Visual marking method and device based on fusion positioning and medium

By integrating positioning technology and GPS/RTK or UWB positioning, the problem of recognition accuracy of airport surface monitoring systems under adverse weather or obstructed conditions has been solved, achieving high-precision real-time monitoring and dynamic adaptability, and improving the reliability of airport surface monitoring and human-computer interaction.

CN121746702APending Publication Date: 2026-03-27浪潮智慧科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing airport surface monitoring systems suffer from decreased recognition accuracy in adverse weather or when objects obstruct the view, have high computing power requirements, poor real-time monitoring adaptability, and difficulty in obtaining accurate geographical location information.

Method used

By employing fusion positioning technology, the system converts camera intrinsic and extrinsic parameters into matrix values ​​and combines this with GPS/RTK or UWB positioning to generate labeled display information in video footage, thereby identifying and labeling occluded targets.

Benefits of technology

It improves the reliability and positioning accuracy of the monitoring system, reduces hardware costs and latency, and enhances system adaptability and human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746702A_ABST
    Figure CN121746702A_ABST
Patent Text Reader

Abstract

The invention discloses a visual marking method and equipment based on fusion positioning, and a medium, belongs to the technical field of airport scene monitoring and management, and aims to solve the problems that the recognition accuracy is reduced, the high computing power requirement is needed, the real-time monitoring adaptability is poor, and the recognition efficiency is high due to the fact that existing airport scene monitoring is easily shielded by vision. And accurate geographical location information is difficult to obtain. The method comprises the following steps: performing coverage calculation of a maximum monitoring area on a visible area on an airport map to obtain a polygonal visible area; acquiring positioning terminal signals of key equipment and workers in an airport range; carrying out conversion processing of related camera point locations on a real-time geographic coordinate system in the real-time position information to obtain a camera coordinate system; performing conversion processing on the camera coordinate system to obtain a two-dimensional image coordinate system; mapping the two-dimensional image coordinate system into a polygonal visible area; and performing shielding labeling display on part of shielding targets in the initial labeling display information to obtain fused labeling display information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of airport surface monitoring and management, and in particular to a visual annotation method, device and medium based on fusion positioning. Background Technology

[0002] As a transportation-intensive location, the safety and efficiency of airport operations directly affect overall operational effectiveness. Current airport surface monitoring primarily relies on the following technologies: (1) pure visual monitoring systems, which capture images through cameras and use computer vision algorithms to identify and track targets; (2) radar- or ADS-B-based surveillance systems, which provide location information but lack visual intuitiveness. However, these existing technologies have significant limitations: 1) Visual occlusion problem: The recognition accuracy of pure vision systems drops significantly in adverse weather conditions (rain, snow, fog) or when objects obstruct each other. For example, when an aircraft engine is partially obstructed by ground service equipment, the vision system may be unable to recognize the engine status, posing a potential safety hazard.

[0003] 2) High computing power requirements: Deep learning-based visual recognition algorithms require a large amount of training data and computing resources, especially when they need to identify various types of devices, vehicles and people, the model complexity increases exponentially.

[0004] 3) Poor adaptability: The airport environment changes dynamically, and the addition of new equipment or temporary changes (such as construction areas) require retraining of the model, resulting in system updates being delayed and unable to meet real-time monitoring requirements.

[0005] 4) Limited accuracy: It is difficult to obtain the precise geographical location information of objects by relying solely on visual recognition, and it cannot be effectively integrated with airport digital twin systems or high-precision maps. Summary of the Invention

[0006] This application provides a visual annotation method, device, and medium based on fusion positioning to solve the following technical problems: existing airport surface monitoring is easily affected by visual occlusion, resulting in a decrease in recognition accuracy, and requires high computing power, has poor real-time monitoring adaptability, and is difficult to obtain accurate geographical location information.

[0007] The embodiments of this application adopt the following technical solutions: On one hand, this application provides a visual annotation method based on fusion positioning, including: calculating the coverage of the maximum monitoring area of ​​the visible area on the airport map based on the cameras pre-deployed in the airport to obtain a polygonal visible area; collecting positioning terminal signals of key equipment and personnel within the airport area and generating real-time location information; converting the real-time geographic coordinate system in the real-time location information according to a unified high-precision map coordinate system to obtain a camera coordinate system; converting the camera coordinate system to obtain a two-dimensional image coordinate system through a camera intrinsic parameter matrix; mapping the two-dimensional image coordinate system onto the polygonal visible area through a perspective projection model to obtain initial annotation display information in the video display screen based on the camera; and displaying partial occlusion targets in the initial annotation display information as occlusion annotations to obtain fused annotation display information in the video display screen.

[0008] This application embodiment, by fusing positioning data, enables the system to correctly label targets in video footage based on their geographical location even when the target is partially or completely obscured, thus improving the reliability of airport surface monitoring. Compared to deep learning-based visual recognition systems, this system can operate stably on a lower-cost hardware platform. Furthermore, employing GPS / RTK or UWB positioning technology, the positioning accuracy of key equipment can reach centimeter level, far exceeding the accuracy of pure visual positioning (typically meter level). Simultaneously, since complex image processing algorithms are not required, the system significantly reduces the latency from acquiring location data to completing labeling, meeting the requirements of real-time airport monitoring. Moreover, newly added equipment only needs to be equipped with a positioning terminal to be identified and tracked by the system, without the need to retrain the visual model, enabling rapid adaptation to dynamic changes in the airport environment and reducing maintenance costs and complexity. Furthermore, monitoring personnel can intuitively see the labeling information of various equipment and vehicles on the video footage, including location, number, and status, reducing workload and improving situational awareness.

[0009] In one feasible implementation, based on the cameras pre-deployed in the airport, the maximum monitoring area coverage calculation is performed on the visible area on the airport map to obtain a polygonal visible area. Specifically, this includes: acquiring and processing the geographical range image of the airport map under the maximum monitoring area based on the parameter information and installation location of the cameras to obtain the maximum monitoring area image; performing ray mapping processing on the airport surface model corresponding to the visible area on the airport map through the four corner points in the maximum monitoring area image to obtain several camera visible area boundary points; and sequentially connecting the camera visible area boundary points to construct the polygonal visible area.

[0010] In one feasible implementation, the method involves collecting positioning terminal signals of key equipment and personnel within the airport area and generating real-time location information. Specifically, this includes: identifying key equipment and personnel within the currently visible airport area based on the electronic fence rules of the polygonal visible area; wherein the key equipment includes at least baggage carts, refueling trucks, and towing vehicles; collecting corresponding equipment positioning terminal signals and personnel positioning terminal signals in real time based on the equipment positioning terminals of the key equipment and the personnel positioning terminals of the personnel; performing location data parsing processing on the equipment positioning terminal signals and personnel positioning terminal signals using the airport's geographic information system to obtain real-time equipment location information and real-time personnel location information; and determining both the real-time equipment location information and the real-time personnel location information as the real-time location information; wherein the real-time location information includes at least latitude and longitude information, speed information, and direction information.

[0011] In one feasible implementation, based on a unified high-precision map coordinate system, the real-time geographic coordinate system in the real-time location information is transformed for camera point locations to obtain the camera coordinate system. Specifically, this includes: based on... The camera coordinate system is obtained. ;in, R is the real-time geographic coordinate system in the real-time location information; R is the rotation matrix in the camera extrinsic matrix; and t is the translation vector in the camera extrinsic matrix.

[0012] In one feasible implementation, the camera coordinate system is transformed according to the camera intrinsic parameter matrix to obtain a two-dimensional image coordinate system. Specifically, this includes: based on... The two-dimensional image coordinate system is obtained. Where K is the camera intrinsic parameter matrix, is the projection transformation matrix in the camera coordinate system.

[0013] In one feasible implementation, the two-dimensional image coordinate system is mapped to the polygonal visible area using a perspective projection model to obtain initial annotation display information in the video display screen based on the camera. Specifically, this includes: constructing a transformation matrix from the world coordinate system to the camera coordinate system based on the coordinates of the four vertices of the polygonal visible area in the world coordinate system and the camera's extrinsic parameter matrix under spatial calibration; multiplying the camera's intrinsic parameter matrix with the transformation matrix to obtain a perspective projection matrix; transforming the world coordinates in the polygonal visible area using the perspective projection matrix to obtain homogeneous coordinates within the two-dimensional image coordinate system; performing perspective division on the homogeneous coordinates and normalizing them to standard device coordinates in the screen coordinate system of the video display screen; and obtaining the initial annotation display information in the video display screen based on the correspondence between the standard device coordinates and the two-dimensional image coordinate system.

[0014] In one feasible implementation, partially occluded targets in the initial annotation display information are annotated to obtain fused annotation display information in the video display. Specifically, this includes: based on device location information and a scene 3D model of the visible area, partially occluded targets in the initial annotation display information are judged for their outline integrity to obtain occlusion result information; if the occlusion result information indicates occlusion, the partially occluded targets are annotated semi-transparently and / or highlighted to generate occlusion annotation display information; the occlusion annotation display information is fused with the initial annotation display information to obtain the fused annotation display information in the video display.

[0015] In one feasible implementation, before calculating the coverage of the maximum monitoring area on the airport map based on the pre-deployed cameras in the airport to obtain the polygonal visible area, the method further includes: performing deployment processing on the cameras under a full-coverage monitoring network based on the visible area on the airport map to obtain the visible area distribution location of the cameras; performing visible range correction processing on each camera based on internal and external parameters, installation location, and orientation angle to obtain camera configuration information for the maximum monitoring area; and generating a camera deployment strategy based on the visible area distribution location of the cameras and the camera configuration information.

[0016] Secondly, embodiments of this application also provide a visual annotation device based on fusion positioning, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to execute a visual annotation method based on fusion positioning as described in any of the above embodiments.

[0017] Thirdly, embodiments of this application also provide a non-volatile computer storage medium, which is a non-volatile computer-readable storage medium storing at least one program, each program including instructions, which, when executed by a terminal, cause the terminal to execute a visual annotation method based on fusion positioning as described in any of the above embodiments.

[0018] This application provides a visual annotation method, device, and medium based on fusion positioning. Compared with the prior art, the embodiments of this application have the following beneficial technical effects: 1. Solve the problem of visual occlusion: By integrating positioning data, even if the target is partially or completely occluded, the system can still correctly mark it in the video frame according to its geographical location, improving the reliability of airport scene monitoring.

[0019] 2. Reduced computational requirements: Eliminating the need for complex visual recognition models reduces reliance on high-performance computing resources such as GPUs. Compared to deep learning-based visual recognition systems, this system can run stably on a lower-cost hardware platform.

[0020] 3. Improved positioning accuracy: By using GPS / RTK or UWB positioning technology, the positioning accuracy of key equipment can reach the centimeter level, which is far higher than the accuracy of pure visual positioning (usually the meter level).

[0021] 4. Enhanced system response speed: Since no complex image processing algorithms are required, the system's delay from acquiring location data to completing annotation is significantly reduced, meeting the requirements of real-time airport monitoring.

[0022] 5. Enhanced system adaptability: Newly added equipment only needs to be equipped with a positioning terminal to be recognized and tracked by the system. There is no need to retrain the visual model, which can quickly adapt to the dynamic changes in the airport environment and reduce maintenance costs and complexity.

[0023] 6. Improve human-computer interaction experience: Monitoring personnel can intuitively see the labeling information of various devices and vehicles on the video screen, including location, number, status, etc., reducing workload and improving situational awareness. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart of a visual annotation method based on fusion localization provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a visual annotation device based on fusion positioning, provided in an embodiment of this application. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0026] This application provides a visual annotation method based on fusion localization, such as... Figure 1 As shown, the visual annotation method based on fusion localization specifically includes steps S101-S106: S101. Based on the cameras pre-deployed in the airport, calculate the coverage of the maximum monitoring area on the visible area of ​​the airport map to obtain the polygonal visible area.

[0027] Specifically, it is necessary to first deploy the cameras under the relevant full-coverage area monitoring network based on the visible area on the airport map, so as to obtain the distribution location of the cameras in the visible area.

[0028] Furthermore, each camera undergoes visual range correction processing based on its internal and external parameters, installation location, and orientation angle to obtain the camera configuration information for the maximum monitoring area.

[0029] Furthermore, a camera deployment strategy is generated based on the distribution of the camera's visible area and the camera configuration information.

[0030] In one embodiment, multiple high-definition cameras are deployed in the visible areas of the airport that need to be monitored, forming a full-coverage surveillance network. Each camera needs to be precisely calibrated to determine its intrinsic and extrinsic parameters (focal length, optical center, distortion coefficient, etc.), as well as its installation location (latitude and longitude) and orientation (yaw, pitch, roll angle).

[0031] Furthermore, by utilizing the camera parameter information and installation location in the camera deployment strategy, the airport map is processed to acquire images of the geographical range under the maximum monitoring area, thus obtaining the image of the maximum monitoring area.

[0032] Furthermore, by using the four corner points in the image of the largest monitoring area, ray mapping is performed on the airport surface model corresponding to the visible area on the airport map to obtain several boundary points of the visible area of ​​the cameras.

[0033] Furthermore, the boundary points of the camera's visible area are connected sequentially to construct a polygonal visible area.

[0034] In one embodiment, the viewport mapping uses a reverse projection method to determine the geographical extent of the camera's viewport. This involves radiating rays from the four corner points of the maximum monitored area image, intersecting them with the airport surface model to obtain the boundary points of the camera's viewport, and connecting these boundary points to form a viewport polygon.

[0035] S102. Collect positioning terminal signals of key equipment and personnel within the airport area and generate real-time location information.

[0036] Specifically, it is also necessary to identify key equipment and personnel within the current visible airport area based on the electronic fence rules of the polygonal visible area. Key equipment includes at least: baggage carts, refueling trucks, and towing vehicles.

[0037] Furthermore, based on the equipment positioning terminals of key equipment and the personnel positioning terminals of staff, the corresponding equipment positioning terminal signals and personnel positioning terminal signals are collected in real time.

[0038] Furthermore, it is necessary to use the airport's geographic information system to perform location data parsing and processing on the equipment positioning terminal signals and personnel positioning terminal signals to obtain real-time equipment location information and real-time personnel location information.

[0039] Furthermore, both real-time device location information and real-time personnel location information are defined as real-time location information. This real-time location information includes at least: latitude and longitude information, speed information, and direction information.

[0040] As a feasible implementation method, key equipment (baggage carts, refueling trucks, towing vehicles, etc.) and personnel within the airport area are equipped with positioning terminals (GPS / RTK or UWB) to upload location data (latitude, longitude, altitude, speed, direction, etc.) in real time.

[0041] S103. Based on the unified high-precision map coordinate system, the real-time geographic coordinate system in the real-time location information is transformed for the relevant camera points to obtain the camera coordinate system.

[0042] It should be noted that the geographical coordinates in the physical world are converted into pixel coordinates in the camera image coordinate system. This conversion is achieved through a perspective projection model, with the core formula as follows: [u,v,1] = K * [R|t] * [X,Y,Z,1] Where [u,v] are the image pixel coordinates, K is the camera intrinsic parameter matrix, [R|t] is the extrinsic parameter matrix (rotation and translation), and [X,Y,Z] are the geographic coordinates.

[0043] Specifically, coordinate system design and transformation are fundamental to this application. Airports need to establish a unified high-precision map coordinate system (such as the UTM coordinate system), with the positions of all cameras and equipment based on this coordinate system. Then, coordinate transformation is performed. In the transformation from geographic coordinates to the camera coordinate system: points in the world coordinate system are transformed to the camera coordinate system using camera extrinsic parameters (rotation matrix R and translation vector t), that is: according to... Obtain the camera coordinate system .in, R is the real-time geographic coordinate system in the real-time location information; R is the rotation matrix in the camera extrinsic matrix; t is the translation vector in the camera extrinsic matrix.

[0044] S104. Using the camera intrinsic parameter matrix, the camera coordinate system is transformed to obtain the two-dimensional image coordinate system by converting the pixel positions of the two-dimensional image.

[0045] Specifically, in the transformation from the camera coordinate system to the image coordinate system: using the camera intrinsic parameter matrix K, the 3D camera coordinates are projected onto the 2D image coordinates, that is, according to... To obtain the two-dimensional image coordinate system Where K is the camera intrinsic parameter matrix, This is the projection transformation matrix in the camera coordinate system.

[0046] S105. By using the perspective projection model, the two-dimensional image coordinate system is mapped onto the polygonal visible area to obtain the initial annotation display information in the video display screen based on the camera.

[0047] Specifically, based on the coordinates of the four vertices of the visible polygon region in the world coordinate system and the external parameter matrix of the camera under spatial calibration, a transformation matrix from the world coordinate system to the camera coordinate system is first constructed.

[0048] Furthermore, the internal parameter matrix of the camera is multiplied with the transformation matrix to obtain the perspective projection matrix.

[0049] As a feasible implementation, a transformation matrix [R|t] from the world coordinate system to the camera coordinate system is constructed based on the coordinates of the four vertices of the polygonal visible region in the world coordinate system and the camera's extrinsic parameters (rotation matrix R and translation vector t) obtained through spatial calibration. Multiplying the intrinsic parameter matrix K by [R|t] yields the complete perspective projection matrix P.

[0050] Furthermore, the world coordinates in the visible polygon area are transformed using a perspective projection matrix to obtain homogeneous coordinates in the two-dimensional image coordinates.

[0051] Furthermore, the homogeneous coordinates are subjected to perspective division and normalized to standard device coordinates in the screen coordinate system of the video display.

[0052] As a feasible implementation method, the world coordinates of the polygonal visible area are transformed using a perspective projection matrix P to obtain its homogeneous coordinates in a two-dimensional image coordinate system. Then, perspective division is performed on the homogeneous coordinates to normalize them to standard device coordinates in the screen coordinate system, thereby accurately determining the boundary of the polygonal visible area in the video display.

[0053] Furthermore, based on the correspondence between the standard device coordinates and the two-dimensional image coordinate system, the initial annotation display information in the video display is obtained.

[0054] As a feasible implementation, for initial annotation information (such as the center point (u, v) of a target box) that needs to be displayed within the visible area of ​​the polygon, the coordinates of the two-dimensional image are back-projected onto the world coordinate system plane defined by the visible area of ​​the polygon using the inverse of the perspective projection matrix P or by solving for it. Finally, based on the coordinates obtained from the back projection and combined with a preset annotation style (such as box size and arrow length), a complete annotation geometry is generated in the world coordinate system, and then reprojected onto the two-dimensional image using the perspective projection matrix P to form initial annotation display information consistent with the scene perspective.

[0055] S106. Occlusion marking is performed on some of the occluded targets in the initial annotation display information to obtain the fused annotation display information in the video display.

[0056] Specifically, by using the device location information and the scene 3D model of the visible area, partial occlusion judgment is made on the marked targets in the initial annotation display information based on the integrity of the target outline, and occlusion result information is obtained.

[0057] Furthermore, if the occlusion result information indicates that an occlusion state exists, then the partially occluded target will be displayed in a semi-transparent annotation and / or outline highlighting mode to generate occlusion annotation display information.

[0058] Furthermore, the occlusion label display information is fused with the initial label display information to obtain the fused label display information in the video display.

[0059] In one embodiment, it can also be determined whether the target is occluded based on the device position and the 3D model of the scene within the visible area. For partially occluded targets, a semi-transparent label or outline highlighting method is used to display them, ensuring that the operator can perceive the existence of the occluded device.

[0060] In addition, embodiments of this application also provide a visual annotation device based on fusion positioning, such as... Figure 2 As shown, the visual annotation device 200 based on fusion positioning specifically includes: At least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein the memory 202 stores instructions executable by the at least one processor 201 to enable the at least one processor 201 to execute: Based on the cameras pre-deployed in the airport, the maximum coverage area of ​​the visible area on the airport map is calculated to obtain the polygonal visible area; Collect positioning terminal signals from key equipment and personnel within the airport area and generate real-time location information; Based on the unified high-precision map coordinate system, the real-time geographic coordinate system in the real-time location information is transformed for the relevant camera points to obtain the camera coordinate system; By using the camera intrinsic parameter matrix, the camera coordinate system is transformed to obtain the two-dimensional image coordinate system by transforming the camera coordinate system for the pixel positions of the two-dimensional image. By using a perspective projection model, the two-dimensional image coordinate system is mapped onto the polygonal visible area to obtain the initial annotation display information in the video display based on the camera; The partially occluded targets in the initial annotation display information are further occluded and annotated to obtain the fused annotation display information in the video display.

[0061] This application, by fusing location data, enables the system to correctly label targets in video footage based on their geographical location even when the targets are partially or completely obscured, thus improving the reliability of airport surface monitoring. Compared to deep learning-based visual recognition systems, this system can operate stably on a lower-cost hardware platform. Furthermore, employing GPS / RTK or UWB positioning technology, the positioning accuracy of key equipment can reach centimeter level, far exceeding the accuracy of pure visual positioning (typically meter level). Simultaneously, since complex image processing algorithms are not required, the system significantly reduces the latency from acquiring location data to completing labeling, meeting the requirements of real-time airport monitoring. Moreover, newly added equipment only needs to be equipped with a positioning terminal to be identified and tracked by the system, eliminating the need to retrain the visual model. This allows for rapid adaptation to dynamic changes in the airport environment, reducing maintenance costs and complexity. Furthermore, monitoring personnel can intuitively see the labeling information of various equipment and vehicles on the video footage, including location, number, and status, reducing workload and improving situational awareness.

[0062] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0063] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0065] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0068] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0069] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0070] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0071] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0072] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of this specification.

Claims

1. A visual annotation method based on fusion localization, characterized in that, The method includes: Based on the cameras pre-deployed in the airport, the maximum coverage area of ​​the visible area on the airport map is calculated to obtain the polygonal visible area; Collect positioning terminal signals from key equipment and personnel within the airport area and generate real-time location information; Based on the unified high-precision map coordinate system, the real-time geographic coordinate system in the real-time location information is transformed for the relevant camera points to obtain the camera coordinate system; By using the camera intrinsic parameter matrix, the camera coordinate system is transformed for the pixel positions of the two-dimensional image to obtain the two-dimensional image coordinate system. By using a perspective projection model, the two-dimensional image coordinate system is mapped onto the polygonal visible area to obtain the initial annotation display information in the video display image based on the camera; The partially occluded targets in the initial annotation display information are occluded and annotated to obtain the fused annotation display information in the video display screen.

2. The visual annotation method based on fusion localization according to claim 1, characterized in that, Based on the pre-deployed cameras in the airport, the maximum coverage area of ​​the visible area on the airport map is calculated to obtain the polygonal visible area, which specifically includes: Based on the parameter information and installation location of the camera, the airport map is processed to acquire images of the geographical range under the maximum monitoring area, thus obtaining the image of the maximum monitoring area. By using the four corner points in the image of the maximum monitoring area, ray mapping is performed on the airport surface model corresponding to the visible area on the airport map to obtain several camera visible area boundary points. The polygonal visible area is constructed by sequentially connecting the boundary points of the camera's visible area.

3. The visual annotation method based on fusion localization according to claim 1, characterized in that, Collect positioning terminal signals from key equipment and personnel within the airport area and generate real-time location information, specifically including: Based on the electronic fence rules of the polygonal visible area, key equipment and personnel within the current visible airport range are identified; wherein, the key equipment includes at least: baggage carts, refueling trucks, and towing vehicles; Based on the equipment positioning terminal of the key equipment and the personnel positioning terminal of the staff, the corresponding equipment positioning terminal signals and personnel positioning terminal signals are collected in real time. The location data of the equipment positioning terminal signal and the personnel positioning terminal signal are parsed and processed by the airport's geographic information system to obtain the real-time equipment location information and the real-time personnel location information. The real-time device location information and the real-time personnel location information are both determined as the real-time location information; wherein, the real-time location information includes at least: latitude and longitude information, speed information, and direction information.

4. The visual annotation method based on fusion localization according to claim 1, characterized in that, Based on a unified high-precision map coordinate system, the real-time geographic coordinate system in the real-time location information is transformed for the relevant camera points to obtain the camera coordinate system, specifically including: according to The camera coordinate system is obtained. ;in, R is the real-time geographic coordinate system in the real-time location information; R is the rotation matrix in the camera extrinsic matrix; and t is the translation vector in the camera extrinsic matrix.

5. The visual annotation method based on fusion localization according to claim 1, characterized in that, By using the camera intrinsic parameter matrix, the camera coordinate system is transformed for the relevant two-dimensional image pixel positions to obtain a two-dimensional image coordinate system, specifically including: according to The two-dimensional image coordinate system is obtained. Where K is the camera intrinsic parameter matrix, is the projection transformation matrix in the camera coordinate system.

6. The visual annotation method based on fusion localization according to claim 1, characterized in that, By using a perspective projection model, the two-dimensional image coordinate system is mapped onto the polygonal visible area to obtain initial annotation display information in the video display image based on the camera, specifically including: Based on the coordinates of the four vertices of the polygonal visible area in the world coordinate system, and based on the external parameter matrix of the camera under spatial calibration, a transformation matrix from the world coordinate system to the camera coordinate system is constructed. The perspective projection matrix is ​​obtained by multiplying the camera's internal parameter matrix with the transformation matrix. The perspective projection matrix is ​​used to transform the world coordinates in the visible area of ​​the polygon to obtain the homogeneous coordinates in the two-dimensional image coordinates. The homogeneous coordinates are subjected to perspective division and normalized to standard device coordinates in the screen coordinate system of the video display. Based on the correspondence between the standard device coordinates and the two-dimensional image coordinate system, the initial annotation display information in the video display screen is obtained.

7. The visual annotation method based on fusion localization according to claim 1, characterized in that, The partially occluded targets in the initial annotation display information are annotated to obtain the fused annotation display information in the video display image, specifically including: Based on the device location information and the scene 3D model of the visible area, the marked targets in the initial annotation display information are partially occluded under the condition of target outline integrity, and the occlusion result information is obtained. If the occlusion result information indicates that there is an occlusion state, then the partially occluded target will be displayed in a semi-transparent marking and / or outline highlighting mode to generate occlusion marking display information; The occlusion label display information is fused with the initial label display information to obtain the fused label display information in the video display.

8. The visual annotation method based on fusion localization according to claim 1, characterized in that, Before calculating the coverage of the maximum surveillance area on the airport map based on the pre-deployed cameras in the airport to obtain the polygonal visible area, the method further includes: Based on the visible area on the airport map, the cameras are deployed under the relevant full-coverage area monitoring network to obtain the visible area distribution location of the cameras. Each camera is corrected for its internal and external parameters, installation position, and orientation angle to obtain camera configuration information for the maximum monitoring area. A camera deployment strategy is generated based on the location of the camera's visible area and the camera's configuration information.

9. A visual annotation device based on fusion positioning, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform a visual annotation method based on fusion localization according to any one of claims 1-8.

10. A non-volatile computer storage medium, characterized in that, The storage medium is a non-volatile computer-readable storage medium that stores at least one program, each program including instructions that, when executed by a terminal, cause the terminal to perform a visual annotation method based on fusion localization according to any one of claims 1-8.