Positioning method based on unmanned aerial vehicle image and infrared two-dimensional code

By using infrared QR code markers and satellite positioning modules on the drone, the high-precision positioning problem of infrared images in complex environments is solved, and the precise positioning and measurement of infrared heat source targets is achieved.

CN120544072APending Publication Date: 2025-08-26SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510583190.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In complex environments, the imaging quality of color images is insufficient, which makes it difficult for traditional visible light images to meet the needs of object detection and precise positioning. Infrared images lack texture information and spatial reference, limiting high-precision positioning applications.

Method used

Infrared QR code markers are used, and the radiation characteristics of low emissivity materials are used in the infrared band. Combined with the satellite positioning module, infrared image data is obtained through drones, image enhancement processing is performed, QR code features are identified, pixel-geographically spatial conversion is established, and high-precision target positioning is achieved.

Benefits of technology

The precise positioning of infrared heat source targets is achieved under complex conditions, improving the safety, real-time and accuracy of measurement, and providing a convenient measurement system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544072A_ABST
    Figure CN120544072A_ABST
Patent Text Reader

Abstract

The invention discloses a positioning method based on an unmanned aerial vehicle image and an infrared two-dimensional code, which can effectively solve the problem of unmanned aerial vehicle visual angle target positioning under the condition that a color mode is limited. The method comprises the following steps: S1, selecting an infrared low-emissivity material to manufacture a two-dimensional code marker, and embedding a satellite positioning module into the two-dimensional code marker; s2, acquiring an infrared image of an experimental site by using an unmanned aerial vehicle, and performing enhancement processing; s3, detecting whether a marker exists in the image frame or not, and extracting positioning information; s4, establishing a local coordinate system by using the positioning information, and calculating a pixel-distance proportionality coefficient and a marker deflection angle; s5, segmenting a target area by using a detail-enhanced attention network, and representing the position of a target point by using a geometric centroid; and S6, calculating a coordinate value based on a local coordinate system, and resolving a global coordinate through a satellite positioning module on the marker. Compared with the prior art, the cooperative two-dimensional code is introduced into the infrared mode for the first time, and the method has the advantages of low cost, light weight and the like, and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV remote sensing image processing and visual positioning, and specifically relates to a positioning method based on UAV images and infrared two-dimensional codes. Background Art

[0002] As a product of the convergence of aviation and remote sensing technologies, drones have become a widely used sensor-carrying platform across various fields due to their high flexibility and low operating costs. Optical sensors onboard drones can efficiently and cost-effectively acquire information on the spatial structure and attitude changes of ground targets. In recent years, with the continuous advancement of flight control and remote sensing technologies, drone technology has been widely applied in smart agriculture, intelligent inspections, military reconnaissance, and logistics and transportation.

[0003] However, in practical applications, complex environmental factors often severely impact the quality of color images, making traditional visible light images inadequate for target detection and precise positioning. This is particularly true in low-light or heavily obscured smoke and dust scenarios, such as nighttime patrols, fire rescue, and forest fire monitoring. Color modality images often struggle to accurately capture target features due to insufficient brightness, reduced contrast, or visual occlusion, impacting the reliability of perception systems and positioning accuracy.

[0004] Infrared imaging technology detects thermal radiation from an object's surface. It has the advantages of being independent of external light sources and capable of penetrating some obstructions. It can stably capture target features with significant temperature differences even in low-light conditions and smog, making it an important supplementary tool for target identification and positioning in complex environments. However, infrared images often lack rich texture and color information, blurring target boundaries and limiting positioning accuracy. Furthermore, the lack of an effective spatial reference calibration system for the infrared modality limits its application in high-precision spatial positioning tasks.

[0005] Therefore, this paper proposes a positioning method based on drone imagery and infrared QR codes. This method utilizes a drone-mounted infrared camera to perceive the target scene, improving the safety, real-time performance, and accuracy of measurement. The use of infrared-recognizable QR codes as a reference for the conversion between infrared pixel space and real physical space ensures the convenience and speed of the measurement system. Summary of the Invention

[0006] This invention aims to provide a positioning method based on drone imagery and infrared QR codes. This method is suitable for environments with limited color vision modalities and can accurately locate infrared heat source targets under complex conditions. Based on the principle of thermal radiation from objects and the weak radiation properties of low-emissivity materials in the infrared band, this method designs infrared-recognizable QR code markers. This system then constructs a drone remote sensing positioning system based on infrared image recognition and pixel-to-geospace conversion, enabling high-precision, non-contact, real-time target measurement under complex conditions.

[0007] The specific plan is as follows:

[0008] A positioning method based on drone images and infrared QR codes is used to perform infrared heat source remote sensing positioning using drones under color modality-restricted conditions. The method includes the following steps:

[0009] Step S1: constructing an infrared QR code marker, wherein the infrared QR code marker is made of a low-emissivity material, the QR code pattern has structural features that can be recognized by a machine vision system, and the center integrates a satellite positioning module for providing real-time geographic location information;

[0010] Step S2: acquiring real-time infrared image data of the target scene through the UAV device and the wireless data transfer module, and performing image enhancement processing on the data;

[0011] Step S3: Detect whether there is an infrared QR code marker in the image frame. If a marker is detected, extract its corner point position, orientation information and other features, and proceed to step S4; if not, repeat step S3 and continue processing the next frame;

[0012] Step S4: Based on the extracted QR code feature information, a local coordinate system with the marker as the origin is established, and the pixel-distance ratio coefficient and the deflection angle of the QR code are calculated;

[0013] Step S5: Segment the target area using the attention network based on detail enhancement, and use the geometric center of the segmented area as the image pixel position of the target;

[0014] Step S6: Calculate the coordinate position of the target in the local coordinate system based on the pixel-distance ratio coefficient and the deflection angle information, and calculate the global geographic location of the target in combination with the satellite positioning data of the center of the marker.

[0015] Furthermore, the infrared two-dimensional code marker in step S1 is an integral device composed of a low-emissivity material and designed according to the two-dimensional code pattern. The low-emissivity material has a low radiant flux and dark characteristics in the thermal infrared band of 8-14 μm, and the infrared emissivity ε is less than 0.1, which appears as a low-brightness area with approximately zero grayscale in infrared imaging. Preferably, the low-emissivity material is a metal-plated polymer film. The two-dimensional code pattern appears in black and white in visible light mode, and the arrangement of its coding pattern is a fixed format, with structural features that can be decoded by a standard two-dimensional code recognition algorithm. Preferably, the two-dimensional code pattern adopts ArUcoMarker, which can be recognized and detected by the built-in library of OpenCV. The infrared two-dimensional code marker adopts the above-mentioned low-emissivity material and is integrated into an integral structure through an integrated process according to the shape of the two-dimensional code pattern. From the perspective of the drone, the infrared two-dimensional code marker presents a clear pixel intensity boundary, which can be recognized and detected by the built-in library. The actual circumference of the infrared two-dimensional code marker is accurately measured during the production stage and is recorded as L. In order to achieve the global geographic coordinate solution of the target and avoid the need to repeatedly measure the center position of the marker for each deployment, a satellite positioning module is integrated into the center of the QR code marker to obtain the real-time geographic location of the infrared marker. Preferably, the satellite positioning module uses Ublox ZED-F9PRTK real-time dynamic differential positioning, which has centimeter-level accuracy and can provide real-time global position of the infrared QR code marker. where λ O is the measured real-time longitude of the marker, The measured real-time latitude of the marker.

[0016] Furthermore, in step S2, the drone device is equipped with an infrared imaging device and a wireless communication device for real-time acquisition and transmission of infrared image stream data. Preferably, the drone device uses a DJIM30T, equipped with a 640×512 resolution infrared imaging device and a DJIRC Plus remote control terminal, which can realize long-distance real-time image transmission and data return functions. The wireless data flow module includes a remote control terminal, a data acquisition card, and a graphics processing terminal, wherein the remote control terminal is responsible for remote command and data relay; the data acquisition card is responsible for receiving image stream data from the remote control terminal and transmitting it to the graphics processing terminal for preprocessing to support subsequent detection and analysis operations. Preferably, the data acquisition card uses the UGREEN CM410 data acquisition card, and the remote control terminal and the data acquisition card communicate via the HDMI protocol, and the data acquisition card and the graphics processing terminal transmit data via the USB interface. In order to make the markers in the real-time infrared image stream present more significant pixel intensity contrast, so that the built-in algorithm library can identify the markers and obtain positioning auxiliary information, the graphics processing terminal preprocesses the received video data to support subsequent target detection and analysis. Image enhancement processing includes image dedistortion and contrast enhancement. Dedistortion uses the Zhang Zhengyou calibration method combined with the Brown-Conrady model to correct barrel and pincushion distortions, ensuring geometric consistency of the image frame. First, the Zhang Zhengyou calibration method is used to obtain the infrared camera's intrinsic parameter matrix and distortion parameters. Then, radial and tangential distortion terms are calculated based on the Brown-Conrady model. Distortion correction is performed using a back-mapping interpolation algorithm to ensure geometric consistency of the image frame and improve subsequent detection accuracy. Contrast enhancement combines histogram equalization with adaptive contrast enhancement to improve both local and global contrast in infrared images, thereby enhancing the visibility of infrared QR codes and targets. Histogram equalization is first performed on the global pixel distribution to enhance the overall dynamic range of brightness. Adaptive contrast adjustment is then applied to local regions. The image is segmented into multiple sub-blocks, and contrast-constrained histogram equalization is performed on each sub-block to prevent excessive noise enhancement while improving the contrast between the infrared marker and the background. Finally, smoothing is performed using Gaussian filtering to reduce image noise and optimize the quality of the enhanced infrared image.

[0017] Furthermore, in step S3, the image frame enhanced by step S2 is subjected to recognition and detection of infrared QR code markers using the OpenCV algorithm library. The image frame after image enhancement is recognized by the visual recognition system. If the QR code pattern is not detected in the current image frame, the system determines that the frame is not recognized successfully, automatically continues to process the next frame, and repeats the detection process; if the QR code pattern is detected, it is determined that the recognition is successful, and the key feature information of the QR code is extracted. The visual recognition system is constructed based on the QR code pattern algorithm library; the feature information includes: the pixel position of the corner point (u p1 ,vp1 ),(u p2 ,v p2 ),(u p3 ,v p3 ),(u p4 ,v p4 ), the marker's deflection angle α, the marker's pixel perimeter l p .

[0018] Furthermore, the pixel-distance ratio coefficient k in step S4 is calculated by the actual perimeter of the marker and the pixel perimeter, which reflects the scale mapping relationship between the infrared image space and the physical world space. p The calculation formula is the actual perimeter L of the marker in step S1:

[0019]

[0020] The scale factor k is used to establish a local coordinate system that reflects the relationship between the image pixels and the actual distance to the ground. The global coordinates of the target can then be obtained through the satellite positioning module built into the infrared QR code marker. The local coordinate system is used to describe the position relationship of the target relative to the infrared marker, including the coordinate origin and the x-axis, y-axis direction, the pixel position of the coordinate origin (u OL ,v OL ) Select the geometric center of the infrared two-dimensional code marker (the position is determined by calculating the pixel coordinates of the corner point extracted in step S3), and its calculation formula is:

[0021]

[0022] The y-axis is the marker's orientation, deflected by an angle α (i.e., the marker's positive direction) as described in step S3, and the x-axis is the y-axis rotated 90° clockwise. Using the pixel-to-distance ratio k and the parameters of the local coordinate system, the real-space positional relationship between the infrared image pixel and the infrared QR code marker can be calculated.

[0023] Furthermore, the detail-enhanced attention network in step S5 is a convolutional neural network for semantic segmentation tasks, which sequentially includes a convolutional layer, a downsampling module, a fusion module, and a detail-enhanced attention block. The convolutional layer uses a 3×3 convolution kernel and is placed before the detail-enhanced attention block to extract basic features of the image. The downsampling module is composed of a convolutional layer with a stride of 2, which is used to downsample the input feature map by two times. The fusion module is composed of a channel-level splicing and a convolutional layer cascade. First, the feature maps of the upsampling and downsampling stages are channel-spliced, and then the shallow and deep features are aggregated through the convolution layer. The detail-enhanced attention block is the main feature extraction module of the detail-enhanced attention network, including a pixel addition module, a differential convolution block, and an attention block. The pixel addition module adds two feature maps of the same shape pixel by pixel according to pixel position to perform feature fusion. The differential convolution block extracts brightness-gradient level features and includes a convolutional layer, a pixel addition module, and a differential convolution. Among them, differential convolution is an improved convolution operation. The convolution kernel is modified according to the operator in the field of machine vision, and the extraction of image gradient information is introduced into the traditional convolution operation. The differential convolution block inputs the feature map, enters the convolution layer branch and the differential convolution branch respectively, and then passes the two features into the pixel addition module to extract the brightness-gradient level features. The attention block extracts image features and adaptively weights the importance of feature channels. Among them, the average pooling operation calculates the average value of each channel feature to generate channel attention weights. The channel multiplication module multiplies this weight with the original feature map channel by channel to achieve weighted enhancement of important channels. The attention block uses convolutional layers to extract features, then uses average pooling to extract channel-level attention weights from the feature map, which are then weighted using a channel-wise multiplication module. The detail enhancement attention block first passes the input feature map through a differential convolution block to extract brightness-gradient features. This is then summed pixel-by-pixel with the original input feature map to preserve gradient information. The fused features are then extracted through convolutional layers, and the attention block performs channel-wise weighting. Finally, the output feature map is summed with the original input map to further enhance key details and maintain gradient stability. The overall structure of the detail enhancement-based attention network is as follows: the input infrared image first enters the convolutional layer to extract shallow features, followed by feature extraction and semantic enhancement through the stacked detail enhancement attention blocks. A downsampling module then extracts 2x downsampled features, and the above process is repeated to extract 4x downsampled features. The 4x downsampled features are then further processed through convolution and the detail enhancement attention block. These features are then merged with the downsampled feature map through a fusion module and restored to 2x resolution through an upsampling module. Finally, they are processed through the stacked detail enhancement attention blocks and fused with the 2x downsampled features to restore to the original resolution. The final output is the semantic segmentation result map, where the target area corresponds to the pixel area output by the semantic segmentation network, and the pixel position of the target (u E ,v E) is determined by the geometric center of the output region and is defined as the geometric center of the pixels in that region.

[0024] Furthermore, the pixel position (u E ,v E ) and the pixel-distance ratio k from step S4 are used to calculate the real-world distance between the target and the marker. Using the local coordinate system and its x- and y-axis directions established in step S5, the projected lengths of the real-world distance between the target and the marker in the x- and y-axis directions of the local coordinate system are calculated, respectively, to obtain the relative coordinate position of the target in the local coordinate system of the marker. Finally, the real-world geographic location of the infrared marker obtained by the satellite positioning module in step S1 and the relative coordinates of the target in the local coordinate system are combined to calculate the geographic coordinates of the target in real geographic space.

[0025] In step S4, the local coordinate system parameters are defined, including the pixel position of the coordinate origin (u OL ,v OL ) and the x-axis, y-axis direction, the Euclidean distance Dist between the target and the coordinate origin in the pixel space p for:

[0026]

[0027] The angle β between the line connecting the target and the coordinate origin and the u-axis of the pixel coordinate system is:

[0028]

[0029] Based on this, the projection of the line connecting the target and the coordinate origin on the x and y axes of the local coordinate system can be calculated:

[0030]

[0031] Among them, Proj x 、Proj y Respectively represent its projection on the x and y axes of the pixel space. Then, according to the pixel-distance ratio relationship k, the local coordinates (x L ,y L ):

[0032]

[0033] Finally, according to the data of the satellite positioning module in the infrared two-dimensional code marker in step S1 Determine the global world position of the target. To facilitate global positioning calculation, preferably, the infrared QR code marker in step S1 is facing north. L,y L ), its offset in the east-west longitude direction:

[0034]

[0035] Where R is the radius of the Earth, approximately 6378 km. The offset in the north-south latitude direction is:

[0036]

[0037] Therefore, the longitude and latitude of the global world position of the target can be obtained

[0038]

[0039] Beneficial effects of the present invention:

[0040] 1. This invention applies QR code markers in visible light mode to infrared image measurement and positioning tasks for the first time. It is used for measurement and positioning tasks under color modality-restricted conditions, which is conducive to solving the problem of multi-target perception in complex working environments and ensuring the speed, accuracy and safety of measurement tasks.

[0041] 2. This paper designs an attention network based on detail enhancement to detect target areas, introduces differential convolution to fully utilize the gradient features of infrared images, combines the channel attention mechanism to screen key feature patterns, and makes a lightweight design for edge devices with limited computing power.

[0042] 3. The present invention constructs a global coordinate positioning system based on infrared two-dimensional codes, providing a complete solution for visual remote sensing measurement systems under color modality-restricted conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a specific implementation flow chart of the positioning method described in the present invention.

[0044] Figure 2 The present invention relates to the production process of infrared ArUcoMarker and the effect display of color and infrared modes.

[0045] Figure 3 This is the attention network structure diagram based on detail enhancement proposed in this invention.

[0046] Figure 4 It is the principle of using the information provided by the infrared ArUco Marker to locate the drone. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0048] like Figure 1 The figure shows an execution flow chart of an embodiment of the present invention. First, an infrared QR code marker is made, and a low-emissivity material with an infrared emissivity ε<0.1 and emitting low-flux thermal infrared radiation is used as the substrate. The two-dimensional structural pattern of the ArUco Marker is used for sewing to form a graphic mark that can be clearly identified in the infrared image. A satellite positioning module is embedded in the geometric center of the marker to provide a global geographic location reference. Secondly, the internal and external parameters of the drone's onboard camera are calibrated using the Zhang Zhengyou calibration method in the offline stage, and the infrared image obtained in the subsequent steps is dedistorted using the dedistortion parameters. Subsequently, the infrared QR code marker is placed in the experimental site, and the infrared image stream data of the training ground is collected in real time using the drone. The collected data is read using the Python OpenCV library, and the ArUcoMarker library is used to detect whether there is a marker in the image. If no marker is detected in the current image, the current frame image is skipped; if a marker is detected, the coordinates of its four positioning corner points are extracted, and a local coordinate system is established based on this. Next, a detection algorithm is used to detect whether the target region exists in the image. If no target region is detected, the process returns to the marker detection process and continues to detect the next frame. If the target region is detected, the corresponding image region is extracted and the geometric center coordinates of this region are calculated. Finally, the local coordinate system parameters are used to calculate the physical coordinates relative to the marker. Combined with the global latitude and longitude data provided by the satellite positioning module in the QR code marker, the global geographic coordinates of the target are finally obtained, completing accurate target positioning.

[0049] like Figure 2 As shown in the figure, it is a schematic diagram of the structure of the infrared two-dimensional code marker used in an embodiment of the present invention. The marker is constructed into a specific pattern from a low-emissivity material to ensure its recognizability in infrared images. In this embodiment, the ArUcoMarker two-dimensional code is selected. The two-dimensional code was originally proposed by S. Garrido-Jurado et al. from the University of Cordoba in Spain. It is a two-dimensional visual identification system based on binary coding design. Its core structure consists of a clear black border and an internal binary coding matrix, which can effectively increase the amount of information required when mapping a two-dimensional image to a three-dimensional space, and is suitable for visual recognition and positioning scenarios. According to the theory of thermal radiation, all objects in nature with a temperature above absolute zero will generate electromagnetic wave radiation through molecular thermal motion, which is called thermal radiation. Its radiation spectrum characteristics follow Planck's blackbody radiation law.

[0050] Under normal ambient temperature conditions, thermal radiation energy is mainly concentrated in the mid-to-far infrared band of 8-14μm, which coincides with the infrared transmission window of the atmosphere, providing good observation conditions for thermal infrared imaging systems. Infrared low-emissivity materials are special materials with low radiation emissivity in the infrared band. Their thermal radiation is almost imperceptible to infrared sensors during imaging and usually appears as a black area on infrared images. After using this material to construct a QR code marker, the QR code pattern can be clearly displayed in the infrared image, realizing the QR code positioning and recognition function under infrared mode. In this embodiment, a waterproof material is used to sew the low-emissivity infrared ArUco pattern into a whole, forming a large-size infrared QR code marker with a side length of about 3m. Experiments show that in infrared images of drones flying at an altitude greater than 100m, the infrared ArUco pattern can still be accurately recognized by the Opencv library.

[0051] The parameter calibration and image dedistortion method used in this embodiment is based on Zhang Zhengyou's calibration method and the Brown-Conrady model. Camera calibration is to establish the relationship between the camera image pixel position and the scene point position. According to the camera imaging model, it is necessary to calibrate the external parameters including the rotation matrix and translation vector, and the internal parameters including the focal length, principal point, tilt factor, and distortion coefficient. Since the actual camera lens has different degrees of distortion, the pixel coordinates calculated by the perspective projection model will be distorted, resulting in a plane coordinate offset. The main distortion error is caused by the radial distortion of the lens. The model description coordinates are different from the real coordinates (u', v', 1) T The relationship is as follows:

[0052]

[0053] where δ u ,δ v is the offset value in the u,v direction, which can be described as:

[0054]

[0055] Where (u0, v0) is the coordinate of the center of the pixel plane, r is the distance from the real coordinate to the plane center, and k1, k2, ... are the coefficients of radial distortion deviation. Two values ​​are usually used for ease of calculation. The internal and external parameters determined using Zhang Zhengyou's calibration method can correct the distorted image captured by the camera. During the calibration process, a black and white checkerboard image is selected as the calibration image. Calibration images are captured at multiple angles using an infrared camera mounted on a drone. After extracting feature corners, the OpenCV calibration tool is used to solve the parameters and output the camera's intrinsic and extrinsic parameters. During use, the calibration parameter values ​​are input, and the aforementioned principle can be used to eliminate camera distortion, thereby preventing the problem of camera lens distortion causing the pixel distribution of the captured image to differ from the true perspective. In a specific example of the present invention, a 9×6 calibration checkerboard with an actual size of 90 cm × 60 cm is used to obtain the infrared camera's internal and extrinsic parameters. Image distortion elimination is achieved using the OpenCV library and a camera transmission imaging model.

[0056] The satellite positioning module used in this example is the Ublox ZED-F9P RTK real-time dynamic differential positioning module, which provides the system with centimeter-level global positioning information for infrared QR code markers. The drone uses a DJI M30T equipped with a 640×512 resolution infrared imaging device and a DJI RC Plus remote control terminal for long-distance real-time image transmission and control. The data acquisition card uses the UGREEN CM410 data acquisition card. The remote control terminal and the data acquisition card communicate using the HDMI protocol, while the data acquisition card and the graphics processing terminal communicate using a USB interface. The graphics processing terminal is a computer equipped with an RTX 4090 graphics computing unit, an Intel I913900K CPU, and a Windows 10 operating system. During the operation of this method, the drone flies to the experimental site, switches to infrared mode, and transmits the real-time infrared image stream to the graphics processing terminal via the image transmission system. Simultaneously, the position information of the satellite positioning module is sent to the graphics processing terminal via the serial port for calculating the global position of the target. The implementation steps involved in this embodiment use Python language to write scripts, use pyserial library and OpenCV library to read sensor data, and use the network written by pytorch library to realize the detection of target area. Figure 4 The mathematical model shown solves the local relative position and uses the satellite positioning module to obtain the global position of the target.

[0057] The detection network proposed in this embodiment adopts an attention mechanism network based on detail enhancement, and its overall structure is as follows: Figure 3As shown. The network is a U-Net-like architecture that realizes the extraction of shallow image features and deep image features, and gradually upsamples the low-resolution deep features, fuses them with shallow detail features, and embeds rich semantic information into shallow features with detailed texture and spatial position, thereby obtaining the semantic segmentation results of the target area. Due to the large difference between the target intensity and background in the infrared modality, the semantic depth required for the segmentation task is significantly reduced compared to the requirements in traditional natural environments. The gradient of the regional edge image changes significantly, and the foreground and background segmentation mainly relies on shallow edge features. Therefore, the U-Net structure network selected in this embodiment only performs two downsamplings, ensuring the segmentation accuracy while ensuring the lightweight model. Furthermore, the present invention uses differential convolution to extract image gradient features and fully utilize the prior information of the infrared image. It also introduces a channel-level attention module to screen the key semantic feature patterns in the feature map. During the training phase, this embodiment uses pre-collected infrared perspective video images, extracts frames from the video to screen key frames, and annotates the two-category semantic segmentation dataset in LabelMe, with a scale of approximately 200 infrared images with pixel-level fine annotations. During training, the model augments images by randomly flipping and cropping them to a size of 256×256. The network was built using the PyTorch framework, with detail-enhancing attention blocks stacked in four layers at both the upsampling and downsampling stages, and eight layers at the bottom feature extraction layer of the U-Net. The dataset was split into training and test sets with an 8:2 ratio, and the model was trained for 100 epochs. Inference was performed on the test set and on 640×512 resolution infrared images collected by a drone.

[0058] After obtaining the pixel position of the target, this embodiment uses Figure 4 The mathematical geometry principle shown in the figure is used to solve the local coordinates. This process involves the characteristics of the infrared QR code marker in physical space and pixel space. During the production stage of the infrared QR code marker, its actual physical perimeter is first measured and recorded as L. When this embodiment is running, the drone flies to the experimental site and obtains real-time infrared images. The OpenCV library is used to detect the infrared marker and obtain its corner pixel position (u p1 ,v p1 ),(u p2 ,v p2 ),(u p3 ,v p3 ),(u p4 ,v p4 ), the marker's deflection angle α, the marker's pixel perimeter l p Based on this, the conversion coefficient k from pixel to actual distance is calculated, and according to the pixel position of the measured target, the angle β between the line connecting the center point of the marker and the u axis of the image coordinate system and the pixel distance Dist are calculated. pProject the line to the defined x-axis and y-axis directions in the pixel space to obtain the pixel projection value Proj x 、Proj y , and calculate the relative coordinates (x in the local coordinate system) based on the pixel-distance ratio relationship k L ,y L ). Finally, the global positioning information of the infrared QR code marker is transmitted in real time based on the satellite positioning module The latitude and longitude coordinates of the target to be measured can be calculated

[0059] The foregoing is an example of the best mode of carrying out the present invention. Any portion not described in detail herein is common knowledge within the skill of one of ordinary skill in the art. The scope of protection of the present invention is determined by the claims. Any equivalent transformation based on the technical teachings of the present invention is also within the scope of protection of the present invention.

Claims

1. A positioning method based on drone images and infrared QR codes, characterized in that: Under the condition of color mode limitation, the positioning of infrared heat source targets using drones includes the following steps: Step S1: constructing an infrared QR code marker, wherein the infrared QR code marker is made of a low-emissivity material, the QR code pattern has structural features that can be recognized by a machine vision system, and the center integrates a satellite positioning module for providing real-time geographic location information; Step S2: acquiring real-time infrared image stream data of the target scene through the UAV device and the wireless data transfer module, and performing image enhancement processing on the data; Step S3: Detect whether there is an infrared QR code marker in the image frame. If a marker is detected, extract its corner position and orientation information features and proceed to step S4; if not, repeat step S3 and continue processing the next frame; Step S4: Based on the extracted QR code feature information, a local coordinate system with the marker as the origin is established, and the pixel-distance ratio coefficient and the deflection angle of the QR code are calculated; Step S5: Segment the target area using the attention network based on detail enhancement, and use the geometric center of the segmented area as the image pixel position of the target; Step S6: Calculate the coordinate position of the target in the local coordinate system based on the pixel-distance ratio coefficient and the deflection angle information, and calculate the global geographic location of the target in combination with the satellite positioning data of the center of the marker.

2. A positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: The infrared QR code marker in step S1 is an integral device composed of a low-emissivity material and designed according to a QR code pattern; the low-emissivity material is a material with an emissivity ε<0.1 and can emit low-flux thermal infrared radiation, which appears as a low-brightness area with approximately zero grayscale in infrared imaging; the QR code pattern appears in black and white in visible light mode, and the arrangement of its coding pattern is a fixed format, with structural characteristics that can be decoded by a standard QR code recognition algorithm; the satellite positioning module is used to obtain the real-time geographic location of the infrared marker.

3. The positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: The drone device in step S2 is equipped with an infrared imaging device and a wireless communication device for real-time acquisition and transmission of infrared image stream data; the wireless data transfer module is composed of a remote control terminal, a data acquisition card and a graphics processing terminal of the drone device, wherein the remote control terminal is responsible for remote command and data relay, the data acquisition card is used to receive and transfer infrared image stream data, and the graphics processing terminal pre-processes the received image data to support subsequent detection and analysis operations; the image enhancement processing includes image dedistortion and contrast enhancement, wherein the image dedistortion process is based on the camera calibration method of the Zhang Zhengyou model, and is combined with the Brown-Conrady model to correct barrel distortion and pincushion distortion to ensure the geometric consistency of the image frame; contrast enhancement is based on histogram equalization and adaptive contrast adjustment methods to improve the local and global contrast of the infrared image, thereby enhancing the visibility of the infrared QR code and the target area.

4. The positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: The specific method of detecting whether there is an infrared QR code marker in the image frame in step S3 includes: using a visual recognition system to identify the image frame after image enhancement processing, if the infrared QR code is successfully identified, it is determined that the detection is successful; if not, it is determined that the detection fails; the visual recognition system is constructed based on a QR code pattern algorithm library; the feature information includes the pixel position of the marker corner point, the marker orientation deflection angle, and the pixel perimeter of the marker.

5. The positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: In step S4, the pixel-to-distance ratio coefficient k is calculated from the actual perimeter of the marker and the pixel perimeter. The local coordinate system is used to describe the positional relationship of the target relative to the infrared marker, and its coordinate origin is set to the center position of the marker pixel corner point, which is determined by the corner point pixel coordinates extracted in step S3. The deflection angle of the marker is determined by the deflection angle information extracted in step S3, where the positive direction of the marker is defined as the y-axis of the local coordinate system, and the x-axis is the direction obtained by rotating the y-axis 90 degrees clockwise.

6. The positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: The detail enhancement-based attention network described in step S5 is a convolutional neural network for semantic segmentation tasks, including a convolution layer, a downsampling module, a fusion module, and a detail enhancement attention block; wherein the convolution layer is used to extract features of the input image and generate a multi-level feature representation; the fusion module is used to splice the feature maps of the upsampling and downsampling stages, and aggregate shallow features and deep features; the detail enhancement attention block includes a pixel addition module, a differential convolution block, and an attention block, which are used to highlight the detail areas in the image, extract the brightness-gradient features in the feature map, and integrate the features using a channel attention weighted mechanism; wherein the target area corresponds to the pixel area output by the semantic segmentation network, and the target pixel position is determined by the geometric center of the output area.

7. The positioning method based on drone images and infrared two-dimensional codes according to claim 1, characterized in that: In step S6, the pixel positions of the target and the marker are calculated, and the real space distance between the target and the marker is calculated using the pixel-distance ratio k in step S5; And using the local coordinate system established in step S5 and its x-axis and y-axis directions, respectively calculate the real distance projection length between the target and the marker in the x-axis and y-axis directions of the local coordinate system, so as to obtain the relative coordinate position of the target in the local coordinate system of the marker; combined with the real-time geographic location of the infrared marker obtained by the satellite positioning module in step S1 and the relative coordinates of the target in the local coordinate system, calculate the geographic coordinates of the target in the real geographic space.