Take-off and landing control method and platform for vehicle-mounted duplex fire-extinguishing rescue unmanned aerial vehicle

By constructing a visual positioning marking area on the vehicle-mounted drone take-off and landing platform and using a multimodal visual perception module for feature fusion, the problem of precise take-off and landing and dual-drone collaboration of traditional drones in complex environments has been solved. This has enabled efficient drone rotation and safe docking, meeting the control requirements for efficient fire fighting and rescue.

CN120872011APending Publication Date: 2025-10-31CONTINENTAL UNIION CHAOLU TECH BEIJING CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511017602.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, traditional drone control methods are difficult to achieve precise take-off and landing and dual-drone coordination in complex fire fighting and rescue scenarios, resulting in problems such as positioning deviation and hose entanglement, which cannot meet the needs of efficient fire fighting and rescue.

Method used

By deploying a container-structured vehicle-mounted UAV take-off and landing platform, activating LED visible light coded beacons to construct a visual positioning identification area, and using a multimodal visual perception module to acquire images and perform self-attention and mutual attention fusion to generate dynamic fused visual features, combined with flight attitude correction control commands and water hose tension adjustment parameters, the UAV can achieve precise take-off and landing and efficient rotation.

Benefits of technology

It enables precise take-off and landing and efficient rotation of drones in complex environments, ensures safe docking of tethered hoses and mounting devices, and meets the control requirements for efficient fire fighting and rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872011A_ABST
    Figure CN120872011A_ABST
Patent Text Reader

Abstract

The invention discloses a take-off and landing control method and platform for a vehicle-mounted duplex fire-extinguishing rescue unmanned aerial vehicle, and relates to the technical field of unmanned aerial vehicle control, and the method comprises the steps: firstly unfolding a vehicle-mounted unmanned aerial vehicle take-off and landing platform, and activating LED beacons at four corners to construct a positioning identification area; when the single-machine rotation mode is obtained, the first unmanned aerial vehicle is controlled to take off, images are collected, and visual features are fused; calculating an offset to generate a flight attitude correction control instruction during homeward voyage; and controlling the first unmanned aerial vehicle to land and synchronously recover the water hose, and then triggering a second unmanned aerial vehicle take-off instruction. The technical problem that in a fire extinguishing and rescue scene, a traditional unmanned aerial vehicle control means is difficult to deal with precise take-off and landing and double-unmanned aerial vehicle cooperation in a complex environment and cannot meet the control requirement for efficient rotation operation of the vehicle-mounted duplex unmanned aerial vehicle is solved, precise take-off and landing and efficient rotation of the vehicle-mounted duplex unmanned aerial vehicle in the complex fire extinguishing and rescue scene are achieved, and the control requirement for efficient rotation operation of the vehicle-mounted duplex unmanned aerial vehicle is met. And the technical effect of high-efficiency fire extinguishing and rescue control requirements is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control technology, and in particular to a take-off and landing control method and platform for a vehicle-mounted dual-unit firefighting and rescue UAV. Background Technology

[0002] In the field of fire and rescue, drones play a significant role in firefighting operations due to their flexibility, especially vehicle-mounted dual-drone drones, which can improve the continuity of rescue efforts. Current technologies mostly use traditional remote control or single navigation and positioning methods to control drone takeoff and landing. These methods are effective in open and stable environments, but they reveal significant limitations in complex rescue scenarios. Due to the complex environment at fire scenes, including smoke and electromagnetic interference, traditional control methods struggle to achieve precise drone takeoff and landing. When rotating between two drones, problems such as positioning deviations and hose entanglement can easily occur, leading to inaccurate data, low coordination efficiency, and failing to meet the demands of efficient firefighting and rescue. Summary of the Invention

[0003] This application provides a take-off and landing control method and platform for a vehicle-mounted dual-unit firefighting and rescue drone, which solves the technical problem that traditional drone control methods are difficult to cope with precise take-off and landing and dual-drone coordination in complex environments in firefighting and rescue scenarios, and cannot meet the control requirements of efficient rotation operations of vehicle-mounted dual-unit drones.

[0004] The first aspect of this application provides a take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone. The method includes: deploying a container-structured vehicle-mounted drone take-off and landing platform; activating a set of LED visible light coded beacons located at the four corners of the platform to construct a visual positioning identification area; acquiring a dual-drone alternating take-off and landing mode; when the dual-drone alternating take-off and landing mode is a single-drone alternating take-off and landing mode, controlling the first drone to take off, and acquiring RGB images and infrared images according to a preset acquisition cycle through a multimodal visual perception module installed on the lower part of the first drone; extracting modal features respectively, and performing multi-stage attention feature fusion to generate dynamic fused visual features; when the first drone returns, calculating the offset between the center of the first drone and the center of the visual positioning identification area based on the spatial relationship between the dynamic fused visual features and the visual positioning identification area, and generating a flight attitude correction control command based on the offset; controlling the first drone to land in conjunction with the flight attitude correction control command, and synchronously controlling the hose recovery device to uniformly retract the hose in a guided manner according to dynamic hose tension adjustment parameters to complete the landing of the first drone, and triggering the take-off command of the second drone.

[0005] A second aspect of this application provides a take-off and landing control platform for a vehicle-mounted dual-unit firefighting and rescue drone. The platform includes: a visual positioning identification area construction module for deploying a container-structured vehicle-mounted drone take-off and landing platform, activating LED visible light coded beacons located at the four corners of the platform, and constructing a visual positioning identification area; and a dynamic fusion visual feature acquisition module for acquiring a dual-drone alternating take-off and landing mode. When the dual-drone alternating take-off and landing mode is a single-drone alternating take-off and landing mode, the module controls the first drone to take off and acquires RGB and infrared images according to a preset acquisition cycle through a multimodal visual perception module installed on the lower part of the first drone. After extracting modal features, multi-stage attention is performed. Force feature fusion generates dynamic fused visual features; a flight attitude correction control command acquisition module is used to calculate the offset between the center of the first UAV and the center of the visual positioning mark area based on the spatial relationship between the dynamic fused visual features and the visual positioning mark area when the first UAV returns, and generate flight attitude correction control commands according to the offset; a second UAV take-off command execution module is used to control the first UAV to perform landing in combination with the flight attitude correction control commands, and synchronously control the hose recovery device to uniformly retract the hose in a guided manner according to the dynamic hose tension adjustment parameters, so as to complete the landing of the first UAV and trigger the take-off command of the second UAV.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application utilizes a container-structured vehicle-mounted UAV takeoff and landing platform and activates an LED visible light coded beacon to construct a visual positioning identification area. It then acquires a dual-drone alternating takeoff and landing mode. After takeoff, the UAV uses a multimodal visual perception module to acquire images and generates dynamic fused visual features through self-attention and mutual attention fusion. Upon return, these features are combined with beacon signals to calculate the offset and generate attitude correction commands. Simultaneously, the hose reel is controlled based on dynamic hose tension adjustment parameters, achieving precise alternating takeoff and landing of the two drones. This ensures safe docking of the tethered hose and mounting device, making the takeoff and landing control of the vehicle-mounted dual-drone firefighting and rescue UAVs more efficient and reliable. The application achieves the technical effect of precise takeoff and landing and efficient alternation of vehicle-mounted dual-drone UAVs in complex firefighting and rescue scenarios, meeting the control requirements for efficient firefighting and rescue. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1This is a flowchart illustrating the take-off and landing control method of a vehicle-mounted dual-unit firefighting and rescue drone provided in an embodiment of this application; Figure 2 This is a schematic diagram of the take-off and landing control platform of a vehicle-mounted dual-unit firefighting and rescue drone provided in an embodiment of this application.

[0009] Figure labeling: Visual positioning identification area construction module 1, dynamic fusion visual feature acquisition module 2, flight attitude correction control command acquisition module 3, second UAV take-off command execution module 4. Detailed Implementation

[0010] This application provides a take-off and landing control method and platform for a vehicle-mounted dual-unit firefighting and rescue drone, which solves the technical problem that traditional drone control methods are difficult to cope with precise take-off and landing and dual-drone coordination in complex environments in firefighting and rescue scenarios, and cannot meet the control requirements of efficient rotation operations of vehicle-mounted dual-unit drones.

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0012] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0013] Example 1, as Figure 1 As shown, a take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone is disclosed, wherein the method includes: Step A100: Deploy the container-structured vehicle-mounted drone take-off and landing platform, activate the LED visible light coded beacon sets located at the four corners of the platform, and construct a visual positioning identification area.

[0014] In this embodiment, the vehicle-mounted drone take-off and landing platform is a container-type retractable and deployable structure placed on the rear of a fire truck or the top of a fire water tanker. It is equipped with a hydraulic or electric drive system and is used to support the take-off, landing, parking, and coordinated operation of dual firefighting and rescue drones. The visible light coded beacon is a positioning device deployed at the four vertices of the vehicle-mounted drone take-off and landing platform, consisting of a control circuit, a coding modulation circuit, a signal amplification circuit, and an LED driving circuit. The visual positioning marker area is the visible light coded beacon at the four vertices of the vehicle-mounted drone take-off and landing platform, which constructs an optical sensing range in space carrying the platform's position coordinate information through high-frequency flashing coded signals.

[0015] Specifically, once the fire water tanker arrives at the work site, the power system of the vehicle-mounted drone take-off and landing platform (i.e., the power compartment, located at the rear of the container, measuring 0.5 meters long * 2 meters wide * 1 meter high, containing the vehicle-mounted power interface, charger, hydraulic power unit, or regulated power supply) drives the container structure take-off and landing platform to initiate the deployment procedure. Relying on a hydraulic linkage mechanism, under preset control logic, the platform gradually unfolds from its folded state, forming a working surface with a horizontal span that meets the take-off and landing requirements of the drone.

[0016] Next, the LED visible light coded beacons at the four corners of the platform are activated: the control circuit inside the beacon first initiates a start command, and the encoding and modulation circuit then generates a coded signal in a specific format, namely, a 2-byte FF as the start field; a 1-byte information field containing a 4-bit vehicle code to distinguish different rescue vehicles; a 2-bit beacon sequence number corresponding to the spatial position at the four corners of the platform and 2 check bits to ensure signal integrity; and a 1-byte end field. The generated signal is amplified and then drives the LEDs to flash at a high frequency of ≥50Hz, continuously broadcasting the position code into the air to construct a visual positioning identification area. In this process, the vehicle code identifies the platform, the beacon sequence number maps spatial orientation, and the high-frequency flashing overcomes interference in the rescue scenario, such as smoke and electromagnetic noise, allowing the drone to stably capture the signal.

[0017] Through a series of actions involving mechanical deployment and signal encoding broadcasting, the platform constructs a visual positioning area carrying spatial coordinate information in vertical space, providing an optical reference for the UAV's subsequent attitude perception and precise docking. Utilizing standardized signal design and hardware integration, the platform's positioning markers were constructed, laying the perceptual foundation for the UAV's precise take-off and landing.

[0018] Step A200: Obtain the dual-drone rotation take-off and landing mode. When the dual-drone rotation take-off and landing mode is single-drone rotation take-off and landing, control the first UAV to take off, and acquire RGB images and infrared images according to a preset acquisition cycle through the multimodal visual perception module installed on the lower part of the first UAV. After extracting modal features respectively, perform multi-stage attention feature fusion to generate dynamic fused visual features.

[0019] In this embodiment of the application, single-drone rotation take-off and landing means that the second drone takes off after the first drone lands, and landing is the reverse process.

[0020] Optionally, those skilled in the art can determine the dual-drone rotation take-off and landing mode based on rescue mission parameters, such as estimated operation time and fire spread rate, and dual-drone status, such as battery level and payload integrity. When it is determined to be a single-drone rotation take-off and landing mode, a full-state self-check is first performed on the first UAV: ​​the battery level is checked and must be ≥80% to meet the power consumption for take-off, landing, and initial operation; the sealing pressure of the tethered hose interface of the mounting device is ≥0.3MPa; and the communication delay between the UAV flight control module and the vehicle-mounted UAV take-off and landing platform is ≤50ms.

[0021] After passing the self-test, a takeoff command is sent to the first UAV. The flight control module drives the multi-rotor motor to perform vertical takeoff with an initial ascent speed of ≤2m / s. Simultaneously, the tethered hose deployment and retrieval mechanism adjusts the hose length in real time according to the UAV's altitude changes, maintaining hose tension within the safe range of 5–10N to prevent tangling or slack. During takeoff, the GNSS module continuously collects and records the takeoff point coordinates. The GNSS module is a global satellite navigation system coordinate system used to record the UAV's takeoff point, providing an initial positioning reference for the UAV's takeoff. During landing, it combines with ground ranging to preliminarily determine whether the UAV has entered the landing zone. Its error is ≤0.5m, serving as the reference data for subsequent return-to-home positioning.

[0022] Finally, the multimodal visual perception module at the bottom of the first UAV acquires RGB and infrared image sequences according to a preset acquisition cycle. After extracting modal features, it first performs a first-stage self-attention fusion on each of the two types of feature sequences, and then performs a second-stage mutual attention fusion on the fusion results to generate dynamic fused visual features. The specific steps are explained in detail in A210-A240.

[0023] Through a seamless process of mode determination, status self-check, and coordinated control of flight control and water hose, the smooth take-off of the UAV in single-unit rotation mode was achieved, ensuring the synchronous matching of the tethered water hose and flight actions, and laying a reliable start-up foundation for the continuous firefighting operation.

[0024] Step A300: When the first UAV returns, based on the spatial relationship between the dynamically fused visual features and the visual positioning marker area, calculate the offset between the center of the first UAV and the center of the visual positioning marker area, and generate a flight attitude correction control command according to the offset.

[0025] In one embodiment of this application, when the first UAV returns, it first matches the preset landing area features based on the dynamic fusion visual features and the visual positioning mark area. After successful matching, it activates the bottom desktop optical landing positioning sensor, combines the phased array detection and processing of the LED beacon flashing signal to obtain the decoding result, determines the current position through least squares fitting, calculates the offset from the center of the area and sends it to the flight control module to generate flight attitude correction control commands. The specific steps are described in detail in A310-A340.

[0026] Step A400: Combine the flight attitude correction control command to control the first UAV to perform landing, and synchronously control the water hose recovery device to uniformly retract the water hose in a guided manner according to the dynamic water hose tension adjustment parameters, so as to complete the landing of the first UAV and trigger the take-off command of the second UAV.

[0027] Specifically, the first UAV adjusts its attitude to land according to flight attitude correction control commands. Simultaneously, the hose recovery device, based on dynamic hose tension adjustment parameters, guides the hose through guide wheels to retract at a constant speed matching the UAV's descent speed, maintaining stable hose tension. Detailed steps are explained in A410-A420. When the first UAV fully touches the ground and completes its landing, the takeoff command for the second UAV is automatically triggered, achieving seamless operation between the two drones.

[0028] Furthermore, step A200 in the method provided in this application embodiment includes: A210: The multimodal visual perception module acquires RGB and infrared images according to a preset acquisition cycle, and obtains RGB image sequences and infrared image sequences.

[0029] A220: Modal feature extraction is performed on the RGB image sequence and the infrared image sequence respectively to obtain the RGB image feature sequence and the infrared image feature sequence.

[0030] A230: Perform a first-stage self-attention feature fusion on the RGB image feature sequence and the infrared image feature sequence respectively to obtain RGB image fusion features and infrared image feature fusion features.

[0031] A240: Perform a second-stage mutual attention feature fusion on the RGB image fusion features and the infrared image fusion features to obtain dynamic fused visual features.

[0032] Specifically, after the first UAV takes off, the multimodal visual perception module installed on its underside is activated. This module includes an RGB imaging component and an infrared sensing component, and begins to work according to a preset acquisition cycle (e.g., 100ms / frame). The RGB imaging component captures scene information within the visible light range, records the texture and color features of the take-off and landing platform and the surrounding environment, and continuously acquires data to form an RGB image sequence; the infrared sensing component penetrates interference such as smoke and dust, captures the thermal radiation distribution of objects in the environment, and simultaneously generates an infrared image sequence.

[0033] The two image sequences are synchronized over time. After each acquisition cycle, the module adds one RGB image and one infrared image, continuously accumulating to form a set of RGB and infrared image sequences containing temporal information. The RGB image sequence preserves detailed features of the visually identifiable location area, such as the brightness variations of the LED beacon, while the infrared image sequence supplements the contour information of the platform and surrounding objects in low-visibility environments. Utilizing synchronized acquisition and sequence accumulation, the multimodal visual perception module provides a continuous visual data foundation covering both visible and infrared frequency bands for subsequent modal feature extraction.

[0034] Next, a multimodal feature extractor is pre-constructed. This extractor is used to extract modal features from RGB image sequences and infrared image sequences respectively to obtain RGB image feature sequences and infrared image feature sequences. The specific steps are explained in detail in A221.

[0035] Then, when performing the first-stage self-attention feature fusion on the RGB image feature sequence and the infrared image feature sequence respectively, the two are used as inputs. The context dependency relationship is constructed through the self-attention encoder module, and multi-head self-attention calculation is performed. After splicing, linear transformation and residual connection and normalization processing, the RGB image fusion feature and the infrared image fusion feature are obtained. The specific steps are explained in detail in A231-A233.

[0036] Finally, when performing the second-stage mutual attention feature fusion of the RGB image fusion features and the infrared image fusion features, the feature similarity set of the two is obtained, and after normalization, a multimodal adjacency matrix is ​​constructed. This matrix is ​​then used to convolve the two fusion features respectively, and the mean of the results is processed to obtain the dynamic fusion visual features. The specific steps are explained in detail in A241-A243.

[0037] Furthermore, step A220 in the method provided in this application embodiment includes: A221: A pre-constructed multimodal feature extractor is used to extract modal features from the RGB image sequence and the infrared image sequence respectively, so as to obtain the RGB image feature sequence and the infrared image feature sequence.

[0038] Optionally, before applying the multimodal visual perception module to acquire RGB and infrared image sequences, a multimodal feature extractor needs to be pre-built. When building the multimodal feature extractor, a dual-branch parallel architecture needs to be designed to adapt to the feature extraction requirements of RGB and infrared image sequences respectively. Both the RGB and infrared branches contain an input layer, convolutional block groups, pooling layers, and a feature output layer. The branch structures are symmetrical but the parameters are independent. The input layer size matches the image resolution, such as 256×256 pixels, converting the original image into a three-dimensional tensor. The RGB branch has 3 channels, while the infrared branch has 1 channel. The convolutional block group consists of 3-5 cascaded convolutional units. Each unit contains convolutional layers that alternate between 3×3 and 5×5 convolutional kernels, with the number increasing with the layer level. For example, the first layer has 32 kernels, and the second layer has 64 kernels. The ReLU activation function and batch normalization layer capture multi-scale features through convolutional kernels of different sizes. The 3×3 kernel focuses on local details, such as the edges of LED beacons, while the 5×5 kernel captures global contours, such as the corners of a platform. The pooling layer uses 2×2 max pooling with a stride of 2, compressing the feature dimension while retaining key information. The feature output layer maps the features to a fixed 256-dimensional dimension using a 1×1 convolutional kernel, forming a standardized feature vector.

[0039] After determining the architecture, the convolutional kernel parameters were initialized using a He normal distribution, and the bias term was set to 0. Training components were configured: an Adam optimizer with a learning rate of 0.001 was selected, and the feature reconstruction error (mean squared error between the input image and the image reconstructed by the extractor and decoder) was used as the loss function. The training set was divided into 80% and the validation set into 20%. Subsequently, over 10,000 frames of RGB and infrared image samples covering scenes such as smoke, strong light, and nighttime were loaded, input into the dual-branch system in batches of 32 frames each, and iterated for 50-100 rounds. Validation was performed every 5 rounds. When the validation set loss dropped below 0.01 and showed no significant fluctuation for 3 consecutive rounds, training was stopped, the model parameters were saved, and the construction of the multimodal feature extractor was completed.

[0040] Next, the RGB image sequence and the infrared image sequence are input into the multimodal feature extractor. The multimodal feature extractor processes each frame of the image independently: for the RGB image sequence, pixel spatial correlation features are extracted through convolutional layers, and the output feature dimension is 256 dimensions / frame. After pooling compression, key texture information is retained; for the infrared image sequence, contour features are extracted based on the thermal radiation intensity distribution through a dedicated convolutional kernel (such as an optimized version of the Sobel operator that focuses on edge detection), and the output feature dimension is 256 dimensions / frame, forming an infrared image feature sequence that corresponds to the temporal sequence of the RGB image feature sequence. The RGB image feature sequence retains the visible light details of the visual positioning marker area, while the infrared image feature sequence enhances the environmental contour information under low visibility conditions.

[0041] By pre-constructing a multimodal feature extractor adapted to rescue scenarios, feature extraction was performed on RGB and infrared image sequences respectively, resulting in feature sequences that accurately reflect the key information of the two modalities, providing high-quality input data for subsequent feature fusion.

[0042] Furthermore, step A230 in the method provided in this application embodiment includes: A231: Using RGB image feature sequences and infrared image feature sequences as inputs, the self-attention encoder module constructs contextual dependencies to obtain RGB image feature dependencies and infrared image feature dependencies.

[0043] A232: Based on the feature dependencies of RGB images and infrared images, multi-head self-attention calculation is performed on the feature sequences of RGB images and infrared images respectively to obtain the attention feature sequences of RGB images and infrared images.

[0044] A233: After concatenating the attention feature sequences of RGB images and infrared images, a linear transformation is performed, and residual connections and normalization are performed with the RGB image feature sequences and infrared image feature sequences to obtain RGB image fusion features and infrared image fusion features.

[0045] In this embodiment, the self-attention encoder module is used to process the input RGB image feature sequence and infrared image feature sequence to construct contextual dependencies. Contextual dependencies refer to the temporal or spatial relationships between features at different locations in the RGB image feature sequence or infrared image feature sequence. Image feature dependencies refer to the dependencies between features within the sequence constructed by the self-attention encoder module for the RGB image feature sequence or infrared image feature sequence.

[0046] Specifically, after the RGB image feature sequence and infrared image feature sequence are generated, the contextual association within the modality needs to be strengthened to address potential issues such as dynamic blurring, background interference, and signal flickering that may occur during the drone's return journey.

[0047] First, when processing RGB image feature sequences and infrared image feature sequences, the self-attention encoder module divides the two types of sequences into continuous local processing windows of 10 frames each to adapt to the temporal correlation of dynamic scenes during the UAV's return journey, such as the continuity of beacon flashing and the stability of the platform outline. For RGB image feature sequences, the module calculates the correlation weight between any two frames within the window, such as frame t and frames t+1 to t+9, for each frame's 256-dimensional feature vector. By quantifying the magnitude of brightness feature changes caused by LED beacon flashing, if the difference in brightness features between adjacent frames is ≤0.1, the weight is increased to above 0.8, emphasizing the dependence between the flashing frame and the preceding and following frames. For example, frame t (the bright beacon frame) and frame t+1 (the dark beacon frame) have complementary brightness features, so their weights are set to the highest value within the window, thereby highlighting the temporal continuity of the beacon in a dynamically blurred environment.

[0048] For the infrared image feature sequence, the module focuses on the differences between the platform's thermal radiation profile and its surrounding environment: it calculates the similarity of the platform's thermal profile features, such as temperature gradient vectors, in each frame within the window. When the overlap of platform profile features between two frames is ≥90%, a high correlation weight is assigned, such as 0.7. Conversely, when the overlap with the profiles of frames containing interference such as flames and smoke is ≤30%, a low weight is assigned, such as below 0.2. This strengthens the continuous correlation of the platform's thermal features in complex backgrounds. This constructs RGB image feature dependencies and infrared image feature dependencies, highlighting the feature correlations of key regions (such as beacons and platform corners) in the sequence.

[0049] Based on the constructed feature dependencies, multi-head self-attention computation is performed on the two types of feature sequences. Eight attention heads are set up, each assigning weights from different feature subspaces (such as the color channels of RGB and the thermal intensity gradient of infrared) to enhance the response intensity to key regions. For example, for the high-frequency flickering feature of LED beacons in the RGB sequence, multi-head self-attention increases the weight ratio of flickering frames in the sequence; for the infrared bright targets of the platform in the infrared sequence, its distinction from the background is enhanced, ultimately generating RGB image attention feature sequences and infrared image attention feature sequences.

[0050] For the attention feature sequence of RGB images, the multi-frame features (each frame being a 256-dimensional vector) are first concatenated intra-frame in chronological order to form a continuous feature matrix. The same operation is performed on the attention feature sequence of infrared images, concatenating them into corresponding feature matrices. Subsequently, the two concatenated feature matrices are processed through a 128-dimensional linear transformation matrix. The weight parameters in the matrix are trained and optimized to map the 256-dimensional high-dimensional features to a 128-dimensional low-dimensional space, reducing redundant information while retaining key features, such as beacon blinking features in RGB and platform thermal contour features in infrared, thus achieving dimensional unification between the two types of attention feature sequences.

[0051] After dimensional unification is completed, the transformed RGB attention features are performed with the original RGB image feature sequence using residual connection: that is, the 128-dimensional attention features are superimposed on the corresponding original 256-dimensional features frame by frame, and the attention feature dimensions are matched with the original features through zero padding. While preserving the basic information of the original features, the key region features highlighted by the self-attention mechanism are strengthened. The infrared image attention features and the original infrared image feature sequence are also residually connected in the same way.

[0052] Finally, LayerNorm normalization is performed on the residual concatenated RGB and infrared features respectively: the mean and variance of the features in each frame are calculated, and the feature values ​​are standardized using the formula (feature value - mean) / standard deviation, so that the mean of the RGB fused features and the variance of the infrared fused features are both 0 and 1, eliminating scale differences between features in different frames and ensuring the stability of feature distribution in subsequent feature fusion processes. Through this process, the final RGB image fused features and infrared image fused features are obtained, which contain both the original feature information and enhance the correlation of key regions.

[0053] Through the above process, the self-attention mechanism effectively improves the feature modeling capability of key regions within a modality. After residual connection and normalization processing, it not only preserves the original feature information but also strengthens the effective correlation in the sequence, providing high-quality feature input for subsequent cross-modal mutual attention interaction.

[0054] Furthermore, step A240 in the method provided in this application embodiment includes: A241: Obtain the feature similarity set between the RGB image fusion feature and the infrared image feature fusion feature.

[0055] A242: Normalize the feature similarity set and construct a multimodal adjacency matrix based on the normalization result.

[0056] A243: The RGB image fusion feature and the infrared image fusion feature are convolved using the multimodal adjacency matrix respectively, and the convolution results are averaged to obtain the dynamic fusion visual feature.

[0057] Specifically, in the second stage of mutual attention feature fusion, the feature similarity set between the RGB image fusion features and the infrared image fusion features is first calculated. Each frame's RGB image fusion feature and the corresponding frame's infrared image fusion feature are both 256-dimensional vectors. The cosine similarity formula is used to calculate the cosine of the angle between the two vectors, obtaining the single-frame similarity, with a value ranging from 0 to 1. The same calculation is performed on a feature sequence containing 30 frames, forming a feature similarity set containing 30 values. The frame similarity for the platform region and LED beacon is typically ≥0.7, because the visible light details of RGB and the thermal outline of infrared are highly matched in this region, while the frame similarity for the background interference region is typically ≤0.3.

[0058] Subsequently, the feature similarity set is normalized: the Softmax function maps all similarity values ​​to the 0-1 range, making the sum equal to 1, thus strengthening the weight of high-similarity frames. For example, the weight of a frame with an original similarity of 0.8 increases to 0.25 after normalization, while the weight of a frame with an original similarity of 0.2 decreases to 0.05. A multimodal adjacency matrix is ​​constructed based on the normalization results. The matrix is ​​initially empty with a dimension of 30×30, consistent with the length of the feature sequence. The normalized similarity values ​​are filled into the corresponding positions according to the frame index. For example, the similarity between the 5th frame and the 5th frame is filled into the 5th row and 5th column of the matrix, forming a matrix reflecting the temporal correlation strength between RGB and infrared features. The diagonal and adjacent frame positions show dense high values ​​due to temporal continuity.

[0059] The multimodal adjacency matrix is ​​used to perform convolution operations on the RGB image fusion features and the infrared image fusion features respectively: a 1×1 convolution kernel with 256 channels is used, and the dot product operation between the matrix and the feature vector is used to strengthen the closely related feature components in the two modalities, such as suppressing background noise features that are independent of the platform. After convolution, the RGB convolution results and the infrared convolution results are averaged over time frames to obtain a dynamic fusion visual feature with a dimension of 256. This feature not only preserves the color texture details of RGB (such as beacon blinking frequency) but also integrates the anti-interference thermal contour information of infrared (such as platform edges).

[0060] Through the above process, the mutual attention mechanism effectively explores the complementary correlation between the two modal features. The dynamically fused visual features generated after convolution and mean processing take into account both detail recognition and anti-interference capabilities in complex rescue environments, providing a more comprehensive visual basis for the calculation of UAV positioning offset.

[0061] Furthermore, step A300 in the method provided in this application embodiment includes: A310: Based on the dynamic fusion visual features, match them with the preset landing area features of the visual positioning marker area. If the match is successful, activate the desktop optical landing positioning sensor installed on the bottom of the first UAV, combine it with phased array detection to capture the flashing signal of the LED visible light coded beacon set and process it to obtain the signal decoding result.

[0062] A320: Based on the signal decoding results, perform least squares fitting to determine the current position of the first UAV.

[0063] A330: Calculate the offset between the current position and the center position of the visual positioning mark area to obtain the offset amount.

[0064] A340: The offset is sent to the flight control module to generate flight attitude correction control commands.

[0065] Specifically, when the first UAV returns to the preset altitude, it begins matching the preset landing area features based on dynamically fused visual features and visual positioning markers. The preset landing area features include baseline features such as the relative positions of the LED beacons at the four corners of the platform and the geometric contours of the platform's edges (4 meters long × 2 meters wide); the dynamically fused visual features integrate the beacon flashing details in the RGB image with the platform's thermal outline in the infrared image. The overlap between the two types of features is calculated using a feature point matching algorithm (such as SIFT). When the percentage of successfully matched feature points is ≥80%, the match is considered successful, indicating that the UAV has entered a range where precise positioning is possible.

[0066] After successful matching, the desktop optical landing and positioning sensor installed on the bottom of the UAV is activated. This sensor, combined with phased array detection technology, captures the signal of the LED visible light coded beacons flashing at a frequency of ≥50Hz at the four corners of the platform in real time. Through photoelectric conversion, filtering, amplification, and differential calculation of the signal, the signal decoding result containing the incident angle and relative distance of each beacon is obtained. The specific steps are explained in detail in A311-A313.

[0067] Based on the signal decoding results, the least squares fitting method is used for data optimization: using the preset known absolute coordinates of the four beacons as a reference, the detected relative distances and angles are converted into three-dimensional coordinate observations of the UAV, and the deviation between the observations and theoretical values ​​is minimized through the fitting algorithm, so as to finally determine the current position of the UAV.

[0068] Next, the current position is offset from the center position of the visual positioning mark area, which is preset as the geometric center coordinates of the platform (2 meters, 1 meter, 0 meters). The offset in three dimensions (Δx, Δy, Δz) is calculated. For example, Δx = 0.3 meters, Δy = -0.2 meters, and Δz = 0.1 meters, which represent the deviation from the center in the x, y, and z axis directions, respectively.

[0069] Finally, the offset is sent to the flight control module, which generates corresponding flight attitude correction control commands based on the offset: for example, for Δx=0.3 meters, the roll angle is adjusted to make the UAV move to the left; for Δy=-0.2 meters, the pitch angle is adjusted to make the UAV move forward, gradually reducing the offset to within ±0.05 meters to ensure accurate docking.

[0070] Through the above process, from feature matching to confirm the landing range, to signal analysis and positioning calculation, and then to the generation of attitude correction commands, a coherent control from coarse positioning to fine alignment is achieved during the UAV's return process, providing precise attitude guidance for the safe docking of the tethered hose and the mounted device.

[0071] Furthermore, step A310 in the method provided in this application embodiment includes: A311: Phased array detection captures the set of flash signals of the LED visible light coded beacon set.

[0072] A312: The flickering signal is combined with photoelectric conversion and component filtering is performed using a bandpass filter to obtain a set of filtered photoelectric signals.

[0073] A313: The filtered photoelectric signal is amplified using a low-noise amplifier and differential calculation is performed to determine the set of beacon incident angles and the set of relative distances. The signal decoding result is then obtained by summarizing the results.

[0074] In this embodiment, the bandpass filter is a device for component filtering of the LED visible light coded beacon flashing signal after photoelectric conversion. It can retain the effective signal component matching the beacon flashing frequency (≥50Hz) and filter out interference signals below or above this frequency band, obtaining a pure set of filtered photoelectric signals. The low-noise amplifier is a device for amplifying the photoelectric signal filtered by the bandpass filter. While increasing the signal amplitude to a processable range, it can control the introduced noise at a low level, avoiding signal distortion.

[0075] In one embodiment, after activating the desktop optical landing and positioning sensor mounted on the bottom of the first UAV, the phased array detection unit of the sensor begins to operate. The phased array consists of multiple optical receiving elements arranged in a specific array. By adjusting the phase response of each element in real time, a directional receiving beam is formed to accurately capture the signal of the LED visible light coded beacon set at the four corners of the platform flashing at a frequency of ≥50Hz, forming a signal set containing the timing flashing characteristics of the four beacons. The signal of each beacon carries its position coding information and exhibits periodic pulse characteristics on the time axis, such as flashing once every 20ms. The LED beacon signal processing parameters are shown in Table 1.

[0076] Next, the captured set of flickering signals is processed: first, the optical signal is converted into a corresponding electrical signal by the photodiode array inside the sensor. At this time, the electrical signal is mixed with low-frequency interference from ambient light such as sunlight and flame light, and high-frequency interference from electronic noise. Then, it is connected to a bandpass filter whose center frequency matches the beacon flicker frequency (e.g., 50Hz). The filter bandwidth is set to 10Hz, retaining only the effective signal component of about 50Hz, and filtering out interference signals below 40Hz and above 60Hz, resulting in a pure set of filtered photoelectric signals.

[0077] The filtered photoelectric signal has a relatively small amplitude, typically in the mV range, and needs to be amplified by a low-noise amplifier. The amplifier gain is set to 30dB to increase the signal amplitude to the V level, while the introduced noise figure is controlled below 2dB to avoid signal distortion. The amplified signal is then sent to a differential calculation unit. By comparing the phase difference of the signals received by different receiving elements of the phased array and combining this with the array geometric parameters, the incident angle of each beacon is calculated.

[0078] Simultaneously, based on the signal strength attenuation model, the relative distance between the UAV and each beacon is calculated. The signal strength attenuation model is constructed based on the known transmission power and signal propagation characteristics of the LED beacons. By pre-collecting received signal strength data at different distances (e.g., 1 to 10 meters) and in different environments (e.g., smoke, no smoke), a curve relating received power to distance is fitted. The core principle uses the basic formula of the free-space propagation model, where received power is inversely proportional to the square of the distance. An environmental attenuation coefficient is introduced, which is dynamically adjusted by those skilled in the art based on factors such as smoke concentration. During calculation, the received power value of each beacon is first extracted from the signal output of the low-noise amplifier, which is converted from the electrical signal amplitude. Combined with the preset beacon transmission power and the environmental attenuation coefficient in the model, the distance is calculated by substituting these values ​​into the formula. Then, the average of the calculation results for 10 consecutive sampling periods is taken to reduce the impact of instantaneous interference, so that the relative distance error between the UAV and each beacon is controlled within ≤0.2 meters.

[0079] Finally, the incident angles and relative distances of all beacons are summarized to form a signal decoding result containing spatial orientation and distance information, which fully reflects the relative positional relationship between the UAV and the four corner beacons of the platform.

[0080] Through the above process, the flashing signal of the LED visible light coded beacon was accurately captured, purified, amplified, and analyzed, providing high-precision raw data support for the subsequent determination of the current position of the UAV.

[0081] Table 1: LED Beacon Signal Processing Parameters

[0082] Furthermore, step A400 in the method provided in this application embodiment includes: A410: Acquire the flight speed and altitude information of the first UAV.

[0083] A420: Based on the flight speed and altitude information, identify the water belt tension adjustment parameters and determine the dynamic water belt tension adjustment parameters.

[0084] In this embodiment, the hose tension adjustment parameter is used to synchronously control the hose retrieval device to uniformly retract the hose in a guided manner, ensuring that the hose maintains appropriate tension during the drone's descent, avoiding tangling or excessive stretching, and achieving safe and synchronous deployment and retrieval of the hose and the drone during takeoff and landing.

[0085] Optionally, when controlling the first UAV to perform a landing in conjunction with flight attitude correction control commands, the flight speed and altitude information of the first UAV are first acquired in real time. The flight speed is collected by a Hall effect velocity sensor on the UAV, with a sampling frequency of 10Hz and an accuracy of ±0.1m / s, covering a landing speed range of 0-5m / s; the altitude information is acquired by a laser rangefinder, also with a sampling frequency of 10Hz, a measurement range of 0-10m, and an error ≤0.05m. The two data are processed synchronously to form a time-series input pair (speed v, altitude h).

[0086] Next, based on this information, the dynamic hose tension adjustment parameters need to be determined, requiring the construction and training of a neural network model. The model adopts a three-layer fully connected neural network structure: the input layer contains 2 neurons, corresponding to flight speed and altitude; the hidden layer has two layers, the first layer has 64 neurons with the ReLU activation function, and the second layer has 32 neurons with the ReLU activation function; the output layer contains 2 neurons, corresponding to the hose retrieval speed coefficient k (range 0.8-1.2) and the tension threshold T (unit N, range 5-20N), where k is used to adjust the base rotation speed of the retrieval device, and T is used to limit the maximum hose tension.

[0087] The model training data came from simulations and actual tests: 10,000 samples were collected under different wind speeds (0-3 m / s) and hose lengths (5-30 m). Each sample included (v, h) and the corresponding optimal k and T. The hose state was recorded using a high-speed camera, and parameters were manually labeled to ensure no tangling and stable tension. The samples were divided into training and validation sets in an 8:2 ratio. An Adam optimizer with a learning rate of 0.001 and mean squared error (MSE) as the loss function was used for iterative training for 80 Lenz iterations. Training was stopped when the validation set MSE dropped below 0.005, and the model parameters were saved.

[0088] In practical applications, the real-time collected (v, h) data is normalized (velocity divided by 5 m / s and altitude divided by 10 m) and then input into the trained model. The model outputs k and T as dynamic hose tension adjustment parameters. The hose recovery device calculates the actual recovery speed based on k: recovery speed = UAV descent speed × k, and drives the winding mechanism via a servo motor. Simultaneously, a tension sensor monitors the hose tension in real time. When the tension approaches T, the controller fine-tunes the motor speed based on the deviation. For example, if the tension exceeds T by 10%, the speed is reduced by 5%. Combined with the guide wheel assembly, this ensures the hose is wound along a fixed path, achieving uniform and stable recovery. When the first UAV touches the ground, the recovery device stops working, completing the landing.

[0089] By constructing a neural network model with flight speed and altitude as inputs and dynamic hose tension adjustment parameters as outputs, precise synchronization between hose retrieval and UAV landing was achieved, ensuring that the hose does not tangle and that the tension is stable, thus ensuring a safe and efficient landing process.

[0090] Furthermore, step A400 in the method provided in this application embodiment includes: A510: When the dual-drone alternating take-off and landing mode is in the dual-drone alternating take-off and landing mode, the first UAV and the second UAV are controlled to take off in sequence and land in reverse order.

[0091] In this embodiment, the dual-drone alternating take-off and landing is the take-off and landing control method of the vehicle-mounted dual-drone fire-fighting and rescue drone in the alternating mode. Specifically, the first drone and the second drone are controlled to take off in a preset order. After the mission is completed, they land in the reverse order of take-off. That is, the second drone lands first according to the flight attitude command, and the hose recovery device simultaneously retracts the hose. After the hose is fully touched down, the landing command of the first drone is triggered, and the first drone completes the landing according to the same process.

[0092] In one embodiment, when the dual-drone takeoff and landing mode is in operation, the first and second drones are controlled to take off sequentially. Before takeoff, both drones undergo a comprehensive self-check to ensure that key components such as the power system, flight control system, and fire extinguishing equipment are functioning normally. The first drone starts first, ascends to an altitude of 10 meters at a vertical climb rate of 2 m / s, then switches to horizontal cruise mode and flies towards the fire extinguishing area at a speed of 8 m / s. Thirty seconds after the first drone takes off, the second drone initiates the same takeoff procedure, with the two drones maintaining a safe distance of more than 50 meters.

[0093] During the descent phase, the drones are controlled to land in reverse order. When the second drone returns after completing its mission, the ground control station calculates parameters such as its remaining battery power and mission completion rate. If the landing conditions are met, such as remaining battery power ≥ 30%, a landing command is sent. Upon receiving the command, the second drone enters the descent process at a descent speed of 1.5 m / s. Simultaneously, the hose recovery device adjusts the hose according to dynamic hose tension adjustment parameters, such as recovery speed coefficient k = 1.0 and tension threshold T = 15 N, to synchronously retract the hose. Once the second drone has completely landed and locked onto the vehicle-mounted platform, it immediately triggers the landing command for the first drone. The first drone performs its descent at the same descent speed and with the same hose recovery control method, ultimately completing the alternating takeoff and landing of the two drones.

[0094] The reverse-order design of the first and second drones taking off and the second drone landing avoids airspace conflicts between the two drones during takeoff. This sequential takeoff of the first and second drones is based on mission priority allocation. The first drone usually carries the main fire extinguishing equipment and enters the core rescue area first, while the second drone carries auxiliary equipment or spare fire extinguishing agents and takes off later to cooperate with the operation. The reverse-order landing of the second and first drones corresponds to the landing order of the second drone before the first drone. That is, when both drones need to return, the second drone, which completes its mission first, receives the landing instruction first due to its later takeoff time and relatively shorter mission cycle. It returns according to the preset path and completes the landing through dynamic fusion of visual feature matching and precise positioning. The hose recovery device simultaneously retracts its attached hose. After the second drone has completely landed and locked onto the vehicle platform, it sends a landing instruction to the first drone, which then performs the landing according to the same procedure.

[0095] By using this control method of taking off in sequence and landing in reverse sequence, combined with dynamic water hose tension adjustment and adaptive adjustment of environmental parameters, efficient connection and safety assurance of dual-aircraft operation are achieved, which improves the operational efficiency of the fire fighting and rescue drone system while reducing operational risks.

[0096] In summary, the take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone provided in this application has the following technical effects: This application constructs a visual positioning marker area by deploying a vehicle-mounted UAV take-off and landing platform and activating an LED visible light coded beacon. Modal features are extracted from the RGB and infrared images acquired by the UAV, and multi-stage attention feature fusion is performed to generate dynamic fused visual features. Based on these features, the offset between the UAV and the center of the visual positioning marker area is calculated, and flight attitude correction control commands are generated. These commands are combined to control the UAV's descent and synchronously retract the hose according to dynamic hose tension adjustment parameters. During dual-drone rotation, the UAVs take off in sequence and land in reverse order, thereby achieving precise take-off and landing and collaborative operation of the vehicle-mounted dual-drone firefighting and rescue system. This makes UAV take-off and landing control safer and more efficient, achieving the technical effect of precise take-off and landing and efficient rotation of vehicle-mounted dual-drone systems in complex firefighting and rescue scenarios, meeting the control requirements of efficient firefighting and rescue.

[0097] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a take-off and landing control platform for a vehicle-mounted dual-unit firefighting and rescue drone, the platform comprising: Visual positioning identifier area construction module 1 is used to deploy a container-structured vehicle-mounted UAV take-off and landing platform, activate the LED visible light coded beacon sets located at the four corners of the platform, and construct a visual positioning identifier area.

[0098] The dynamic fusion visual feature acquisition module 2 is used to acquire the dual-aircraft rotation take-off and landing mode. When the dual-aircraft rotation take-off and landing mode is single-aircraft rotation take-off and landing, the first UAV is controlled to take off, and RGB images and infrared images are acquired by the multimodal visual perception module installed on the lower part of the first UAV according to the preset acquisition cycle. After extracting the modal features respectively, multi-stage attention feature fusion is performed to generate dynamic fusion visual features.

[0099] The flight attitude correction control command acquisition module 3 is used to calculate the offset between the center of the first UAV and the visual positioning mark area based on the spatial relationship between the dynamic fused visual features and the visual positioning mark area when the first UAV returns, and generate flight attitude correction control commands according to the offset.

[0100] The second UAV takeoff command execution module 4 is used to control the first UAV to land in conjunction with the flight attitude correction control command, and synchronously control the water hose recovery device to uniformly retract the water hose in a guided manner according to the dynamic water hose tension adjustment parameters, so as to complete the landing of the first UAV and trigger the second UAV takeoff command.

[0101] Furthermore, the dynamic fusion visual feature acquisition module 2 is used to perform the following steps: A multimodal visual perception module acquires RGB and infrared images according to a preset acquisition cycle to obtain RGB image sequences and infrared image sequences. Modal features are extracted from the RGB and infrared image sequences respectively to obtain RGB image feature sequences and infrared image feature sequences. A first-stage self-attention feature fusion is performed on the RGB and infrared image feature sequences respectively to obtain RGB image fusion features and infrared image feature fusion features. A second-stage mutual attention feature fusion is performed on the RGB and infrared image feature fusion features to obtain dynamic fused visual features.

[0102] Furthermore, the dynamic fusion visual feature acquisition module 2 is used to perform the following steps: A pre-constructed multimodal feature extractor is used to extract modal features from the RGB image sequence and the infrared image sequence respectively, thereby obtaining the RGB image feature sequence and the infrared image feature sequence.

[0103] Furthermore, the dynamic fusion visual feature acquisition module 2 is used to perform the following steps: Using RGB image feature sequences and infrared image feature sequences as inputs, a self-attention encoder module is used to construct context dependencies, obtaining RGB image feature dependencies and infrared image feature dependencies. Based on these dependencies, multi-head self-attention computation is performed on the RGB and infrared image feature sequences respectively, yielding RGB image attention feature sequences and infrared image attention feature sequences. The RGB and infrared image attention feature sequences are then concatenated and linearly transformed, and residual connections and normalization are performed with the RGB and infrared image feature sequences to obtain RGB image fusion features and infrared image fusion features.

[0104] Furthermore, the dynamic fusion visual feature acquisition module 2 is used to perform the following steps: Obtain the feature similarity set of the RGB image fusion feature and the infrared image feature fusion feature; normalize the feature similarity set and construct a multimodal adjacency matrix based on the normalization result; use the multimodal adjacency matrix to convolve the RGB image fusion feature and the infrared image feature fusion feature respectively, and perform mean processing on the convolution result to obtain the dynamic fused visual feature.

[0105] Furthermore, the flight attitude correction control command acquisition module 3 is used to perform the following steps: The system matches the dynamically fused visual features with the preset landing area features of the visual positioning marker area. If the match is successful, the desktop optical landing positioning sensor installed on the bottom of the first UAV is activated. It combines phased array detection to capture and process the flashing signal of the LED visible light coded beacon set to obtain the signal decoding result. Based on the signal decoding result, least squares fitting is performed to determine the current position of the first UAV. The current position is offset from the center position of the visual positioning marker area to obtain the offset amount. The offset amount is sent to the flight control module to generate flight attitude correction control commands.

[0106] Furthermore, the flight attitude correction control command acquisition module 3 is used to perform the following steps: The phased array detects and captures the set of flashing signals of the LED visible light coded beacon set; the flashing signals are combined and photoelectrically converted, and component filtering is performed using a bandpass filter to obtain a set of filtered photoelectric signals; the filtered photoelectric signals are amplified using a low-noise amplifier and differential calculation is performed to determine the set of beacon incident angles and the set of relative distances, and the signal decoding result is obtained by summarizing them.

[0107] Furthermore, the second UAV takeoff command execution module 4 is used to perform the following steps: The flight speed and altitude information of the first UAV are acquired; based on the flight speed and altitude information, the water belt tension adjustment parameters are identified, and the dynamic water belt tension adjustment parameters are determined.

[0108] Furthermore, the second UAV takeoff command execution module 4 is used to perform the following steps: When the dual-drone alternating take-off and landing mode is in the dual-drone alternating take-off and landing mode, the first UAV and the second UAV are controlled to take off in sequence and land in reverse order.

[0109] The take-off and landing control platform for a vehicle-mounted dual-unit firefighting and rescue drone provided in this embodiment of the invention can execute the take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0110] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0111] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for controlling the take-off and landing of a vehicle-mounted dual-unit firefighting and rescue drone, characterized in that, The method includes: Deploy the container-structured vehicle-mounted drone take-off and landing platform, activate the LED visible light coded beacon sets located at the four corners of the platform, and construct a visual positioning identification area; The dual-drone rotation take-off and landing mode is obtained. When the dual-drone rotation take-off and landing mode is single-drone rotation take-off and landing, the first drone is controlled to take off. The multimodal visual perception module installed on the lower part of the first drone acquires RGB images and infrared images according to a preset acquisition cycle. After extracting modal features respectively, multi-stage attention feature fusion is performed to generate dynamic fused visual features. When the first UAV returns, based on the spatial relationship between the dynamically fused visual features and the visual positioning marker area, the offset between the first UAV and the center of the visual positioning marker area is calculated, and a flight attitude correction control command is generated according to the offset. The flight attitude correction control command is combined with the first UAV to perform landing, and the water hose recovery device is synchronously controlled to uniformly retract the water hose in a guided manner according to the dynamic water hose tension adjustment parameters, so as to complete the landing of the first UAV and trigger the take-off command of the second UAV.

2. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 1, characterized in that, The multimodal visual perception module installed on the underside of the first UAV acquires RGB and infrared images according to a preset acquisition cycle. After extracting modal features, multi-stage attention feature fusion is performed to generate dynamically fused visual features, including: The multimodal visual perception module acquires RGB and infrared images according to a preset acquisition cycle, thus obtaining RGB image sequences and infrared image sequences. Modal features are extracted from the RGB image sequence and the infrared image sequence respectively to obtain the RGB image feature sequence and the infrared image feature sequence; The first-stage self-attention feature fusion is performed on the RGB image feature sequence and the infrared image feature sequence respectively to obtain the RGB image fusion feature and the infrared image feature fusion feature; A second-stage mutual attention feature fusion is performed on the RGB image fusion features and the infrared image fusion features to obtain dynamic fused visual features.

3. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 2, characterized in that, A pre-constructed multimodal feature extractor is used to extract modal features from the RGB image sequence and the infrared image sequence respectively, thereby obtaining the RGB image feature sequence and the infrared image feature sequence.

4. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 2, characterized in that, The first-stage self-attention feature fusion is performed on the RGB image feature sequence and the infrared image feature sequence respectively to obtain RGB image fusion features and infrared image feature fusion features, including: Using RGB image feature sequences and infrared image feature sequences as inputs, the context dependency relationship is constructed through a self-attention encoder module to obtain the RGB image feature dependency relationship and the infrared image feature dependency relationship; Based on the feature dependencies of RGB images and infrared images, multi-head self-attention calculation is performed on the feature sequences of RGB images and infrared images respectively to obtain the attention feature sequences of RGB images and infrared images. The attention feature sequences of RGB images and infrared images are concatenated and then linearly transformed. The resulting sequences are then residually connected and normalized with the RGB and infrared image feature sequences to obtain the RGB image fusion features and the infrared image fusion features.

5. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 4, characterized in that, A second-stage mutual attention feature fusion is performed on the RGB image fusion features and the infrared image fusion features to obtain dynamic fused visual features, including: Obtain the feature similarity set between the RGB image fusion features and the infrared image feature fusion features; The feature similarity set is normalized, and a multimodal adjacency matrix is ​​constructed based on the normalization result. The RGB image fusion feature and the infrared image fusion feature are convolved using the multimodal adjacency matrix, and the convolution results are averaged to obtain the dynamic fused visual features.

6. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 1, characterized in that, When the first UAV returns, based on the spatial relationship between the dynamically fused visual features and the visual positioning marker area, the offset between the center of the first UAV and the center of the visual positioning marker area is calculated, and a flight attitude correction control command is generated according to the offset, including: The dynamic fusion visual features are matched with the preset landing area features of the visual positioning marker area. If the match is successful, the desktop optical landing positioning sensor installed on the bottom of the first UAV is activated. The phased array detection captures the flashing signal of the LED visible light coded beacon set and processes it to obtain the signal decoding result. Based on the signal decoding results, least squares fitting is performed to determine the current position of the first UAV; The offset is calculated by comparing the current position with the center position of the visual positioning mark area; The offset is sent to the flight control module to generate flight attitude correction control commands.

7. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 6, characterized in that, The desktop optical landing and positioning sensor installed on the bottom of the first UAV is then activated. Combined with phased array detection, it captures and processes the flashing signals of the LED visible light coded beacon set to obtain the signal decoding results, including: The phased array detector captures the set of flash signals of the LED visible light coded beacon set; The flickering signals are combined and photoelectrically converted, and then component filtering is performed using a bandpass filter to obtain a set of filtered photoelectric signals. The filtered photoelectric signal is amplified using a low-noise amplifier and differential calculation is performed to determine the set of beacon incident angles and the set of relative distances. The signal decoding result is then obtained by summarizing these results.

8. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 1, characterized in that, Combining the flight attitude correction control commands with the first UAV's landing control, and synchronously controlling the hose recovery device to uniformly and guide the hose reeling in according to the dynamic hose tension adjustment parameters, the first UAV's landing is completed, including: Obtain the flight speed and altitude information of the first drone; Based on the flight speed and altitude information, the water belt tension adjustment parameters are identified, and the dynamic water belt tension adjustment parameters are determined.

9. The take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in claim 8, characterized in that, When the dual-drone alternating take-off and landing mode is in the dual-drone alternating take-off and landing mode, the first UAV and the second UAV are controlled to take off in sequence and land in reverse order.

10. A take-off and landing control platform for a vehicle-mounted dual-unit firefighting and rescue drone, characterized in that, The platform is used to implement the take-off and landing control method for a vehicle-mounted dual-unit firefighting and rescue drone as described in any one of claims 1-9, the platform comprising: The visual positioning identification area construction module is used to deploy a container-structured vehicle-mounted UAV take-off and landing platform, activate the LED visible light coded beacon sets located at the four corners of the platform, and construct the visual positioning identification area; The dynamic fusion visual feature acquisition module is used to acquire the dual-aircraft rotation take-off and landing mode. When the dual-aircraft rotation take-off and landing mode is single-aircraft rotation take-off and landing, the module controls the first UAV to take off and acquires RGB images and infrared images according to a preset acquisition cycle through the multimodal visual perception module installed on the lower part of the first UAV. After extracting modal features respectively, multi-stage attention feature fusion is performed to generate dynamic fusion visual features. The flight attitude correction control command acquisition module is used to calculate the offset between the center of the first UAV and the visual positioning mark area based on the spatial relationship between the dynamic fused visual features and the visual positioning mark area when the first UAV returns to home, and generate flight attitude correction control commands according to the offset. The second UAV takeoff command execution module is used to control the first UAV to land in conjunction with the flight attitude correction control command, and synchronously control the hose recovery device to uniformly retract the hose in a guided manner according to the dynamic hose tension adjustment parameters, so as to complete the landing of the first UAV and trigger the second UAV takeoff command.

Citation Information

Patent Citations

  • Precise landing control system and method based on multi-information fusion

    CN109407708A

  • Radar and infrared fused unmanned aerial vehicle landing control method and device

    CN111176323A

  • Target detection method and device based on multi-modal image fusion

    CN114694001A

  • Document image entity extraction method and device and storage medium

    CN116486420A

  • Multi-ship tracking method and device based on multi-modal information fusion

    CN118379328A