Unmanned aerial vehicle end side-based intelligent identification model adaptive updating method

CN122821284APending Publication Date: 2026-09-25XIDIAN UNIV HANGZHOU RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611005688.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这些方案在一定程度上推动了无人机智能化的发展,但普遍存在对通信条件依赖较强、模型适应性不足或数据安全风险较高等问题,难以同时满足复杂应用场景下对实时性、适应性与安全性的综合要求

Benefits of technology

[0004]为了解决现有技术中所存在的上述问题,本发明提供了一种基于无人机端侧的智能识别模型自适应更新方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821284A_ABST
    Figure CN122821284A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on unmanned aerial vehicle end side's intelligent identification model self-adapting updating method, applied to unmanned aerial vehicle, method includes: collection field perception data;With initial visual language recognition model based on field perception data carries out target identification, obtains confidence;If confidence is less than preset threshold, then utilize field perception data as training sample, update target parameter in initial visual language recognition model by back propagation algorithm, obtain final visual language recognition model;With final visual language recognition model carries out target identification.The method of the present application, when the confidence of target identification is insufficient, the target parameter in initial visual language recognition model is updated using the collected field perception data, so that the model learns the current scene information, and then the updated final visual language recognition model is used for target identification, which can improve the real-time performance, stability and adaptability of unmanned aerial vehicle in multiple scene conditions, and improve the target identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, specifically relating to an adaptive update method for an intelligent recognition model based on the UAV terminal. Background Technology

[0002] As the application of drones in security patrols, emergency rescue, environmental monitoring, and other fields continues to expand, drones are gradually evolving from simple image acquisition devices into intelligent terminals with perception, analysis, and decision-making capabilities. In practical applications, drones typically need to operate for extended periods in complex and changing environments, continuously identifying targets and assessing their status. This places higher demands on the adaptability, operational efficiency, and data security of edge-side intelligent recognition models.

[0003] Existing technologies for intelligent identification of drones mainly include cloud-based centralized identification, edge-side fixed model identification, and edge-cloud collaborative model updates. While these solutions have promoted the development of drone intelligence to some extent, they generally suffer from problems such as strong dependence on communication conditions, insufficient model adaptability, or high data security risks, making it difficult to simultaneously meet the comprehensive requirements of real-time performance, adaptability, and security in complex application scenarios. Summary of the Invention

[0004] To address the aforementioned problems in the existing technology, this invention provides an adaptive update method for an intelligent recognition model based on the UAV terminal.

[0005] The technical problem to be solved by this invention is achieved through the following technical solution: On the one hand, an adaptive update method for an intelligent recognition model based on the UAV edge is provided. The method is applied to UAVs and includes: Collect on-site sensing data; The initial visual language recognition model is used to identify targets based on the on-site perception data to obtain confidence scores; If the confidence level is less than a preset threshold, the on-site perception data is used as training samples, and the target parameters in the initial visual language recognition model are updated through the backpropagation algorithm to obtain the final visual language recognition model. The final visual language recognition model is used for target recognition.

[0006] In one embodiment, the initial visual language recognition model includes general fixed parameters and scene adaptation parameters; Updating the target parameters in the initial visual language recognition model using the backpropagation algorithm includes: updating the scene adaptation parameters in the initial visual language recognition model using the backpropagation algorithm.

[0007] In one embodiment, the output of the initial visual language recognition model is represented as follows: Out = WX + ABX; Where Out represents the output of the initial visual language recognition model, X represents the input scene perception data, W represents the general fixed parameters, and A and B represent low-rank matrices, and the low-rank matrices contain the scene adaptation parameters.

[0008] In one embodiment, the method further includes: The target parameters are sent to the main control platform.

[0009] In one embodiment, the method further includes: Send the current status information of the drone to the main control platform; The current status information includes at least one of the following: the average confidence level, the update record of the target parameter, the current target parameter, and the resource usage parameter.

[0010] In one embodiment, collecting on-site sensing data includes: Determine the data acquisition parameters based on the task requirements; The field sensing data is collected based on the data acquisition parameters.

[0011] In one embodiment, the data acquisition parameters include at least one of the following: acquisition frequency, resolution of the field sensing data, and type of the field sensing data.

[0012] In one embodiment, target recognition is performed based on the on-site perception data using an initial visual language recognition model, including: The field-sensed data is subjected to size standardization processing; Target recognition is performed using the initial visual language recognition model based on the size-normalized on-site perception data.

[0013] On the other hand, an electronic device is provided, comprising a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory, and the processor being used to execute program data to implement the steps in the above-mentioned method for adaptive updating of intelligent recognition models based on UAV edge devices.

[0014] On another aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps in the above-mentioned method for adaptive updating of intelligent recognition models based on UAV end-side.

[0015] This invention provides an adaptive update method for an intelligent recognition model based on an unmanned aerial vehicle (UAV) edge. The method, applied to a UAV, includes: collecting on-site perception data; performing target recognition using an initial visual language recognition model based on the on-site perception data to obtain a confidence score; if the confidence score is less than a preset threshold, using the on-site perception data as training samples, updating the target parameters in the initial visual language recognition model through a backpropagation algorithm to obtain a final visual language recognition model; and performing target recognition using the final visual language recognition model. This method, when the confidence score for target recognition is insufficient, updates the target parameters in the initial visual language recognition model using the collected on-site perception data, enabling the model to learn current scene information. Then, the updated final visual language recognition model is used for target recognition, improving the real-time performance, stability, and adaptability of the UAV under multiple scene conditions, and enhancing target recognition accuracy.

[0016] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating an adaptive update method for an intelligent recognition model based on a UAV terminal, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0019] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the intelligent recognition model adaptive update method based on the UAV edge of this application. The method of this embodiment is applied to a UAV and specifically includes: Step S1: Collect on-site sensing data.

[0020] Specifically, the drone is equipped with an image acquisition module, such as a visible light camera or a high-definition video sensor.

[0021] The drone flies along a preset cruise route. During the flight, the data acquisition module continuously collects on-site perception data and continuously caches the raw frame data to the onboard local flash memory to facilitate target identification.

[0022] It should be noted that high-definition camera sensors acquire visible light images of natural scenes in real time. The original output formats include camera native encoding formats such as RAW, YUV, and BMP. Different camera models have different output bitrates and pixel formats. The RAW, YUV, and BMP native camera formats are batch-converted to the model's standard input RGB and JPG formats, removing redundant header information and invalid auxiliary parameters to reduce data volume.

[0023] During the collection of on-site sensing data, different data acquisition parameters are used for different task requirements. Therefore, in one embodiment, the data acquisition parameters are determined according to the task requirements; the on-site sensing data is then collected based on the data acquisition parameters. The data acquisition parameters include at least one of the following: acquisition frequency, resolution of the on-site sensing data, and type of on-site sensing data (e.g., image or video).

[0024] For example, in power line inspection tasks, the acquisition frequency is set to 2 frames per second. The drone cruises along the tower at a constant speed and samples at intervals to avoid repeatedly collecting a large number of homogeneous images in a short period of time, which would cause storage waste. 4K high-definition resolution imaging is used to identify minor defects, including broken insulators, broken strands of conductors, and floating objects entangled in the line, based on high-precision images. Static high-definition images are the main focus, and short segments of on-site video are recorded only when abnormal points on the line are identified for preservation and evidence collection.

[0025] For example, in urban security patrol missions, the acquisition frequency is set to 10 frames per second for high-frequency acquisition, continuously capturing instantaneous dynamic scenes such as people climbing over walls, gathering in groups, and open flames; 1080P resolution is used to balance recognition processing speed and target outline clarity, and is adapted to real-time inference on the airborne end side; continuous video stream output throughout the process is pushed to the model frame by frame in real time for target recognition.

[0026] For example, in the task of monitoring floating objects in rivers and lakes, the acquisition frequency is set to 1 frame / 3 seconds, and drones are used to sweep the river at high altitudes and low altitudes over a wide area, reducing the sampling frequency and enabling rapid inspection of large water areas; 2K resolution is adopted to meet the needs of large-scale screening of floating garbage and abnormal sewage outlets; all images are stored as static aerial images.

[0027] Step S2: Use the initial visual language recognition model to identify the target based on the on-site perception data and obtain the confidence level.

[0028] In one specific embodiment, the field perception data is subjected to size standardization processing; the initial visual language recognition model is then used to perform target recognition based on the size-standardized field perception data. Specifically, original images of different resolutions are uniformly scaled and cropped to the model's fixed input size (preferably 640×640), removing invalid black borders and redundant edge areas to adapt to the basic model's input specifications; the formatted image data circulates only in onboard memory, with one path fed into the initial visual language recognition model for real-time target recognition, and the other path archived locally as a training data source for subsequent model fine-tuning.

[0029] Specifically, the initial visual language recognition model synchronously calculates the classification confidence score Conf for each detected target. The confidence score ranges from [0, 1], with values ​​closer to 1 indicating higher accuracy in the model's target recognition. The confidence score is obtained by normalizing the classification output logistic value using the softmax function, representing the probability that the object within the current candidate box belongs to the corresponding labeled category.

[0030] Step S3: If the confidence level is less than a preset threshold, the on-site perception data is used as training samples, and the target parameters in the initial visual language recognition model are updated through the backpropagation algorithm to obtain the final visual language recognition model.

[0031] If the confidence level is lower than a preset threshold, the initial visual language recognition model is updated. For example, under normal, clear weather conditions, the recognition confidence level can reach 0.92. However, when the weather changes, such as in the event of heavy fog, the confidence level will drop significantly when performing target recognition based on the collected on-site perception images with fog, falling below the preset threshold. When the confidence level is lower than the preset threshold, the on-site perception images with fog are used as training samples, and the target parameters in the initial visual language recognition model are updated through the backpropagation algorithm to obtain the final visual language recognition model. This allows the model to learn the ability to recognize targets in foggy scenes, and using the final visual language model to perform target recognition in foggy scenes can greatly improve the confidence level. This improves the real-time performance, stability, and adaptability of the UAV under multiple scene conditions, and enhances the accuracy of target recognition.

[0032] In one specific embodiment, the initial visual language recognition model includes general fixed parameters and scene adaptation parameters; when updating the target parameters, the scene adaptation parameters in the initial visual language recognition model are updated through a backpropagation algorithm.

[0033] Specifically, the output of the initial visual language recognition model is represented as follows: Out = WX + ABX; Where Out represents the output of the initial visual language recognition model, X represents the input scene perception data, W represents the general fixed parameters, and A and B represent low-rank matrices, and the low-rank matrices contain the scene adaptation parameters.

[0034] Specifically, when updating the scene adaptation parameters in the initial visual language recognition model, only the low-rank matrices A and B need to be updated. This significantly reduces the number of model parameters and computational complexity. During training, by updating only the parameters in the low-rank matrices, direct updates to the original large-scale matrices are avoided, thereby reducing the computational burden and maintaining the model's performance level while reducing computational complexity and resource consumption.

[0035] Step S4: Use the final visual language recognition model to perform target recognition.

[0036] It should be noted that the confidence level mentioned in the above embodiments can be the confidence level of a single target recognition, or the average confidence level of multiple consecutive target recognitions. When the confidence level is greater than or equal to a preset threshold, it is determined that the current environment matches the model's pre-training scene well, the recognition effect is qualified, and model fine-tuning is not initiated; the model maintains its original low-rank matrices A and B unchanged. When the confidence level is less than the preset threshold, it is determined that the environment has changed significantly, such as fog, rain, or snow, and the initial visual language recognition model cannot adapt to the current environmental conditions, resulting in insufficient recognition effect. At this time, model fine-tuning is initiated, using the currently collected on-site perception data to adjust the model's low-rank matrices A and B, enabling the model to learn the ability to perform target recognition in the current environment, thereby improving the accuracy of target recognition.

[0037] The intelligent recognition model adaptive update method based on the UAV terminal of this invention can meet the recognition requirements of different application scenarios and provide a data foundation for subsequent model inference and parameter adaptive updates. Furthermore, basic tasks such as target detection and target recognition are implemented directly on the UAV terminal, without needing to be transmitted to the main control platform, thus ensuring data security.

[0038] Furthermore, the method of this application, when the model's adaptability to different scenarios is insufficient, utilizes real-time collected on-site perception data to update the parameters of the deployed model, enabling the model to adapt to different application scenarios. Additionally, when updating the model parameters, only the low-rank matrix is ​​updated. This low-rank matrix contains scenario-adaptive parameters, thus reducing computational complexity and storage overhead, enabling the model to respond quickly to scene changes, ensuring model performance stability, and preventing performance fluctuations caused by frequent parameter changes during long-term system operation, thereby improving overall operational reliability.

[0039] In one embodiment of this application, the method further includes: sending the current status information of the UAV to the main control platform; the current status information includes at least one of: the average confidence level, the update record of the target parameter, the current target parameter, and the resource usage parameter.

[0040] The method includes: sending the target parameters to the main control platform.

[0041] Specifically, sending the average confidence score of the drone to the main control platform allows the platform to remotely assess the drone's on-site recognition performance. If the overall confidence score is high, it indicates that the model is well adapted to the current environment and no manual intervention is needed. If the confidence score is significantly low, it indicates a severe change in the on-site environment, and control commands can be remotely issued to adjust the flight path. The recognition quality can be assessed without retrieving the original on-site images.

[0042] Sending the drone's target parameters (i.e., the low-rank matrix) and their update records to the main control platform allows the platform to save records of all parameter modifications, distinguishing between specific low-rank parameters for different scenarios such as heavy fog, backlight, rain, and snow, thus forming a scenario parameter database. When the drone returns to base, its onboard storage is formatted, and it flies back to the same work area, the main control platform directly sends the historically backed-up low-rank matrix, which the drone loads directly, eliminating the 1-2 minute adaptation process of on-site fine-tuning.

[0043] Sending the drone's resource usage parameters to the main control platform allows the platform to understand the drone's CPU usage, memory consumption, and storage space availability, enabling real-time monitoring of the model's operational load. In cases of resource overload, the image acquisition frame rate and resolution can be remotely reduced to prevent recognition lag and crashes caused by exhausting onboard computing power. Conversely, when resources are plentiful, sampling accuracy can be temporarily increased and recognition details optimized.

[0044] In other applications, R&D personnel can also rely on the massive scene parameters and confidence change data collected by the main control platform to iteratively optimize the pre-trained backbone model and continuously upgrade the factory basic model. This is an offline optimization in the background and does not affect the real-time operation of the drone on site.

[0045] From a system architecture perspective, the method described in this application eliminates the need for drones to transmit collected field perception data back to the main control platform, reducing the risk of data leakage and ensuring data security and privacy protection capabilities.

[0046] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, such as... Figure 2 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. Memory 603 is used to store computer programs; When the processor 601 executes the program stored in the memory 603, it implements the steps of the above-described adaptive update method for the intelligent recognition model based on the UAV terminal.

[0047] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0048] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0049] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0050] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0051] The present invention also provides a computer-readable storage medium. A computer program is stored in this computer-readable storage medium, and when executed by a processor, the computer program implements the steps of the above-described adaptive update method for the intelligent recognition model based on the UAV edge.

[0052] Optionally, the computer-readable storage medium may be non-volatile memory (NVM), such as at least one disk storage device.

[0053] Optionally, the computer-readable storage medium may also be at least one storage device located remotely from the aforementioned processor.

[0054] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of the above-described adaptive update method for the intelligent recognition model based on the UAV terminal.

[0055] It should be noted that, for the embodiments of electronic devices / storage media / computer programs, since they are basically similar to the method embodiments, the description is relatively simple. For relevant parts, please refer to the description of the method embodiments. All embodiments of the above-described intelligent recognition model adaptive update method based on UAV terminal side are applicable to the electronic device and storage medium, and can achieve the same or similar beneficial effects.

[0056] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0057] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0058] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0059] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus (devices), or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects, all of which are collectively referred to herein as "modules" or "systems." Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The computer program may be stored / distributed in a suitable medium, provided with or as part of other hardware, or may take other distribution forms, such as via the Internet or other wired or wireless telecommunications systems.

[0060] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0061] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0063] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. An adaptive update method for an intelligent recognition model based on an unmanned aerial vehicle (UAV) terminal, characterized in that, The method is applied to a drone, and the method includes: Collect on-site sensing data; The initial visual language recognition model is used to identify targets based on the on-site perception data to obtain confidence scores; If the confidence level is less than a preset threshold, the on-site perception data is used as training samples, and the target parameters in the initial visual language recognition model are updated through the backpropagation algorithm to obtain the final visual language recognition model. The final visual language recognition model is used for target recognition.

2. The method according to claim 1, characterized in that, The initial visual language recognition model includes general fixed parameters and scene adaptation parameters; Updating the target parameters in the initial visual language recognition model using the backpropagation algorithm includes: updating the scene adaptation parameters in the initial visual language recognition model using the backpropagation algorithm.

3. The method according to claim 2, characterized in that, The output of the initial visual language recognition model is represented as follows: Out = WX + ABX; Where Out represents the output of the initial visual language recognition model, X represents the input scene perception data, W represents the general fixed parameters, and A and B represent low-rank matrices, and the low-rank matrices contain the scene adaptation parameters.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The target parameters are sent to the main control platform.

5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Send the current status information of the drone to the main control platform; The current status information includes at least one of the following: the average confidence level, the update record of the target parameter, the current target parameter, and the resource usage parameter.

6. The method according to any one of claims 1 to 3, characterized in that, Collect on-site sensing data, including: Determine the data acquisition parameters based on the task requirements; The field sensing data is collected based on the data acquisition parameters.

7. The method according to claim 6, characterized in that, The data acquisition parameters include at least one of the following: acquisition frequency, resolution of the field sensing data, and type of the field sensing data.

8. The method according to claim 1, characterized in that, Target recognition is performed based on the on-site perception data using an initial visual language recognition model, including: The field-sensed data is subjected to size standardization processing; Target recognition is performed using the initial visual language recognition model based on the size-normalized on-site perception data.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor coupled to each other. The processor is used to execute program instructions stored in the memory and to execute program data to implement the steps in the adaptive update method of the intelligent recognition model based on the UAV end-side as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the adaptive update method for intelligent recognition models based on UAV end-side as described in any one of claims 1 to 8.