Unmanned aerial vehicle autonomous navigation method and storage medium

By acquiring global mission information and combining it with visual language models for environmental perception, the problem of insufficient intelligence and efficiency in autonomous navigation of drones is solved, achieving more intelligent and accurate navigation effects.

CN120669725APending Publication Date: 2025-09-19CHINA WEST NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510870172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing drone autonomous navigation technology lacks intelligent and efficient navigation capabilities when performing target tasks, making it difficult to accurately complete designated tasks.

Method used

By acquiring global task information, outputting local task information at a first frequency based on the global task information, and outputting action control instructions at a second frequency, combined with onboard sensors and image acquisition devices, and using visual language models for environmental perception and reasoning, autonomous navigation of the UAV is achieved.

Benefits of technology

It improves the intelligence and accuracy of drone autonomous navigation, enabling it to complete tasks in a way that is closer to human thinking patterns, and improves flight efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669725A_ABST
    Figure CN120669725A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an unmanned aerial vehicle autonomous navigation method and a storage medium, and the method comprises the steps: obtaining global task information; the global task information comprises destination information; outputting one or more pieces of local task information at a first frequency at least based on the global task information, and outputting one or more action control instructions at a second frequency; the local task information comprises intermediate position information in a process from a preset initial position to a destination, and the action control instruction is used for controlling the state of the unmanned aerial vehicle; wherein the second frequency is not less than the first frequency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of drone navigation technology, and in particular to a drone autonomous navigation method and storage medium. Background Art

[0002] Drones (or unmanned aerial vehicles) have been widely used in various fields due to their flexibility, efficiency, and low cost. In some applications, drones are required to fly to a specific location to complete a target task. For example, drones can be flown over disaster areas to deliver relief supplies, or they can be flown over farmland to spray pesticides.

[0003] The completion of drone target missions depends on drone navigation. In order to complete target missions more efficiently and accurately, drones are expected to have more intelligent and efficient autonomous navigation and flight capabilities.

[0004] Therefore, it is necessary to provide a method and storage medium for autonomous navigation of a UAV. Summary of the Invention

[0005] One or more embodiments of the present specification provide a method for autonomous navigation of an unmanned aerial vehicle (UAV), which is executed by one or more processors, and the method includes: obtaining global mission information; the global mission information includes destination information; outputting one or more local mission information at a first frequency and outputting one or more action control instructions at a second frequency based at least on the global mission information; the local mission information includes intermediate position information in the process of reaching the destination from a preset starting position, and the action control instructions are used to control the state of the UAV.

[0006] One or more embodiments of the present specification provide an autonomous navigation system for an unmanned aerial vehicle (UAV), the system comprising: an acquisition module for acquiring global mission information; the global mission information including destination information; an output module for outputting one or more local mission information at a first frequency and one or more action control instructions at a second frequency based at least on the global mission information; the local mission information including intermediate position information in the process of reaching the destination from a preset starting position, and the action control instructions being used to control the state of the UAV.

[0007] One or more embodiments of this specification provide a drone autonomous navigation device, including a processor, wherein the processor is configured to execute the above-mentioned drone autonomous navigation method.

[0008] One or more embodiments of this specification provide a storage medium storing computer instructions. When a processor executes at least part of the computer instructions in the storage medium, the above-mentioned drone autonomous navigation method can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein: Figure 1 is a schematic diagram of an application scenario of a drone autonomous navigation system according to some embodiments of this specification; Figure 2 is an exemplary flow chart of a method for autonomous navigation of a drone according to some embodiments of this specification; Figure 3 is an exemplary flow chart of a round of cyclic processing according to some embodiments of this specification; Figure 4 is an exemplary flow chart of determining environmental image information in a round of loop processing according to some embodiments of this specification; Figure 5 is an exemplary module diagram of a drone autonomous navigation system according to some embodiments of this specification; Figure 6 is an exemplary schematic diagram of an attention area according to some embodiments of the present specification. DETAILED DESCRIPTION

[0010] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0011] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.

[0012] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0013] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0014] Figure 1 This is a schematic diagram of an application scenario of a drone autonomous navigation system according to some embodiments of this specification.

[0015] like Figure 1 As shown, the application scenario 100 of the drone autonomous navigation system may include a drone 110, a processor 120, a terminal device 130, a storage device 140 and a network 150.

[0016] UAV 110 refers to an aircraft capable of autonomous flight or remote control without the need for direct human control. UAV 110 can automate specific tasks by integrating sensors, control systems, power units, and communication modules. Sensors can include various types of sensors, such as visual sensors, radar sensors, and infrared sensors. For example, an onboard image acquisition device can be configured within the UAV based on the visual sensor to capture images. In some embodiments, UAV 110 may include multi-rotor (e.g., quadcopter), fixed-wing (long-endurance), vertical take-off and landing (VTOL) hybrid UAVs, and others. In some embodiments, UAV 110 can be used in a variety of fields, including civilian applications such as geographic surveying and mapping, agricultural plant protection, power inspections, aerial filming for film and television, and emergency supply delivery; military applications such as battlefield reconnaissance, electronic countermeasures, and precision strikes; and scientific applications such as meteorological observation, wildlife tracking, and polar expeditions.

[0017] Processor 120 can process data and / or information obtained from drone 110, terminal device 130, storage device 140, or other components of the system's application scenario 100. For example, processor 120 can obtain global mission information from terminal device 130; the global mission information includes destination information. Based at least on the global mission information, processor 120 can output one or more local mission information at a first frequency and one or more motion control instructions at a second frequency. The local mission information includes intermediate position information from a preset starting position to a destination, and the motion control instructions are used to control the drone's state. The second frequency is no less than the first frequency. In some embodiments, processor 120 can be local or remote. For example, processor 120 can access information and / or data from drone 110, terminal device 130, and / or storage device 140 via network 150. In some embodiments, processor 120 can be integral to drone 110, for example, integrated into a control system.

[0018] Terminal device 130 may be a user-operated terminal. Terminal device 130 can input user commands and operations, or output drone status. In some embodiments, a user can input global mission information through terminal device 130 to control drone 110. Throughout this specification, "user" and "user terminal" are used interchangeably. In the embodiments of this specification, terminal device 130 may include a mobile device 130-1, a tablet computer 130-2, a laptop computer 130-3, or any combination thereof. In some embodiments, terminal device 130 may also be a dedicated control device for drone 110. In this specification, a user may be the operator of drone 110, such as a technician or a general user.

[0019] Storage device 140 can store data, instructions, and / or any other information. In some embodiments, storage device 140 can store data obtained from drone 110 and / or processor 120. For example, various input and output data related to autonomous drone navigation, as well as drone control instructions, can be stored. In some embodiments, storage device 140 can include one or more storage components, each of which can be a standalone device or part of another device. In some embodiments, storage device 140 can include random access memory (RAM), read-only memory (ROM), mass storage, removable storage, volatile read-write memory, or any combination thereof. Exemplary mass storage devices include magnetic disks, optical disks, solid-state disks, and the like. In some embodiments, storage device 140 can be implemented on a cloud platform.

[0020] Network 150 may comprise any suitable network capable of facilitating information and / or data exchange. In some embodiments, at least one component of the drone autonomous navigation system application scenario 100 (e.g., drone 110, processor 120, terminal device 130, storage device 140) may exchange information and / or data with at least one other component of the system application scenario 100 via network 150. For example, processor 120 may obtain images output by an onboard image acquisition device from drone 110 via network 150.

[0021] It should be noted that the application scenario 100 of the autonomous drone navigation system is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art will readily appreciate that various modifications and variations can be made based on the description of this specification. For example, the application scenario 100 of the autonomous drone navigation system may also include a database. For another example, the application scenario 100 of the autonomous drone navigation system may implement similar or different functions on other devices. However, such modifications and variations do not deviate from the scope of this specification.

[0022] Figure 2 This is an exemplary flow chart of the autonomous navigation method of a drone according to some embodiments of this specification. Figure 2 As shown, the process 200 includes the following steps. In some embodiments, the process 200 may be executed by a processor (such as the processor 120).

[0023] Step 202: Obtain global task information.

[0024] Global mission information is information related to a user-specified mission. In some embodiments, global mission information includes destination information. For example, a global mission might be a user specifying that a drone fly to Mountain Area A (latitude: 28.6543° N, longitude: 112.3456° E, altitude: 1,856 meters) to drop supplies. The global mission information would then include the drone's destination (Mountain Area A).

[0025] Destination information can be a three-dimensional spatial area or a precise three-dimensional spatial point. Destination information can be represented by three-dimensional spatial coordinate values ​​(such as longitude, latitude, and altitude), or it can be a place name. When the destination information is a place name, the three-dimensional spatial coordinate information of the destination can be further determined by combining it with map data.

[0026] In some embodiments, global task information can be input by the user through an input device (such as through an external input device, a terminal device used by the user, etc.). For example, the user can specify the global task information through text input, voice input, or picture input; it can also be input in advance by the user and stored in a storage device (such as a database, etc.), and obtained by reading from the storage device.

[0027] Step 204 : output one or more local task information at a first frequency and output one or more motion control instructions at a second frequency based at least on the global task information.

[0028] Local mission information refers to the intermediate tasks that need to be completed in the process of completing the global mission, such as the intermediate location information that the drone will fly to in the next period or at the next moment.

[0029] In some embodiments, the local mission information includes intermediate position information during the process of reaching the destination from the preset starting position. The preset starting position can be the initial position of the drone's takeoff or the current position of the drone.

[0030] The intermediate position information may be represented by a precise three-dimensional spatial coordinate value, or may be a relative value, such as 50 meters to the right, 100 meters to the northeast, or 25 meters down in height.

[0031] Action control commands are instructions / information used to control the drone's state. Controlling a drone's state can involve executing a targeted action to achieve a specific state. A targeted action can be the action the drone must perform to reach its destination, such as dropping supplies, conducting reconnaissance, landing, or hovering. For example, a user might specify that the drone fly to Mountain Area A (latitude: 28.6543° N, longitude: 112.3456° E, altitude: 1,856 meters) to complete a supplies drop. Another example of controlling a drone's state is executing intermediate actions related to the current local mission, such as backing up, rolling, pitching, obstacle avoidance, and deceleration.

[0032] The drone status includes flight altitude (if the drone has not taken off, the flight altitude can be 0), heading, posture (pitch angle, yaw angle and / or roll angle), IMU acceleration and angular acceleration, satellite navigation coordinates (the position of an object on the earth's surface or in space determined by the global navigation satellite system GNSS, such as GPS, Beidou, GLONASS, Galileo, etc.), and the drone's position (such as the current position of the drone determined by measuring changes in velocity vector and direction).

[0033] In some embodiments, the motion control instructions may specifically include speed control instructions and posture control instructions, etc. The speed control instructions and posture control instructions may be the same as the control instructions issued by the remote control to the drone controller when operating the drone remote control. Such a design can quickly and concisely realize the control coupling of the navigation system and the drone without understanding the underlying design or control principles of the drone.

[0034] In some embodiments, the first frequency may be a constant, ie, a constant frequency, or may be a variable, such as the first frequency is f1 during a period of time, and changes to f2 during another period of time.

[0035] In some embodiments, the second frequency may be a constant or a variable.

[0036] The magnitude of the first frequency and the second frequency may depend on the computing frequency f of the computing device that executes the drone autonomous navigation method (or navigation system) disclosed in the embodiments of this specification. The second frequency does not exceed the computing frequency f, and the second frequency is not less than the first frequency.

[0037] In some embodiments, the processor may perform at least one round of loop processing based on the global task information, and then output one or more local task information at a first frequency and one or more action control instructions at a second frequency according to the result of the loop processing.

[0038] For detailed description of the loop process, please refer to the following Figure 3 Description.

[0039] In some embodiments of this specification, the UAV autonomous navigation system can plan several local task information based on global task information, and plan UAV motion control instructions at a higher frequency based on the global task information and local task information, thereby enabling the UAV to perform autonomous navigation in a fast and slow "thinking" coupling manner, which is closer to the human thinking rules and improves the intelligence of autonomous navigation and flight.

[0040] Figure 3 This is an exemplary flow chart of a round of cyclic processing according to some embodiments of this specification. Figure 3 As shown, the process 300 includes the following steps. In some embodiments, the process 300 may be executed by a processor (such as the processor 120).

[0041] Step 302: Obtain an input sequence.

[0042] The input sequence is the input data of the planning model. The input sequence can be a text sequence with a preset format. For example, the input sequence can be text that is classified and arranged according to different data information.

[0043] In some embodiments, the input sequence includes the following information: global task information, local task information calculated from the previous loop processing, and motion planning information. In some embodiments, the input sequence includes at least the global task information. If the current round is the first loop processing, the local task and motion planning information are initialized.

[0044] Action planning information is the basic information for generating action control instructions. In some embodiments, the action planning information can be relatively rough relative information, such as left translation, right translation, forward translation, backward translation, ascent, descent, rotation (clockwise / counterclockwise rotation around the vertical axis of the drone), pitch, roll, and hover. Correspondingly, action control instructions are control instructions that are more directly related to components such as the drone's rotor motors. For example, if a drone has one or more rotors, the action control instructions can be to control the speed of one or more of the rotor motors, ultimately achieving the aforementioned planned actions such as left translation, rotation, and roll.

[0045] In some embodiments, the input sequence may also include drone status information of the drone and environmental image information.

[0046] Drone status information refers to various types of information related to the current state of the drone, including but not limited to flight altitude, heading, attitude (pitch, yaw and / or roll angles), IMU acceleration and angular acceleration, satellite navigation coordinates, and aircraft position information.

[0047] In some embodiments, drone status information may be determined based on output data from one or more onboard sensors on the drone. Onboard sensors refer to sensors mounted on the drone. These sensors may include altitude sensors, posture sensors, acceleration sensors, position sensors, and image sensors. This embodiment does not limit the specific sensor types.

[0048] Environmental image information may include information such as the spatial location of the drone and objects within that location. In some embodiments, environmental image information may be determined based on images captured by an onboard image acquisition device on the drone. Images output by the onboard image acquisition device on the drone may be images captured in real time by the onboard image acquisition device. Image types include, but are not limited to, natural images and infrared images.

[0049] Optionally, the environmental image information can be further determined based on a sub-image corresponding to the attention area in the image captured by the onboard image acquisition device. In some embodiments, a round of cyclic processing also includes: obtaining an image output by the onboard image acquisition device on the drone; determining the attention area in the image and obtaining one or more sub-images located in the attention area; and determining the environmental image information based on the one or more sub-images. For detailed descriptions, see Figure 4 Related description.

[0050] The attention region refers to one or more attention regions in an image, and the sub-image corresponding to the attention region refers to the portion of the image corresponding to the one or more attention regions.

[0051] For example, the sub-image corresponding to the attention area can be as follows Figure 6 As shown, Figure 6 The following is an exemplary schematic diagram of an attention region according to some embodiments of this specification. The black dashed circle illustrates the attention region in the image. In some alternative embodiments, the shape of the region may be rectangular, triangular, elliptical, pentagonal, or irregular. In some embodiments, the image output by the onboard image acquisition device may be divided into multiple grids, with the grids touched by the black dashed circle forming the sub-image corresponding to the attention region. Alternatively, the portion of the image output by the onboard image acquisition device captured by the black dashed circle may constitute the sub-image corresponding to the attention region.

[0052] In some embodiments, the attention region information can be determined based on an output sequence, which is the output result of the loop processing. For example, the attention region information can be determined based on the output sequence of the previous loop processing. For more information about the output sequence, please refer to the description of step 304 below.

[0053] Exemplarily, a round of cyclic processing may also include the following operations: acquiring an image output by an onboard image acquisition device on the drone; intercepting one or more sub-images located in the attention area in the image based on the attention area information; and determining the environmental image information based on the one or more sub-images.

[0054] In some embodiments, the environmental image information is determined based on the one or more sub-images, and the sub-image corresponding to the attention area itself can be used as the environmental image information, or the vector or string obtained by processing the attention area sub-image based on the encoding or convolution algorithm can be used as the environmental image information.

[0055] In this embodiment, the planning model's input sequence includes drone state information and environmental imagery, enabling the navigation system to further perceive the environment and perform reasoning based on this perceived environmental information, further improving the accuracy of autonomous navigation. For example, the planning model can accurately reason and calculate the drone's action plan based on the drone's state, environmental imagery, and local mission information.

[0056] Step 304: Process the input sequence through the planning model to obtain an output sequence.

[0057] The output sequence is the output result obtained by the planning model after processing the input sequence. The output sequence is also the output result of a round of cyclic processing. In some embodiments, the output sequence can include at least one of the following information: local task information calculated during the current round of cyclic processing and action planning information.

[0058] In some embodiments, the processor may input an input sequence into a planning model for processing, thereby obtaining an output sequence.

[0059] In some embodiments, the planning model's input data and output data are both sequences of pre-formatted language symbols. The planning model's input data, or input sequence, includes a global mission information segment, a local mission information segment, an action plan information segment, a drone state information segment, and an environmental image information segment. The planning model's output data, or output sequence, includes a flag bit, a reuse information segment, and an attention region information segment.

[0060] Language symbols refer to symbols with semantics, and symbols can be numbers, letters, Chinese characters, etc. When language symbols are letters or numbers, specific semantics can be agreed upon in advance for each symbol. For example, the letter L is agreed to represent left, and for another example, the number 1 is agreed to represent the exploration environment. When the language symbols are Chinese characters, they can use the original semantics of Chinese characters, or they can agree on other semantics similar to letters and numbers. In some embodiments, language symbols can even be formatted natural languages. Formatted natural language refers to natural language that has been processed according to a certain format. For example, the time format xx year xx month xx day is a formatted natural language; for another example, "longitude: aa, latitude: bb" is also a formatted natural language.

[0061] A language symbol sequence is a combination of one or more language symbols. For example, the first part of the planning model's input sequence (the global mission information segment) might be "Latitude: 28.6543° N, Longitude: 112.3456° E, 0," where Latitude: 28.6543° N, Longitude: 112.3456°, E represents the destination location, and "0" indicates the drop of supplies. Language symbols format the input / output sequences. By permuting and combining symbols, different symbol sequences can be generated, expressing different semantic information. This effectively compresses the planning model's input space (also known as the state space) and output space (also known as the action space), significantly improving model training efficiency.

[0062] In the input sequence: The global task information segment is used to store global task information and is the first part of the input sequence.

[0063] The local task information segment is used to store local task information and is the second part of the input sequence. The content of the local task information segment in the input sequence is determined based on the second part of the output sequence from the previous loop. If the second part of the output sequence from the previous loop contains local task information, the local task information segment in the input sequence from the current loop is updated. If the second part of the output sequence from the previous loop contains action planning information, the local task information segment in the input sequence from the current loop is not updated.

[0064] The action planning information segment is used to store action planning information and is the third part of the input sequence. The content of the action planning information segment in the input sequence is determined based on the second part of the output sequence from the previous loop. If the second part of the output sequence from the previous loop contains action planning information, the action planning information segment in the input sequence from the current loop is updated. If the second part of the output sequence from the previous loop contains local task information, the action planning information segment in the input sequence from the current loop is not updated.

[0065] The drone status information segment is used to store drone status information and is the fourth part of the input sequence.

[0066] The environmental image information segment is used to store environmental image information and is the fifth part of the input sequence.

[0067] In the output sequence, the value of the flag bit is used to indicate whether the value of the multiplexed information segment is local task information or action planning information.

[0068] The flag bit is the first part of the output sequence. When the value is the first value, such as "0", "L", "low frequency", etc., it indicates that the multiplexed information segment stores local task information. When the value is the second value, such as "1", "H", "high frequency", etc., it indicates that the multiplexed information segment stores action planning information.

[0069] The multiplexed information segment is the second part of the output sequence and is used to store local task information or action planning information.

[0070] The attention region information segment is the third part of the output sequence and is used to store the attention region information.

[0071] In some embodiments, the planning model can be a machine learning model, such as a deep learning model. Furthermore, the planning model can be a visual language model. A visual language model (VLM) is an artificial intelligence model that can simultaneously understand and process visual (image / video) and textual information. Using deep learning techniques, the VLM combines computer vision (CV) with natural language processing (NLP) to achieve cross-modal semantic understanding and generation. Examples of VLMs include, but are not limited to, CLIP, Flamingo, BLIP / BLIP-2, and GPT-4V.

[0072] In some embodiments, the planning model may be trained based on the first training sample with the identifier. Specifically, the first training sample with the identifier is input into the planning model, and the parameters of the planning model are updated through training.

[0073] In some embodiments, the first training sample may be a pre-collected sample input sequence. The specific content of the first training sample can refer to the description of the relevant input sequence.

[0074] In some embodiments, the first training sample can be obtained by extracting it from a data set related to autonomous navigation of the drone.

[0075] In some embodiments, the identifier may be a corresponding output sequence. In some embodiments, the identifier may be obtained by manual annotation.

[0076] In some embodiments, training can be performed based on the first training sample using various methods, such as a gradient descent method.

[0077] In some embodiments, training ends when a preset condition is met. The preset condition may include the number of iterations reaching a preset number (e.g., 10,000) or the convergence of the loss function used for model training. In some embodiments, when the output sequence does not include action planning information, the processor may directly process the output sequence to obtain the input sequence for the next round of loop processing and then proceed to the next round of loop processing. When the output sequence includes action planning information, the processor may execute step 306 to obtain action control instructions based on the action planning information.

[0078] Step 306: When the output sequence includes the action planning information, obtain the action control instruction based on the action planning information through a translation model.

[0079] In some embodiments, the input data of the translation model includes an action plan information segment and an environment image information segment. The input data of the translation model is a language symbol sequence in a preset format. For a detailed description of the language symbol sequence, please refer to the description of step 304 above.

[0080] In some embodiments, the processor may determine input data for the translation model based on the action planning information and the environmental image information. The action planning information segment within the translation model's input data is used to store the action planning information output by the planning model. The environmental image information segment within the translation model's input data is used to store environmental image information acquired by a peripheral program based on the attention region information output by the planning model.

[0081] In some embodiments, the processor may perform language formatting based on the action planning information and the environmental image information, thereby obtaining a language symbol sequence in a preset format.

[0082] The processor may input the input data of the translation model into the translation model for processing, so as to obtain the action control instruction by processing the input data through the translation model.

[0083] In some embodiments, the translation model can be a machine learning model, such as a deep learning model. Furthermore, the translation model can also be a visual language model. The specific type of the translation model can be the same as or different from the specific type of the planning model, and this embodiment does not limit this.

[0084] In some embodiments, the translation model may be trained based on the second training sample with the identifier. Specifically, the second training sample with the identifier is input into the translation model, and the parameters of the translation model are updated through training.

[0085] In some embodiments, the second training sample may be a pre-collected sequence of language symbols in a preset format, including a sample action planning information segment and a sample environment image information segment. The specific content of the second training sample can be found in the description of the input data of the translation model.

[0086] In some embodiments, the second training sample can be obtained by manually formatting the action planning information segment and the environment image information segment during the autonomous navigation of the UAV.

[0087] In some embodiments, the identifier may be a corresponding action control instruction. In some embodiments, the identifier may be obtained by manual annotation.

[0088] In some embodiments, training can be performed based on the second training sample using various methods, such as a gradient descent method.

[0089] In some embodiments, training ends when a preset condition is met, which may include the number of iterations reaching a preset number (e.g., 10,000) or the loss function used for model training converging.

[0090] In some embodiments, the planning model and translation model can be trained independently as described in the above embodiments, or jointly as described in the following embodiments. The joint training method can be reinforcement learning, supervised training, or a combination of supervised training and reinforcement learning. The following exemplifies three training methods.

[0091] Training method 1: Joint training combined with supervised training.

[0092] In some embodiments, the sample input sequence of the planning model and the sample output data of the translation model can be artificially constructed, and the sample input sequence is input into the planning model. Figure 3 After the process shown, the translation model outputs the action control instruction. The difference between the action control instruction output by the translation model and the sample output data is determined, and the parameters of the planning model and the translation model are adjusted to reduce the difference between the action control instruction output by the translation model and the sample output data.

[0093] Training method 2 - joint training combined with reinforcement learning.

[0094] During training, the navigation system can be made to perform Figure 3 The illustrated process operates a drone to obtain drone state information at multiple time points. The multiple time points can be determined based on the computation frequency f of the computing device in the aforementioned drone autonomous navigation method (or navigation system). The drone state information can be used as action (a) in reinforcement learning.

[0095] For each time point t, a reward value or action-value function value (qt) is determined based on the drone’s state information and the input sequence of the planning model, where the input sequence can be used as the state (st) in reinforcement learning.

[0096] During the training process, the model parameters of the planning model and the translation model are adjusted to maximize the reward value or action value function.

[0097] In some embodiments, the reward value or action-value function value can be obtained through a judgment model. The judgment model can be obtained based on neural network training. The input data of the judgment model is the drone state information of the drone and the input sequence of the planning model, and the output is the reward value or action-value function value. The training method of the judgment model can refer to the separate training method for the planning model or translation model mentioned above. The difference lies in the different training samples used. The training samples of the judgment model are sample drone state information and sample input sequences, and the output is manually annotated reward value or action-value function value.

[0098] In some embodiments, a reward value or action-value function value can also be manually assigned based on the current actual environment and the drone's flight state. For example, if the drone successfully avoids an obstacle using a smaller avoidance path, a larger reward value is assigned; if the drone collides with an obstacle, a negative reward value is assigned. For another example, if the drone's current action allows its flight path to better cover a farmland area, a positive reward value is assigned; if the drone's current action prevents its flight path from better covering the farmland area, a negative reward value is assigned.

[0099] Reinforcement training can be performed during the actual flight of the drone, or a virtual drone and a virtual environment can be constructed in a simulation environment (such as AirSim software) and reinforcement learning can be performed in the simulation. This embodiment does not limit this.

[0100] Training method 3 — Joint training combines supervised training and reinforcement learning.

[0101] The first stage is supervised training, and the training process can be similar to the above training method 1. After a certain degree of supervised training, the navigation system will have a certain degree of autonomous navigation capability.

[0102] The second phase is the reinforcement learning phase, and its training process is similar to the above training method 2. The second phase can be to reinforce the navigation system through actual flight scenarios, which will help improve its performance.

[0103] In this embodiment, the planning model and the translation model are implemented based on the visual language model, so that both can process language and images simultaneously, thereby improving the efficiency and accuracy of navigation reasoning.

[0104] Figure 4 This is an exemplary flow chart of determining environmental image information in a round of cyclic processing according to some embodiments of this specification. Figure 4 As shown, the process 400 includes the following steps. In some embodiments, the process 400 may be executed by a processor (such as the processor 120).

[0105] Step 402: Acquire an image output by an onboard image acquisition device on the UAV.

[0106] Step 404: determine an attention area in the image, and obtain one or more sub-images located in the attention area.

[0107] In some embodiments, the processor may determine the attention area in the image through the operations shown in the following embodiments.

[0108] The processor may determine a center point position of the attention area in the image; and determine an offset based on a flight speed and / or a flight altitude of the drone.

[0109] The center point position may be the position where the center of the attention area is located. In some embodiments, the center point position may be represented by an image coordinate value (uv coordinate value), or by a grid where the center point is located (e.g. Figure 6 The serial number or ID representation of the division method shown in the figure) is used.

[0110] In some embodiments, the center point can be a constant position, such as being located at the center of the image regardless of the captured image, or being divided using the same partitioning method for different images, with the grid number or ID of the grid where the center point is located being fixed. The center point can also be variable, i.e., the center point varies for different captured images, and the center point position can be predicted by a planning model. For example, an image is input into the planning model, and the planning model outputs the predicted center point position coordinates or coordinate range.

[0111] The offset can be a distance value, such as 10 pixels (one pixel corresponds to a certain distance, which can be related to the drone's flight altitude), 5cm, or the number of grids, such as 3. When the attention area is circular, the offset is a unique value and can be considered the length of the radius. When the attention area is rectangular, the offset can contain two components: one component is the offset in the v-axis direction, and the other is the offset in the u-axis direction. For other shapes such as ellipses and pentagons, the number and value of the offset components can be determined by the specific geometry of the area.

[0112] The offset reflects the distance from the center point to the boundary of the attention area, wherein the offset is negatively correlated with the flight speed and / or positively correlated with the flight altitude.

[0113] The flight speed / flight altitude can be derived from the drone status information or directly determined based on the output data of the corresponding onboard sensors.

[0114] In some embodiments, a flight altitude-offset function can be pre-fitted. This function can be a monotonically increasing function or a piecewise function, with each piecewise function having a different slope. For example, the slope of a function segment corresponding to a higher flight altitude is smaller than the slope of a function segment corresponding to a lower flight altitude. The offset can be determined based on the flight altitude-offset function and the flight altitude. For example, the offset can be calculated by substituting the flight altitude into the function. Similarly, a monotonically decreasing flight speed-offset function, flight altitude, and a flight speed-offset function can be fitted.

[0115] In some embodiments, a flight altitude-offset mapping table may be established, and the offset may be obtained by looking up the table based on the flight altitude. Similarly, a flight speed-offset mapping table, a flight altitude and flight speed-offset mapping table may be constructed.

[0116] Step 406: Determine the environmental image information based on the one or more sub-images.

[0117] about Figure 4 For details of other steps, such as steps 402 and 406, please refer to Figure 2 The relevant description of step 204 is shown in FIG.

[0118] In this embodiment, determining the size of the attention zone for captured images based on the drone's flight altitude and / or speed aligns with how humans observe the environment while in motion. This accurately captures the visual information required for decision-making or model reasoning, improving model prediction accuracy. Furthermore, a higher flight speed corresponds to a smaller attention zone, reducing the amount of environmental image information processed by the planning model. This, based on the response characteristics of language models (response speed is negatively correlated with the amount of input information), helps improve the planning model's response speed, making it more suitable for intelligent and accurate autonomous navigation of drones at high speeds.

[0119] It should be noted that the above descriptions of the various processes are for illustrative purposes only and do not limit the scope of this specification. Those skilled in the art may, under the guidance of this specification, make various modifications and alterations to the various processes. However, such modifications and alterations remain within the scope of this specification. For example, a storage step may be added.

[0120] Figure 5 5 is an exemplary module diagram of a drone autonomous navigation system according to some embodiments of this specification. In some embodiments, the drone autonomous navigation system 500 may include an acquisition module 510 and an output module 520.

[0121] The acquisition module 510 is used to acquire global task information; the global task information includes destination information.

[0122] The output module 520 is used to output one or more local task information at a first frequency based on at least the global task information, and to output one or more action control instructions at a second frequency; the local task information includes intermediate position information in the process of reaching the destination from a preset starting position, and the action control instructions are used to control the state of the drone; wherein the second frequency is not less than the first frequency.

[0123] It should be understood that Figure 5 The illustrated systems and their modules can be implemented in various ways. For example, in some embodiments, the systems and their modules can be implemented using hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic, while the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated hardware. Those skilled in the art will appreciate that the methods and systems described above can be implemented using computer-executable instructions and / or contained in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and their modules described herein can be implemented not only using hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips or transistors, or programmable hardware devices such as field programmable gate arrays or programmable logic devices, but can also be implemented using software, such as executed by various types of processors, or a combination of such hardware circuits and software (e.g., firmware).

[0124] It should be noted that the above description of the UAV autonomous navigation system and its modules is for convenience only and does not limit this specification to the scope of the embodiments. It is understandable that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the modules or form subsystems connected to other modules without deviating from the principles. In some embodiments, Figure 5 The acquisition module 510 and output module 520 disclosed in the disclosure may be different modules in a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.

[0125] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.

[0126] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.

[0127] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.

[0128] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.

[0129] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may vary according to the required features of the individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values ​​are as accurate as possible within the feasible range.

[0130] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This excludes any application history documents that are inconsistent with or conflicting with the content of this specification, as well as any documents (currently or subsequently appended to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.

[0131] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.

Claims

1. A method for autonomous navigation of a drone, characterized in that: Executed by one or more processors, the method includes: Acquire global task information; the global task information includes destination information; At least based on the global task information, one or more local task information is output at a first frequency, and one or more action control instructions are output at a second frequency; the local task information includes intermediate position information in the process of reaching the destination from a preset starting position, and the action control instructions are used to control the state of the drone.

2. The method according to claim 1, characterized in that The method includes more than one round of cyclic processing, wherein one round of cyclic processing includes: Obtain an input sequence; the input sequence includes global task information, local task information calculated in the previous round of loop processing, and action planning information; if the current round is the first round of loop processing, the local task information and the action planning information are initial values; Processing the input sequence through the planning model to obtain an output sequence; the output sequence includes the following information: local task information and action planning information obtained by the current round of loop processing; When the output sequence includes the action planning information, the action control instruction is obtained based on the action planning information through a translation model.

3. The method according to claim 2, characterized in that The input sequence further includes drone state information of the drone and environmental image information; wherein the drone state information is determined based on output data of one or more onboard sensors on the drone.

4. The method according to claim 3, characterized in that One of the cycle processes further includes: Acquire an image output by an onboard image acquisition device on the UAV; determining an attention region in the image, and obtaining one or more sub-images located in the attention region; The environmental image information is determined based on the one or more sub-images.

5. The method according to claim 4, characterized in that Determining the attention area in the image includes: Determining a center point position of the attention area in the image; An offset is determined based on a flight speed and / or a flight altitude of the drone, where the offset reflects a distance from the center point to a boundary of the attention area; wherein the offset is negatively correlated with the flight speed and / or positively correlated with the flight altitude.

6. The method according to claim 3, characterized in that The output sequence also includes attention area information; the attention area information includes the center point position and offset of the attention area; One of the cycle processes further includes: Acquire an image output by an onboard image acquisition device on the UAV; intercepting one or more sub-images located in the attention area in the image based on the attention area information; The environmental image information is determined based on the one or more sub-images.

7. The method according to claim 3, characterized in that The step of processing the action planning information through a translation model to obtain the flight control instruction includes: Determining input data of a translation model based on the action planning information and the environment image information; The input data is processed by the translation model to obtain the action control instruction.

8. The method according to claim 2, characterized in that The second frequency is not less than the first frequency; The input data and output data of the planning model are both language symbol sequences in a preset format, and the input data of the translation model is also a language symbol sequence in a preset format; The input data of the planning model includes a global task information segment, a local task information segment, an action plan information segment, a drone state information segment, and an environment image information segment; the output data of the planning model includes a flag bit, a multiplexed information segment, and an attention region information segment, wherein the value of the flag bit is used to indicate whether the value of the multiplexed information segment is local task information or action plan information; The input data of the translation model includes an action planning information segment and an environment image information segment.

9. The method according to claim 2, characterized in that The planning model and the translation model are both visual language models.

10. A storage medium storing computer instructions, wherein when a processor executes at least part of the computer instructions in the storage medium, the method for autonomous navigation of a drone as described in any one of claims 1 to 9 can be implemented.