Unmanned aerial vehicle intelligent decision-making method and device based on multi-modal large model
Through multimodal large models, the drone image, text and data information are integrated, and preprocessing, feature extraction and fusion are performed, which solves the shortcomings of single modal decision-making of drones, achieves more accurate and comprehensive decision-making, and improves the drone's mission execution capabilities in complex environments.
Patent Information
- Application Number
- CN202510183549.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the process of autonomous decision-making, single modal data analysis cannot fully utilize the advantages of multiple data sources, resulting in limited decision-making accuracy and reliability, especially in complex environments or emergencies.
The multimodal big model is adopted to integrate drone images, text information and data information, and through preprocessing, feature extraction and multimodal information fusion, it uses convolutional neural networks, recurrent neural networks and long-term memory networks to perform feature extraction and fusion, and finally makes intelligent decisions in the multimodal big model.
It improves the accuracy and robustness of decision-making in complex environments, enhances the ability to respond to various tasks, reduces the risk of misjudgment under the limitation of single modal data, and improves the safety and success rate of task execution.
Smart Images

Figure CN120180349A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent decision-making method, device, storage medium and electronic device for drones based on a multimodal large model. Background Art
[0002] The applications of drones in fields such as aerial photography, agriculture, plant protection, micro selfies, express delivery, disaster relief, observing wild animals, monitoring infectious diseases, surveying and mapping, news reporting, power line inspection, disaster relief, and film and television shooting have greatly expanded their uses. Developed countries are also actively expanding industry applications and developing drone technology.
[0003] In the autonomous decision-making process of drones, multimodal data collected by sensors (such as drone images, videos, radar, lidar, audio, etc.) provides a rich information source.
[0004] Currently, the intelligent decision-making methods of drones usually rely on the analysis of single-modal data. In the face of complex environments or emergencies, single-modal decision-making methods often cannot fully utilize the advantages of all data sources, resulting in limitations in the accuracy and reliability of decisions. Summary of the Invention
[0005] Embodiments of this application provide an intelligent decision-making method, device, storage medium and electronic device for drones based on a multimodal large model, which can enable drones to make more comprehensive and accurate decisions in complex or changing environments and improve their decision-making ability to handle various tasks.
[0006] Embodiments of this application provide an intelligent decision-making method for drones based on a multimodal large model, including: Obtain drone images, text information and data information; Preprocess the drone images, the text information and the data information respectively; Extract features from the preprocessed drone images, text information and data information respectively to obtain drone image features, text features and data features; Perform multimodal information fusion on the drone image features, the text features and the data features to obtain multimodal features; Input the multimodal features into a trained multimodal large model to obtain an intelligent decision result.
[0007] Further, in the above intelligent decision-making method for drones based on a multimodal large model, the drone images include RGB drone images of the on-board camera of the drone, depth camera drone images, thermal imaging drone image information and video stream drone image information.
[0008] Further, in the above-mentioned UAV intelligent decision-making method based on a multimodal large model, the text information includes UAV position text information, map text information, environmental text information, and task text information; The data information includes flight parameter data information, environmental data information, battery status data information, and communication signal strength data information.
[0009] Further, in the above-mentioned UAV intelligent decision-making method based on a multimodal large model, the preprocessing of the UAV image, the text information, and the data information respectively includes: Cropping the UAV image and normalizing the pixel values of the cropped UAV image; Performing text cleaning, word segmentation, and stop word removal on the text information; Sequentially performing data cleaning and completion, denoising processing, and data standardization and normalization on the data information.
[0010] Further, in the above-mentioned UAV intelligent decision-making method based on a multimodal large model, the feature extraction of the preprocessed UAV image to obtain UAV image features includes: Inputting the preprocessed UAV image into a convolutional neural network, and performing local feature extraction on the UAV image through convolutional operations to obtain an output feature map; Among them, the convolutional operation is expressed as:
[0011] Among them, is a certain position of the output feature map, X is the input UAV image, W is the convolutional kernel weight, and b is the bias; Sequentially passing the output feature map through the activation layer, pooling layer, and fully connected layer of the convolutional neural network to obtain UAV image features.
[0012] Further, in the above-mentioned UAV intelligent decision-making method based on a multimodal large model, the feature extraction of the preprocessed text information to obtain text features includes: Inputting the preprocessed text information into a recurrent neural network to obtain text features; Among them, the current state is expressed as:
[0013] Among them, is the current hidden state, is the current input, and are weight matrices, b is the bias, and f is the activation function.
[0014] Further, in the above-mentioned UAV intelligent decision-making method based on a multimodal large model, the multimodal information fusion of the UAV image features, the text features, and the data features to obtain multimodal features includes: Input the text features into the first-layer LSTM to obtain the first hidden layer state of each neuron; After splicing the data features with the first hidden layer state, input them into the second-layer LSTM to obtain the second hidden layer state of each neuron in the second layer; After splicing the UAV image features with the third hidden layer state, input them into the third-layer LSTM to obtain the third hidden layer state of each neuron in the third layer, and use the third hidden layer state as the multimodal features.
[0015] An embodiment of the present application also provides a UAV intelligent decision-making device based on a multimodal large model, including: An acquisition module, configured to acquire UAV images, text information, and data information; A preprocessing module, configured to preprocess the UAV images, the text information, and the data information respectively; A feature extraction module, configured to extract features from the preprocessed UAV images, text information, and data information respectively to obtain UAV image features, text features, and data features; A multimodal fusion module, configured to perform multimodal information fusion on the UAV image features, the text features, and the data features to obtain multimodal features; A decision-making module, configured to input the multimodal features into a trained multimodal large model to obtain an intelligent decision-making result.
[0016] An embodiment of the present application also provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned UAV intelligent decision-making methods based on a multimodal large model.
[0017] An embodiment of the present application also provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above-mentioned UAV intelligent decision-making methods based on a multimodal large model.
[0018] The present application provides a method, apparatus, storage medium, and electronic device for intelligent decision-making of unmanned aerial vehicles (UAVs) based on multimodal large models. The present application collects multimodal information for intelligent decision-making of UAVs. Compared with the existing UAV intelligent decision-making methods that collect single-modal information, multimodal information fusion can capture UAV status, environmental, and task information from multiple perspectives, compensating for the deficiencies of single-modal data under specific conditions. This enables UAVs to make more comprehensive and accurate decisions in complex or changing environments, improving their decision-making ability to handle various tasks. The embodiments of the present application utilize multimodal large models for information processing, which can enhance the robustness and adaptability of the UAV intelligent decision-making system. In the face of challenges such as bad weather, poor lighting conditions, or electromagnetic interference, multimodal data can complement each other, reducing the risk of misjudgment caused by the limitations of single-modal data. This allows UAVs to perform tasks in more diverse environments, enhancing the safety and success rate of UAV task execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of the present application, will make the technical solutions and other beneficial effects of the present application obvious.
[0020] Figure 1 It is a flowchart of the method for intelligent decision-making of UAVs based on multimodal large models provided by the embodiments of the present application.
[0021] Figure 2 It is a schematic diagram of the UAV flight parameter data information provided by the embodiments of the present application.
[0022] Figure 3 It is a flowchart of multimodal information fusion provided by the embodiments of the present application.
[0023] Figure 4 It is a schematic diagram of the structure of the UAV intelligent decision-making device based on multimodal large models provided by the embodiments of the present application.
[0024] Figure 5 It is a schematic diagram of the structure of the electronic device provided by the embodiments of the present application.
[0025] Figure 6 It is another schematic diagram of the structure of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will, in conjunction with the drawings in the embodiments of the present application, clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0027] The embodiments of this application provide a method, device, storage medium, and electronic device for intelligent decision-making of drones based on multi-modal large models. The intelligent decision-making device for drones based on multi-modal large models provided by the embodiments of this application can be integrated into an electronic device, which can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro-processing box, or other devices, etc.
[0028] The method, device, storage medium, and electronic device for intelligent decision-making of drones based on multi-modal large models provided by this application can be widely applied to scenarios such as post-disaster rescue, agricultural monitoring, and urban traffic management. In the post-disaster rescue scenario, the method of the present invention can use drones equipped with multi-modal sensors such as high-definition cameras and thermal imagers to quickly generate disaster maps, accurately locate trapped people, and improve rescue efficiency. In the agricultural monitoring scenario, the method of the present invention can monitor the growth status and pest and disease conditions of crops in real time through multi-spectral sensors, analyze the data using deep learning models, and help farmers apply fertilizers and pesticides accurately to increase crop yields. In the urban traffic management scenario, the method of the present invention can build an intelligent traffic management system based on the traffic flow and speed data collected by drones in real time, dynamically adjust traffic lights and publish traffic information, thereby optimizing traffic flow, reducing congestion, and improving traffic efficiency.
[0029] Please refer to Figure 1 , Figure 1 is a flowchart of the method for intelligent decision-making of drones based on multi-modal large models provided by the embodiments of this application. It is applied to an electronic device. The method for intelligent decision-making of drones based on multi-modal large models includes the following steps: S1, Obtain drone images, text information, and data information.
[0030] Among them, the drone images include RGB drone images of the drone's on-board camera, depth camera drone images, thermal imaging drone image information, and video stream drone image information; the text information includes drone position text information, map text information, environmental text information, and task text information; the data information includes flight parameter data information, environmental data information, battery status data information, and communication signal strength data information. Among them, Figure 2 is a schematic diagram of the drone flight parameter data information provided by the embodiments of this application. It can be referred to Figure 2 .
[0031] S2, Preprocess the drone images, text information, and data information respectively.
[0032] Step S2 can include the following steps: S21. Crop the UAV image and normalize the pixel values of the cropped UAV image.
[0033] Specifically, first, perform operations such as scaling, cropping, rotating, and flipping on the UAV image to unify the pixel size of the image.
[0034] Then, perform normalization. Normalize the pixel values of the UAV image with the cropped size, and convert the pixel values from the range of 0 - 255 to the range of -1 to 1. The image normalization formula is as follows:
[0035] Among them, represents the value of the image pixel point, min(x) , max(x) respectively represent the maximum and minimum values of the image pixels.
[0036] S22. Perform text cleaning, word segmentation, and stop word removal on the text information.
[0037] Among them, text cleaning includes removing unnecessary symbols, punctuation marks, special characters, and extra spaces in the UAV text information. Word segmentation and stop word removal include splitting the cleaned UAV text information into words or phrases and removing common words that are meaningless for analysis to reduce noise.
[0038] S23. Perform data cleaning and completion, noise removal processing, and data standardization and normalization on the data information in sequence.
[0039] Data cleaning and completion include removing duplicate, incorrect, or incomplete UAV data and using linear interpolation method and polynomial interpolation method to complete the UAV data. The linear interpolation method and polynomial interpolation calculation formulas are as follows: Linear interpolation method:
[0040] Polynomial interpolation method:
[0041] Noise removal processing includes using mean filtering to remove noise from the cleaned and completed UAV data information. The mean filtering formula is as follows:
[0042] Data standardization and normalization include using the Z - score standardization method and the minimum - maximum normalization method to standardize and normalize the denoised UAV data information.
[0043] S3. Perform feature extraction on the pre - processed UAV image, text information, and data information respectively to obtain UAV image features, text features, and data features.
[0044] Step S3 includes the following steps: S31. Input the preprocessed drone image into a convolutional neural network (CNN), and perform local feature extraction on the drone image through convolutional operations to obtain an output feature map.
[0045] Among them, the convolutional operation is expressed as:
[0046] Among them, is a certain position of the output feature map, X is the input drone image, W is the convolutional kernel weight, and b is the bias; Pass the output feature map through the activation layer (introducing non-linearity through the activation function), pooling layer (reducing the feature dimension through pooling operations), and fully connected layer of the convolutional neural network in sequence to obtain the drone image features.
[0047] S32. Input the preprocessed text information into a recurrent neural network (RNN) to obtain text features, and the RNN processes the input sequence (text information) step by step.
[0048] Among them, the current state is expressed as:
[0049] Among them, is the current hidden state, is the current input, and are weight matrices, b is the bias, and f is the activation function.
[0050] After that, obtain the hidden state through backpropagation through time as the drone text information feature.
[0051] S4. Perform multimodal information fusion on the drone image features, text features, and data features to obtain multimodal features.
[0052] Figure 3 This is the flowchart of multimodal information fusion provided by the embodiments of the present application. As Figure 3 shown, step S4 is specifically as follows: Input the text features into the first layer of LSTM to obtain the first hidden layer state of each neuron; splice the data features with the first hidden layer state and input them into the second layer of LSTM to obtain the second hidden layer state of each neuron in the second layer; splice the drone image features with the third hidden layer state and input them into the third layer of LSTM to obtain the third hidden layer state of each neuron in the third layer, and use the third hidden layer state as the multimodal feature.
[0053] S5. Input the multimodal features into the trained multimodal large model to obtain an intelligent decision result.
[0054] Among them, the intelligent decision result is the recognition result or prediction result of the multi-modal large model for multi-modal features. In different scenarios, the intelligent decision result is different, depending on the input and training process of the multi-modal large model.
[0055] In the embodiments of the present application, intelligent decision-making for drones is carried out by collecting multi-modal information. Compared with the existing method of intelligent decision-making for drones that collects single-modal information, multi-modal information fusion can capture drone status, environmental, and task information from multiple perspectives, making up for the deficiencies of single-modal data under specific conditions. This enables drones to make more comprehensive and accurate decisions in complex or changing environments, improving their decision-making ability to handle various tasks.
[0056] The embodiments of the present application use a multi-modal large model for information processing, which can improve the robustness and adaptability of the drone intelligent decision-making system. In the face of challenges such as bad weather, poor lighting conditions, or electromagnetic interference, multi-modal data can complement each other, reducing the risk of misjudgment caused by the limitation of single-modal data. This allows drones to perform tasks in more diverse environments, enhancing the safety and success rate of drone task execution.
[0057] According to the method described in the above embodiments, this embodiment will further describe from the perspective of a drone intelligent decision-making device based on a multi-modal large model. The drone intelligent decision-making device based on a multi-modal large model can be specifically implemented as an independent entity, or integrated in an electronic device. The electronic device can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro-processing box, or other devices, etc.
[0058] Please refer to Figure 4 , Figure 4 which specifically describes the drone intelligent decision-making device based on a multi-modal large model provided by the embodiments of the present application, applied to an electronic device. The drone intelligent decision-making device based on a multi-modal large model may include: An acquisition module, configured to acquire drone images, text information, and data information; A preprocessing module, configured to preprocess the drone images, text information, and data information respectively; A feature extraction module, configured to extract features from the preprocessed drone images, text information, and data information respectively, to obtain drone image features, text features, and data features; A multi-modal fusion module, configured to perform multi-modal information fusion on the drone image features, text features, and data features to obtain multi-modal features; A decision module, configured to input the multi-modal features into a trained multi-modal large model to obtain an intelligent decision result.
[0059] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference can be made to the method embodiments described above. For the beneficial effects that can be specifically achieved, please also refer to the beneficial effects in the method embodiments described above, which will not be elaborated here.
[0060] In addition, an embodiment of the present application also provides an electronic device, which can be a device such as a computer or a tablet computer. As Figure 5 shown, the electronic device 400 includes a processor 401 and a memory 402. Among them, the processor 401 is electrically connected to the memory 402.
[0061] The processor 401 is the control center of the electronic device 400, connecting various parts of the entire electronic device through various interfaces and lines. By running or loading application programs stored in the memory 402, and by calling data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device.
[0062] In this embodiment, the processor 401 in the electronic device 400 will load the instructions corresponding to the processes of one or more application programs into the memory 402 according to the following steps, and the processor 401 will run the application programs stored in the memory 402 to implement various functions: Obtain drone images, text information, and data information; Preprocess the drone images, the text information, and the data information respectively; Extract features from the preprocessed drone images, text information, and data information respectively to obtain drone image features, text features, and data features; Perform multi-modal information fusion on the drone image features, the text features, and the data features to obtain multi-modal features; Input the multi-modal features into a trained multi-modal large model to obtain an intelligent decision result.
[0063] This electronic device can implement the steps in any embodiment of the drone intelligent decision method based on a multi-modal large model provided by the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any drone intelligent decision method based on a multi-modal large model provided by the embodiments of the present invention. For details, please refer to the previous embodiments, which will not be elaborated here.
[0064] Figure 6The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. This electronic device can be used to implement the intelligent decision-making method for drones based on the multimodal large model provided in the above embodiment. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0065] The RF circuit 510 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions. For example, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, enterprise intranets, wireless networks, or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network, or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols, and technologies, including but not limited to the Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as the Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.
[0066] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizes functions such as taking pictures with the front camera, processing the captured drone images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0067] The input unit 530 can be used to receive input digital or character information, and generate keyboards and mice related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).
[0068] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may further include an earphone jack to provide communication between an external peripheral earphone and the electronic device 500.
[0069] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to an essential component of the electronic device 500, and it can be completely omitted within the scope of not changing the essence of the invention according to needs.
[0070] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and calling the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0071] The electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0072] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. One or more programs include instructions for performing the following operations: Obtain drone images, text information, and data information; Preprocess the drone images, the text information, and the data information respectively; Extract features from the preprocessed drone images, text information, and data information respectively to obtain drone image features, text features, and data features; Perform multi-modal information fusion on the drone image features, the text features, and the data features to obtain multi-modal features; Input the multi-modal features into a trained multi-modal large model to obtain an intelligent decision result.
[0073] In specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. The specific implementation of the above-mentioned each module can refer to the method embodiments described above and will not be elaborated here.
[0074] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any of the embodiments of the method for intelligent decision-making of an unmanned aerial vehicle based on a multimodal large model provided by the embodiments of the present invention.
[0075] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0076] Since the instructions stored in the storage medium can execute the steps in any of the embodiments of the method for intelligent decision-making of an unmanned aerial vehicle based on a multimodal large model provided by the embodiments of the present invention, the beneficial effects achievable by any of the methods for intelligent decision-making of an unmanned aerial vehicle based on a multimodal large model provided by the embodiments of the present invention can be realized. For details, see the previous embodiments and will not be elaborated here.
[0077] The above has introduced in detail a method, device, storage medium, and electronic device for intelligent decision-making of an unmanned aerial vehicle based on a multimodal large model provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A UAV intelligent decision-making method based on a multimodal large model, characterized in that: The method comprises: Obtain drone images, text information, and data information; Preprocessing the drone image, the text information, and the data information respectively; The preprocessed UAV images, text information and data information are respectively subjected to feature extraction to obtain UAV image features, text features and data features; Performing multimodal information fusion on the drone image features, the text features, and the data features to obtain multimodal features; The multimodal features are input into a trained multimodal large model to obtain an intelligent decision result.
2. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 1 is characterized in that: The drone images include drone onboard camera RGB drone images, depth camera drone images, thermal imaging drone image information, and video stream drone image information.
3. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 1 is characterized in that: The text information includes drone location text information, map text information, environment text information and mission text information; The data information includes flight parameter data information, environmental data information, battery status data information and communication signal strength data information.
4. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 1 is characterized in that: The preprocessing of the drone image, the text information and the data information respectively comprises: Cropping the drone image, and normalizing pixel values of the cropped drone image; Performing text cleaning, word segmentation, and stop word removal on the text information; The data information is sequentially cleaned and completed, denoised, and standardized and normalized.
5. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 1 is characterized in that: The preprocessed UAV image is subjected to feature extraction to obtain the UAV image features, including: Inputting the preprocessed UAV image into a convolutional neural network, performing local feature extraction on the UAV image through a convolution operation, and obtaining an output feature map; Wherein, the convolution operation is expressed as: in, is a certain position of the output feature map, X is the input drone image, W is the convolution kernel weight, and b is the bias; The output feature map is sequentially passed through the activation layer, pooling layer and fully connected layer of the convolutional neural network to obtain the drone image features.
6. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 5 is characterized in that: The preprocessed text information is subjected to feature extraction to obtain text features, including: Input the preprocessed text information into the recurrent neural network to obtain text features; Among them, the current state is expressed as: in, is the current hidden state, is the current input, and is the weight matrix, b is the bias, and f is the activation function.
7. The intelligent decision-making method for unmanned aerial vehicles based on a multimodal large model according to claim 6 is characterized in that: The multimodal information fusion of the drone image feature, the text feature and the data feature to obtain the multimodal feature includes: Input the text features into the first layer of LSTM to obtain the first hidden layer state of each neuron; The data features are concatenated with the first hidden layer state and then input into the second layer LSTM to obtain the second hidden layer state of each neuron in the second layer; The drone image features are concatenated with the third hidden layer state and input into the third layer LSTM to obtain the third hidden layer state of each neuron in the third layer, and the third hidden layer state is used as the multimodal feature.
8. An intelligent decision-making device for unmanned aerial vehicles based on a multimodal large model, characterized in that: include: Acquisition module, used to acquire drone images, text information and data information; A preprocessing module, used to preprocess the drone image, the text information and the data information respectively; A feature extraction module is used to extract features from the preprocessed drone image, text information and data information to obtain drone image features, text features and data features; A multimodal fusion module, used for performing multimodal information fusion on the drone image features, the text features and the data features to obtain multimodal features; The decision module is used to input the multimodal features into the trained multimodal large model to obtain an intelligent decision result.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the intelligent decision-making method for unmanned aerial vehicle based on a multimodal large model as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the drone intelligent decision-making method based on a multimodal large model as described in any one of claims 1 to 7.