Unmanned aerial vehicle identification method based on multi-modal three-dimensional data and related equipment

By constructing a recognition method for multimodal 3D data of UAVs, acquiring and mapping it to a discrete symbol space, and generating structured language results, the problem of poor recognition performance of UAVs in diverse application scenarios is solved, and the intelligent recognition capability and robustness are improved.

CN120932129APending Publication Date: 2025-11-11HUBEI WUCHANG XINGYUTONG SKY TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510951489.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing UAV intelligent recognition technology is not adaptable to diverse application scenarios, has high module coupling, low intelligence level, and poor robustness, especially in the detection of small targets in complex environments.

Method used

By acquiring multimodal 3D data of the drone within a preset range, constructing 3D geometric features, mapping them to a preset discrete symbol space, obtaining temporal and spatial correlations, and finally generating structured language results, intelligent recognition of the environment is achieved.

Benefits of technology

It enhances the intelligent recognition capabilities of drones in diverse application scenarios, improves recognition results, reduces reliance on VL models, and enhances the efficiency and accuracy of image understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932129A_ABST
    Figure CN120932129A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle identification method based on multi-modal three-dimensional data and related equipment. The method comprises the steps that the multi-modal three-dimensional data in a preset range where an unmanned aerial vehicle is located is acquired; three-dimensional geometric features are constructed based on the multi-modal three-dimensional data; mapping the three-dimensional geometric features to a preset discrete symbol space to obtain logic symbols corresponding to the three-dimensional geometric features; obtaining a time association relationship and a space association relationship among the three-dimensional geometric features; and obtaining a structured language result based on the logic symbols, the time association relationship and the space association relationship. The method can improve the capability of intelligent recognition of the unmanned aerial vehicle to adapt to diversified application scenes, thereby improving the recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and related equipment for the identification of unmanned aerial vehicles (UAVs) based on multimodal 3D data. Background Technology

[0002] In the field of intelligent recognition technology for drones, there are two relatively typical technologies, each with its own advantages and disadvantages. One is multimodal fusion based on timestamp interpolation, represented by Intel RealSense T265, which is widely used in consumer drones and industrial inspection scenarios and supports basic module expansion. However, it has significant drawbacks: high module coupling, requiring deep modifications to the flight control firmware when adding new functions, and strong inter-module dependencies; it also has low intelligence levels, relying on preset rules, such as PID parameter tuning, and cannot learn from dynamic environmental changes online. The other is YOLO-based image recognition technology, which relies on a single-stage detection architecture to achieve real-time obstacle detection of vehicles, pedestrians, trees, etc., on drones. However, it also has shortcomings: poor performance in small target detection, poor robustness in complex environments, and weak generalization ability, because it relies on data annotation and lacks contextual understanding capabilities.

[0003] Therefore, how to improve the ability of drones to intelligently identify and adapt to diverse application scenarios, thereby improving the recognition effect, is an urgent problem to be solved. Summary of the Invention

[0004] To address the aforementioned issues, embodiments of this application provide a method and apparatus for identifying unmanned aerial vehicles (UAVs) based on multimodal 3D data, as well as an UAV, electronic device, computer-readable storage medium, and computer program product.

[0005] Firstly, in order to solve the above-mentioned technical problems, this application provides a method for UAV recognition based on multimodal 3D data, including: Acquire multimodal 3D data of the drone within a preset range; Three-dimensional geometric features are constructed based on the multimodal three-dimensional data; The three-dimensional geometric features are mapped to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features; the temporal and spatial correlation relationships between the three-dimensional geometric features are obtained. The structured language result is obtained based on the logical symbols, the temporal associations, and the spatial associations.

[0006] The beneficial effects are: In the technical solution provided by the embodiments of this application, multimodal 3D data of the UAV within a preset range is acquired; 3D geometric features are constructed based on the multimodal 3D data; the 3D geometric features are mapped to a preset discrete symbol space to obtain logical symbols corresponding to the 3D geometric features; the temporal and spatial correlations between the 3D geometric features are acquired; and structured language results are obtained based on the logical symbols, temporal correlations, and spatial correlations. In this way, this application uniformly maps information from multiple sensors to a preset discrete symbol space, and obtains structured language results based on various correlations within it. For tasks involving temporal concepts and dependencies, it can output structured natural language descriptions, facilitating understanding and reasoning by large models. It does not require reliance on a VL (Vision-Language Model) model for image understanding, thereby enhancing the UAV's ability to intelligently recognize and adapt to diverse application scenarios, thus improving recognition performance.

[0007] Furthermore, acquiring multimodal 3D data within a preset range of the UAV includes: The spatial reference system of multiple sensors of the UAV is unified by coordinate system projection transformation to obtain multiple unified sensors; Using the aforementioned unified sensors, multimodal three-dimensional data within a preset spatial range of the UAV is acquired; The multimodal 3D data is spatiotemporally aligned using a soft timestamp interpolation algorithm to obtain aligned multimodal 3D data, which is then used to construct the corresponding 3D geometric features.

[0008] Furthermore, the construction of three-dimensional geometric features based on the multimodal three-dimensional data includes: The preset range space is voxelized to obtain multiple cubic units; Obtain the three-dimensional coordinates of the multimodal three-dimensional data, and assign the multimodal three-dimensional data to the cube unit corresponding to the three-dimensional coordinates; The multimodal 3D data is geometrically represented within the cubic unit to obtain the corresponding 3D geometric features.

[0009] Furthermore, the step of mapping the three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features includes: Obtain a preset discrete symbol space, which includes pre-constructed logical predicates and spatial predicates under the BeiDou grid code; The three-dimensional geometric features are mapped based on the logical predicate and the spatial predicate using a neural network to obtain logical symbols corresponding to the three-dimensional geometric features.

[0010] Furthermore, obtaining the temporal and spatial correlations between the three-dimensional geometric features includes: obtaining time point information of the three-dimensional geometric features, and obtaining event correlation information between the three-dimensional geometric features based on the time point information; Obtain the position information of the three-dimensional coordinates corresponding to the three-dimensional geometric features in the BeiDou grid code; At each time point, the spatial correlation of the three-dimensional geometric features is obtained based on the location information.

[0011] Furthermore, obtaining the structured language result based on the logical symbols, the temporal associations, and the spatial associations includes: Based on the logical symbols, the association symbol results are obtained through the temporal and spatial association relationships. The associated symbol results are imported into a preset template format to form a structured language result.

[0012] Secondly, the present invention provides a drone, characterized in that it uses the drone-based multimodal three-dimensional data recognition method described above to achieve environmental detection.

[0013] Thirdly, the present invention provides a UAV recognition device based on multimodal 3D data, comprising: The acquisition unit is used to acquire multimodal 3D data of the UAV within a preset range; Geometric feature units are used to construct three-dimensional geometric features based on the multimodal three-dimensional data; A mapping unit is used to map the three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features; The association unit is used to obtain the temporal and spatial correlation relationships between the three-dimensional geometric features; The language result unit is used to obtain structured language results based on the logical symbols, the temporal correlation, and the spatial correlation.

[0014] Thirdly, this application also provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the aforementioned UAV recognition method based on multimodal 3D data.

[0015] Fourthly, this application also provides a computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a computer's processor, cause the computer to perform the UAV recognition method based on multimodal 3D data as described above.

[0016] Fifthly, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the UAV recognition method based on multimodal 3D data provided in the various alternative embodiments described above.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a flowchart illustrating an exemplary embodiment of the present application of a method for identifying unmanned aerial vehicles (UAVs) based on multimodal 3D data; Figure 2 This is a schematic diagram of the BeiDou grid code in a real-world scenario; Figure 3 This refers to the temporal and spatial correlations between multimodal three-dimensional data in an exemplary embodiment of this application; Figure 4 This is a schematic diagram illustrating the transformation of associated symbol results into structured language results using a prompt template in an exemplary embodiment of this application; Figure 5 This is a block diagram illustrating an unmanned aerial vehicle (UAV) recognition device based on multimodal 3D data, as shown in an exemplary embodiment of this application. Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing the electronic devices of the present application embodiments. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0020] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0021] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0022] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0023] To address the issue of poor intelligent recognition performance of drones, embodiments of this application propose a method, apparatus, electronic device, and computer-readable storage medium for drone recognition based on multimodal 3D data. These embodiments primarily relate to drone recognition technology based on multimodal 3D data, which is included in data processing technologies. The following will provide a detailed description of these embodiments.

[0024] Please refer to the following first. Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of a UAV recognition method based on multimodal 3D data. The method can be executed by a server, which can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. No restrictions are imposed here. Figure 1 As shown in an exemplary embodiment, the UAV recognition method based on multimodal 3D data may include steps S101 to S105, which are described in detail below: Step S101: Obtain multimodal 3D data of the UAV within a preset range.

[0025] Step S102: Construct three-dimensional geometric features based on multimodal three-dimensional data.

[0026] Step S103: Map the three-dimensional geometric features to a preset discrete symbol space to obtain the logical symbols corresponding to the three-dimensional geometric features.

[0027] Step S104: Obtain the temporal and spatial relationships between the three-dimensional geometric features.

[0028] Step S105: Obtain the structured language result based on logical symbols, temporal relationships, and spatial relationships.

[0029] As can be seen from the above, in the method provided in this embodiment, multimodal 3D data of the UAV within a preset range is acquired; 3D geometric features are constructed based on the multimodal 3D data; and the 3D geometric features are mapped to a preset discrete symbol space to obtain logical symbols corresponding to the 3D geometric features. This allows information from multiple sensors to be uniformly mapped to a preset discrete symbol space, facilitating management and subsequent calculations.

[0030] Then, the temporal and spatial relationships between 3D geometric features are obtained. Based on logical symbols, temporal and spatial relationships, structured language results are obtained. The structured output has the advantages of being easy to remember (small space occupation) and fast retrieval (fast retrieval of key information) of environmental information. In this way, structured language results are obtained based on multiple relationships in discrete symbol space. For tasks that perform tasks with temporal concepts and dependencies, structured natural language descriptions can be output, which are easy for large models to understand and reason about. It does not require reliance on VL models for image understanding, thereby improving the ability of UAV intelligent recognition to adapt to diverse application scenarios and thus improving recognition performance.

[0031] In an exemplary embodiment of this application, the specific steps for acquiring multimodal three-dimensional data of the UAV within a preset range may include: By unifying the spatial reference systems of multiple sensors of the UAV through coordinate system projection transformation, multiple unified sensors are obtained. By using multiple unified sensors, multimodal 3D data of the UAV within a preset spatial range can be acquired. Spatiotemporal alignment of multimodal 3D data is performed using a soft timestamp interpolation algorithm to obtain aligned multimodal 3D data, which is then used to construct corresponding 3D geometric features.

[0032] In this embodiment, during the acquisition of multimodal 3D data within a preset range of the UAV, the sensors collecting the data and the acquired data are preprocessed to improve data processing efficiency. Preferably, the spatial reference systems of multiple sensors of the UAV are unified through coordinate system projection transformation to obtain multiple unified sensors, eliminating spatiotemporal misalignment of cross-modal data; using multiple unified sensors, multimodal 3D data within the preset range of the UAV is acquired; and the multimodal 3D data is spatiotemporally aligned using a soft timestamp interpolation algorithm to obtain aligned multimodal 3D data, thereby constructing corresponding 3D geometric features. The multimodal 3D data may include multi-source signals such as visual frames, lidar point clouds, millimeter-wave radar ranging information, ultrasonic sensor near-range obstacle detection information, IMU (Inertial Measurement Unit) acceleration information, and GNSS (Global Navigation Satellite System) positioning information.

[0033] Thus, through the above embodiments, this application eliminates the problem of uncontrolled inconsistencies between sensors and data by unifying the spatial reference system of sensors and the spatiotemporal alignment of data, providing a consistent spatiotemporal representation basis for the construction of three-dimensional geometric features.

[0034] In an exemplary embodiment of this application, the specific steps for constructing three-dimensional geometric features based on multimodal three-dimensional data may include: The preset range space is voxelized to obtain multiple cubic units; The three-dimensional coordinates of the multimodal three-dimensional data are obtained, and the multimodal three-dimensional data are assigned to the cube cells corresponding to the three-dimensional coordinates. The multimodal three-dimensional data are geometrically represented within the cube cells to obtain the corresponding three-dimensional geometric features.

[0035] In this embodiment, the multimodal 3D data all include 3D coordinates and intensity information of the corresponding type. The preferred step in constructing the 3D geometric features is to perform voxelization on a preset range space to obtain multiple cubic units, that is, to divide the preset range space into tiny cubic units; obtain the 3D coordinates of the multimodal 3D data; assign the multimodal 3D data to the cubic units corresponding to the 3D coordinates; and perform geometric representation of the multimodal 3D data within the cubic units, for example, as lines, surfaces, or spheres, thereby obtaining the corresponding 3D geometric features.

[0036] Thus, through the above embodiments, this application constructs three-dimensional geometric features based on modal three-dimensional data through feature transformation, and uses a unified data format for subsequent data processing, which facilitates the handling of complex tasks.

[0037] In an exemplary embodiment of this application, the specific steps of mapping three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features may include: Obtain a preset discrete symbol space, which includes pre-constructed logical predicates and spatial predicates under the BeiDou grid code; By using a neural network, three-dimensional geometric features are mapped based on logical predicates and spatial predicates to obtain logical symbols corresponding to the three-dimensional geometric features.

[0038] In this embodiment, the preset discrete symbol space includes pre-constructed logical predicates and spatial predicates under the BeiDou grid code, such as object recognition predicates (object(drone,obejct), object(tree,obstacle_2), and relational predicates (distance(drone,obstacle_2,near)). The neural network used for mapping can be a differentiable symbolic network (such as Neural Logic Machines).

[0039] Spatial predicates are modeled using spatiotemporal constraints under the BeiDou grid code. The BeiDou grid code provides binding location information, thereby forming spatial associations. In discrete symbolic space mapping, using the BeiDou grid code for location calculation and representation can improve the efficiency of path planning in complex spatial domains. For example... Figure 2 As shown, Figure 2 This is a schematic diagram of the BeiDou grid code in a real-world scenario. The BeiDou grid code divides three-dimensional space into a three-dimensional grid and combines it with the time dimension, which can significantly simplify the path planning problem of UAVs in complex airspace. Compared with traditional floating-point operations based on latitude and longitude, the binary operation of the grid code improves the computational efficiency by more than 10 times. No matter how much the number of objects in the airspace increases, the computational efficiency only changes linearly.

[0040] Thus, through the above embodiments, this application maps three-dimensional geometric features to a discrete symbol space to obtain corresponding logical symbols, which facilitates reasoning, facilitates natural language description, and improves generalization ability.

[0041] In an exemplary embodiment of this application, the specific steps for obtaining the temporal and spatial correlation relationships between three-dimensional geometric features may include: Acquire time point information of 3D geometric features, and obtain event correlation information between 3D geometric features based on the time point information; obtain the position information of the 3D coordinates corresponding to the 3D geometric features in the BeiDou grid code; At each time point, the spatial correlation of three-dimensional geometric features is obtained based on location information.

[0042] In this embodiment, after acquiring multimodal 3D data, the information from multiple sensor sources is concatenated using time (i.e., temporal correlation), and then the spatial changes of the object at each time point are reflected (i.e., spatial correlation). For example... Figure 3 As shown, Figure 3 This refers to the temporal and spatial correlations between multimodal three-dimensional data in an exemplary embodiment of this application.

[0043] Therefore, after constructing the three-dimensional geometric features, it is possible to obtain the time point information of the three-dimensional geometric features based on the correlation between data, obtain the event association information between the three-dimensional geometric features based on the time point information, obtain the position information of the three-dimensional coordinates corresponding to the three-dimensional geometric features in the Beidou grid code, and obtain the spatial association relationship of the three-dimensional geometric features based on the position information at each time point.

[0044] In another exemplary embodiment, the specific steps for obtaining structured language results based on logical symbols, temporal relationships, and spatial relationships may include: By using temporal and spatial relationships, the associated symbol results are obtained based on logical symbols. Import the results of the association symbols into a preset template format to form a structured language result.

[0045] In this embodiment, by using temporal and spatial relationships, association symbol results are obtained based on logical symbols, so that the final output is a natural language description displayed chronologically. Then, the association symbol results are imported into a preset template format (such as a prompt template) to form a structured language result. Figure 4 As shown, Figure 4 This is a schematic diagram illustrating how the prompt template is used to transform the associated symbol result into a structured language result in an exemplary embodiment of this application.

[0046] Thus, through the above embodiments, this application associates these mapped discrete logical symbols by combining time point information and Beidou grid code location information, and then imports the association symbol results into a preset template format to form a structured language result, outputting a natural language description displayed in a timeline.

[0047] In an exemplary embodiment of this application, the UAV recognition method based on multimodal 3D data is applied to a UAV. After multiple sensors of the UAV acquire multimodal 3D data within a preset range, the information from the multiple sensors is uniformly mapped to a preset discrete symbol space through a server or other processing module. Structured language results are then obtained based on various relationships within this space, facilitating memorization (small space occupation) and rapid retrieval (fast retrieval of key information) of environmental information. Furthermore, for tasks involving sequential concepts and dependencies, structured natural language descriptions can be output, facilitating understanding and reasoning by large models without relying on VL models for image understanding. This enhances the UAV's ability to intelligently recognize and adapt to diverse application scenarios, thereby improving the UAV's recognition performance of its surrounding environment.

[0048] Figure 5 This is a block diagram illustrating an exemplary embodiment of this application of a drone-based multimodal 3D data recognition device 500. Figure 5 As shown, the device includes: Acquisition unit 501 is used to acquire multimodal three-dimensional data of the UAV within a preset range; Geometric feature unit 502 is used to construct three-dimensional geometric features based on multimodal three-dimensional data; The mapping unit 503 is used to map three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features; Association unit 504 is used to obtain the temporal and spatial correlation relationships between three-dimensional geometric features; Language result unit 505 is used to obtain structured language results based on logical symbols, temporal relationships and spatial relationships.

[0049] This device utilizes the UAV recognition method based on multimodal 3D data provided in this application. The acquisition unit 501 acquires multimodal 3D data within a preset range of the UAV; the geometric feature unit 502 constructs 3D geometric features based on the multimodal 3D data; the mapping unit 503 maps the 3D geometric features to a preset discrete symbol space, obtaining logical symbols corresponding to the 3D geometric features; the association unit 504 acquires the temporal and spatial relationships between the 3D geometric features; and the language result unit 505 obtains structured language results based on logical symbols, temporal relationships, and spatial relationships. In this way, this application uniformly maps information from multiple sensors to a preset discrete symbol space, and obtains structured language results based on various relationships within it. For tasks involving temporal concepts and dependencies, it can output structured natural language descriptions, facilitating understanding and reasoning by large models without relying on VL models for image understanding. This enhances the UAV's ability to intelligently recognize diverse application scenarios, thereby improving recognition performance.

[0050] In another exemplary embodiment, the acquisition unit 501 is further configured to perform unified processing on the spatial reference system of multiple sensors of the UAV through coordinate system projection transformation to obtain multiple unified sensors; use the multiple unified sensors to acquire multimodal three-dimensional data within a preset range of space where the UAV is located; and perform spatiotemporal alignment on the multimodal three-dimensional data through a soft timestamp interpolation algorithm to obtain aligned multimodal three-dimensional data in order to construct the corresponding three-dimensional geometric features.

[0051] In another exemplary embodiment, the geometric feature unit 502 is further configured to perform voxelization on the preset range space to obtain multiple cube units; acquire the three-dimensional coordinates of the multimodal three-dimensional data, assign the multimodal three-dimensional data to the cube units corresponding to the three-dimensional coordinates; and perform geometric representation on the multimodal three-dimensional data within the cube units to obtain the corresponding three-dimensional geometric features.

[0052] In another exemplary embodiment, the mapping unit 503 is further configured to obtain a preset discrete symbol space, which includes pre-constructed logical predicates and spatial predicates under the BeiDou grid code; and to map the three-dimensional geometric features based on the logical predicates and spatial predicates through a neural network to obtain logical symbols corresponding to the three-dimensional geometric features.

[0053] In another exemplary embodiment, the association unit 504 is further configured to obtain time point information of the three-dimensional geometric features, obtain event association information between the three-dimensional geometric features based on the time point information; obtain the position information of the three-dimensional coordinates corresponding to the three-dimensional geometric features in the Beidou grid code; and obtain the spatial association relationship of the three-dimensional geometric features based on the position information at each time point.

[0054] In another exemplary embodiment, the language result unit 505 is further configured to obtain associated symbol results based on logical symbols through temporal and spatial association relationships; and import the associated symbol results into a preset template format to form a structured language result.

[0055] It should be noted that the UAV recognition device based on multimodal 3D data provided in the above embodiments and the UAV recognition method based on multimodal 3D data provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the UAV recognition device based on multimodal 3D data provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0056] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, enable the electronic device to implement the UAV recognition method based on multimodal 3D data provided in the above embodiments.

[0057] Figure 6 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 6 The computer system 600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0058] like Figure 6 As shown, the computer system 600 includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in Read-Only Memory (ROM) 602 or a program loaded from Storage Section 608 into Random Access Memory (RAM) 603. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.

[0059] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.

[0060] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs various functions defined in the system of this application.

[0061] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0063] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0064] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned UAV recognition method based on multimodal 3D data. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently without being assembled into the electronic device.

[0065] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the UAV recognition method based on multimodal 3D data provided in the various embodiments above.

[0066] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for identifying unmanned aerial vehicles (UAVs) based on multimodal 3D data, characterized in that, The method includes: Acquire multimodal 3D data of the drone within a preset range; Three-dimensional geometric features are constructed based on the multimodal three-dimensional data; The three-dimensional geometric features are mapped to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features; the temporal and spatial correlation relationships between the three-dimensional geometric features are obtained. The structured language result is obtained based on the logical symbols, the temporal associations, and the spatial associations.

2. The method according to claim 1, characterized in that, The acquisition of multimodal 3D data within a preset range of the UAV includes: The spatial reference system of multiple sensors of the UAV is unified by coordinate system projection transformation to obtain multiple unified sensors; Using the aforementioned unified sensors, multimodal three-dimensional data within a preset spatial range of the UAV is acquired; The multimodal 3D data is spatiotemporally aligned using a soft timestamp interpolation algorithm to obtain aligned multimodal 3D data, which is then used to construct the corresponding 3D geometric features.

3. The method according to claim 1, characterized in that, The three-dimensional geometric features constructed based on the multimodal three-dimensional data include: The preset range space is voxelized to obtain multiple cubic units; Obtain the three-dimensional coordinates of the multimodal three-dimensional data, and assign the multimodal three-dimensional data to the cube unit corresponding to the three-dimensional coordinates; The multimodal 3D data is geometrically represented within the cubic unit to obtain the corresponding 3D geometric features.

4. The method according to claim 1, characterized in that, The step of mapping the three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features includes: Obtain a preset discrete symbol space, which includes pre-constructed logical predicates and spatial predicates under the BeiDou grid code; The three-dimensional geometric features are mapped based on the logical predicate and the spatial predicate using a neural network to obtain logical symbols corresponding to the three-dimensional geometric features.

5. The method according to claim 1, characterized in that, The acquisition of the temporal and spatial correlations between the three-dimensional geometric features includes: Obtain the time point information of the three-dimensional geometric features, and obtain the event association information between the three-dimensional geometric features based on the time point information; Obtain the position information of the three-dimensional coordinates corresponding to the three-dimensional geometric features in the BeiDou grid code; At each time point, the spatial correlation of the three-dimensional geometric features is obtained based on the location information.

6. The method according to claim 1, characterized in that, The process of obtaining structured language results based on the logical symbols, the temporal relationships, and the spatial relationships includes: Based on the logical symbols, the association symbol results are obtained through the temporal and spatial association relationships. The associated symbol results are imported into a preset template format to form a structured language result.

7. A drone, characterized in that, The method for identifying unmanned aerial vehicles (UAVs) based on multimodal three-dimensional data as described in any one of claims 1 to 6 is used to achieve environmental detection.

8. A recognition device for unmanned aerial vehicles (UAVs) based on multimodal three-dimensional data, characterized in that, include: The acquisition unit is used to acquire multimodal 3D data of the UAV within a preset range; Geometric feature units are used to construct three-dimensional geometric features based on the multimodal three-dimensional data; A mapping unit is used to map the three-dimensional geometric features to a preset discrete symbol space to obtain logical symbols corresponding to the three-dimensional geometric features; The association unit is used to obtain the temporal and spatial correlation relationships between the three-dimensional geometric features; The language result unit is used to obtain structured language results based on the logical symbols, the temporal correlation, and the spatial correlation.

9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the UAV recognition method based on multimodal three-dimensional data as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, It stores computer-readable instructions, which, when executed by the computer's processor, cause the computer to perform the UAV recognition method based on multimodal three-dimensional data as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Spatial intelligent multi-mode scene memory method and device and medium

    CN121600196A