Low-altitude object identification method, system and device, and storage medium

By extracting audio signal features in real time through a distributed sound source acquisition network and combining them with a voiceprint classification model, the problem of low accuracy in recognizing slow, small targets at low altitudes is solved. This enables accurate recognition and real-time positioning of low-altitude targets, and has the advantages of being radiation-free, low-cost, low-latency, and self-evolving.

CN121662081APending Publication Date: 2026-03-13CHENGDU GONGDING TECHNOLOGY CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies are easily affected by terrain occlusion and environmental interference when identifying low-altitude, slow-moving, small targets, resulting in low recognition accuracy. Furthermore, traditional systems struggle to capture their dynamic trajectories in real time.

Method used

By constructing a distributed sound source acquisition network, the power spectral density and harmonic frequency characteristics of audio signals are extracted in real time using a microphone array. This is then combined with a voiceprint classification model for identification, and front-end processing and cloud model updates are performed at edge nodes to achieve passive detection and real-time positioning of low-altitude targets.

Benefits of technology

It achieves accurate identification and real-time positioning of low-altitude targets, avoids the risk of electromagnetic exposure, breaks through the limitations of traditional technologies, and has the capabilities of being radiation-free, low-cost, low-latency, and self-evolving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662081A_ABST
    Figure CN121662081A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a low-altitude object identification method, system and device, and a storage medium, the method is applied to any node in a distributed sound source acquisition end network, and an original audio signal is acquired by acquiring a microphone; performing feature extraction on the original audio signal to obtain target voiceprint information; the target voiceprint information is input into a voiceprint classification model to obtain a target object category, the voiceprint classification model is used for performing comparison processing with a voiceprint library based on the target voiceprint information, and the voiceprint library comprises multiple pieces of voiceprint information and the object category corresponding to each piece of voiceprint information; wherein the corresponding voiceprint information comprises power spectral density and harmonic frequency. According to the scheme, key voiceprint features such as the power spectral density and the harmonic frequency in the audio signals are extracted in real time through the distributed nodes, rapid comparison is carried out by means of the voiceprint classification model and the voiceprint library, accurate recognition of the low-altitude target is achieved, and the defect that the recognition precision is poor due to the terrain influence is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a method, system, device, and storage medium for identifying low-altitude objects. Background Technology

[0002] With the rapid expansion of demand for low-altitude economy (such as drone logistics, agricultural plant protection, and urban security), the threat posed by low-altitude, slow-moving, small targets (such as consumer drones, micro aircraft, and cruise missiles) is becoming increasingly severe.

[0003] Currently, the identification of low-altitude, slow-moving, small targets is mainly based on air defense systems, primarily using radar, radio detection, and optical / infrared surveillance.

[0004] However, current technologies are hampered by terrain occlusion and other factors when identifying such small targets, making accurate identification difficult. Summary of the Invention

[0005] This application provides a method, system, device, and storage medium for identifying low-altitude objects, in order to improve the technical effect of improving the identification efficiency and accuracy of low-altitude targets.

[0006] In a first aspect, embodiments of this application provide a method for identifying low-altitude objects, applied to any node in a distributed sound source acquisition network, including:

[0007] Acquire the raw audio signal captured by the microphone;

[0008] Feature extraction is performed on the original audio signal to obtain the target voiceprint information;

[0009] The target voiceprint information is input into the voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with the voiceprint database. The voiceprint database includes: multiple voiceprint information and the object category corresponding to each voiceprint information.

[0010] The corresponding voiceprint information includes: power spectral density and harmonic frequency.

[0011] In one or more embodiments, the target object category includes: at least one categorized object and a confidence level for each categorized object;

[0012] Accordingly, the method further includes:

[0013] If the confidence level of a first classification object in at least one classification object is greater than the confidence threshold, first key information is generated. The first key information includes: the acquisition timestamp corresponding to the original audio signal, sound pressure intensity information, the identifier of the first classification object, and the azimuth angle of the sound source.

[0014] The first key information is sent to the first device, which is used to determine the trajectory and perform alarm processing based on the first key information.

[0015] In one or more embodiments, the method further includes:

[0016] If the confidence level of the target object category is empty or there is no category object is greater than the confidence threshold, the original audio signal is desensitized and encrypted.

[0017] The processed original audio signal is uploaded to a cloud server, which is used to update the voiceprint classification model and the voiceprint database based on the processed original audio signal.

[0018] In one or more embodiments, before performing feature extraction on the original audio signal to obtain target voiceprint information, the method further includes:

[0019] The original audio signal is filtered and spectral subtraction is performed to obtain the processed original audio signal.

[0020] Secondly, embodiments of this application provide a low-altitude object identification device, applied to any node in a distributed sound source acquisition network, the device comprising:

[0021] The acquisition module is used to acquire the raw audio signal captured by the microphone;

[0022] The first processing module is used to extract features from the original audio signal to obtain target voiceprint information;

[0023] The second processing module is used to input the target voiceprint information into the voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with the voiceprint database. The voiceprint database includes: multiple voiceprint information and the object category corresponding to each voiceprint information.

[0024] The corresponding voiceprint information includes: power spectral density and harmonic frequency.

[0025] In one or more embodiments, the target object category includes: at least one categorized object and a confidence level for each categorized object;

[0026] Accordingly, the second processing module is also used for:

[0027] If the confidence level of a first classification object in at least one classification object is greater than the confidence threshold, first key information is generated. The first key information includes: the acquisition timestamp corresponding to the original audio signal, sound pressure intensity information, the identifier of the first classification object, and the azimuth angle of the sound source.

[0028] The first key information is sent to the first device, which is used to determine the trajectory and perform alarm processing based on the first key information.

[0029] In one or more embodiments, the second processing module is further configured to:

[0030] If the confidence level of the target object category is empty or there is no category object is greater than the confidence threshold, the original audio signal is desensitized and encrypted.

[0031] The processed original audio signal is uploaded to a cloud server, which is used to update the voiceprint classification model and the voiceprint database based on the processed original audio signal.

[0032] In one or more embodiments, before performing feature extraction on the original audio signal to obtain target voiceprint information, the second processing module is further configured to:

[0033] The original audio signal is filtered and spectral subtraction is performed to obtain the processed original audio signal.

[0034] Thirdly, embodiments of this application provide a low-altitude object identification system, the system comprising: a distributed sound source acquisition terminal network and a first device, the distributed sound source acquisition terminal network comprising: multiple nodes;

[0035] Each node is used to execute the method described in the first aspect and any one of its components;

[0036] After the first device acquires first key information sent by at least one node, when the first category object determined based on each first key information is the same target and the number of nodes sending the first key information is greater than a preset number threshold, the device determines the running trajectory and alarm information of the first category object based on at least one first key information.

[0037] In one or more embodiments, the first device is configured to determine the trajectory of the first classified object based on at least one first key piece of information, including:

[0038] The first device is used to determine multiple location information of the first classified object based on the acquisition timestamp and the azimuth angle of the sound source in the at least one first key information;

[0039] Based on the multiple location information, the running trajectory of the first classified object is constructed.

[0040] In one or more embodiments, the first device is further configured to:

[0041] Obtain multi-source detection information for the first category object, wherein the multi-source detection information includes at least one of the following: infrared thermal imaging information, radio frequency detection information, and video surveillance information;

[0042] Based on the multi-source detection information, the first key information is corrected.

[0043] Fourthly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0044] The memory stores computer-executed instructions;

[0045] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0046] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0047] In a sixth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0048] The low-altitude object identification method, system, device, and storage medium provided in this application embodiment are applied to any node in a distributed sound source acquisition network. The method acquires raw audio signals via a microphone; extracts features from the raw audio signals to obtain target voiceprint information; and inputs the target voiceprint information into a voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with a voiceprint database, which includes multiple voiceprint information entries and the object category corresponding to each voiceprint entry. The corresponding voiceprint information includes power spectral density and harmonic frequencies. This scheme extracts key voiceprint features such as power spectral density and harmonic frequencies from audio signals in real time through distributed nodes and performs rapid comparison with the voiceprint database using a voiceprint classification model, achieving accurate identification of target object categories and effectively solving the drawback of poor identification accuracy caused by terrain influence in target acquisition techniques. Attached Figure Description

[0049] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0050] Figure 1 A schematic diagram of the first structure of the low-altitude object identification system provided in this application;

[0051] Figure 2 Flowchart of the low-altitude object identification method provided in this application Figure 1 ;

[0052] Figure 3 Flowchart of the low-altitude object identification method provided in this application Figure 2 ;

[0053] Figure 4 Flowchart of the low-altitude object identification method provided in this application Figure 3 ;

[0054] Figure 5 A schematic diagram of a second structure for the low-altitude object identification system provided in this application;

[0055] Figure 6 A schematic diagram of the structure of the low-altitude object identification device provided in this application;

[0056] Figure 7 A schematic diagram of the structure of the electronic device provided in this application.

[0057] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0058] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0059] With the rapid expansion of the low-altitude economy (such as drone logistics, agricultural plant protection, and urban security) and military needs, the threat posed by low-altitude, slow-moving, small targets (such as consumer drones, micro-aircraft, and cruise missiles) is becoming increasingly serious.

[0060] These targets are characterized by small radar cross-sections (e.g., less than 0.1 m²), low flight altitudes (e.g., less than 1000 meters), slow speeds (e.g., less than 120 km / h), high maneuverability, and high stealth. They are often used in scenarios such as border infiltration, sabotage of key facilities, and illegal surveillance. Currently, air defense systems primarily rely on radar, radio detection, and optical / infrared surveillance, but they suffer from the following significant shortcomings:

[0061] 1) Radar has limited ability to detect low, slow and small targets, is easily blocked by terrain and poses a risk of electromagnetic radiation, is inflexible in deployment and is costly;

[0062] 2) Radio detection relies on the target actively transmitting signals. It fails when faced with frequency hopping or silent mode. For example, once the UAV adopts a preset flight path for autonomous flight, frequency hopping or radio silence mode, this technology will immediately fail.

[0063] 3) Optical / infrared monitoring is severely affected by weather conditions (fog, rain, snow) and complex backgrounds. At night, in severe weather or complex backgrounds such as fog, rain, snow, etc., the detection effect drops sharply and the monitoring range is limited, making it difficult to achieve wide-area seamless coverage.

[0064] 4) Low-speed, small targets often move in groups or use urban buildings for concealment, making it difficult for traditional systems to capture their dynamic trajectories in real time.

[0065] Accordingly, in response to the technical problems existing in the prior art, the inventors of this application have the following concept: By constructing a distributed sound source acquisition network, the passive detection signal of sound can be used as the core information source. The inherent acoustic characteristics of the target aircraft, especially the combination of power spectral density and harmonic frequency with strong recognizability, can be used as the acoustic signature feature. The acoustic signature feature extraction and classification can be completed in real time through edge nodes, thereby forming a non-contact detection of low-altitude slow-moving small targets. This avoids the electromagnetic exposure risk and terrain obstruction defects of traditional radar, breaks through the dependence of radio detection on active signals, and overcomes the limitation of optical monitoring being restricted by ambient light.

[0066] Based on the above technical concepts Figure 1 A first structural schematic diagram of the low-altitude object recognition system provided in this application is shown below. Figure 1 As shown, the low-altitude object identification system may include: a cloud-based Artificial Intelligence (AI) layer (deploying the cloud server described below), a central decision-making layer (deploying the first device described below), and an edge AI layer (i.e., a distributed sound source acquisition network, deploying multiple nodes described below).

[0067] The cloud-based AI layer can provide incremental model training and regular updates; the central decision-making layer uses the Time Difference of Arrival (TDOA) - Angle of Arrival (AOA) algorithm to jointly locate four-dimensional perception interaction (integrated with external infrared, radio frequency, vision and other sensors for analysis); and the edge AI layer can provide front-end intelligent processing, data acquisition and hot model updates.

[0068] The cloud AI layer and the edge AI layer can communicate via a long-distance LoRa self-organizing network.

[0069] Optionally, the distributed sound source acquisition network can consist of multiple (e.g., thousands) passive microphone array units and multiple nodes (each node can be responsible for the audio signals acquired by one or more microphone array units). For example, a microphone array unit configured with an AI chip can also serve as a node.

[0070] In one possible implementation, microphone array units are deployed at high points such as utility poles, high-voltage towers, tall trees, and tall buildings to create a seamless auditory network; the nodes can be powered by solar panels or replaceable batteries to ensure energy supply for long-term field work.

[0071] The above is only a brief introduction to the content involved in this application. The following will elaborate on the parts that were not described in detail above.

[0072] The technical solution of this application and how it solves the above-mentioned technical problems are described in detail below with specific embodiments. The execution subject of the method involved in this technical solution is any node in a distributed sound source acquisition network. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0073] Figure 2 Flowchart of the low-altitude object identification method provided in this application Figure 1 ,like Figure 2 As shown, this method is applied to any node in a distributed sound source acquisition network, including:

[0074] Step 21: Acquire the raw audio signal captured by the microphone;

[0075] In this step, at least one microphone (i.e., a microphone array unit) is deployed on any node in the distributed sound source acquisition network. This microphone continuously listens to the surrounding environment and converts the sound wave vibrations propagating in the air into continuous electrical signals, i.e., the original analog audio signals, denoted as the original audio signals.

[0076] Optionally, the raw audio signal contains mixed information from all sound sources within the detection area (e.g., drones, ambient noise, etc.).

[0077] Step 22: Extract features from the original audio signal to obtain the target voiceprint information;

[0078] The corresponding voiceprint information includes: power spectral density and harmonic frequency.

[0079] In this step, the unique or significant identity features that can be extracted from the original audio signal containing a large amount of redundant information are extracted.

[0080] Among them, power spectral density reveals the distribution of signal power at different frequency components, and can characterize the unique operating spectrum profile of a target (such as a drone motor or rotor); harmonic frequency can reflect the structure of the fundamental frequency and its harmonics in an audio signal, and is directly related to the physical vibration characteristics of the sound source itself.

[0081] Furthermore, by extracting the power spectral density and harmonic frequencies from the original audio signal, a stable and highly recognizable voiceprint can be constructed, which is denoted as the voiceprint DNA of the original audio signal, i.e., the target voiceprint information.

[0082] Optionally, before step 22, the following can also be performed: filtering and spectral subtraction of the original audio signal to obtain the processed original audio signal.

[0083] In this implementation, the original audio signal is preprocessed to improve the accuracy of the aforementioned feature extraction.

[0084] The filtering process can employ a bandpass filter to preserve the typical operating frequency band of the target sound source (e.g., a pre-defined drone motor or rotor) while effectively filtering out low-frequency environmental noise (e.g., wind noise, traffic hum) and irrelevant high-frequency interference, thus initially focusing on the useful signal components. The spectral subtraction process identifies a relatively stable background noise spectrum from the original audio signal and subtracts it from the current original audio signal (which may be the filtered original audio signal) in the spectral domain, thereby enhancing the saliency of the target sound source signal and improving its signal-to-noise ratio. The signal obtained after these two steps is...

[0085] Step 23: Input the target voiceprint information into the voiceprint classification model to obtain the target object category.

[0086] Among them, the voiceprint classification model is used to compare the target voiceprint information with the voiceprint database, which includes multiple voiceprint information and the object category corresponding to each voiceprint information.

[0087] In this step, the extracted target voiceprint feature vector is input into a pre-trained voiceprint classification model (e.g., a machine learning-based classifier or a deep learning neural network).

[0088] The voiceprint classification model encapsulates classification logic, which quickly compares and calculates similarity between the input voiceprint information and the pre-stored, known voiceprint database, and finally outputs the object category that is approximately or consistent with the target object category.

[0089] The voiceprint database is essentially a knowledge database that links specific voiceprint features (such as the typical power spectral density and harmonic frequencies of a certain type of drone) with the corresponding object category (such as drones, cruise missiles, helicopters, background noise, etc.).

[0090] Optionally, the target object categories include: at least one categorized object and the confidence level for each categorized object.

[0091] In this implementation, when the voiceprint classification model processes the target voiceprint information, it compares it with the voiceprint database. The output target object category can also be at least one classified object (i.e., candidate object category, such as drone or helicopter) and its corresponding confidence level.

[0092] The confidence level is a value between 0 and 1, which quantifies the degree of matching between the current target voiceprint information and the object category.

[0093] The low-altitude object identification method provided in this application is applied to any node in a distributed sound source acquisition network. This method acquires raw audio signals via a microphone; extracts features from the raw audio signals to obtain target voiceprint information; and inputs the target voiceprint information into a voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with a voiceprint database, which includes multiple voiceprint information entries and the object category corresponding to each voiceprint entry. The corresponding voiceprint information includes power spectral density and harmonic frequencies. This scheme extracts key voiceprint features such as power spectral density and harmonic frequencies from the audio signal in real time through distributed nodes and performs rapid comparison with the voiceprint database using a voiceprint classification model, achieving accurate identification of the target object category and effectively solving the drawback of poor identification accuracy caused by terrain influence in target acquisition techniques.

[0094] Based on the above embodiments, Figure 3 Flowchart of the low-altitude object identification method provided in this application Figure 2 ,like Figure 3 As shown, the method also includes:

[0095] Step 31: If at least one category contains an object whose confidence level in the first category is greater than the confidence threshold, generate the first key information.

[0096] The first key information includes: the acquisition timestamp of the original audio signal, sound pressure intensity information, the identifier of the first category object, and the azimuth angle of the sound source;

[0097] In this step, the confidence scores of the multiple classified objects output by the voiceprint classification model are compared with a preset confidence threshold. When the confidence score of any classified object exceeds this confidence threshold, it is determined that this is a valid and reliable detection-trigger target detection event. Then, the classified object is identified as the first classified object, such as a drone, which is regarded as a threat object.

[0098] Furthermore, this node does not upload the massive amount of raw audio signal, but generates a highly condensed first key information, which includes the core dimensions of the event: collection timestamp: marking the time when the event occurred; sound pressure level: which can help estimate distance or target size; identification of the first category object: clarifying the identity of the threat object; sound source arrival azimuth: providing the direction of the threat object.

[0099] In one possible implementation, the confidence threshold could be 0.85.

[0100] Step 32: Send the first key information to the first device;

[0101] The first device is used for trajectory determination and alarm processing based on the first key information.

[0102] In this step, the node sends the generated first key information, which is a very small amount of data, to the first device through a low-power wide-area network technology (such as LoRa).

[0103] The first device gathers key information from multiple nodes in the network. By using the collection timestamp, sound pressure intensity information, identification of the first classified object, and azimuth angle of the sound source in this key information, and performing data fusion through algorithms such as TDOA or triangulation, the real-time position and trajectory of the target can be calculated.

[0104] Furthermore, based on the determined trajectory and the target identity of the first category object, the first device also triggers an alarm of the corresponding level and guides the defense device to respond.

[0105] The low-altitude object identification method provided in this application generates first key information by recognizing that at least one of the first classification objects has a confidence level greater than a confidence threshold. The first key information includes: the acquisition timestamp of the original audio signal, sound pressure level information, the identifier of the first classification object, and the azimuth angle of the sound source. The first key information is then sent to a first device, which performs trajectory determination and alarm processing based on the first key information. This scheme avoids network congestion caused by continuous data transmission by only uploading multi-dimensional data containing timestamps, sound pressure levels, target identifiers, and azimuth angles to the first device when the identification confidence level exceeds a threshold. Furthermore, by fusing spatiotemporal dimension information, it provides a data foundation for trajectory tracking. While ensuring low bandwidth usage, it achieves real-time accurate positioning, trajectory reconstruction, and alarming of low-altitude targets, significantly improving the decision-making efficiency and practicality of the distributed sensing system.

[0106] Based on the above embodiments, Figure 4 Flowchart of the low-altitude object identification method provided in this application Figure 3 ,like Figure 4 As shown, the method also includes:

[0107] Step 41: If the confidence level of the target object category is empty or there is no category object, and the confidence level is greater than the confidence level threshold, the original audio signal is desensitized and encrypted.

[0108] In this step, when a node cannot identify the target object category (the voiceprint classification model outputs an empty result) or the confidence of all identification results (at least one classified object) is lower than the confidence threshold, it indicates that an unknown new target or complex interference scenario has been encountered.

[0109] Therefore, to prevent the loss of data value, the original audio signal was uploaded.

[0110] However, strict security measures must be taken before uploading: de-identification will remove or obfuscate sensitive information such as voice and location that may be contained in the audio to protect privacy and compliance; encryption will ensure that the data cannot be stolen or tampered with during transmission, building a security defense for subsequent data circulation and processing.

[0111] Step 42: Upload the processed raw audio signal to the cloud server;

[0112] The cloud server is used to update the voiceprint classification model and voiceprint database based on the processed original audio signal.

[0113] In this step, the processed raw audio signal is uploaded (e.g., periodically transmitted, at 1 a.m. every day) to the cloud server; the cloud server collects difficult samples of such raw audio signals from at least one node in the distributed sound source acquisition network, and at this time, computing resources can be used to perform in-depth batch processing and analysis on these raw audio signals.

[0114] For example, more refined feature extraction, manual annotation, or model training (i.e., updating the voiceprint classification model).

[0115] Accordingly, the cloud server can optimize the parameters of the voiceprint classification model, add new voiceprint information and the corresponding object category to the voiceprint database, and finally distribute the upgraded model and database to the nodes (e.g., periodically, once a week), so that the recognition capability of the entire distributed network can continue to evolve.

[0116] The low-altitude object identification method provided in this application involves desensitizing and encrypting the original audio signal when the confidence level for a target object category being empty or lacking a classification object is greater than a confidence threshold. The processed original audio signal is then uploaded to a cloud server. The cloud server updates the voiceprint classification model and voiceprint database based on the processed original audio signal. This technical solution ensures the security and privacy compliance of sensitive original data by uploading unidentifiable or low-confidence original audio signals to the cloud after desensitization and encryption. It also transforms massive edge nodes into a continuously learning perceptual network, using these challenging samples to optimize the model and voiceprint database in the cloud. This improves the entire network's ability to identify new, camouflaged, or background noise targets, achieving a leap from static protection to dynamic evolution in the early warning system.

[0117] Based on the above method embodiments, Figure 5 A second structural schematic diagram of the low-altitude object identification system provided in this application is shown below. Figure 5 As shown, the low-altitude object identification system includes: a distributed sound source acquisition network and a first device; the distributed sound source acquisition network includes: multiple nodes;

[0118] Each node is used in any of the methods described in the above method embodiments;

[0119] After acquiring first key information sent by at least one node, the first device determines the running trajectory and alarm information of the first category object based on at least one first key information when the first category object determined based on each first key information is the same target and the number of nodes sending the first key information is greater than a preset number threshold.

[0120] In this implementation, multiple key input information are associated with a target. By determining that the first category objects identified are of the same type (e.g., all are "drones"), and combining the timestamp and azimuth information, it is preliminarily determined that they originate from the same target.

[0121] Subsequently, a preset threshold condition was introduced, meaning that subsequent processing is only triggered when the number of nodes sending information reaches a certain scale. This eliminates false alarms caused by misreporting from a single or a few nodes or by accidental detection, ensuring the statistical significance of detection events and thus greatly improving the reliability of system alarms.

[0122] After the above conditions are met, the first device performs data fusion using multiple first key information packets from different nodes. By integrating the timestamps, azimuth angles of sound sources, and sound pressure levels (which can be used to assist in ranging) in these information packets, it can use algorithms such as triangulation and TDOA to calculate the spatial position of the target at different time points, and thus determine its trajectory.

[0123] Furthermore, based on this trajectory, more detailed alarm information, including target location, speed, heading, and threat level, can be generated, providing accurate and reliable situational awareness for final countermeasures and interception.

[0124] Optionally, the first device is used to determine the trajectory of the first classified object based on at least one first key piece of information, including the following steps:

[0125] Step 1: The first device is used to determine multiple location information of the first classified object based on the acquisition timestamp and the azimuth of the sound source in at least one first key information.

[0126] In this implementation, the first device gathers key information from multiple nodes and first utilizes the TDOA positioning principle: by comparing the acquisition timestamps of the same target acoustic signal arriving at different nodes, the time difference between each pair of nodes is calculated. Since the speed of sound is constant, each time difference can define a hyperboloid with the two nodes as foci, and the target must lie on this hyperboloid. Multiple such hyperboloids intersect in space, forming a series of possible positioning points.

[0127] To quickly converge to a unique solution from these ambiguous points, AOA assistance is introduced: the azimuth angle of arrival of the sound source provided by each node adds a directional constraint to the TDOA hyperboloid, just like adding a compass rose to each curve, thereby eliminating false solutions in space and calculating one or more location information corresponding to the target (i.e., the first-classified object) at each time step.

[0128] Step 2: The first device is also used to construct the running trajectory of the first category of objects based on multiple location information.

[0129] In this implementation, after obtaining multiple location information of the first classified object at a series of consecutive time points through TDOA-AOA fusion positioning, the first device associates and sorts these discrete location points according to the order of collection timestamps.

[0130] Subsequently, trajectory filtering algorithms (such as Kalman filtering) are used to process these location points to smooth out positioning jitter caused by measurement errors and to predict the target's motion state (such as speed and orientation). By connecting these processed and filtered continuous location points, a smooth and continuous running trajectory is constructed.

[0131] Optionally, the first device is also used for:

[0132] Obtain multi-source detection information for the first category object, and correct the first key information based on the multi-source detection information.

[0133] The multi-source detection information includes at least one of the following: infrared thermal imaging information, radio frequency detection information, and video surveillance information.

[0134] In this implementation, after generating the first key information (preliminary localization and identification) for the first category object (e.g., drone) based on acoustic detection, multi-source detection information from other heterogeneous sensors can be acquired to increase the identification accuracy.

[0135] For example, infrared thermal imaging information can be used to verify and supplement acoustic positioning at night or in inclement weather, accurately capturing the heat-generating components of a target (such as an engine); radio frequency detection information can sniff and identify the remote control or image transmission signals of a target, providing independent electronic fingerprint evidence for identification; while video surveillance information can provide the most intuitive visual confirmation, effectively distinguishing the target from drones, birds or other objects in complex scenarios.

[0136] Then, data fusion algorithms (such as Kalman filtering and Bayesian inference) are run to align and correlate the first key information with other sensor information in time and space. Through mutual verification and compensation, the first key information is corrected.

[0137] For example, it can more accurately calibrate the target's position coordinates, improve the confidence of identity recognition, and even calculate the speed and heading, thereby ultimately outputting a more accurate, stable, and reliable fusion situation, providing a high-quality information foundation for subsequent alarm and countermeasure decisions.

[0138] In addition, the first device is also used to predict the trajectory of the first classified object. Based on its trajectory, the motion state of the first classified object can be determined, and the sharing path of the target within a first preset time period can be extrapolated and predicted, such as the flight path of 3-10 seconds.

[0139] The low-altitude object identification system provided in this application embodiment achieves passive, wide-area, real-time, and high-precision monitoring and early warning of low-altitude targets, effectively overcoming the limitations of traditional technologies. It has significant advantages such as no radiation, low cost, low latency, high intelligence, and self-evolution. The system can be described from the following perspectives:

[0140] 1) Novelty: The collaborative architecture of local lightweight identification at edge nodes - integrated positioning and decision-making at the command center - continuous evolution of cloud models; TDOA-AOA joint acoustic positioning and fusion of terrain data; incremental learning and hot update mechanism of the model; thousands of distributed sound source acquisition terminals to build an auditory sky network;

[0141] 2) Creativity: The distributed microphone nodes have built-in AI chips for real-time voiceprint recognition; the acoustic positioning results can be fused with infrared, radio frequency, and vision in four dimensions; samples are collected daily and the model is updated incrementally every week;

[0142] 3) Practicality: Easy to deploy; lower cost than radar; no active electromagnetic radiation.

[0143] Based on the above method embodiments, Figure 6 A schematic diagram of the structure of the low-altitude object identification device provided in this application is shown below. Figure 6 As shown, the device is applied to any node in a distributed sound source acquisition network and includes:

[0144] Acquisition module 61 is used to acquire the raw audio signal collected by the microphone;

[0145] The first processing module 62 is used to extract features from the original audio signal to obtain the target voiceprint information;

[0146] The second processing module 63 is used to input the target voiceprint information into the voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with the voiceprint database. The voiceprint database includes: multiple voiceprint information and the object category corresponding to each voiceprint information.

[0147] The corresponding voiceprint information includes: power spectral density and harmonic frequency.

[0148] In one or more embodiments, the target object category includes: at least one categorized object and a confidence level for each categorized object;

[0149] Correspondingly, the second processing module 63 is also used for:

[0150] If at least one of the classification objects has a confidence level greater than the confidence threshold, first key information is generated. The first key information includes: the acquisition timestamp corresponding to the original audio signal, sound pressure intensity information, the identifier of the first classification object, and the azimuth angle of the sound source.

[0151] The first key information is sent to the first device, which is used to determine the trajectory and handle alarms based on the first key information.

[0152] In one or more embodiments, the second processing module 63 is further configured to:

[0153] If the confidence level of the target object category is empty or there is no category object is greater than the confidence threshold, the original audio signal is desensitized and encrypted.

[0154] The processed raw audio signal is uploaded to the cloud server, which then updates the voiceprint classification model and voiceprint library based on the processed raw audio signal.

[0155] In one or more embodiments, before performing feature extraction on the original audio signal to obtain the target voiceprint information, the second processing module 63 is further configured to:

[0156] The original audio signal is filtered and spectral subtraction is performed to obtain the processed original audio signal.

[0157] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical element, or they can be physically separated. Furthermore, these modules can be implemented entirely in software through processing element calls, or entirely in hardware. Alternatively, some modules can be implemented through processing element calls in software, while others can be implemented in hardware. Moreover, these modules can be integrated together or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.

[0158] As can be seen from the above, the low-altitude object identification device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0159] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device provided in this embodiment can be the node described above, and may include at least one processor 71 and a memory 72.

[0160] Optionally, the electronic device also includes a communication component 73.

[0161] The processor 71, memory 72, and communication component 73 are connected via bus 74.

[0162] In a specific implementation, at least one processor 71 executes computer execution instructions stored in memory 72, causing at least one processor 71 to perform the above-described method.

[0163] The specific implementation process of processor 71 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0164] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0165] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0166] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0167] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0168] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0169] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0170] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0171] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0172] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0173] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0174] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0175] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

Claims

1. A method for identifying low-altitude objects, characterized in that, The method, applied to any node in a distributed sound source acquisition network, includes: Acquire the raw audio signal captured by the microphone; Feature extraction is performed on the original audio signal to obtain the target voiceprint information; The target voiceprint information is input into the voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with the voiceprint database. The voiceprint database includes: multiple voiceprint information and the object category corresponding to each voiceprint information. The corresponding voiceprint information includes: power spectral density and harmonic frequency.

2. The method according to claim 1, characterized in that, The target object category includes: at least one categorized object and the confidence level of each categorized object; Accordingly, the method further includes: If the confidence level of a first classification object in at least one classification object is greater than the confidence threshold, first key information is generated. The first key information includes: the acquisition timestamp corresponding to the original audio signal, sound pressure intensity information, the identifier of the first classification object, and the azimuth angle of the sound source. The first key information is sent to the first device, which is used to determine the trajectory and perform alarm processing based on the first key information.

3. The method according to claim 2, characterized in that, The method further includes: If the confidence level of the target object category is empty or there is no category object is greater than the confidence threshold, the original audio signal is desensitized and encrypted. The processed original audio signal is uploaded to a cloud server, which is used to update the voiceprint classification model and the voiceprint database based on the processed original audio signal.

4. The method according to any one of claims 1-3, characterized in that, Before performing feature extraction on the original audio signal to obtain the target voiceprint information, the method further includes: The original audio signal is filtered and spectral subtraction is performed to obtain the processed original audio signal.

5. A low-altitude object identification system, characterized in that, The system includes: a distributed sound source acquisition terminal network and a first device, wherein the distributed sound source acquisition terminal network includes: multiple nodes; Each node is used to execute the method described in any one of claims 1-4; After the first device acquires first key information sent by at least one node, when the first category object determined based on each first key information is the same target and the number of nodes sending the first key information is greater than a preset number threshold, the device determines the running trajectory and alarm information of the first category object based on at least one first key information.

6. The system according to claim 5, characterized in that, The first device is used to determine the trajectory of the first classified object based on at least one first key piece of information, including: The first device is used to determine multiple location information of the first classified object based on the acquisition timestamp and the azimuth angle of the sound source in the at least one first key information; Based on the multiple location information, the running trajectory of the first classified object is constructed.

7. The system according to claim 6, characterized in that, The first device is also used for: Obtain multi-source detection information for the first category object, wherein the multi-source detection information includes at least one of the following: infrared thermal imaging information, radio frequency detection information, and video surveillance information; Based on the multi-source detection information, the first key information is corrected.

8. A device for identifying low-altitude objects, characterized in that, The device, applicable to any node in a distributed sound source acquisition network, comprises: The acquisition module is used to acquire the raw audio signal captured by the microphone; The first processing module is used to extract features from the original audio signal to obtain target voiceprint information; The second processing module is used to input the target voiceprint information into the voiceprint classification model to obtain the target object category. The voiceprint classification model is used to compare the target voiceprint information with the voiceprint database. The voiceprint database includes: multiple voiceprint information and the object category corresponding to each voiceprint information. The corresponding voiceprint information includes: power spectral density and harmonic frequency.

9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-4.