A method of detecting an intrusion target and an image capturing device

CN122842243APending Publication Date: 2026-09-29SHENZHEN ZHILING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611192427.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,红外传感器基于热辐射变化原理工作,对于无人机等非发热或低热源目标难以有效识别和触发,导致系统无法及时响应无人机偷拍、踩点等新型安全威胁

Benefits of technology

[0015]本申请上述技术方案与现有技术相比具有如下优点:本申请提供了一种入侵目标的检测方法及图像捕获设备,本申请在检测是否存在入侵目标时,需要实时采集环境声音信号,若环境声音信号的当前环境音量超过音量阈值,且音频特征与目标飞行物的音频特征数据匹配时,根据视频帧图像检测运动目标;响应于运动目标在连续预定帧图像中均位于预定监测区域内,且帧间位移小于预设位移阈值,将运动目标判定为入侵目标;可见,本申请基于声音匹配的方式可有效检测出发生特定声音的入侵目标,从而对热辐射变化不明显的入侵目标进行有效检测;并且,本申请通过声音初步判定存在入侵目标后,还会通过视频帧图像进行复查,只有复查通过后才最终确定存在入侵目标,确保检测结果的准确性。该双重验证机制通过音频特征匹配进行初步筛选,再由视频帧图像的时空运动特征分析进行二次确认,从技术层面消除了单一传感器对非热源目标漏检及雷达方案高误判的固有缺陷;所述音频处理器与图像处理器采用分级唤醒策略,使高功耗的图像处理器仅在音频特征匹配成功时才被中断信号唤醒,从而实现极低待机功耗下的持续监听。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842243A_ABST
    Figure CN122842243A_ABST
Patent Text Reader

Abstract

This application relates to the field of intrusion target detection, and discloses an intrusion target detection method and image capture device. When detecting the presence of an intrusion target, this application needs to collect ambient sound signals in real time. If the current ambient volume of the ambient sound signal exceeds a volume threshold and the audio characteristics match the audio characteristic data of the target flying object, the moving target is detected based on the video frame image. In response to the moving target being located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement being less than a preset displacement threshold, the moving target is determined to be an intrusion target. It can be seen that this application can effectively detect intrusion targets that produce specific sounds based on sound matching, thereby effectively detecting intrusion targets with insignificant changes in thermal radiation. Furthermore, after initially determining the presence of an intrusion target through sound, this application will also conduct a review through video frame images. Only after the review is passed will the presence of an intrusion target be finally confirmed, ensuring the accuracy of the detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intrusion target detection, and in particular to a method for detecting intrusion targets and an image capture device. Background Technology

[0002] Network cameras (IP cameras, IPCs) are important security monitoring devices widely used in homes, public places, and other locations. Currently, mainstream IPCs typically use passive infrared sensors as the motion detection trigger mechanism. When a moving object is detected, the main control ISP (Image Signal Processor) is activated to determine if there is any suspicious intrusion. However, infrared sensors operate based on the principle of thermal radiation changes, making it difficult to effectively identify and trigger non-heat-generating or low-heat-source targets such as drones. This results in the system being unable to respond promptly to emerging security threats such as drone surveillance and reconnaissance.

[0003] Therefore, how to effectively detect intrusion targets with insignificant changes in thermal radiation is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a method for detecting intrusion targets and an image capture device to effectively detect intrusion targets with insignificant changes in thermal radiation.

[0005] Firstly, this application provides a method for detecting an intrusion target, comprising: Real-time acquisition of ambient sound signals; When the current ambient volume of the ambient sound signal exceeds the volume threshold, an audio segment is acquired, and the audio features of the audio segment are extracted. The audio features are matched with the audio feature data of the target flying object to generate a first matching result; When the first matching result is a successful match, a moving target is detected from the video frame image; In response to the fact that the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold, the moving target is determined to be an intrusion target.

[0006] Optionally, after real-time acquisition of ambient sound signals, the system also includes: The noise reduction process is performed by utilizing the phase difference between various ambient sound signals to obtain a noise-reduced sound signal; wherein, each ambient sound signal is a sound signal acquired in real time by each sound acquisition unit in the sound acquisition array; The ambient volume is determined based on the noise-reduced sound signal.

[0007] Optionally, when the current ambient volume of the ambient sound signal exceeds a volume threshold, an audio segment is acquired, and the audio features of the audio segment are extracted, including: When the audio processor, which is in low-power standby mode, detects that the current ambient volume of the ambient sound signal exceeds the volume threshold, the audio processor is switched from low-power standby mode to active mode. The audio processor then collects audio segments and extracts the audio features of the audio segments.

[0008] Optionally, after extracting the audio features of the audio segment, the method further includes: The audio processor matches the audio features with the audio feature data of each dangerous event to generate a second matching result. When the second matching result indicates that the audio feature successfully matches the audio feature data of the target dangerous event, a dangerous event is determined to have occurred, and an alarm signal is issued.

[0009] Optionally, the audio features are matched with the audio feature data of the target flying object to generate a first matching result; when the first matching result is a successful match, detecting the moving target from the video frame image includes: The audio processor matches the audio features with the audio feature data of the target flying object to generate a first matching result; When the first matching result is a successful match, an interrupt signal is sent to the image processor to switch the image processor from sleep state to working state; The image processor acquires the current frame image and compares it with a pre-stored background image to obtain a comparison result; the pre-stored background image is an environmental image that the image processor stores in advance before entering a sleep state. In response to the comparison result indicating the presence of a new moving target, a threat status monitoring process for the moving target is performed.

[0010] Optionally, the threat status monitoring process for the moving target includes: Determine the position information of the moving target in the current frame image and subsequent consecutive frame images, and determine the inter-frame displacement of two consecutive frame images based on the position information; The location information is compared with the location range of the predetermined monitoring area to obtain a location comparison result; the inter-frame displacement is compared with a preset displacement threshold to obtain a displacement comparison result. When the position comparison result indicates that the moving target is located within a predetermined monitoring area in consecutive predetermined frame images, and the displacement comparison result indicates that the inter-frame displacement of the moving target in consecutive predetermined frame images is less than a preset displacement threshold, it is determined that the target is under UAV threat.

[0011] Optionally, the detection method further includes: In response to the fact that the moving target is not always located within the predetermined monitoring area in consecutive predetermined frame images, or that the inter-frame displacement is not less than a preset displacement threshold, the number of consecutive false alarms is updated. When the number of consecutive false alarms exceeds the threshold, the sound acquisition array and audio processor are powered off. The sound acquisition array and the audio processor are powered on again at fixed intervals so that the sound acquisition array can reacquire the ambient sound signal. The audio processor compares the current ambient volume of the reacquired ambient sound signal with a volume threshold to obtain a volume comparison result. When the volume comparison result indicates that the current ambient volume of the re-acquired ambient sound signal exceeds the volume threshold, the sound acquisition array and the audio processor are powered off.

[0012] Secondly, this application provides an image capture device, comprising: A sound acquisition array is used to acquire ambient sound signals in real time. An audio processor is configured to acquire an audio segment and extract audio features of the audio segment when the current ambient volume of the ambient sound signal exceeds a volume threshold; and to match the audio features with the audio feature data of the target flying object to generate a first matching result. An image processor is configured to detect a moving target from a video frame image when the first matching result is a successful match; and to determine the moving target as an intrusion target when the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold.

[0013] Optionally, the sound acquisition array includes a first sound acquisition unit and a second sound acquisition unit; the first sound acquisition unit and the second sound acquisition unit are respectively disposed on both sides of the image acquisition unit of the image capture device; the first sound acquisition unit is used to acquire a first ambient sound signal, and the second sound acquisition unit is used to acquire a second ambient sound signal; the image acquisition unit is used to acquire video frame images; The audio processor is specifically used to perform noise reduction processing based on the phase difference between the first ambient sound signal and the second ambient sound signal to obtain a noise-reduced sound signal, and to determine the current ambient volume based on the noise-reduced sound signal.

[0014] Optionally, the audio processor and the image processor are connected via an integrated circuit bus, which is used to transmit interactive data; The audio processor and the image processor are also connected via an interrupt signal line, which is used to transmit an interrupt signal sent by the audio processor to switch the image processor from a sleep state to a working state.

[0015] Compared with the prior art, the above-mentioned technical solution of this application has the following advantages: This application provides a method for detecting intrusion targets and an image capture device. When detecting the existence of an intrusion target, this application needs to collect ambient sound signals in real time. If the current ambient volume of the ambient sound signal exceeds the volume threshold and the audio characteristics match the audio characteristic data of the target flying object, the moving target is detected based on the video frame image. In response to the moving target being located within the predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement being less than the preset displacement threshold, the moving target is determined to be an intrusion target. It can be seen that this application can effectively detect intrusion targets that produce specific sounds based on sound matching, thereby effectively detecting intrusion targets with insignificant changes in thermal radiation. Furthermore, after initially determining the existence of an intrusion target through sound, this application will also conduct a review through video frame images. Only after the review is passed will the existence of an intrusion target be finally confirmed, ensuring the accuracy of the detection results. This dual verification mechanism performs initial screening through audio feature matching and secondary confirmation through spatiotemporal motion feature analysis of video frame images. From a technical perspective, it eliminates the inherent defects of single sensors in missing non-heat source targets and high misjudgment of radar schemes. The audio processor and image processor adopt a hierarchical wake-up strategy, so that the high-power image processor is only woken up by the interrupt signal when the audio feature matching is successful, thereby achieving continuous monitoring with extremely low standby power consumption. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0019] Figure 1 This is a schematic diagram of the hardware structure of an image capture device provided in an embodiment of this application; Figure 2This is a schematic diagram of the hardware structure of another image capture device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of another image capture device provided in an embodiment of this application; Figure 4 This is a flowchart of an intrusion target detection method provided in an embodiment of this application. Detailed Implementation

[0020] This specific embodiment is merely an explanation of this application and is not intended to limit it. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of this application.

[0021] It should be noted that, in the optional embodiments of this application, the data related to object information, when applied to specific products or technologies, requires the permission or consent of the object. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this application involve data related to an object, it must be obtained with the object's authorization and consent, the authorization and consent of relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the individual's consent. If sensitive information is involved, the separate consent of the information subject is required. The embodiments also need to be implemented with the object's authorization and consent.

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0024] Image capture devices are those capable of capturing and analyzing images, including network cameras and doorbell devices. Traditional solutions use passive infrared (PIR) sensors as the motion detection trigger mechanism. This method cannot effectively identify and trigger non-heat-generating or low-heat-source targets such as drones. Some traditional solutions also use radar sensors as the motion detection trigger mechanism; however, radar sensors have high power consumption, and when detecting moving objects, they can only sense the physical displacement of the target. Therefore, when a radar sensor detects a displacement signal, the system cannot determine whether the displacement is caused by a drone, a bird, or other interference, leading to a high false positive rate and an inability to effectively identify intruding targets such as drones. Furthermore, battery-powered image capture devices are limited by energy supply; even with solar panels, it is difficult to achieve continuous 24 / 7 recording, making it impossible to detect intruding targets in real time via video.

[0025] In summary, traditional solutions mainly rely on infrared sensors or radar for motion detection, which cannot effectively detect intrusion targets such as drones that do not have obvious thermal characteristics. Therefore, in order to solve the above problems, this application provides an intrusion target detection method and image capture device to effectively detect intrusion targets with insignificant changes in thermal radiation.

[0026] To provide a clear explanation of this application, an image capture device provided in this application will be described first, see [link to relevant documentation]. Figure 1 This is a schematic diagram of the hardware structure of an image capture device provided in an embodiment of this application, such as... Figure 1 As shown, the image capture device 100 includes: The sound acquisition array 110 is used to acquire ambient sound signals in real time. The audio processor 120 is used to acquire audio segments and extract audio features of the audio segments when the current ambient volume of the ambient sound signal exceeds a volume threshold; and to match the audio features with the audio feature data of the target flying object to generate a first matching result. The image processor 130 is used to detect a moving target from a video frame image when the first matching result is a successful match; and to determine the moving target as an intrusion target when the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold.

[0027] In this application, the image capture device 100 can be installed in the area to be monitored, and can be a network camera, video doorbell, or other device for security monitoring of the monitored area. The image capture device 100 includes a sound acquisition array 110, an audio processor 120, and an image processor 130. The sound acquisition array 110 is connected to the audio processor 120. The sound acquisition array 110 is used to acquire ambient sound signals in real time and send them to the audio processor 120 for processing. The audio processor 120 can be a general-purpose audio DSP (Digital Signal Processor) chip added to an existing architecture. The audio processor 120 is connected to the image processor 130 through an independent interrupt signal line, which is used to wake up the image processor 130 from sleep mode via a hardware interrupt.

[0028] In another embodiment, the audio processing function of the audio processor 120 can be implemented by a digital signal processor core integrated within the image processor 130, thereby reducing hardware costs and circuit board area. In this embodiment, the digital signal processor core within the image processor 130 continuously monitors the volume of the ambient sound signal in a low-power standby state. When the volume exceeds a volume threshold, the digital signal processor core switches from the low-power standby state to the active state and performs audio feature extraction and matching operations.

[0029] See Figure 2 The figure shows a schematic diagram of the hardware structure of another image capture device 100 provided in this application embodiment. The sound acquisition array 110 includes a first sound acquisition unit 111 and a second sound acquisition unit 112. The first sound acquisition unit 111 and the second sound acquisition unit 112 are respectively disposed on both sides of the image acquisition unit 140 of the image capture device 100. The first sound acquisition unit 111 is used to acquire a first ambient sound signal, and the second sound acquisition unit 112 is used to acquire a second ambient sound signal. The image acquisition unit 140 is used to acquire video frame images. The audio processor 120 in this application is specifically used to perform noise reduction processing based on the phase difference between the first ambient sound signal and the second ambient sound signal to obtain a noise-reduced sound signal, and to determine the current ambient volume based on the noise-reduced sound signal.

[0030] Specifically, the sound acquisition array 110 in this application includes a first sound acquisition unit 111 and a second sound acquisition unit 112. The first sound acquisition unit 111 and the second sound acquisition unit 112 are physically positioned on opposite sides of the image acquisition unit 140, maintaining sufficient spacing to obtain an effective phase difference. This application can utilize the sound acquisition array 110 to acquire the phase difference of ambient sound signals, allowing the audio processor 120 to effectively filter out co-propagating ambient noise at the front end. In this embodiment, the sound acquisition array 110 can be an array composed of two microphones, the image acquisition unit 140 can be a camera, and the DSP chip is connected to the microphone array. Based on the phase difference of the ambient sound signals acquired by the two microphones, co-propagating ambient noise is effectively filtered out at the front end, improving the target audio signal-to-noise ratio.

[0031] Furthermore, the DSP chip in this application is preloaded with pre-trained predetermined audio feature data, including audio feature data of the target flying object, such as the audio feature data of the drone propeller (DATA_Audio). The audio processor 120 is equipped with a two-level wake-up mechanism to reduce power consumption. Specifically, in the low-power standby state, the audio processor 120 removes environmental noise from the ambient sound signal to obtain a denoised sound signal. Based on the denoised sound signal, it determines the current ambient volume. When the current ambient volume of the ambient sound signal exceeds the volume threshold, the audio processor 120 switches from the low-power standby state to the active state, acquires audio segments, and extracts the audio features of the audio segments. It matches the audio features with the audio feature data of the target flying object to generate a first matching result. When the first matching result is successful, it sends an interrupt signal to the image processor 130 to switch the image processor 130 from the sleep state to the working state. The image processor 130 can be the master control ISP.

[0032] See Figure 3 This is a schematic diagram of the hardware structure of another image capture device 100 provided in an embodiment of this application, as shown below. Figure 3 As shown, the audio processor 120 and the image processor 130 are connected via an integrated circuit bus, which can be an I²C bus (Inter-Integrated Circuit Bus). The audio processor 120 communicates with the image processor 130 via the I²C bus. This integrated circuit bus is used to transmit interactive data, including audio feature data, user-defined volume thresholds, etc. The audio processor 120 and the image processor 130 are also connected via an interrupt signal line. The interrupt signal line is used to transmit the interrupt signal sent by the audio processor 120 to switch the image processor 130 from sleep state to working state.

[0033] Specifically, the main control ISP and the DSP chip in this application are connected in two ways: one is through the standard I²C bus, and the other is through the GPIO (General Purpose Input / Output) interrupt signal line. To reduce power consumption, the main control ISP is normally in a low-power sleep state, the I²C bus is suspended, and the GPIO interrupt is configured as a high-priority interrupt. The I²C bus can only be woken up by this high-priority interrupt. After being woken up, the main control ISP can communicate with the DSP chip through the I²C bus. Specifically, after the image processor 130 is woken up, it is used to acquire video frame images through the image acquisition unit 140, detect moving targets from the video frame images, and determine the moving target as an intrusion target if the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold.

[0034] In summary, this application employs a two-level wake-up mechanism for the audio processor within the image capture device. When the volume threshold is not triggered, it maintains a low-power standby state to reduce power consumption. If the volume threshold is triggered and audio feature matching is performed, the image processor remains in a low-power sleep state. Only after the audio processor determines that the audio feature matching is successful is the image processor awakened via an interrupt signal, further reducing overall device power consumption. The image processor detects moving targets using video frame images. If the moving target is located within a predetermined monitoring area in consecutive predetermined frame images, and the inter-frame displacement is less than a preset displacement threshold, the moving target is identified as an intrusion target. Therefore, this application performs initial screening through audio feature matching, followed by secondary confirmation through spatiotemporal motion feature analysis of video frame images. This method effectively reduces the false alarm rate while ensuring the detection capability for non-thermal radiation change targets such as drones, effectively compensating for the limitations of passive infrared or radar solutions in detecting non-thermal source targets. Furthermore, this application places the first sound acquisition unit and the second sound acquisition unit on both sides of the image acquisition unit and maintains a predetermined distance. This allows for noise reduction processing using the phase difference between the two sound signals, effectively suppressing environmental common-mode noise propagating in the same direction and improving the signal-to-noise ratio of the audio front end.

[0035] It should be noted that this application primarily applies to image capture devices, such as battery-powered network cameras or video doorbells. Unlike traditional professional security systems, consumer-grade image capture devices are typically battery-powered, making frequent battery replacements or power cord connections difficult after installation. In such devices, standby power consumption is the core factor determining battery life; if the average standby current is too high, the device's battery life will be significantly shortened. Therefore, minimizing standby power consumption while maintaining detection capabilities is the most significant technical constraint distinguishing consumer-grade image capture devices from traditional professional security systems.

[0036] In traditional solutions, if a System-on-Chip (SoC) is used to continuously analyze audio signals to detect targets such as drones, the SoC needs to remain constantly operational to run audio processing algorithms. Due to the high power consumption of the SoC, it cannot continuously run audio analysis algorithms while maintaining extremely low power consumption. Even when the SoC enters sleep mode, its audio processing module typically powers down, preventing it from performing any audio signal processing tasks during sleep. To address these technical issues, this application adds a dedicated audio processor to the image capture device. This audio processor is specifically designed for audio signal acquisition, processing, and feature recognition. Since the audio processor is a dedicated chip designed for audio signal processing, its hardware architecture is optimized for audio algorithms, enabling it to extract and match audio features with extremely low power consumption. By offloading the audio analysis task from the SoC to the dedicated audio processor, the SoC can remain in deep sleep mode when no triggering event occurs, only being awakened when the audio processor confirms the presence of a target audio event. This significantly reduces overall system power consumption while achieving continuous audio monitoring.

[0037] To further reduce power consumption, this application incorporates a two-level wake-up mechanism within the audio processor. The first level is volume threshold triggering. In low-power standby mode, the audio processor only monitors the ambient sound volume; at this time, the core processing module is off, resulting in extremely low power consumption. Only when the ambient volume exceeds a preset threshold does the audio processor switch to active mode, activating the audio acquisition and feature extraction module and entering the second level of processing. The second level is audio feature matching. In active mode, the audio processor matches the extracted audio features with pre-loaded target flying object audio feature data; only when a match is successful does the audio processor send an interrupt signal to the main control SoC to wake it up. Compared to a single-level wake-up mechanism, this two-level wake-up mechanism offers a significant quantitative advantage in power consumption. Under a single-level wake-up mechanism, the audio processor needs to remain active at all times to continuously run the feature extraction and matching algorithm. However, under the two-level wake-up mechanism of this application, the audio processor is in a low-power standby state for most of the time, only performing volume detection. Through the above two-level wake-up mechanism, this application can keep the average power consumption of the audio monitoring function at an extremely low level while maintaining effective detection of intrusion targets such as drones, significantly extending the battery life of the battery-powered image capture device.

[0038] See Figure 4 The above is a flowchart of an intrusion target detection method provided in an embodiment of this application. The method specifically includes: S101. Real-time acquisition of ambient sound signals.

[0039] In this application, the ambient sound signals of the monitoring area can be collected in real time by a sound acquisition device. The sound acquisition device can be a microphone or other types of acoustic-electric transducer, as long as it can convert the sound signals in the environment into electrical signals.

[0040] S102. When the current ambient volume of the ambient sound signal exceeds the volume threshold, an audio segment is collected and the audio features of the audio segment are extracted.

[0041] In this application, the current ambient volume needs to be determined based on the currently acquired ambient sound signal. This determination can be achieved by calculating the average amplitude or root mean square value of the ambient sound signal per unit time, and is not specifically limited here. After determining the current ambient volume, it is compared with a preset volume threshold. The volume threshold is a pre-set amplitude limit used to distinguish whether the environment contains a target sound, which is the sound emitted by a dangerous intrusion target to be identified, including drone propellers, etc. Audio feature matching is only performed when the ambient volume exceeds this volume threshold.

[0042] The specific value of the volume threshold can be preset according to the actual application scenario, or automatically determined after identifying the application scenario. Specifically, when setting the volume threshold, users can use the sensitivity level selection function in the APP (Application). Users can adjust the triggered volume threshold according to the installation environment. Different sensitivity levels correspond to different volume thresholds. For example, if the installation environment is on the street, users can set it to a low sensitivity level, which corresponds to a higher volume threshold. If the installation environment is in a courtyard, users can set it to a medium sensitivity level, which corresponds to a moderate volume threshold. If the installation environment is in a quiet indoor environment, users can set it to a high sensitivity level, which corresponds to a lower volume threshold.

[0043] If the current ambient volume exceeds a volume threshold, an audio segment of preset duration is collected. This audio segment can be extracted from the continuously collected ambient sound signal in step S101, or it can be collected again after the volume threshold is exceeded. After the audio segment is collected, audio features are extracted from it. The extraction of audio features can be done using various methods such as Mel-frequency cepstral coefficients and linear prediction coefficients; this application does not specify a particular method for this. If the current ambient volume does not exceed the volume threshold, step S101 is executed to monitor the ambient volume, but the above-mentioned audio segment collection and subsequent steps are not performed.

[0044] S103. Match the audio features with the audio feature data of the target flying object to generate the first matching result.

[0045] In this application, after extracting the audio features of an audio segment, it needs to be matched with pre-loaded predetermined audio feature data to generate a first matching result. This predetermined audio feature data is audio feature data of a specific intrusion target that has been pre-trained and feature-extracted. In this application, the predetermined audio feature data specifically refers to the audio feature data of the target flying object, i.e., the audio feature data of a drone propeller. It should be noted that the predetermined audio feature data in this application can also be audio feature data of threatening wild animals, or audio feature data of abnormal shouts, etc., to identify various types of intrusion targets. The audio processor calculates the first matching degree between the extracted audio features and the audio feature data of the target flying object, and generates a first matching result based on the first matching degree. When calculating the first matching degree between the extracted audio features and the audio feature data of the target flying object, cosine similarity calculation methods, Euclidean distance calculation methods, or deep neural network-based matching methods, etc., can be used; this application does not specifically limit this method.

[0046] The first matching result in this application is used to indicate whether the current audio feature matches the audio feature data of the target flying object successfully. For example, when the first matching degree exceeds a preset first matching degree threshold, the first matching result is a successful match; when the first matching degree does not exceed the preset first matching degree threshold, the first matching result is a failed match. The specific value of the first matching degree threshold can be set according to the requirements of false alarm rate and false negative rate in the actual application scenario. When a high detection sensitivity is required, the first matching degree threshold can be set to a lower value, such as 60%, so that sounds that are partially similar to the audio features of the target flying object can also trigger a successful match; when a low false alarm rate is required, the first matching degree threshold can be set to a higher value, such as 85%, so that only sounds that are highly similar to the audio features of the target flying object can trigger a successful match.

[0047] S104. When the first matching result is a successful match, detect the moving target from the video frame image.

[0048] In this application, when the first matching result is successful, it is necessary to acquire video frame images and detect moving targets from the video frame images. Specifically, when detecting moving targets from video frame images, this application can acquire one or more video frame images, process the video frame images, and detect whether there are moving targets in them. In this application, video frame images can be acquired through an image acquisition device, which can be a camera. The moving target refers to an object detected from the video frame image whose position changes relative to the background environment. If a moving target is detected, the position information of the moving target in the image is acquired, which may include, for example, the pixel coordinates of the target's center point in the image, and then S105 is executed; if no moving target is detected, the trigger is determined to be a false alarm, and no subsequent processing is performed.

[0049] S105. In response to the fact that the moving target is located within the predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than the preset displacement threshold, the moving target is determined to be an intrusion target.

[0050] In this application, after detecting a moving target from a video frame image, it is necessary to track and detect the moving target in order to determine whether the moving target is an intrusion target.

[0051] Specifically, this application obtains the position information of a moving target in the current frame image and performs the same target detection and position tracking processing on subsequent consecutive frames to obtain the position information of the moving target in each subsequent frame image. The number of frames in each subsequent consecutive frame image can be customized according to actual conditions, such as 5 to 15 frames. Then, the changes in the position information of the moving target in different frame images are compared to determine the inter-frame displacement of the moving target, which is the pixel distance that the moving target moves between two adjacent frame images.

[0052] When determining whether a moving target is an intrusion target, the first step is to determine whether the target's position in each frame is within a predetermined monitoring area. This predetermined monitoring area is a pre-defined boundary of the monitoring area, which can be a virtual polygonal region in the field of view of the image acquisition device, such as the boundary of a courtyard or the perimeter of a house monitored by the image capture device. Then, it is determined whether the inter-frame displacement of the moving target between adjacent frames is less than a preset displacement threshold. These consecutive predetermined frames refer to a pre-defined number of consecutive image frames, such as 5 or 10 consecutive frames, used to ensure that the moving target is continuously present within the predetermined monitoring area rather than merely passing through. When the inter-frame displacement is less than the preset displacement threshold, it indicates that the moving target is hovering or flying slowly within the monitoring area; when the inter-frame displacement is greater than or equal to the preset displacement threshold, it indicates that the moving target is rapidly crossing the monitoring area.

[0053] In this application, a moving target is determined to be in a threatening state when it is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement between adjacent frames is less than a preset displacement threshold. This determination method ensures that a moving target is only considered an intrusion target if it is continuously present within the predetermined monitoring area and is hovering or flying slowly, effectively eliminating non-threatening behaviors such as birds flying quickly overhead. For example, when a moving target is continuously and stably flying within the predetermined monitoring area, such as a drone hovering or slowly circling over a courtyard, it can be determined that the drone is located within the predetermined monitoring area in consecutive predetermined frame images, and the inter-frame displacement between adjacent frames is less than a preset displacement threshold. Therefore, its behavior pattern matches the characteristics of aerial reconnaissance or surveillance of the monitoring area, which is a typical security threat scenario, and thus it is determined to be in a threatening state. When a moving target disappears quickly from the image field of view, such as a bird flying rapidly across the monitoring area, this situation does not meet the condition that the moving target is located within the predetermined monitoring area in consecutive predetermined frames of images, nor does it meet the condition that the inter-frame displacement between adjacent frames is less than a preset displacement threshold. Therefore, its behavior pattern does not belong to reconnaissance or intrusion behavior, and is therefore determined to be in a non-threat state.

[0054] When all the above conditions are met, the moving target is identified as an intrusion target, and an alarm message is issued. This alarm message can be sent to the user's terminal device, which may include a mobile phone or tablet. The form of the alarm message may include text push notifications, image captures, short video clips, or a combination of the above forms. The alarm message may also be a voice alarm message issued directly through on-site equipment. If the moving target is not located within the predetermined monitoring area in consecutive predetermined frames of images, or if there is an inter-frame displacement greater than or equal to a preset displacement threshold in at least one frame, the trigger is determined to be a false alarm, and no alarm processing is performed.

[0055] In summary, this embodiment sets up a two-level audio judgment mechanism: volume threshold triggering and audio feature matching. Only after successful audio feature matching can the detection and judgment of moving targets based on video frame images be triggered, avoiding the high power consumption caused by continuous image detection. At the same time, this application can accurately distinguish between hovering reconnaissance drones and fast-moving interference objects such as birds by using the judgment condition that all consecutive predetermined frame images are located within a predetermined monitoring area and the inter-frame displacement is less than a preset displacement threshold. While ensuring the detection capability of non-thermal radiation change intrusion targets, it effectively reduces the false alarm rate.

[0056] In another embodiment of this application, after real-time acquisition of ambient sound signals, the method further includes: The noise reduction process utilizes the phase difference between various ambient sound signals to obtain a noise-reduced sound signal; each ambient sound signal is a sound signal collected in real time by each sound collector in the sound acquisition array; the current ambient volume is determined based on the noise-reduced sound signal.

[0057] In this application, multiple ambient sound signals can be acquired in real time by each sound acquisition unit in the sound acquisition array, and noise reduction processing can be performed using the phase difference between the multiple ambient sound signals to obtain the noise-reduced sound signal, so as to obtain the accurate current ambient volume.

[0058] Specifically, the sound acquisition array in this application includes multiple sound acquisition units. The number of sound acquisition units can be determined according to actual conditions, such as 2, 3, or 4, etc. The specific number can be determined comprehensively based on factors such as equipment size, cost, and computational complexity. In this embodiment, only two sound acquisition units are used as an example. The two sound acquisition units in this application can specifically be a first sound acquisition unit and a second sound acquisition unit. The two sound acquisition units are set on both sides of the image acquisition unit, maintaining a predetermined distance between them. Due to the physical distance between the sound acquisition units, there is a slight difference in the time it takes for the sound emitted from the same target sound source to reach each sound acquisition unit. This time difference causes a phase difference between the environmental sound signals output by each sound acquisition unit. For example, environmental noise such as wind noise or long-distance traffic noise, which is far from the sound acquisition array, has a similar phase and amplitude as the environmental noise signals output by each sound acquisition unit are far from the sound acquisition array.

[0059] Therefore, this application utilizes the phase difference between various ambient sound signals for noise reduction. For example, when two ambient sound signals are subtracted, the ambient noise cancels out because the phases of the two signals are close. However, the target sound, which is closer, is retained after subtraction due to the phase difference. After the above processing, a noise-reduced sound signal is obtained, in which the ambient noise component has been suppressed. Based on the noise-reduced sound signal, the current ambient volume can be accurately determined.

[0060] In summary, this application utilizes the phase difference between ambient sound signals collected by multiple sound acquisition devices for noise reduction processing, which can effectively suppress ambient noise, improve the audio signal-to-noise ratio, provide more accurate audio for subsequent volume judgment, and thus improve detection accuracy.

[0061] In another embodiment of this application, when the current ambient volume of the ambient sound signal exceeds a volume threshold, an audio segment is acquired, and the audio features of the audio segment are extracted, including: When the audio processor, which is in low-power standby mode, detects that the current ambient sound volume exceeds the volume threshold, it switches the audio processor from low-power standby mode to active mode, acquires audio segments, and extracts the audio features of the audio segments.

[0062] In this application, an audio processor is used to process the acquired ambient sound signal. This audio processor can be an audio DSP chip. The audio processor has two states: a low-power standby state and an active state. Correspondingly, the audio processor has a two-level wake-up mechanism: the first level is triggered by a dynamically adjustable volume threshold. In the first level, the audio processor is in a low-power standby state, where it only performs the volume detection process, consuming very little power (e.g., less than 100 microamps) to continuously monitor the current ambient volume of the ambient sound signal acquired by the sound collector. When the volume of the ambient sound signal picked up by the sound collector exceeds the currently set volume threshold, the audio processor wakes up from the ultra-low-power standby state, switches to the active state, and begins acquiring audio segments, entering the second-level feature matching and judgment process. This process requires extracting audio features from the audio segments. Compared to the low-power standby state, the active state of the audio processor consumes more power but possesses complete audio signal processing capabilities. If the current ambient volume does not exceed the volume threshold, the audio processor remains in the low-power standby state, continuing to monitor the ambient volume without switching states or performing audio acquisition or feature extraction.

[0063] In summary, if the main control ISP chip analyzes the ambient sound signal in real time, it will significantly increase power consumption and affect the device's battery life. Therefore, in this application, the collected ambient sound signal is processed by an audio processor. The audio processor is normally in a low-power standby state and only switches to an active state to collect audio and extract features when the ambient volume exceeds the volume threshold, which significantly reduces the device's power consumption during real-time volume monitoring.

[0064] In another embodiment of this application, after extracting the audio features of the audio segment, the method further includes: using the audio processor to match the audio features with the audio feature data of each dangerous event to generate a second matching result; When the second matching result indicates that the audio feature successfully matches the audio feature data of the target dangerous event, a dangerous event is determined to have occurred, and an alarm signal is issued.

[0065] In this application, the audio feature data within the audio processor can be updated via remote firmware upgrades to add audio feature data for various hazardous events, supporting the addition of new recognition scenarios. This audio feature data includes features for sounds like breaking windows and gunshots. The audio feature data for each hazardous event is obtained by pre-training and feature extraction from a large number of real hazardous event audio samples. Specifically, the audio feature data for each hazardous event can be pre-obtained through the following methods: collecting raw audio samples of various hazardous events, with multiple samples collected under different environmental conditions for each type of hazardous event, such as window-breaking sound samples and gunshot samples under different distances, directions, and background noise conditions; labeling the collected audio samples, including event type labels; extracting features from the labeled audio samples; and training and clustering the extracted features to obtain the audio feature data corresponding to each type of hazardous event.

[0066] It should be noted that this application can employ specific audio feature extraction methods for dangerous events when extracting audio features from audio segments to improve the accuracy of dangerous event detection. Specifically, the sound of breaking windows and gunshots in dangerous events are both transient impact sound signals. For this type of sound, this application preferably uses short-time energy and short-time zero-crossing rate as detection features. Short-time energy reflects the intensity change of the audio signal within each frame, effectively characterizing the suddenness and energy concentration of transient impact sounds; short-time zero-crossing rate counts the number of times the signal crosses the zero axis within each frame, effectively distinguishing transient impact sounds from stable background noise. Furthermore, this application can also employ a dedicated sound event detection model for feature extraction; this model is pre-trained using a large number of audio samples labeled with dangerous event tags such as breaking windows and gunshots, and after training, it is deployed as a detection model in the audio processor. During runtime, the current audio segment is input into the detection model, and the model directly outputs the probability value of the audio segment belonging to various dangerous events. By employing the combined features of short-time energy and short-time zero-crossing rate, or by extracting audio features based on a sound event detection model, compared to general audio feature extraction methods, this approach exhibits higher detection sensitivity and a lower false alarm rate in transient impact sound recognition tasks.

[0067] In this application, after the audio processor extracts the audio features of an audio segment, it needs to match the audio features with the audio feature data of each dangerous event to obtain a second matching degree between the audio features and the audio feature data of each dangerous event. A second matching result is generated based on the second matching degree. When the second matching result indicates that the audio features successfully match the audio feature data of the target dangerous event, a dangerous event is determined to have occurred. Specifically, this application can compare the extracted audio features with the audio feature data of each dangerous event separately, and calculate the second matching degree for each. That is, for each dangerous event, the audio processor calculates the second matching degree between the audio feature and the audio feature data of that dangerous event. The audio processor compares each second matching degree with a preset second matching degree threshold to obtain a second matching result. When the second matching degree exceeds the preset second matching degree threshold, the second matching result is a successful match; when the second matching degree does not exceed the preset second matching degree threshold, the second matching result is a failed match. This second matching degree threshold is a pre-set similarity threshold used to determine whether the current audio belongs to a certain type of dangerous event. The specific value of the second matching degree threshold can be set according to the detection sensitivity and false alarm rate requirements for various types of dangerous events.

[0068] For example, for high-security-level events, to ensure the principle of "better to err on the side of false alarms than false alarms," ​​the second matching threshold can be set to a lower value, such as 50%, so that sounds similar to the audio characteristics of dangerous events can also trigger alarms. For events of general security level, the second matching threshold can be set to a higher value, such as 80%, to reduce the false alarm rate. For each type of dangerous event, when its corresponding second matching degree exceeds the corresponding second matching threshold, it is determined that the dangerous event has occurred in the current environment.

[0069] If the second matching result indicates that the audio feature successfully matches the audio feature data of the target dangerous event, a dangerous event is determined to have occurred. The audio processor immediately issues an alarm signal. This alarm signal can be sent by the audio processor to the image processor to wake it up, and the image processor then sends the alarm information externally through the communication module. Alternatively, the audio processor can directly send the alarm information externally through the communication module. The alarm signal can include the type of dangerous event, the corresponding audio data, and possibly on-site video, etc. This application, for high-security-level dangerous events such as broken windows or gunshots, issues an alarm signal directly without image re-evaluation, ensuring timely response.

[0070] In summary, this application addresses high-security-level dangerous events such as broken window sounds and gunshots by directly matching the collected audio features with preset audio feature data for each dangerous event through an audio processor. This allows for effective identification of dangerous events through audio, and an alarm signal is immediately issued upon successful matching, ensuring timely notification and handling of dangerous events.

[0071] In another embodiment of this application, the audio features are matched with the audio feature data of the target flying object to generate a first matching result; when the first matching result is a successful match, detecting the moving target from the video frame image includes: The audio processor matches the audio features with the audio feature data of the target flying object to generate a first matching result; when the first matching result is successful, an interrupt signal is sent to the image processor to switch the image processor from sleep state to working state. The image processor acquires the current frame image and compares it with a pre-stored background image to obtain a comparison result. The pre-stored background image is an environmental image that the image processor stores in advance before entering a sleep state. In response to the comparison result indicating the presence of a new moving target, a threat status monitoring process for the moving target is performed.

[0072] In this application, the audio processor needs to calculate the first matching degree between the extracted audio features and the preloaded audio feature data of the target flying object. When the first matching degree exceeds a preset first matching degree threshold, the generated first matching result is considered a successful match, that is, the audio processor determines that the current audio contains the sound of the drone propeller. Subsequently, the audio processor sends a GPIO interrupt signal to the image processor via the interrupt signal line to notify the image processor of a trigger event. This interrupt signal is a level-transition triggered type, and its logic level state corresponds to a Boolean value, for example, a high level is 1 and a low level is 0. The trigger action is generated when the level transitions from 0 to 1 or from 1 to 0. This interrupt signal is used to switch the image processor from sleep state to working state. If the image processor does not receive the interrupt signal, it remains in sleep state. In sleep state, most of the image processor's functional modules are powered down, with only the interrupt response circuit receiving power. When a valid level transition occurs on the interrupt signal line, the interrupt response circuit of the image processor is triggered, and the image processor wakes up from sleep state and enters working state.

[0073] After being awakened, the image processor executes multi-frame image re-judgment logic. It acquires the current frame image via an image acquisition device, which can be the first frame captured after receiving the image acquisition command from the image processor. The current frame image is compared with a pre-stored background image to determine if any new moving targets exist in the scene. This pre-stored background image is an environmental image of the monitored area that the image processor caches before entering sleep mode. When determining whether a new moving target exists, the image processor can do so based on image differences; that is, it compares the pixels of the current frame image and the pre-stored background image, calculates the differences between the two frames, and obtains the comparison result. When the size and shape of the differing pixels match the shape characteristics of the target flying object, the comparison result indicates the presence of a new moving target. For example, if the size and shape of the differing pixels exhibit a four-axis symmetrical structure or a six-axis symmetrical structure, and their size and proportion match the projection characteristics of the drone from the corresponding viewpoint, then it is determined that a new moving target exists and is identified as a drone. If the size and shape of the differing pixels do not match the above-mentioned drone shape, the comparison result indicates that no new moving target exists, meaning that there is no new moving target or drone in the normal camera view.

[0074] It should be noted that in real-world scenarios, changes in lighting, shadows, and wind-blown leaves can all cause differences between the current frame image and the pre-stored background image. To improve the accuracy of the comparison results, this application employs a Gaussian Mixture Model (GMM) or a Visual Background Extractor (ViBe) algorithm to model the background of the monitored scene, distinguishing between normal background fluctuations and genuine foreground moving targets. The background model is updated using both triggered and periodic updates. Triggered updates mean that after the image processor is activated and completes a moving target detection, if the determination result indicates that there is no moving target or the moving target is in a non-threatening state, the current frame image is used as the observation data input to update the background model. Periodic updates mean that the image processor updates the background model at preset time intervals to keep pace with long-term changes in the scene.

[0075] Furthermore, to address the false alarm problem caused by sudden changes in illumination, this application also includes an illumination change compensation mechanism. For example, before performing background subtraction, the image processor first calculates the overall brightness difference between the current frame image and the background image; when this difference exceeds a preset threshold, it is determined that the scene has undergone an illumination change. At this time, global brightness compensation is performed on the current frame image to match its overall brightness with the background image before performing the background subtraction operation.

[0076] If the comparison result indicates no new moving targets, the image processor re-enters sleep mode. If the comparison result indicates a new moving target, a threat status monitoring process for the moving target is performed. This threat status monitoring process involves detecting whether the moving target meets the following criteria: the moving target is located within a predetermined monitoring area in consecutive predetermined frames, and the inter-frame displacement is less than a preset displacement threshold. If the above criteria are met, the moving target is determined to be an intrusion target. It should be noted that the dangerous event detection process described in the previous embodiment and the target flying object detection process described in this embodiment are two independent detection processes, which can be executed in parallel. The dangerous event detection process is used to identify high-security events such as broken windows and gunshots. These events are characterized by their suddenness and high degree of harm, requiring immediate response upon confirmation. The target flying object detection process is used to identify flying targets such as drones. These targets require a comprehensive judgment combining audio and image features to avoid false alarms.

[0077] When the detection process for a dangerous event and the detection process for a target flying object are triggered simultaneously—that is, when the audio feature successfully matches the audio feature data of both a dangerous event and the target flying object—the dangerous event detection process has priority in response. This priority is based on the premise that the security risk level of the dangerous event is higher than that of the target flying object. With priority in response, the dangerous event detection process is executed first, directly determining the occurrence of a dangerous event and issuing an alarm signal without waiting for the target flying object's detection process to complete. Simultaneously, the target flying object's detection process can continue or terminate, depending on the specific processing strategy in the application scenario. When only the target flying object's detection process is triggered—that is, when the audio feature only matches the target flying object's audio feature data and not any dangerous event's audio feature data—the target flying object's detection process is executed. This involves waking up the image processor through audio feature matching, which then performs moving target detection and threat status determination on the video frame to confirm whether it is an intrusion target.

[0078] In summary, this application matches audio features with the audio feature data of the target flying object, and only wakes up the image processor via an interrupt signal when a match is successful. This keeps the image processor in a dormant state when there is no triggering event, avoiding the power consumption overhead caused by continuous operation. After the image processor is woken up, it monitors the threat status of moving targets. Through a dual verification mechanism of audio and image, the reliability of detecting non-thermal radiation targets such as drones is improved. Furthermore, through the detection processes of dangerous events and target flying objects, this application can issue alarm signals with minimal delay when dangerous events occur, ensuring timely response. When only a flying target is present, the dual verification mechanism of audio and image makes an accurate judgment, effectively reducing the false alarm rate.

[0079] In another embodiment of this application, the process of monitoring the threat status of a moving target includes: Determine the position information of the moving target in the current frame image and subsequent consecutive frames image, and determine the inter-frame displacement between two consecutive frames image based on the position information; compare the position information with the position range of the predetermined monitoring area to obtain the position comparison result; compare the inter-frame displacement with the preset displacement threshold to obtain the displacement comparison result. When the position comparison result indicates that the moving target is located within the predetermined monitoring area in consecutive predetermined frame images, and the displacement comparison result indicates that the inter-frame displacement of the moving target in consecutive predetermined frame images is less than a preset displacement threshold, it is determined that the target is under threat from a drone.

[0080] In this application, when monitoring the threat status of a moving target, it is first necessary to determine the position information of the moving target in the current frame image, and then perform the same target detection and position tracking processing on subsequent consecutive frames to obtain the position information of the moving target in subsequent consecutive frames. This position information can be the pixel coordinates of the center point of the moving target in the image, or it can be the position range of the image pixel area occupied by the moving target. Based on the position information of the moving target in each frame image, the inter-frame displacement of the moving target between two adjacent frames is calculated; the inter-frame displacement can be determined by calculating the pixel distance between the center points of the moving target in two adjacent frames.

[0081] The location information is compared with the location range of a predetermined monitoring area to obtain a location comparison result. This location comparison result indicates whether the position of the moving target in each frame of the image is within the predetermined monitoring area. Specifically, for each frame of the image, it is determined whether the location information of the moving target falls within the location range of the predetermined monitoring area; if so, the location comparison result of that frame indicates that the target object is within the predetermined monitoring area; if not, the location comparison result of that frame indicates that the target object is not within the predetermined monitoring area. Then, the inter-frame displacement is compared with a preset displacement threshold to obtain a displacement comparison result. This displacement comparison result indicates whether the inter-frame displacement between adjacent frames is less than the preset displacement threshold. Specifically, for each pair of adjacent frames, it is determined whether the calculated inter-frame displacement is less than the preset displacement threshold; if so, the displacement comparison result of that inter-frame displacement indicates that the inter-frame displacement is less than the preset displacement threshold; if not, the displacement comparison result of that inter-frame displacement indicates that the inter-frame displacement is greater than or equal to the preset displacement threshold.

[0082] When the position comparison result indicates that the moving target is within the predetermined monitoring area in consecutive predetermined frames, and the displacement comparison result indicates that the inter-frame displacement of the moving target between consecutive adjacent frames is less than a preset displacement threshold, the moving target is determined to be in a drone threat state. When the position comparison result indicates that the moving target is not within the predetermined monitoring area in consecutive predetermined frames, or the displacement comparison result indicates that the inter-frame displacement of at least one frame is greater than or equal to the preset displacement threshold, the moving target is determined to be in a non-threat state. For example, when the moving target moves quickly, its inter-frame displacement will be greater than the preset displacement threshold, indicating that the target is rapidly crossing the monitoring area; when the moving target disappears from the image field of view, the target's position information cannot be detected in subsequent frames, and the position comparison result indicates that it is not within the predetermined monitoring area, indicating that the target has flown out of the monitoring range; the above situations do not constitute threatening behavior and are therefore determined to be non-threat states.

[0083] In summary, this embodiment obtains a position comparison result by comparing the position information with a predetermined monitoring area, and obtains a displacement comparison result by comparing the inter-frame displacement with a preset displacement threshold. Only when the position comparison result indicates that the moving target is located within the predetermined monitoring area in consecutive predetermined frames, and the displacement comparison result indicates that the inter-frame displacement of the moving target in consecutive predetermined frames is less than the preset displacement threshold, is the drone considered a threat. If the above conditions are not met, it indicates the presence of a fast-moving or disappearing moving target, and is therefore considered a non-threat state. This method can accurately distinguish between drone hovering reconnaissance and fast-moving or disappearing interference such as birds, effectively reducing the false alarm rate.

[0084] In another embodiment of this application, the detection method further includes: In response to the fact that the moving target is not always located within the predetermined monitoring area in consecutive predetermined frame images, or the inter-frame displacement is not less than the preset displacement threshold, the number of consecutive false alarms is updated. When the number of consecutive false alarms exceeds the threshold, the sound acquisition array and audio processor are powered off. The sound acquisition array and audio processor are powered on again at fixed intervals so that the sound acquisition array can reacquire ambient sound signals. The audio processor compares the current ambient volume of the reacquired ambient sound signal with a volume threshold to obtain a volume comparison result. When the volume comparison result indicates that the current ambient volume of the reacquired ambient sound signal exceeds the volume threshold, the sound acquisition array and audio processor are powered off.

[0085] In this application, when it is determined that there is no moving target in a threatening state, that is, when the moving target is not located in the predetermined monitoring area in consecutive predetermined frame images, or the inter-frame displacement is not less than the preset displacement threshold, the frame image re-judgment result is determined to be non-threatening, and it is recorded as a false alarm. The number of consecutive false alarms is updated, for example, by incrementing the count value of the false alarm counter by 1 to realize the update of the number of consecutive false alarms.

[0086] After updating the number of consecutive false alarms, it is necessary to determine whether the updated number of consecutive false alarms exceeds a preset threshold N. This threshold is a preset positive integer, and its specific value can be set according to the environmental noise tolerance requirements of the actual application scenario. For example, in scenarios with relatively stable environmental noise, the threshold can be set to a smaller value, such as 3 times, so that the system can quickly respond to changes in environmental noise and enter deep sleep in time to save power. In scenarios with large fluctuations in environmental noise, the threshold can be set to a larger value, such as 10 times, to avoid the system frequently entering and exiting deep sleep due to short-term noise fluctuations. The specific value of the threshold can be pre-stored as a fixed software parameter and cannot be configured by the user, but it can also be adjusted according to the actual situation.

[0087] If the number of consecutive false alarms does not exceed the threshold, the system continues to operate normally without power-off. If the number of consecutive false alarms exceeds the threshold, the system is determined to be too noisy and unsuitable for effective audio analysis. For example, the presence of continuously operating devices such as lawnmowers or soy milk makers that produce audio characteristics similar to drones will generate too many false alarms. In this case, the image processor sends a control command to the power management unit via the integrated circuit bus. Upon receiving the control command, the power management unit performs a power-off operation on the audio processor and sound acquisition array, cutting off their power supply and completely stopping their operation, thus eliminating power consumption. Simultaneously, the image processor can also push a noise exceeding alarm notification to the user via the communication module, informing them that the current ambient noise is too high and the system has temporarily disabled the audio monitoring function. After the audio processor and sound acquisition array are powered off, i.e., enter deep sleep mode, the image processor briefly wakes up at fixed intervals and actively controls the audio processor and sound acquisition array to power on again via the power management unit.

[0088] After power is restored, the sound acquisition array begins collecting ambient sound signals. The audio processor determines whether the current ambient volume exceeds a volume threshold. If it does, it indicates that the ambient noise is still too high and unsuitable for effective audio analysis. The image processor then uses the power management unit to power off both the audio processor and the sound acquisition array. If the volume does not exceed the threshold, it means the environment has returned to quiet. The image processor keeps the audio processor in normal operating mode and controls the sound acquisition array to enter low-power standby mode, resuming normal monitoring functions.

[0089] It should be noted that the process by which the audio processor determines whether the current ambient sound volume exceeds a volume threshold can be performed multiple times. That is, it repeatedly checks whether the environment has returned to quiet. If the environment has returned to quiet after multiple confirmations, the normal function of the sound acquisition array and audio processor is restored. The entire power-on and initialization process has low latency and does not affect the response to subsequent valid events. Through this method, the audio processor and sound acquisition array can be actively and completely powered off when ambient noise is excessive, thus eliminating power consumption. Simultaneously, by periodically powering on and detecting whether the ambient volume has returned to normal, the system automatically recovers after the environment becomes quiet, balancing energy saving and reliability.

[0090] As can be seen from the above embodiments, this application proposes a low-power intrusion detection scheme for image capture devices based on audio feature recognition and dynamic power management. This scheme introduces a dedicated audio DSP chip to process the signals collected by the microphone array, achieving intelligent wake-up based on sound features with extremely low power consumption. Compared with the traditional method where the main control processes the audio throughout, this scheme puts the high-power main control ISP into a deep sleep state, only triggering its operation when the DSP confirms the presence of specific audio events such as drone propeller sounds, significantly reducing the overall system power consumption. In addition, by remotely updating the audio feature database, it can be expanded to support the recognition of various dangerous scenarios such as broken windows and gunshots, breaking through the limitation of PIR / radar only being able to sense physical displacement, enhancing the product's functional diversity and market competitiveness. This solution further integrates dual-microphone phase difference noise reduction to effectively suppress environmental common-mode noise and improve the signal-to-noise ratio of the audio front end; it enables dynamic adjustment of the volume threshold through user-adjustable sensitivity levels, improving product usability and scene adaptability; through multi-frame image hovering and re-judgment, it can accurately distinguish between drones and birds, improving detection accuracy; through a deep sleep strategy that completely shuts down power when noise exceeds the limit, it achieves extreme energy saving and avoids unnecessary power consumption; through a mechanism of periodically waking up and re-powering the DSP, it ensures that the system can automatically resume normal operation after the environment recovers, balancing reliability and energy efficiency. Through the above process, this application forms a complete low-power, highly robust intrusion detection closed loop, achieving effective monitoring of non-heat source intrusion targets while meeting low-power constraints.

[0091] The various embodiments described in this specification are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” used herein may also mean the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a specific order described or illustrated, unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0092] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0093] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for detecting an intrusion target, characterized in that, include: Real-time acquisition of ambient sound signals; When the current ambient volume of the ambient sound signal exceeds the volume threshold, an audio segment is acquired, and the audio features of the audio segment are extracted. The audio features are matched with the audio feature data of the target flying object to generate a first matching result; When the first matching result is a successful match, a moving target is detected from the video frame image; In response to the fact that the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold, the moving target is determined to be an intrusion target.

2. The detection method according to claim 1, characterized in that, After real-time acquisition of ambient sound signals, it also includes: The noise reduction process is performed by utilizing the phase difference between various ambient sound signals to obtain a noise-reduced sound signal; wherein, each ambient sound signal is a sound signal acquired in real time by each sound acquisition unit in the sound acquisition array; The ambient volume is determined based on the noise-reduced sound signal.

3. The detection method according to claim 1, characterized in that, When the current ambient volume of the ambient sound signal exceeds a volume threshold, an audio segment is acquired, and the audio features of the audio segment are extracted, including: When the audio processor, which is in low-power standby mode, detects that the current ambient volume of the ambient sound signal exceeds the volume threshold, the audio processor is switched from low-power standby mode to active mode. The audio processor then collects audio segments and extracts the audio features of the audio segments.

4. The detection method according to claim 3, characterized in that, After extracting the audio features of the audio segment, the process also includes: The audio processor matches the audio features with the audio feature data of each dangerous event to generate a second matching result. When the second matching result indicates that the audio feature successfully matches the audio feature data of the target dangerous event, a dangerous event is determined to have occurred, and an alarm signal is issued.

5. The detection method according to claim 3, characterized in that, The audio features are matched with the audio feature data of the target flying object to generate a first matching result; when the first matching result is successful, detecting the moving target from the video frame image includes: The audio processor matches the audio features with the audio feature data of the target flying object to generate a first matching result; When the first matching result is a successful match, an interrupt signal is sent to the image processor to switch the image processor from sleep state to working state; The image processor acquires the current frame image and compares it with a pre-stored background image to obtain a comparison result; the pre-stored background image is an environmental image that the image processor stores in advance before entering a sleep state. In response to the comparison result indicating the presence of a new moving target, a threat status monitoring process for the moving target is performed.

6. The detection method according to claim 5, characterized in that, The process of monitoring the threat status of the moving target includes: Determine the position information of the moving target in the current frame image and subsequent consecutive frame images, and determine the inter-frame displacement of two consecutive frame images based on the position information; The location information is compared with the location range of the predetermined monitoring area to obtain a location comparison result; the inter-frame displacement is compared with a preset displacement threshold to obtain a displacement comparison result. When the position comparison result indicates that the moving target is located within a predetermined monitoring area in consecutive predetermined frame images, and the displacement comparison result indicates that the inter-frame displacement of the moving target in consecutive predetermined frame images is less than a preset displacement threshold, it is determined that the target is under UAV threat.

7. The detection method according to any one of claims 1 to 6, characterized in that, The detection method further includes: In response to the fact that the moving target is not always located within the predetermined monitoring area in consecutive predetermined frame images, or that the inter-frame displacement is not less than a preset displacement threshold, the number of consecutive false alarms is updated. When the number of consecutive false alarms exceeds the threshold, the sound acquisition array and audio processor are powered off. The sound acquisition array and the audio processor are powered on again at fixed intervals so that the sound acquisition array can reacquire the ambient sound signal. The audio processor compares the current ambient volume of the reacquired ambient sound signal with a volume threshold to obtain a volume comparison result. When the volume comparison result indicates that the current ambient volume of the re-acquired ambient sound signal exceeds the volume threshold, the sound acquisition array and the audio processor are powered off.

8. An image capture device, characterized in that, include: A sound acquisition array is used to acquire ambient sound signals in real time. An audio processor is configured to acquire an audio segment and extract audio features of the audio segment when the current ambient volume of the ambient sound signal exceeds a volume threshold; and to match the audio features with the audio feature data of the target flying object to generate a first matching result. An image processor is configured to detect a moving target from a video frame image when the first matching result is a successful match; and to determine the moving target as an intrusion target when the moving target is located within a predetermined monitoring area in consecutive predetermined frame images and the inter-frame displacement is less than a preset displacement threshold.

9. The image capture device according to claim 8, characterized in that, The sound acquisition array includes a first sound acquisition unit and a second sound acquisition unit; the first sound acquisition unit and the second sound acquisition unit are respectively disposed on both sides of the image acquisition unit of the image capture device; the first sound acquisition unit is used to acquire a first ambient sound signal, and the second sound acquisition unit is used to acquire a second ambient sound signal; the image acquisition unit is used to acquire video frame images; The audio processor is specifically used to perform noise reduction processing based on the phase difference between the first ambient sound signal and the second ambient sound signal to obtain a noise-reduced sound signal, and to determine the current ambient volume based on the noise-reduced sound signal.

10. The image capture device according to claim 8, characterized in that, The audio processor and the image processor are connected via an integrated circuit bus, which is used to transmit interactive data. The audio processor and the image processor are also connected via an interrupt signal line, which is used to transmit an interrupt signal sent by the audio processor to switch the image processor from a sleep state to a working state.