An intelligent target detection and real annotation system based on acoustic-optical imaging
By integrating a microphone array and a visible light camera, and combining deep learning and acoustic imaging algorithms, multimodal target detection and localization are achieved, solving the problem of detecting occluded or hidden targets, and providing realistic human-computer interaction and high-precision detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 陈明鸣
- Filing Date
- 2025-12-04
- Publication Date
- 2026-06-26
AI Technical Summary
Existing target detection technologies based on video signals are not accurate enough when targets are occluded or hidden, and lack realistic human-computer interaction.
It integrates a microphone array with a visible light camera to collect sound wave and image data in real time. It combines deep learning and acoustic imaging algorithms to perform multimodal target detection and localization, and generates realistic annotations through projection display.
It improves the detection accuracy of occluded or hidden targets and realizes realistic human-computer interaction, enhancing the accuracy and real-time response capability of target detection.
Smart Images

Figure CN122289639A_ABST
Abstract
Description
Technical Field
[0001] This application relates to intelligent target detection and annotation systems, and more particularly to an intelligent target detection and real-world annotation system based on acoustic-optical imaging, belonging to the field of acoustic-optical image target detection technology. Background Technology
[0002] The core task of object detection is to identify all objects of interest in an image or video, and determine their position and size. It is one of the most challenging problems in the field of machine vision. Existing video-based object detection techniques mostly employ a single modality, meaning they only utilize image data from the video for object detection. When the target in the video is severely occluded or hidden, the object detection task will fail. When clear acoustic data about the target exists, image-based single-modality object detection techniques cannot fully utilize this acoustic data for detection. Furthermore, existing object detection systems generally only annotate the detection results on a host computer, lacking real-world human-computer interaction. Summary of the Invention
[0003] In view of this, the present invention discloses an intelligent target detection and reality annotation system based on acousto-optic imaging, which solves the problems mentioned in the background art.
[0004] This invention discloses an intelligent target detection and reality annotation system based on acousto-optic imaging, comprising: The multimodal data acquisition module integrates a microphone array and a visible light camera to acquire visible light image data and synchronous acoustic wave data of the target area in real time, forming processed data.
[0005] The data processing circuit module preprocesses the acoustic wave data and visible light image data from each microphone, and then transmits the processed acoustic wave data and image data to the intelligent target detection and localization module for multimodal intelligent target detection and localization.
[0006] The intelligent target detection and localization module detects targets of interest in visible light images based on deep learning algorithms, and combines acoustic imaging algorithms to locate the sound source and image the sound field of the target using synchronous microphone sound wave data. The acoustic imaging results are fused with the target detection results of visible light images to achieve multimodal intelligent detection and localization of targets, which greatly improves the detection accuracy of severely occluded and hidden targets.
[0007] The reality annotation module, also known as the projection display service module, uses a projector or matrix lighting to generate target projection content, which is used to annotate the detection results of the intelligent target detection and positioning module according to user needs.
[0008] Furthermore, the multimodal data acquisition module includes a microphone array and a camera. The microphone array is used to acquire sound wave data from each microphone, and the camera is used to acquire visible light image data.
[0009] Furthermore, the data processing circuit module primarily preprocesses the acoustic wave data from each microphone and the visible light image data. For example, it performs noise reduction, filtering, and amplification on the acoustic wave data from each microphone. Then, through a spatiotemporal synchronization error control matrix, it achieves precise alignment between the acoustic wave data and the visible light image data, forming preprocessed data to improve data quality and accuracy. The preprocessed acoustic wave data and image data are then transmitted to the intelligent target detection and localization module for multimodal intelligent target detection and localization.
[0010] Furthermore, the intelligent target detection and localization module mainly includes a deep learning-based image target detection algorithm module, which learns and extracts features of targets of interest in visible light images to detect specific targets; simultaneously, a matching acoustic imaging algorithm module, which uses acoustic imaging algorithms to locate the sound source and image the sound field of the target of interest based on sound wave data collected by a microphone array; finally, the acoustic imaging results are fused with the visible light image target detection results to achieve multimodal intelligent detection and localization of targets, significantly improving target detection accuracy and reducing the false detection rate, especially for targets that are severely occluded or hidden.
[0011] Furthermore, the reality annotation module, also known as the projection display service module, uses projectors or matrix lights to provide projection display services. It mainly generates target projection content based on user needs using the target detection and positioning results from the intelligent target detection and positioning module, and performs real-world human-computer interactive annotation on the observed actual space. Beneficial effects
[0012] Compared with existing image-based single-modal target detection technologies, this invention adds microphone array acoustic wave data synchronized with video image data. By using acoustic imaging algorithms based on the microphone array data, the sound source of the target of interest is located and the sound field image is output, realizing multimodal detection and localization of targets of interest in real-world scenes. This improves the accuracy of single-modal target detection, especially for severely occluded and hidden targets. It can not only significantly improve the target detection accuracy but also achieve precise target localization.
[0013] Compared to existing target detection technologies, this invention adds a reality annotation module. Existing target detection technologies typically only perform annotation on a host computer after target detection, lacking actual human-computer interaction. This invention adds a reality annotation module, utilizing a projector or matrix lighting for projection display services. The target detection and localization results based on acoustic-optical imaging are used to generate projected target content according to user needs, providing reality annotation of the captured actual space and achieving real-world human-computer interaction. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the external structure of the intelligent target detection and reality annotation system based on acousto-optic imaging in this invention; Wherein: 1—microphone array, 2—camera, 3—projection display device, 4—system back-end panel circuit interface; 5—system internal processing and control circuit; Figure 2 This is a schematic diagram of the hardware circuit module structure of the intelligent target detection and reality annotation system based on acousto-optic imaging in this invention; Among them: 1—microphone array, 2—camera, 3—projection display device, 4—data processing circuit, 5—main control chip; Figure 3 This is a flowchart of the intelligent target detection and reality annotation method based on acousto-optic imaging in this invention; Figure 4 This is a flowchart illustrating the application steps of the intelligent target detection and reality annotation system based on acoustic-optical imaging in intelligent monitoring scenarios. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016]
Example 1
[0017] Specifically, microphone array 1 acquires sound wave data from each microphone, and camera 2 acquires visible light image data. The microphone sound wave data and visible light image data are transmitted to the internal processing and control circuit 5 for processing. Further, specifically, the internal processing and control circuit 5 includes a data processing circuit and a main control chip. The data processing circuit preprocesses the microphone sound wave data and visible light image data and outputs it to the main control chip. The main control chip deploys a deep learning algorithm for target detection based on the visible light image, and simultaneously deploys an acoustic imaging algorithm for sound source localization and sound field imaging based on the sound wave data. The acoustic imaging results are fused with the visible light image target detection results to achieve multimodal intelligent detection and localization of the target of interest. Then, the results of target detection and localization based on acoustic imaging and visible light image are output to the projection display device 3 to generate target projection content and provide real-world annotation of the captured space. The system's back-end panel circuit interface 4 mainly provides power to the system and enables signal input / output.
[0018]
Example 2
[0019]
Example 3
[0020]
Example 4
[0021] In addition to its application in the aforementioned intelligent monitoring scenarios, another possible application of this application is intelligent monitoring of factory production lines. The system of this invention performs real-time audio and video monitoring of factory production lines. When a problem is detected in a product being produced on the production line, it projects the issue directly onto the problematic product using a projection device, providing real-time annotation and information prompts to promptly alert production personnel to take action.
[0022] Another possible application scenario for this application is intelligent monitoring of traffic accident safety. By installing this system in sections of road where traffic accidents frequently occur, the system can detect traffic accidents by collecting audio and video signals from the traffic scene, and use projection equipment to mark the accident area with matrix lights in the form of bright lights, flashing lights, etc., to promptly remind pedestrians and vehicles in the surrounding area to ensure their safety and prevent the further escalation of traffic accidents.
[0023] Another possible application scenario for this application is intelligent safety monitoring of indoor swimming pools. When the system of this invention detects a dangerous situation such as drowning among swimmers, it uses the system's projection equipment to directly project onto the drowning person, providing real-time annotation and information prompts to promptly remind rescue personnel to carry out emergency rescue.
[0024] It should be further noted that in the above-mentioned application scenarios, external audio playback devices can also be connected through the back-end panel circuit interface of this system to simultaneously provide real-time annotation and audio-visual multimodal information prompts with the projection display service module.
[0025] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.
Claims
1. An intelligent target detection and reality annotation system based on acousto-optic imaging, characterized in that, It includes a multimodal data acquisition module, a data processing circuit module, an intelligent target detection and localization module, and a reality annotation module.
2. The intelligent target detection and reality annotation system based on acousto-optic imaging according to claim 1, characterized in that, The multimodal data acquisition module includes a microphone array and a camera. The microphone array is used to acquire sound wave data from each microphone, and the camera is used to acquire visible light image data.
3. The intelligent target detection and reality annotation system based on acousto-optic imaging according to claim 1, characterized in that, The data processing circuit module preprocesses the acoustic data from each microphone and the visible light image data, including but not limited to noise reduction, filtering, and amplification of the acoustic data from each microphone, and precise alignment of the acoustic data and the visible light image data, forming preprocessed data and outputting it to the intelligent target detection and positioning module.
4. The intelligent target detection and reality annotation system based on acousto-optic imaging according to claim 1, characterized in that, The intelligent target detection and localization module mainly includes a deep learning-based image target detection algorithm module, which learns and extracts features of targets of interest in visible light images to detect specific targets; a matching acoustic imaging algorithm module, which uses acoustic imaging algorithms to locate the sound source and image the sound field of the target of interest based on sound wave data collected by a microphone array; and finally, the acoustic imaging results are fused with the visible light image target detection results to achieve multimodal intelligent detection and localization of the target.
5. The intelligent target detection and reality annotation system based on acousto-optic imaging according to claim 1, characterized in that, The real-world annotation module, also known as the projection display service module, uses a projector or matrix lighting to generate target projection content based on user needs from the target detection and positioning results, and performs real-world human-computer interactive annotation on the observed actual space.
6. A method for intelligent target detection and reality annotation based on acousto-optic imaging, used in the intelligent target detection and reality annotation system based on acousto-optic imaging as described in claims 1-5, characterized in that, include, Sound wave data from each microphone is acquired using a microphone array. Visible light image data is acquired through a camera; The acoustic wave data and visible light image data of each microphone are transmitted to the data processing circuit, which preprocesses the acoustic wave data and visible light image data of the microphone array. The preprocessed microphone acoustic data and visible light image data are transmitted to the main control chip. Deep learning algorithms are deployed on the main control chip to perform target detection based on visible light image data; An acoustic imaging algorithm is deployed on the main control chip to obtain acoustic imaging results based on sound wave data, including the direction of the sound source and the sound field image. The visible light image target detection results and acoustic imaging results are fused on the main control chip to provide the detection and localization results of the target of interest; The fused detection and positioning results are output to the projection display service device to generate target projection content and to mark the actual space captured.
7. An application of an intelligent target detection and reality annotation system based on acousto-optic imaging in intelligent surveillance, characterized in that, The method includes using the multimodal data acquisition module of claim 2 to acquire monitoring audio and video signals in real time, using the data processing circuit module of claim 3 to preprocess the acquired multimodal sound wave data and image data, using the intelligent target detection and positioning module of claim 4 to realize the detection and positioning of the target of interest, and using the real-time annotation module of claim 5 to directly annotate the detected target and provide audio-visual multimodal information prompts on the monitoring screen using a projection device.