Visual task device based on inertia-dynamic vision-single photon sensor fusion

Through the fusion of inertial-dynamic vision-single-photon sensors, the problem of insufficient accuracy of intelligent vision equipment in complex scenes is solved, and efficient scene perception and task execution are achieved, which is suitable for visual tasks in a variety of complex scenes.

CN120651224APending Publication Date: 2025-09-16INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717956.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing intelligent vision devices perform poorly in complex scenarios, especially under high-speed and low-light conditions, making it difficult to achieve high-precision target recognition and scene perception.

Method used

Combining inertial sensors, dynamic vision sensors and single-photon sensors, information is complemented through information processors. Inertial sensors are used to obtain motion posture information, dynamic vision sensors to obtain dynamic information, and single-photon sensors to obtain static detail information. Dynamic and grayscale imaging is performed through binocular lenses to achieve data fusion.

Benefits of technology

It improves the accuracy of visual tasks in complex scenarios and is suitable for tasks such as target recognition, drone navigation, and SLAM, achieving comprehensive scene perception and processing sensor signals in real time without the need for external computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120651224A_ABST
    Figure CN120651224A_ABST
Patent Text Reader

Abstract

The invention provides a visual task device based on inertia-dynamic vision-single photon sensor fusion, relates to the technical field of visual information processing, and aims to solve the technical problem that existing intelligent visual equipment is poor in performance in a complex scene. The device comprises an inertial sensor used for determining motion attitude information of a target according to speed information; the dynamic vision sensor is used for acquiring dynamic information of a target in a specified area scene; the single-photon sensor is used for acquiring static detail information of a target in a specified area scene; and the information processor is used for calculating and processing the moving posture information, the dynamic information and the static detail information and outputting a calculation result. According to the device, the inertia-dynamic vision-single photon sensor is combined, information complementation is achieved through the processor, then comprehensive scene perception under the complex condition is achieved, the accuracy of the device for conducting the intelligent vision task is improved, and the device can be suitable for the vision task under the complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual information processing technology, and more specifically, to a visual task device based on inertial-dynamic vision-single-photon sensor fusion. Background Art

[0002] Faced with various visual tasks in complex scenes, the device needs to accurately and comprehensively perceive the surrounding scene, and at the same time obtain the real-time status of the device itself in the scene. This requires a variety of different sensors suitable for the scene to obtain different information and realize the coordinated processing of multiple visual or non-visual information.

[0003] Dynamic Vision Sensors (DVS) can capture scene brightness changes in the form of event pulses, thereby accurately perceiving the contours of objects or scenes moving relative to the device. They have the advantages of high temporal resolution and sensitive perception. However, due to the imaging characteristics of the sensor, it is relatively difficult to obtain detailed grayscale information of objects or scenes, making it unsuitable for tasks such as high-precision target recognition.

[0004] Single Photon Avalanche Diodes (SPAD) sensors have the ability to detect single photons and have a high signal-to-noise ratio in darker scenes. Their temporal resolution is also significantly better than that of traditional complementary metal oxide semiconductor (CMOS) cameras. This sensor can obtain detailed information about objects or scenes in complex scenes such as low light and high speed.

[0005] Currently, most vision systems involving inertial sensors (IMUs), such as SLAM (Simultaneous Localization and Mapping), use a single CMOS camera. This causes the device to suffer from motion blur in high-speed scenarios and poor image quality in low-light scenes, making it difficult to work in complex scenarios. Summary of the Invention

[0006] In view of this, the present invention provides a visual task device based on inertial-dynamic vision-single-photon sensor fusion, which aims to solve the technical problem that existing intelligent vision equipment performs poorly in complex scenarios.

[0007] One aspect of the present invention provides a visual task device based on inertial-dynamic vision-single-photon sensor fusion, comprising: an inertial sensor for determining the motion posture information of a target based on velocity information; a dynamic vision sensor for acquiring dynamic information of the target in a specified area scene; a single-photon sensor for acquiring static detail information of the target in a specified area scene; and an information processor for calculating and processing the motion posture information, dynamic information, and static detail information and outputting the calculation results.

[0008] According to an embodiment of the present invention, the device also includes: a binocular lens, wherein the binocular lens includes: a first lens, arranged on the photosensitive surface of the dynamic vision sensor, for focusing and imaging the dynamic vision sensor; a second lens, arranged on the photosensitive surface of the single photon sensor, for focusing and imaging the single photon sensor.

[0009] According to an embodiment of the present invention, the information processor includes: a processor array, including multiple computing units, wherein each computing unit is deployed with multiple neural network models, and the neural network models are used to calculate and process motion posture information, dynamic information and static detail information; a memory, used to store the calculation data of multiple neural network models.

[0010] According to an embodiment of the present invention, the inertial sensor, the dynamic vision sensor and the single photon sensor are arranged on the same printed circuit board; wherein the printed circuit board is provided with an interface for introducing data from the inertial sensor, the dynamic vision sensor and the single photon sensor.

[0011] According to an embodiment of the present invention, the inertial sensor is connected to the printed circuit board by means of bolts.

[0012] According to an embodiment of the present invention, the dynamic vision sensor and the single photon sensor are connected to the printed circuit board by welding.

[0013] According to an embodiment of the present invention, the information processor is connected to the inertial sensor, the dynamic vision sensor, and the single photon sensor via a bus protocol.

[0014] According to an embodiment of the present invention, the processor array is configured as a field programmable logic gate array.

[0015] According to an embodiment of the present invention, the memory is configured as a double data rate synchronous dynamic random access memory.

[0016] According to an embodiment of the present invention, the device is provided with a high-speed serial computer expansion bus standard interface for connecting to an external host computer; wherein the device is configured to be able to perform function configuration and signal reading through the host computer.

[0017] Compared with the prior art, the visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the present invention has at least the following beneficial effects:

[0018] (1) The visual task device based on inertial-dynamic vision-single photon sensor fusion provided by the embodiment of the present invention combines inertial sensors, dynamic vision sensors and single photon sensors, wherein the inertial sensor obtains the motion posture information of the target, the dynamic vision sensor obtains the dynamic information of the surrounding scene objects, and the single photon sensor obtains the detailed grayscale information of the surrounding scene objects. Then, the information is complemented by the processor, thereby achieving comprehensive scene perception in complex situations, improving the accuracy of the device in performing intelligent visual tasks, and can be applied to visual tasks in various complex scenarios such as target recognition and positioning, automatic navigation of drones, and simultaneous localization and mapping (SLAM).

[0019] (2) The visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the embodiment of the present invention can simultaneously perform dynamic imaging and grayscale imaging of the scene through binocular lenses, and can better perform sensor data fusion to achieve more comprehensive scene perception.

[0020] (3) The visual task device based on inertial-dynamic vision-single photon sensor fusion provided by the embodiment of the present invention integrates the memory and the processor, so that the device can process sensor signals in real time at the edge without the need for external devices such as PCs (Personal Computers), making the use of the device more flexible.

[0021] (4) In the visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the embodiment of the present invention, each device is connected through the AXI (Advanced eXtensible Interface) bus protocol, which can provide efficient data transmission and processing and meet the device's requirements for bus bandwidth and processing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0023] Figure 1 The structural block diagram of a visual task device based on inertial-dynamic vision-single-photon sensor fusion according to an embodiment of the present invention is schematically shown.

[0024] Figure 2 The diagram schematically shows the internal data flow interaction of a visual task device based on inertial-dynamic vision-single-photon sensor fusion according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concept of the present invention.

[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0028] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0029] In the embodiments of the present invention, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken to prevent unauthorized access to user personal information data and maintain the security of user personal information and network security.

[0030] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0031] Figure 1 The structural block diagram of a visual task device based on inertial-dynamic vision-single-photon sensor fusion according to an embodiment of the present invention is schematically shown.

[0032] like Figure 1As shown, the visual task device based on inertial-dynamic vision-single photon sensor fusion of this embodiment may include, for example: an inertial sensor (IMU), a dynamic vision sensor (DVS), a single photon sensor (SPAD) and an information processor.

[0033] Among them, the inertial sensor is used to determine the target's motion posture information based on the speed information.

[0034] In this embodiment, for example, an inertial sensor may be used to first obtain the acceleration and angular velocity information of the target, then the velocity information is obtained by integrating the acceleration, and then the real-time path and posture of the target are obtained in combination with the angular velocity information.

[0035] Dynamic vision sensors are used to obtain dynamic information of targets in a specified area scene.

[0036] In this embodiment, for example, a dynamic vision sensor can be used to obtain dynamic information of the target's surrounding scene, and a spiking neural network (SNN) can be used to implement intelligent vision tasks.

[0037] Single-photon sensors are used to obtain static detail information of the target in a specified area scene.

[0038] In this embodiment, for example, a single-photon sensor can be used to obtain detailed grayscale information of the target's surrounding scene, and a convolutional neural network (CNN) can be used to implement intelligent vision tasks.

[0039] The information processor is used to calculate and process motion posture information, dynamic information and static detail information, and output calculation results.

[0040] In this embodiment, for example, an information processor may be used to process the data of the three sensors to obtain a processing result.

[0041] The visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by an embodiment of the present invention combines inertial sensors, dynamic vision sensors and single-photon sensors. Among them, the inertial sensor obtains the motion posture information of the target, the dynamic vision sensor obtains the dynamic information of the surrounding scene objects, and the single-photon sensor obtains the detailed grayscale information of the surrounding scene objects. The information is then complemented by the processor to achieve comprehensive scene perception in complex situations, thereby improving the accuracy of the device in performing intelligent visual tasks. It can be applied to visual tasks in various complex scenarios such as target recognition and positioning, drone automatic navigation, and simultaneous localization and mapping (SLAM).

[0042] According to an embodiment of the present invention, the visual task device based on inertial-dynamic vision-single-photon sensor fusion may further include: a binocular lens.

[0043] The binocular lens includes a first lens and a second lens.

[0044] The first lens is arranged on the photosensitive surface of the dynamic vision sensor and is used for focusing and imaging the dynamic vision sensor.

[0045] The second lens is arranged on the photosensitive surface of the single-photon sensor and is used for focusing and imaging the single-photon sensor.

[0046] The visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the embodiment of the present invention simultaneously performs dynamic imaging and grayscale imaging of the scene through binocular lenses, which can better fuse sensor data to achieve more comprehensive scene perception.

[0047] According to an embodiment of the present invention, the inertial sensor, dynamic vision sensor and single-photon sensor are arranged on the same printed circuit board (PCB), and the printed circuit board is provided with an interface for introducing data from the inertial sensor, dynamic vision sensor and single-photon sensor.

[0048] In this embodiment, the inertial sensor can be connected to the printed circuit board by means of bolts, or can be connected to the printed circuit board by other physical fixing methods other than bolts, such as riveting, etc. This embodiment does not limit this.

[0049] In this embodiment, the dynamic vision sensor and the single photon sensor can be connected to the printed circuit board by welding, for example.

[0050] According to an embodiment of the present invention, the information processor may include, for example: a processor array and a memory.

[0051] The processor array may include multiple computing units, each computing unit is deployed with multiple neural network models, and the neural network model is used to calculate and process motion posture information, dynamic information and static detail information.

[0052] The memory is used to store the calculation data of multiple neural network models.

[0053] In this embodiment, the memory and the processor array can be soldered on the same printed circuit board as mentioned above, or can be soldered on separate printed circuit boards, and data communication can be achieved through interface connection.

[0054] The visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the embodiment of the present invention integrates the memory and the processor, so that the device can process sensor signals in real time at the edge without the need for external devices such as PCs (Personal Computers), making the use of the device more flexible.

[0055] According to an embodiment of the present invention, the information processor is connected to the inertial sensor, the dynamic vision sensor, and the single photon sensor via a bus protocol.

[0056] In this embodiment, the information processor and the inertial sensor, the dynamic vision sensor, and the single photon sensor may be connected via an AXI bus protocol, for example.

[0057] AXI (Advanced eXtensible Interface) is an on-chip bus designed for high performance, high bandwidth, and low latency. Its address / control and data phases are separated, supporting unaligned data transfers. In burst transfers, only the first address is required. Separate read and write data channels support both burst access and out-of-order access, making timing closure easier.

[0058] The visual task device based on inertial-dynamic vision-single-photon sensor fusion provided by the embodiment of the present invention has each device connected through an AXI protocol bus, which can provide efficient data transmission and processing and meet the device's requirements for bus bandwidth and processing power.

[0059] Figure 2 The diagram schematically shows the internal data flow interaction of a visual task device based on inertial-dynamic vision-single-photon sensor fusion according to an embodiment of the present invention.

[0060] like Figure 2 As shown, in the visual task device based on inertial-dynamic vision-single-photon sensor fusion of this embodiment, the processor array is configured as a field programmable gate array (FPGA), and the memory is configured as a double data rate synchronous dynamic random access memory, specifically DDR4 memory, that is, the fourth generation double data rate synchronous dynamic random access memory, which is used to store preprocessed signals and intermediate data of the neural network.

[0061] The entire device is equipped with a high-speed serial computer expansion bus standard interface (Peripheral Component Interconnect Express, PCIe) for connecting to an external host computer. The host computer can be used to configure the device's functions and read signals, making the device easier to use.

[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0063] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A visual task device based on inertial-dynamic vision-single photon sensor fusion, characterized in that: The device comprises: Inertial sensor, used to determine the target's motion posture information based on speed information; Dynamic vision sensor, used to obtain dynamic information of the target in a specified area scene; Single-photon sensor, used to obtain static detail information of the target in a specified area scene; An information processor is used to calculate and process the motion posture information, the dynamic information and the static detail information, and output a calculation result.

2. The device according to claim 1, characterized in that The device further comprises: a binocular lens, wherein the binocular lens comprises: A first lens is provided on the photosensitive surface of the dynamic vision sensor and is used to focus and image the dynamic vision sensor; The second lens is provided on the photosensitive surface of the single-photon sensor and is used for focusing and imaging the single-photon sensor.

3. The device according to claim 1, characterized in that The information processor includes: A processor array comprising a plurality of computing units, wherein each computing unit is deployed with a plurality of neural network models, wherein the neural network model is used to perform computational processing on the motion posture information, the dynamic information, and the static detail information; Memory, used to store calculation data of multiple neural network models.

4. The device according to claim 3, characterized in that The inertial sensor, the dynamic vision sensor and the single photon sensor are arranged on the same printed circuit board; The printed circuit board is provided with an interface for introducing data from the inertial sensor, the dynamic vision sensor and the single photon sensor.

5. The device according to claim 4, characterized in that The inertial sensor is connected to the printed circuit board through bolts.

6. The device according to claim 4, characterized in that The dynamic vision sensor and the single photon sensor are connected to the printed circuit board by welding.

7. The device according to claim 1, characterized in that The information processor is connected to the inertial sensor, the dynamic vision sensor and the single photon sensor via a bus protocol.

8. The device according to claim 3, characterized in that The processor array is configured as a field programmable logic gate array.

9. The device according to claim 3, characterized in that The memory is configured as a double data rate synchronous dynamic random access memory.

10. The device according to claim 1, characterized in that The device is provided with a high-speed serial computer expansion bus standard interface for connecting to an external host computer; The device is configured to be capable of function configuration and signal reading via the host computer.