Multimodal Image Processing via Sensor Time Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for processing multimodal images are inefficient and costly due to the need for manual recording of object identifications and specific acquisition scenarios, limiting their application in biometrics and security control.

Innovation Solution

A method and apparatus that utilize multiple vision sensors in a preset identity recognition scenario to automatically acquire and process multimodal images, where each sensor performs image acquisition based on a preset strategy, and object identification information is determined using acquisition time information, reducing manual intervention and improving efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual recording of object identification is used in specific acquisition scenarios, then object identification can be obtained, but processing efficiency is low and acquisition cost is high

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtime consumption
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs automatic object identification and recording without manual intervention. The processor automatically compares images from multiple sensors, determines object identities, and records the results, eliminating the need for manual recording operations and significantly improving processing efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-establishes correspondence between different sensor types and objects through automatic comparison. By performing image matching and object association in advance based on acquisition time and spatial position, the system prepares identification results before they are needed, reducing processing time

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple vision sensors are deployed for multimodal image acquisition, then image acquisition capability is enhanced, but system complexity and processing cost increase

Engineering Contradiction:
Improveimage acquisition capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The processor performs multiple functions using a unified approach: it handles image acquisition from different sensor types (infrared, visible light, depth sensors), performs automatic comparison and matching, determines object identities, and records results. This multi-functional design enhances acquisition capability while avoiding the need for separate complex processing systems for each sensor type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges images from multiple different sensor types into a unified identification result. By combining infrared images, visible light images, and depth sensor data through automatic comparison based on acquisition time and spatial position, the system achieves comprehensive object identification while simplifying the overall processing architecture

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If automatic image acquisition and identification is implemented, then processing efficiency is improved, but acquisition time coordination complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtime coordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses acquisition time information as feedback to coordinate multiple sensors. By continuously monitoring and comparing acquisition timestamps from different sensors, the system automatically determines which images correspond to the same object and adjusts the identification process accordingly, simplifying time coordination through automated feedback-based matching

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3944135B1Method for processing multimodal images, apparatus, device and storage medium
Publication Date: 2023.09.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP3944135B1 patent drawingFigure 1
  • EP3944135B1 patent drawingFigure 2
  • EP3944135B1 patent drawingFigure 3

AI summary

The present application discloses a method for processing multimodal images, an apparatus, a device and a storage medium, and relates to the technical fields of computer vision and deep learning in artificial intelligence. A specific implementation is that: multiple types of vision sensors are disposed in a first preset identity recognition scenario, and the method includes: if it is determined that a first vision sensor detects a biometric part of a target object, controlling each vision sensor to separately perform image acquisition for the biometric part in accordance with a preset acquisition strategy to obtain a target visual image of a corresponding type and acquisition time information of the target visual image; performing identity recognition for the target object according to a first target visual image to determine object identification information corresponding to the first target visual image; determining object identification information corresponding to a target visual image of other type other than the first target visual image according to the acquisition time information of each target visual image and the object identification information corresponding to the first target visual image.