A multi-modal image data generation method and system for small target recognition

CN122780718APending Publication Date: 2026-09-18BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610895833.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

缺陷一:缺乏多视角的“真值”素材,导致合成数据空间特征失真

Benefits of technology

1.双轴联动的半球空间扫描装置结构:区别于传统的云台或单轴转台,本结构能保证传感器模组始终以等地距(球半径)对位于球心的待测目标进行全方位的半球空间覆盖。双轴正交的机械结构布局,以及双光谱传感器(可见光+红外)在末端的刚性同轴安装方式。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780718A_ABST
    Figure CN122780718A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal image data generation methods and systems for small target identification, belong to computer vision and artificial intelligence technical field;The system includes by base support component, horizontal orientation adjusting mechanism, vertical pitch adjusting mechanism and dual-spectrum acquisition terminal constitute dual-shaft linkage acquisition device, and host computer control system.The method uses the nested loop strategy of pitch fixed axis-azimuth scanning to drive A1 pitch axis and A2 azimuth axis linkage, executes "stop-shoot-store" timing control to realize the synchronous acquisition of visible light and infrared thermal imaging, and the pitch angle and azimuth angle are encoded into image filename, the mapping relationship between physical space coordinates and image is established;When data enhancement, according to the angle index in filename, call the corresponding view angle foreground material, respectively implant in visible light and infrared background map with the same space mapping relationship, generate dual-mode synthetic image.The application can safely and automatically obtain full-azimuth, dual-mode strictly aligned small target sample in hemispherical space, provide high-quality training data for deep learning model, significantly improve small target recognition precision and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and artificial intelligence technology, specifically relating to a method and system for generating multimodal image data for small target recognition. Background Technology

[0002] With the development of deep learning technology, target detection algorithms based on convolutional neural networks (CNNs) have been widely used in industrial and military fields. However, the performance of deep learning models is highly dependent on the quantity and quality of training data. In the field of mine detection and clearance, obtaining real small target image data (especially field data containing multiple models and angles) is extremely dangerous and costly, resulting in extremely scarce samples.

[0003] To address the problem of insufficient samples, data augmentation has become a standard solution in the industry.

[0004] (1) General image-level data augmentation

[0005] Early augmentation methods primarily focused on image-level transformations. For example, the paper "YOLOv4: Optimal Speed ​​and Accuracy of Object Detection" by Bochkovskiy et al. elaborated on data augmentation strategies such as Mosaic and Mixup. These methods effectively improved the robustness of the detector by mixing multiple images or cropping and scaling them. However, such methods typically operate on the entire scene, making it difficult to preserve fine-grained features for specific targets.

[0006] (2) Object-level enhancement based on "copy-paste" To more accurately expand the sample of a specific target, the academic community has proposed a method to extract the target from the source image and paste it into a new background.

[0007] The paper "Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation" published by Ghiasi et al. demonstrates that the simple "copy-paste" mechanism has a strong effect on instance segmentation and object detection tasks, and can significantly improve model performance.

[0008] For small targets, Kisantal et al.'s paper, "Augmentation for small object detection," points out that small targets occupy a small percentage of pixels in an image, making detection difficult. The paper proposes that oversampling and copying and pasting the small target into the image can effectively solve the problem of low detection accuracy for small targets.

[0009] (3) Considering the advanced enhancement of consistency and spatial relationships To address the issue of pasted images appearing "unnatural" or lacking physical consistency, existing technologies have been further explored.

[0010] The paper "Cut-Paste Consistency Learning for Semi-Supervised Lesion Segmentation" published by Yap et al. proposes a cut-paste method based on consistency learning in the field of medical image segmentation, which emphasizes the consistency of feature representation between the foreground and background.

[0011] Furthermore, in the field of autonomous driving, relevant literature such as "Object-level Data Augmentation for Visual 3D Object Detection in Autonomous Driving" has begun to attempt to introduce 3D spatial information during the enhancement process in order to avoid situations where the object's pasting position does not conform to physical laws (such as a vehicle floating in the air).

[0012] Although the aforementioned literature has demonstrated at the algorithmic level that "image synthesis" is an effective means of addressing data scarcity, when applied to multimodal small target recognition tasks, the existing data acquisition and generation processes still suffer from the following insurmountable hardware and methodological deficiencies: Lack of high-quality source material from multiple perspectives: The "copy-paste" methods mentioned in the aforementioned literature (such as the schemes of Ghiasi and Kisantal) typically assume that 2D images of the target already exist. However, in practical applications, small targets may appear in arbitrary poses. Existing technology lacks an automated hardware device to batch acquire high-resolution close-ups of small targets at arbitrary angles (pitch, azimuth) within hemispherical space. If only single-view images are pasted, the resulting training set will severely lack 3D spatial features.

[0013] Dual-modal data is difficult to align and simulate: Existing literature (such as YOLOv4 or Cut-Paste Consistency) mainly focuses on visible light images. In small target detection, infrared thermal imaging (IR) is a key method. Infrared features are related to object material, temperature, and ambient radiation, making them difficult to perfectly forge using simple software algorithms (such as GANs). Current technologies lack a mechanism that can simultaneously acquire strictly aligned visible light and infrared thermal imaging data.

[0014] Low acquisition efficiency and fragmented processes: Currently, acquiring "source material" mainly relies on manual shooting, which is extremely inefficient and yields inaccurate perspectives. There is a lack of a system-level solution that integrates "automated hardware acquisition" with "data augmentation algorithms."

[0015] Therefore, the present invention aims to provide an automated acquisition and data enhancement system based on dual-axis linkage and dual-spectral camera, which solves the problem of acquiring high-quality multimodal small target samples from a physical source.

[0016] Although the aforementioned literature (such as Ghiasi's Copy-Paste method and Kisantal's small target augmentation method) provides effective data augmentation strategies at the algorithm level, the existing data acquisition and generation processes have the following significant shortcomings in practical applications for multimodal small target recognition: Defect 1: Lack of multi-view "ground truth" material leads to distortion of spatial features in the synthesized data. Existing "copy-paste" techniques are typically based on static 2D image libraries. However, small targets, as three-dimensional objects, exhibit significantly different appearance features depending on the viewing angle (top view, eye view, oblique view). Current technology lacks a hardware device capable of automatically acquiring omnidirectional views of the target within hemispherical space. Using only single-angle material for enhancement (e.g., forcibly pasting a side view onto a top-down background) results in images that violate perspective principles, causing deep learning models to learn incorrect geometric features.

[0017] Defect 2: Inaccurate alignment of infrared and visible light dual-modal data. Small target identification heavily relies on infrared thermal imaging to help eliminate camouflage interference. Existing data augmentation methods are mostly designed for single-modal (visible light) images. Because infrared thermal imagers and visible light cameras have different fields of view (FOV), resolutions, and optical axes, current technologies lack a mechanism for simultaneous triggering of both cameras, alignment of the field of view centers, and automatic filename association at the physical acquisition end. This lack of a mechanism results in spatial mismatch in the acquired dual-modal data, making it unsuitable for training multimodal fusion detection networks.

[0018] Defect 3: The sample collection process is dangerous and uncontrollable. To obtain high-quality, realistic small target samples, traditional methods require close-range manual photography, posing significant safety risks. Furthermore, manual photography makes it difficult to ensure a uniform distribution of angles (e.g., it's difficult to precisely control the timing of each shot every 5 degrees), resulting in a long-tailed distribution in the dataset and affecting the model's recognition rate for specific angles. Summary of the Invention

[0019] In view of this, the purpose of this invention is to provide a multimodal image data generation method and system for small target recognition, which can efficiently acquire small target samples with omnidirectional visual coverage and strict alignment of visible light and infrared thermal imaging under the premise of safety and automation, thereby providing high-quality, multimodal standard "foreground" material for image synthesis (Copy-Paste) based data augmentation algorithms, and ultimately improving the generalization ability and robustness of small target recognition models.

[0020] A method for generating multimodal image data for small target recognition includes the following steps: Step S1: System initialization and parameter configuration, establish communication connection between the host computer and the motion controller and dual-spectrum acquisition terminal, and set the boundary and step size parameters of the scanning space; Step S2: Based on nested loop hemispherical space trajectory planning, a nested loop control strategy of pitch axis fixation and azimuth scanning is adopted to drive the A1 pitch axis to move to the target pitch angle and lock it, and drive the A2 azimuth axis to perform azimuth scanning at the target pitch angle by step angle. Step S3: Motion-static coordination and dual-spectral synchronous acquisition, execute the "stop-shoot-save" timing control logic, and perform a delay after the A1 axis and / or A2 axis are in place to suppress mechanical aftershocks. Then, the visible light camera and infrared thermal imager in the dual-spectral acquisition terminal are triggered in parallel to synchronously acquire dual-modal images of the current viewpoint. Step S4: Encode and store data based on spatial coordinate mapping. Read the current pitch angle of the A1 axis and azimuth angle of the A2 axis at the moment of acquiring the dual-modal image, encode the pitch angle and azimuth angle information into the file name of the two modal images, and establish a one-to-one mapping relationship between physical spatial coordinates and image data. Step S5: Based on spatially aware multimodal data enhancement, using the pitch and azimuth indexes in the image file name, retrieve the visible light foreground and infrared foreground that match the target background viewpoint from the material library, and embed the visible light foreground and infrared foreground into the visible light background image and infrared background image respectively with the same spatial mapping relationship to generate the composite images of the two modal images, thereby achieving pixel-level alignment of the visible light and infrared modal in spatial coordinates.

[0021] Preferably, the visible light camera communicates with the host computer via a GigE or USB interface, and the infrared thermal imager communicates with the host computer via the RTSP protocol.

[0022] Preferably, the boundary and step size parameters of the scanning space in step S2 include the total travel and step angle of the pitch angle of the A1 axis, and the total travel and step angle of the azimuth angle of the A2 axis. The combination of the pitch angle and the azimuth angle constitutes a discrete sampling grid covering the hemisphere.

[0023] Preferably, the parallel triggering method described in step S3 is as follows: two acquisition threads are triggered in parallel at the same time point. The main thread calls the visible light industrial camera SDK to send a soft trigger command to acquire visible light frames, and the auxiliary thread synchronously captures the current frame of the infrared RTSP stream.

[0024] Preferably, the encoding format of the image file name in step S4 is: {BatchID}.{Elevation}.{Azimuth}, where BatchID is the global task identifier, Elevation is the pitch angle, and Azimuth is the azimuth angle.

[0025] Preferably, before calling the visible light foreground and infrared foreground that match the target background viewpoint from the material library in step S5, the method further includes: using the background conditions under the controlled acquisition environment, performing pixel-level foreground extraction on the original acquired image using adaptive threshold segmentation or chroma keying technology to obtain a visible light foreground and infrared foreground material library with a transparency channel.

[0026] Preferably, in step S5, the visible light foreground and the infrared foreground are respectively embedded into the visible light background image and the infrared background image with the same spatial mapping relationship, so that the images of the two modes of visible light and infrared are aligned at the pixel level in spatial coordinates.

[0027] A multimodal image data generation system for small target recognition includes: The dual-axis linkage acquisition device includes a base support assembly (1), a horizontal orientation adjustment mechanism (2), a vertical pitch adjustment mechanism (3), and a dual-spectrum acquisition terminal (4). The base support assembly (1) includes a base platform and vertical support arms installed on both sides of the base; The horizontal orientation adjustment mechanism (2) is located at the center of the base platform and includes a rotating platform driven by a stepper motor. The rotation axis of the rotating platform is perpendicular to the ground and is used to carry the target to be measured and rotate horizontally around the vertical axis. The vertical pitch adjustment mechanism (3) spans and is disposed above the rotating platform, including an arc-shaped guide rail. The two ends of the arc-shaped guide rail are fixedly connected to the vertical support arm. The geometric center of the arc-shaped guide rail coincides with the rotation center of the rotating platform in space. The dual-spectrum acquisition terminal (4) is installed at the end of the arc-shaped guide rail and includes a visible light camera and an infrared thermal imager. The optical axes of the visible light camera and the infrared thermal imager both point to the rotation center of the rotating stage. The host computer control system is communicatively connected to the dual-axis linkage acquisition device and is used to execute the method according to any one of claims 1 to 7.

[0028] Preferably, the geometric center of the arc-shaped guide rail coincides with the rotation center of the rotating stage in space, so that the straight-line distance between the dual-spectrum acquisition terminal (4) and the target to be measured remains constant when the arc-shaped guide rail moves along it; the slider assembly of the arc-shaped guide rail is driven by a stepper motor through a synchronous belt or gear transmission mechanism, and moves in a stepping manner between 0 degrees and 180 degrees along the arc-shaped trajectory.

[0029] Preferably, the host computer control system establishes a connection with the slave motion controller via a serial communication protocol, the visible light camera communicates with the host computer control system via a GigE or USB interface, and the infrared thermal imager communicates with the host computer control system via the RTSP protocol; the host computer control system includes: The scanning control module is used to drive the vertical pitch adjustment mechanism (3) and the horizontal azimuth adjustment mechanism (2) using a nested loop control strategy of pitch fixed-axis scanning. The synchronous acquisition module is used to perform delayed image stabilization after the motion is in place, and to trigger the visible light camera and the infrared thermal imager to acquire data synchronously in parallel. The coordinate encoding module is used to read the current elevation and azimuth angles and encode them into the image file name; The data enhancement module is used to call the foreground material of the corresponding viewpoint based on the angle index in the image file name to perform multimodal image synthesis.

[0030] The present invention has the following beneficial effects: 1. Dual-axis linkage hemispherical spatial scanning device structure: Unlike traditional pan-tilt units or single-axis turntables, this structure ensures that the sensor module always provides omnidirectional hemispherical spatial coverage of the target located at the center of the sphere at an equal distance (sphere radius). This is achieved through a dual-axis orthogonal mechanical structure layout and a rigid coaxial mounting of the dual-spectrum sensors (visible light + infrared) at the end.

[0031] 2. An automated acquisition method based on "stop-shoot-save" timing control: This method solves the blurring problem caused by mechanical motion aftershocks, as well as the time synchronization and frame alignment problems between cameras with different protocols (industrial camera SDK vs. network stream RTSP).

[0032] 3. Encoding method for mapping physical spatial coordinates to image data: It abandons meaningless serial number naming and realizes "data is coordinates", providing a spatial index that does not require manual annotation for subsequent data augmentation.

[0033] 4. A multimodal data augmentation method based on real-viewpoint indexing: This method overcomes perspective errors caused by "blind pasting" in existing technologies and achieves pixel-level strict alignment of visible light and infrared thermal imaging data during the synthesis process. It includes a small-scale target sample data augmentation method encompassing the entire "acquisition-indexing-synthesis" workflow.

[0034] 5. This application solves the problem of "perspective distortion" in synthetic data, significantly improving the physical realism of training samples. Through a linked scanning structure of the A1 pitch axis and A2 azimuth axis, this application can automatically acquire realistic close-ups of small targets at any angle (e.g., pitch 0°~90°) within a hemispherical space. During subsequent data augmentation, the system can directly call the corresponding "ground truth" material for synthesis based on the perspective requirements of the background image, using the filename index. This ensures that the synthesized image fully conforms to physical laws in terms of geometric perspective, allowing the model to learn true 3D spatial features.

[0035] 6. This invention achieves strict "pixel-level" alignment between visible light and infrared thermal imaging data, filling a gap in multimodal small target samples. Through coaxial mounting of dual-spectral sensors and concurrent triggering logic at the software level (i.e., multi-threaded synchronous capture in the code), this invention ensures that both sets of sensors simultaneously acquire the scene from the same viewpoint at the instant physical motion stops. The resulting dual-modal dataset exhibits strict spatial consistency, enabling the AI ​​model to accurately learn the key features of small targets—"a specific thermal radiation distribution accompanying a camouflaged appearance"—significantly improving anti-interference capabilities.

[0036] 7. The entire process from data collection to annotation is fully automated, significantly improving sample construction efficiency while ensuring personnel safety. This application adopts a technical solution of "automated control by a host computer + spatial coordinate encoding naming." In terms of safety, operators only need to remotely set scanning parameters to complete the entire data collection process without contact with hazardous materials. Regarding efficiency and standardization, the system can automatically traverse all viewpoints at set step sizes (e.g., every 5 degrees) and automatically write the angle information into the filename. This not only improves collection efficiency several times over, but more importantly, it generates a uniformly distributed standard dataset with built-in spatial annotations, saving the enormous cost of manual screening and annotation. Attached Figure Description

[0037] Figure 1 A three-dimensional schematic diagram of the multimodal image data generation system for small target recognition provided by the present invention; Figure 2 A cross-sectional schematic diagram of the multimodal image data generation system for small target recognition provided by the present invention.

[0038] Among them, 1-base support assembly, 2-horizontal orientation adjustment mechanism (A2 axis), 3-vertical pitch adjustment mechanism (A1 axis), 4-dual spectrum acquisition terminal. Detailed Implementation

[0039] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0040] 1. Dual-axis linkage acquisition device This invention provides a dual-axis linkage acquisition device. This device employs a spherical scanning frame structure design, aiming to achieve omnidirectional, equidistant imaging of a centrally located object (such as a small target sample) within a hemispherical space. Its overall mechanical structure is mainly composed of four tightly connected parts: a base support assembly, a horizontal azimuth adjustment mechanism (A2 axis), a vertical pitch adjustment mechanism (A1 axis), and a dual-spectrum acquisition terminal (e.g., ...). Figure 1 and 2 (As shown).

[0041] Specifically, the base support assembly 1 of the device constitutes the physical foundation of the entire system. It includes a stable base platform and vertical support arms mounted on both sides of the base. The base platform is constructed of high-strength metal to ensure that the entire device does not wobble during motor movement, thereby guaranteeing imaging stability.

[0042] At the center of the base platform is a horizontal orientation adjustment mechanism 2 (i.e., axis A2 in the control system). The core of this mechanism is a rotating stage driven by a precision stepper motor. The stage's rotation axis is perpendicular to the ground, and its upper surface is used to place small target samples to be collected. Driven by the motor, this stage can rotate 360 ​​degrees horizontally around its center without any blind spots, allowing a stationary camera to capture the side features of the object at different azimuth angles.

[0043] Spanning above the rotating stage is the vertical pitch adjustment mechanism 3 (i.e., the A1 axis in the control system). This mechanism employs a specially designed arc-shaped guide rail (or gantry structure), with both ends fixedly connected to support arms on either side of the base, forming an arched structure covering the stage. The geometric center of this arc-shaped guide rail and the rotation center of the rotating stage are spatially perfectly aligned, forming a spherical structure with an isofocal plane. A dual-spectrum acquisition terminal 4 is installed at the end of the arc-shaped swing arm. The swing arm is driven by a stepper motor via a synchronous belt or gear transmission mechanism, enabling continuous or stepwise movement along the arc-shaped trajectory from 0 degrees (horizontal viewpoint) to 180 degrees. This design ensures that the straight-line distance between the sensor and the object remains constant regardless of the camera's angle of movement, avoiding focusing issues caused by changes in depth of field.

[0044] The dual-spectrum acquisition terminal 4 is equipped with two sets of sensors: one is a high-resolution visible light industrial camera used to capture the texture details of objects; the other is an infrared thermal imager used to capture the thermal radiation characteristics of objects. To ensure the spatial consistency of multimodal data, both sets of sensors undergo precise optical axis calibration, ensuring that their lenses both point towards the center of the stage, and that the centers of their fields of view overlap as much as possible. Through the combined superposition of pitch motion on the A1 axis and rotation motion on the A2 axis, this acquisition terminal can traverse any coordinate point on the hemispherical surface centered on the object, thereby completing automated data acquisition across the entire field of view.

[0045] 2. Automated data acquisition and data augmentation methods The method of this invention relies on the aforementioned dual-axis linkage acquisition device to achieve automated full-space scanning and dual-modal data enhancement of small target samples through a host computer control system. The core logic of this method lies in establishing a one-to-one mapping relationship between physical space coordinates (pitch angle, azimuth angle) and digital image data, and using this mapping relationship to achieve high-fidelity data synthesis. The specific implementation steps are as follows: Step S1: System Initialization and Parameter Configuration Before the data acquisition task begins, the control system first performs a hardware handshake and parameter initialization.

[0046] Communication establishment: The system establishes a connection with the lower-level motion controller through a serial communication protocol (such as RS232), initializes the visible light camera SDK through the GigE / USB interface, and pulls the real-time video stream from the infrared thermal imager through the RTSP protocol.

[0047] Scanning parameter settings: Users set the boundaries and step size of the scanning space according to the sampling density requirements. Specifically, this includes the total travel and step angle of the A1 axis (tilt) and the total travel and step angle of the A2 axis (azimuth), which constitute a discrete sampling grid covering a hemispherical surface.

[0048] Step S2: Hemispherical trajectory planning based on nested loops To traverse the entire hemispherical space, the system employs a nested loop control strategy of "pitch axis fixation - azimuth scan".

[0049] Outer loop (pitch control): The control system first drives the A1 axis to the first pitch angle and keeps it locked.

[0050] Inner loop (azimuth scan): With the A1 axis fixed, the A2 axis is driven to start rotating. The system enters the "acquisition cycle" every time the A2 axis rotates by one step angle.

[0051] This strategy ensures that the sensor can scan small target samples layer by layer in a manner similar to the Earth's latitude and longitude grid, avoiding any omissions in the field of view.

[0052] Step S3: Coordinated Motion and Static Acquisition with Dual-Spectrum Synchronization To eliminate motion ambiguity caused by mechanical movement and ensure real-time synchronization of dual-modal data, the system executes strict "stop-capture-save" timing control logic: Displacement command transmission: The system sends position commands (such as CJXCG...) to the lower-level computer to drive the A2 axis to the target coordinate.

[0053] Mechanical aftershock suppression: Upon reaching the target position, the program enforces a preset delay. This delay is used to allow the minor vibrations of the robotic arm to completely decay, ensuring image clarity.

[0054] Concurrent triggering mechanism: The system triggers two acquisition threads in parallel at the same time. The main thread calls the industrial camera SDK to send a soft trigger command to acquire a high-resolution visible light frame; the auxiliary thread synchronously captures the current frame of the infrared RTSP stream.

[0055] Return and Reset: After the inner loop (A2 axis) completes one scan, the system drives the A1 axis to the next level. After all levels have been scanned, the system automatically sends a reset command to return the A1 and A2 axes to zero, awaiting the next task.

[0056] Step S4: Encoding and data storage based on spatial coordinate mapping This invention achieves digital storage of spatial dimensions through coordinate mapping naming. At the instant the sensor (visible light / infrared) acquires the raw image, the system synchronously reads the current motor coordinate values ​​(including elevation and azimuth) of the motion control unit in real time. Using a preset string conversion logic, the system formats and concatenates the aforementioned three-dimensional spatial vector information with the global task identifier (BatchID) to generate a unique corresponding filename. Visible light image: {BatchID}.{Elevation}.{Azimuth} Infrared image: {BatchID}.{Elevation}.{Azimuth} This scheme enables zero-cost retrieval by directly encoding the three-dimensional spatial location attributes of the physical world into the file system metadata (filename) of the two-dimensional image file. In subsequent processing stages, there is no need to parse file header information or call external databases; the spatial distribution index of sampling points can be directly reconstructed from the filename. Multi-source data alignment: It provides an accurate and efficient indexing benchmark for heterogeneous data fusion of visible light and infrared images, spatial consistency calibration, and subsequent automated data augmentation.

[0057] Step S5: Context-Aware Multimodal Data Augmentation 1. High-fidelity automated foreground extraction: Utilizing the natural "green screen" conditions provided by a controlled acquisition environment (such as a high-absorption black velvet background), the algorithm employs adaptive threshold segmentation or chroma keying techniques to decouple the original image at the pixel level. The system automatically extracts the fine edge mask of the target, thereby obtaining a library of visible light and infrared foreground materials with alpha channels.

[0058] 2. Physically Driven Viewpoint Matching: The core of this step lies in utilizing the filename space metadata encoded in the S4 stage. When the system needs to simulate a specific task scenario (e.g., a drone's 45° overhead view), the algorithm no longer performs random image stacking. Instead, it precisely retrieves and calls the corresponding target samples from the material library by parsing the altitude and azimuth information in the filenames. This ensures that the perspective of the target in the synthesized image is completely consistent with the physical logic of the background environment, eliminating the "floating" and "proportionality imbalance" problems commonly found in traditional data augmentation.

[0059] 3. Cross-Modal Alignment: The system executes a parallel synthesis strategy: mapping the selected visible light foreground to specified coordinates (x, y) of the background image; simultaneously, embedding the corresponding infrared foreground into the infrared background image with the same spatial mapping relationship. This process ensures absolute pixel-level alignment of the visible light and infrared modes in spatial coordinates, fundamentally solving the pain point of manually labeling and aligning heterogeneous sensor data.

[0060] 4. Through this step, the system constructs a controllable synthetic image pipeline. It can not only produce high-fidelity samples that conform to the laws of physical perspective in batches, but also solve the core contradictions in the field of target detection—the scarcity of infrared feature samples and the high cost of multimodal annotation—through the "one-time extraction, two-way alignment" method.

[0061] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating multimodal image data for small target recognition, characterized in that, Includes the following steps: Step S1: System initialization and parameter configuration, establish communication connection between the host computer and the motion controller and dual-spectrum acquisition terminal, and set the boundary and step size parameters of the scanning space; Step S2: Based on nested loop hemispherical space trajectory planning, a nested loop control strategy of pitch axis fixation and azimuth scanning is adopted to drive the A1 pitch axis to move to the target pitch angle and lock it, and drive the A2 azimuth axis to perform azimuth scanning at the target pitch angle by step angle. Step S3: Motion-static coordination and dual-spectral synchronous acquisition, execute the "stop-shoot-save" timing control logic, execute the delay after the A1 axis and / or A2 axis are in place to suppress mechanical aftershocks, and then trigger the visible light camera and infrared thermal imager in the dual-spectral acquisition terminal to synchronously acquire dual-modal images of the current viewpoint. Step S4: Encode and store data based on spatial coordinate mapping. Read the current pitch angle of the A1 axis and azimuth angle of the A2 axis at the moment of acquiring the dual-modal image, encode the pitch angle and azimuth angle information into the file name of the two modal images, and establish a one-to-one mapping relationship between physical spatial coordinates and image data. Step S5: Based on spatially aware multimodal data enhancement, using the pitch and azimuth indexes in the image file name, retrieve the visible light foreground and infrared foreground that match the target background viewpoint from the material library, and embed the visible light foreground and infrared foreground into the visible light background image and infrared background image respectively with the same spatial mapping relationship to generate the composite images of the two modal images, thereby achieving pixel-level alignment of the visible light and infrared modal in spatial coordinates.

2. The method according to claim 1, characterized in that, The visible light camera communicates with the host computer via a GigE or USB interface, and the infrared thermal imager communicates with the host computer via the RTSP protocol.

3. The method according to claim 1, characterized in that, The boundary and step size parameters of the scanning space mentioned in step S2 include the total travel and step angle of the pitch angle of the A1 axis, and the total travel and step angle of the azimuth angle of the A2 axis. The combination of the pitch angle and the azimuth angle constitutes a discrete sampling grid covering the hemisphere.

4. The method according to claim 1, characterized in that, The parallel triggering method described in step S3 is as follows: two acquisition threads are triggered in parallel at the same time point. The main thread calls the visible light industrial camera SDK to send a soft trigger command to acquire visible light frames, and the auxiliary thread synchronously captures the current frame of the infrared RTSP stream.

5. The method according to claim 1, characterized in that, The image file name encoding format in step S4 is: {BatchID}.{Elevation}.{Azimuth}, where BatchID is the global task identifier, Elevation is the pitch angle, and Azimuth is the azimuth angle.

6. The method according to claim 1, characterized in that, Before calling the visible light foreground and infrared foreground that match the target background perspective from the material library in step S5, the method further includes: using the background conditions under the controlled acquisition environment, performing pixel-level foreground extraction on the original acquired image using adaptive threshold segmentation or chroma keying technology to obtain a visible light foreground and infrared foreground material library with transparency channels.

7. The method according to claim 1, characterized in that, In step S5, the visible light foreground and the infrared foreground are respectively embedded into the visible light background image and the infrared background image with the same spatial mapping relationship, so that the images of the two modes of visible light and infrared are aligned at the pixel level in spatial coordinates.

8. A multimodal image data generation system for small target recognition, characterized in that, include: The dual-axis linkage acquisition device includes a base support assembly (1), a horizontal orientation adjustment mechanism (2), a vertical pitch adjustment mechanism (3), and a dual-spectrum acquisition terminal (4). The base support assembly (1) includes a base platform and vertical support arms installed on both sides of the base; The horizontal orientation adjustment mechanism (2) is located at the center of the base platform and includes a rotating platform driven by a stepper motor. The rotation axis of the rotating platform is perpendicular to the ground and is used to carry the target to be measured and rotate horizontally around the vertical axis. The vertical pitch adjustment mechanism (3) spans and is disposed above the rotating platform, including an arc-shaped guide rail. The two ends of the arc-shaped guide rail are fixedly connected to the vertical support arm. The geometric center of the arc-shaped guide rail coincides with the rotation center of the rotating platform in space. The dual-spectrum acquisition terminal (4) is installed at the end of the arc-shaped guide rail and includes a visible light camera and an infrared thermal imager. The optical axes of the visible light camera and the infrared thermal imager both point to the rotation center of the rotating stage. The host computer control system is communicatively connected to the dual-axis linkage acquisition device and is used to execute the method according to any one of claims 1 to 7.

9. The system according to claim 8, characterized in that, The geometric center of the arc-shaped guide rail coincides with the rotation center of the rotating stage in space, so that the straight-line distance between the dual-spectrum acquisition terminal (4) and the target to be measured remains constant when the arc-shaped guide rail moves along it; the slider assembly of the arc-shaped guide rail is driven by a stepper motor through a synchronous belt or gear transmission mechanism, and moves in a stepping manner between 0 degrees and 180 degrees along the arc-shaped trajectory.

10. The system according to claim 8, characterized in that, The host computer control system establishes a connection with the slave motion controller via a serial communication protocol; the visible light camera communicates with the host computer control system via a GigE or USB interface; and the infrared thermal imager communicates with the host computer control system via the RTSP protocol. The host computer control system includes: The scanning control module is used to drive the vertical pitch adjustment mechanism (3) and the horizontal azimuth adjustment mechanism (2) using a nested loop control strategy of pitch fixed-axis scanning. The synchronous acquisition module is used to perform delayed image stabilization after the motion is in place, and to trigger the visible light camera and the infrared thermal imager to acquire data synchronously in parallel. The coordinate encoding module is used to read the current elevation and azimuth angles and encode them into the image file name; The data enhancement module is used to call the foreground material of the corresponding viewpoint based on the angle index in the image file name to perform multimodal image synthesis.