Multi-source sensing data augmentation method and system
By simulating perturbations in multi-source sensing data using a pre-trained neural network model, augmented data under harsh environments is generated, solving the security and cost issues of data acquisition in harsh environments and achieving spatiotemporal consistency of data and improved recognition performance.
Patent Information
- Application Number
- CN202511125529.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-25
Smart Images

Figure CN121010852A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of perception and computer science, and in particular to a method and system for augmenting multi-source perception data. Background Technology
[0002] In future applications such as automated search and rescue using intelligent robots, unmanned vehicles, and drones, these intelligent machines will need to be equipped with multiple sensors, such as RGB and infrared cameras, lasers, and millimeter-wave radars. These sensors will collect multi-source sensing information for tasks such as target recognition. The multi-source sensing data obtained by these sensors in normal environments (such as clear skies, no wind, and stationary conditions) differs significantly from that in harsh environments (heavy rain, strong winds, and high-speed movement). Specifically, multi-source sensing data obtained in harsh environments is affected by rain streaks, cloud cover, wind-induced vibrations, and blurring caused by rapid movement, resulting in different physical perturbations for image-based and radar point cloud-based sensing data. This severe physical perturbation can significantly degrade the performance of tasks such as target recognition.
[0003] Target recognition in harsh environments is crucial for emergency rescue and other scenarios, and acquiring multi-source sensing datasets for such environments is a critical step. However, collecting multi-source sensing data in harsh environments is dangerous, costly, and can easily lead to personal injury and property damage. Therefore, a multi-source sensing data augmentation method is of great significance for improving the performance of target recognition tasks in harsh environments at a low cost. Summary of the Invention
[0004] This application proposes a multi-source sensing data augmentation method and system to solve the problems of physical distortion and cross-modal perturbation inconsistency in existing data augmentation techniques.
[0005] In a first aspect, embodiments of this application provide a method for augmenting multi-source sensing data, comprising the steps of: Acquire first sensing data; the first sensing data is multi-source sensing data under normal conditions; The first sensing data is input into a pre-trained neural network model to generate a second sensing dataset; the second sensing dataset is a collection of multi-source sensing data with different perturbations generated by the neural network model performing spatiotemporally consistent perturbation simulation on the first sensing data.
[0006] Furthermore, the pre-training process of the neural network model includes the following steps: Acquire first-sensory data and third-sensory data; the third-sensory data is environmental scrambling data collected synchronously with the first-sensory data, and the two are spatiotemporally aligned and have the same scene. The neural network model is trained using the first perception data as input and the third perception data as a supervision signal.
[0007] In one embodiment, the neural network model is a generative model.
[0008] In one embodiment, the perturbation simulation includes: Add optical perturbations to the image data; Add electromagnetic disturbances to radar point cloud data; Maintain the spatial correspondence between image data and point cloud data.
[0009] In one embodiment, before the first perceptual data is input into the neural network model, the step further includes: Geometric transformations are performed on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
[0010] In one embodiment, the third sensing data is multi-source sensing data of the multi-source sensors being physically disturbed and / or the sensed environment and target being physically disturbed.
[0011] Secondly, embodiments of this application also provide a multi-source sensing data augmentation system for implementing the multi-source sensing data augmentation method described in any embodiment of the first aspect, comprising: a multi-source sensor module and a scrambling module. The multi-source sensor module is used to acquire first sensing data. The scrambling module is used to perform spatiotemporally consistent perturbation simulation on the first sensing data, generating a set of multi-source sensing data with different perturbations.
[0012] Furthermore, it also includes a transformation module for performing geometric transformations on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
[0013] This application also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in any of the embodiments of the first aspect.
[0014] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any of the embodiments of the first aspect.
[0015] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: This application utilizes a data augmentation method based on generative neural networks to effectively transform multi-source sensing data (such as images and radar point cloud data) collected under normal conditions into augmented data with characteristics specific to harsh environments. This method avoids the safety risks and costs associated with collecting data in real-world hazardous environments while ensuring the physical authenticity of the generated data and the spatiotemporal consistency among the multi-source data. The augmented data generated using this method can be used to train various machine learning models, significantly improving their performance in tasks such as target recognition under harsh environments, and providing reliable data support for the application of intelligent machines in complex environments. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a multi-source sensing data augmentation method according to an embodiment of this application; Figure 2 This is a flowchart illustrating the multi-source sensing data augmentation process in an embodiment of this application. Figure 3 This is a schematic diagram of the separate and fused multi-source sensor structures according to embodiments of this application; Figure 4 This is a schematic diagram of the pre-training dataset acquisition device and system according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0019] Figure 1 This is a flowchart of a multi-source sensing data augmentation method according to an embodiment of the present application, including steps 110-120.
[0020] Step 110: Acquire the first sensing data; the first sensing data is multi-source sensing data under normal conditions; Data Acquisition: Acquire multi-source sensing data under normal conditions, including image data (RGB / infrared) and / or radar point cloud data (LiDAR / millimeter-wave radar).
[0021] The data structure constructed in this application contains a complete multi-source sensing data representation system, in which a hierarchical coding system is used to identify environmental features, including: environmental type identification code, physical parameter set and / or time reference.
[0022] For example, the environment type identifier code is 0x00 for normal environment and 0x01 for harsh environment.
[0023] For example, the set of physical parameters includes quantitative indicators such as rainfall (unit: mm / h), wind speed (unit: m / s), and vibration amplitude (unit: g).
[0024] For example, the time base uses a 64-bit nanosecond-level timestamp. This identifier is embedded as metadata in the header of all sensed data, ensuring the traceability of the data acquisition environment.
[0025] Multi-source sensing data includes image data and lightning point cloud data.
[0026] The feature requirements for the multi-sensor data include sensor type identification or data storage structure.
[0027] The sensor type identifier of the image data includes RGB images marked as 0x01 or infrared images marked as 0x02.
[0028] The data storage structure of the image data includes a pixel matrix with dimensions H×W×C, where C is the number of channels, with C=3 for RGB and C=1 for infrared. The sensor type identifier for the image data has an enumerated value containing either the binary value corresponding to RGB or infrared.
[0029] The image data also includes auxiliary information: including sensor pose matrix and environmental labels. The sensor type identifier of the radar point cloud data includes lidar marked as 0x11 or millimeter-wave radar marked as 0x12.
[0030] The sensor type identifier of the radar point cloud data has an enumerated value that includes binary values corresponding to lidar or millimeter-wave radar.
[0031] The data storage structure of the radar point cloud data includes a point cloud coordinate array, the intensity value of the corresponding coordinate point, and the corresponding timestamp or sensor type identifier.
[0032] The point cloud coordinate array contains a coordinate matrix: an N×3 floating-point array (x, y, z coordinates).
[0033] For example, the point cloud coordinate array is a three-dimensional Cartesian coordinate system.
[0034] It also includes an intensity vector: an N-dimensional normalized array (range [0,1]). For example, the intensity value of the corresponding coordinate point has a normalized floating-point range of 0 to 1.
[0035] The radar point cloud data also includes calibration data, specifically calibration parameters such as radar intrinsic and extrinsic parameters.
[0036] Identifying the environment involves appending the environment status as a binary encoded or JSON string identifier to each set of image / radar point cloud data.
[0037] The so-called normal environment refers to a situation where there is no physical disturbance to the multi-source sensing devices or the sensed environment and target.
[0038] A normal environment is defined as an environment free from physical disturbances: the sensor is free from physical vibrations caused by wind, heavy rain, etc.; there are no optical disturbances caused by precipitation, fog, haze, lightning, strong light, etc.; and there is no electromagnetic interference caused by strong electromagnetic interference to sensors such as infrared and radar.
[0039] The criteria for determining physical disturbances in the normal environment include, for example, no precipitation meaning rainfall ≤ 0 mm / h; no wind meaning wind speed below a set threshold, such as ≤ 1 m / s; no strong light meaning stable illumination, such as illuminance variation ≤ 10%; physical vibration also includes no active motion ambiguity meaning the relative velocity of the sensor or the measured target is ≤ 1 m / s; and no strong electromagnetic interference including electromagnetic noise power intensity of 0 W generated by the scrambler.
[0040] Specifically, no scrambler simulates physical disturbances such as precipitation, wind, and strong light being applied to the multi-source sensors or the sensed environment and targets.
[0041] The methods described above for linking multimodal data together to form a dataset, such as achieving cross-modal data association through a spatiotemporal coupling matrix, include: Primary key identifier: UUIDv4 is used to ensure uniqueness.
[0042] Time alignment: The timestamp deviation of multi-source data is controlled within ±1ms.
[0043] Spatial mapping: Establish a coordinate transformation matrix (4×4 homogeneous matrix) between sensors.
[0044] Set the Euclidean distance error threshold (≤0.05m).
[0045] This association system ensures that multi-source data in the same scenario have strict spatiotemporal consistency.
[0046] This application achieves the acquisition, processing, and application of multi-source sensing data in three stages: In the data acquisition stage, the system injects environmental identifiers into the sensor data stream in real time, synchronously triggers data acquisition from multiple sensors, and automatically records the pose parameters of each sensor; In the data processing stage, high-precision timestamps are used to achieve strict alignment of multi-source data frames, spatial transformation matrices are calculated using pre-calibrated sensor parameters, and a cross-modal data association index table is constructed to ensure the spatiotemporal consistency of heterogeneous data such as images and point clouds; In the data application stage, efficient cross-modal queries are achieved based on the association index table, training datasets that meet the requirements are dynamically filtered through environmental identifiers, and the physical rationality of generative augmented data is verified using spatiotemporal constraints, thereby providing high-quality, multimodal collaborative augmented data for sensing tasks in harsh environments.
[0047] The input data for the scrambler is input in a structured data format (such as JSON or binary encoding), including: a perturbation type identifier, an intensity parameter, or a duration.
[0048] The scrambler, as the core physical perturbation unit, receives input control commands via a structured data protocol (JSON / binary encoding).
[0049] The control commands include: a disturbance type identifier, an intensity parameter, or a duration.
[0050] The disturbance type identifier has an enumerated value that includes optical disturbance code 101 or electromagnetic disturbance code 102.
[0051] The intensity parameters include a simulated rainfall intensity of 50 mm / h, a simulated wind speed of 10 m / s, and a simulated light intensity of 100 lux.
[0052] The duration is in milliseconds, for example, duration=5000ms.
[0053] It should be noted that the scrambler adjusts the generated environmental characteristics such as precipitation, wind force, and light intensity according to the input parameters, but does not output data itself.
[0054] The control commands are generated by the central controller and issued in the form of API calls to ensure strict synchronization with the data acquisition from the multi-source sensors.
[0055] The control commands for the scrambler are generated by the central controller and sent to the scrambler via API calls or configuration files. Commands are generated in formats such as JSON strings or binary instruction sets. Core commands include: scrambling activation / stop commands, the aforementioned parameter configuration commands, and synchronization commands.
[0056] Triggering timing: Driven by user commands, triggered when data needs to be collected under disturbed conditions.
[0057] The specific scrambling functions of the scrambling device include weather simulation, optical interference, or motion simulation.
[0058] The meteorological simulation function includes: Rain / Fog Disturbance: Use an adjustable rainfall intensity spray system (e.g., simulating 0-100 mm / h rainfall) to add rain streaks and point cloud visibility reduction to the image.
[0059] Wind disturbance: A controllable wind speed (0-20m / s) is generated through a wind tunnel to simulate sensor jitter and target motion blur.
[0060] The optical interference function includes: Strong light / backlight: Use high-brightness floodlights or polarized light sources to add overexposure or glare effects to the image.
[0061] Low-light simulation: Adjust the light intensity to reproduce nighttime or hazy environments.
[0062] Sensor vibration: Simulate bumps (such as vehicle movement) using a vibration table to add image blur and point cloud position drift.
[0063] Step 120: Input the first sensing data into a pre-trained neural network model to generate a second sensing dataset; the second sensing dataset is a collection of multi-source sensing data with different perturbations generated by the neural network model performing spatiotemporally consistent perturbation simulation on the first sensing data.
[0064] The first sensing data is input into a pre-trained generative neural network model. The neural network model performs spatiotemporally consistent perturbation simulation on the first sensing data to generate augmented sensing data under harsh environments. The augmented sensing data is used for training machine learning models. It should be noted that the harsh environment described in this application refers to the environment after the normal environment has been subjected to physical disturbance. In terms of data, it can refer to the second perception data obtained after simulating various physical disturbances by perturbing the first perception data, or it can refer to the third perception data after the multi-sensor device is subjected to real physical disturbances by a disturbance device in reality, or the target area collected by the multi-sensor device receives real physical disturbances.
[0065] The term "time consistency" refers to the synchronicity of the evolution of multi-source data (images / point clouds) over time.
[0066] For example, timestamps of image / radar point cloud data are synchronized at the millisecond level, and the duration of disturbances is consistent (such as raindrop trajectories maintaining the same appearance / disappearance duration in image frame sequences and point cloud time series).
[0067] Spatial consistency refers to the consistency of the physical location in three-dimensional space during the image / radar point cloud perception process.
[0068] For example, the relative position of the image / radar sensor and the observed target is the same or the relative position error is in the centimeter to decimeter range, and the geometry of the disturbance is consistent (such as the 2D position of a raindrop in the image corresponding to a 3D reflection point in the point cloud).
[0069] The datasets collected under scrambling conditions are defined in the same way as those defined in step 110, including the same data format, representation, and data content. Data association is established by directly mapping the same target data collected for the same duration under both the normal environment in step 110 and the scrambling environment in this step.
[0070] The second perceptual dataset described in this application embodiment is an augmented dataset.
[0071] The disturbance simulation includes at least one of rain streaks, cloud cover, and motion blur.
[0072] It should be noted that in the embodiments of this application, both the scrambler and the active feed participate in the scrambling process, and the power intensity of the aforementioned electromagnetic interference is added by the active feed. Disturbances such as weather, wind, and glare are added by the scrambler.
[0073] Furthermore, prior to steps 110-120, the pre-training process of the neural network model includes the following steps: Step 100-1: Obtain first perception data and third perception data; the third perception data is environmental scrambling data collected synchronously with the first perception data, and the two are spatiotemporally aligned and have the same scene; It should be noted that the first sensing data in step 100-1 and the first sensing data in step 110 are both multi-source sensing data under normal conditions. The difference is that the first sensing data in step 110 is any newly collected multi-source sensing data under normal conditions, while the first sensing data in step 100-1 must be strictly paired with the third sensing data for collection. For example, the same camera first takes a clear image in the laboratory, and then turns on the rainfall simulator to take a rainstorm image.
[0074] The first sensing data is data collected under normal conditions.
[0075] The normal environment is defined as an environment free from physical disturbances: the sensor is free from physical vibrations caused by wind, heavy rain, etc., the environment is free from optical disturbances caused by precipitation, fog, haze, lightning, strong light, etc., and the environment is free from electromagnetic interference caused by strong electromagnetic interference to sensors such as infrared and radar.
[0076] The digital indicators are described as follows: no precipitation (rainfall ≤ 0 mm / h), wind speed below the threshold (e.g., ≤ 1 m / s), stable illumination (illuminance change ≤ 10%), no active motion blur (relative velocity of sensor or target ≤ 1 m / s), and electromagnetic noise power intensity generated by the scrambler is 0 W.
[0077] For example, use binary encoding or JSON string identifiers to identify the normal environment, such as 0x00 representing the normal environment.
[0078] The third type of sensing data is the sensing data collected after environmental scrambling.
[0079] For example, harsh environments can be identified using binary encoding or JSON string identifiers, such as 0x01 representing a harsh environment. The scrambled values for rainfall, wind speed, light intensity, vibration amplitude, etc., are then appended as strings to the collected sensor data.
[0080] The first perception data can be directly input into the pre-trained neural network model, but the first perception data needs to be spatiotemporally aligned to ensure that it is valid as a ground truth.
[0081] The third sensing data refers to the raw data collected by the sensor under real physical disturbance conditions, obtained through a pre-trained dataset acquisition device.
[0082] For example, without enabling the scrambler, the multi-source sensor is activated in real time to collect and store multi-source sensor sensing data (first sensing data) obtained under normal conditions. This dataset is denoted as [dataset name missing]. .
[0083] The environmental characteristics to be simulated determine the environmental disturbances that need to be added, such as precipitation, wind force, and direct interference light intensity level, and determine whether an active feed source is needed to simulate ambient light intensity and radar signals. In this way, the working environment of multi-source sensors under various climates and conditions can be simulated.
[0084] With the scrambler enabled, the multi-source sensors are activated, and multi-source sensor data obtained under the set environment is collected and stored. This dataset is denoted as [dataset name missing]. .
[0085] Use the obtained dataset and Training an AI-based scrambling module. The scrambling module is trained based on a commonly used generative model, where the dataset... As training input As a truth value.
[0086] In one embodiment, the third sensing data is multi-source sensing data of the multi-source sensors being physically disturbed and / or the sensed environment and target being physically disturbed.
[0087] Furthermore, the third sensing data is acquired in a controlled, harsh environment such as an artificial rain laboratory or a vibration platform. It includes paired data from the same scene, containing real physical disturbance effects and multi-source sensing data from a normal environment, with strict spatiotemporal alignment.
[0088] For example, point cloud data collected by lidar in real rainstorms includes raindrop noise and visibility reduction effects.
[0089] Step 100-2: Train the neural network model using the first perception data as input and the third perception data as a supervision signal.
[0090] For example, using normal environment perception data as input and corresponding harsh environment ground truth data as supervision signals, the model is trained by minimizing the loss function.
[0091] In one embodiment, the neural network model is a generative model.
[0092] For example, the scrambling module can be trained based on commonly used generative models, such as generative adversarial networks, diffusion models, etc.
[0093] For example, the generative neural network model is a generative adversarial network (GAN), which includes: Generator: Takes normal sensing data as input and outputs augmented data (i.e., second sensing data).
[0094] Discriminator: Determines whether the input data is third-party sensory data or generated data.
[0095] For example, the generator uses a U-Net structure, and the discriminator uses a PatchGAN structure.
[0096] The third sensing data refers to the raw data collected by the sensor under real physical disturbance conditions.
[0097] The generated data refers to the third-party sensory data simulated by the generator based on normal data. It is obtained by simulating perturbations in normal data using a neural network model, without any real physical perturbations, and is synthesized through an algorithm.
[0098] For example, inputting a sunny day image into the generator will output an image with simulated rain streaks added by an algorithm.
[0099] The third sensing data refers to sensor data collected in a controlled environment using a scrambling device (such as a rainfall simulator). The generated data refers to perturbation-containing data simulated by a neural network based on normal data.
[0100] In one embodiment, the perturbation simulation includes: Optical perturbations are added to the image data. For example, these perturbations are caused by motion blur, weather conditions such as rain, fog, or haze.
[0101] Electromagnetic disturbances are added to the radar point cloud data. For example, the electromagnetic disturbances may include signal attenuation noise or background interference.
[0102] Maintain the spatial correspondence between image data and point cloud data.
[0103] In one embodiment, before the first perceptual data is input into the neural network model, the step further includes: Geometric transformations are performed on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
[0104] The geometric transformations include linear transformations and nonlinear transformations.
[0105] The linear transformations include rotation, translation, scaling, mirroring, or shearing.
[0106] The nonlinear transformations include surface distortion, radial distortion, non-uniform noise injection, or point cloud topology reconstruction.
[0107] Figure 2 This application also provides a multi-source sensing data augmentation system structure diagram for implementing the multi-source sensing data augmentation method described in any embodiment of the first aspect, including: a multi-source sensor module 21 and a scrambling module 22.
[0108] The multi-source sensor module is used to acquire the first sensing data.
[0109] For example, it is specifically used to obtain RGB, infrared, and other image data and / or point cloud data from lidar and millimeter-wave radar.
[0110] Multi-source sensor modules can employ either a discrete or fused design. A discrete design involves acquiring the sensor output data from multiple independent sensors and storing it in the same memory; a fused design uses a unified power supply, signaling, and sensor data output interface, all stored in the same memory. Examples of different designs are as follows: Figure 3 As shown.
[0111] Figure 3 The sensors 1 to n in the text can be any one of the following sensors: RGB, infrared camera, lidar, millimeter-wave radar, SAR radar, etc.
[0112] The scrambling module is used to perform spatiotemporally consistent perturbation simulation on the first sensing data, generating a set of multi-source sensing data with different perturbations.
[0113] Artificial intelligence-based scrambling modules use techniques such as generative neural network models to physically scramble the transformed perceptual data, adding effects such as cloud cover, rain streaks, and jitter. These environmental perturbations are significant factors affecting the performance of applications such as target recognition under certain harsh conditions.
[0114] Furthermore, it also includes a transformation module 23, which is used to perform geometric transformations on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
[0115] For example, the transformation module can perform linear transformations such as rotation, cropping, and scaling, or nonlinear transformations such as image and point cloud distortion.
[0116] The transformation module can use standard processing techniques such as random cropping, flipping, rotating, translating, and twisting to process the sensor data, increasing the quantity and diversity of the dataset.
[0117] In one embodiment, the AI-based scrambling module requires pre-training. The pre-training process consists of two steps and can be designed using a pre-training dataset acquisition device, including: Pre-training dataset collection.
[0118] The collected dataset is used to train an AI-based scrambling module.
[0119] To achieve the collection of pre-trained datasets, a pre-trained dataset collection method is designed as follows: Figure 4 As shown.
[0120] In one embodiment, the pre-training dataset acquisition device of this application includes a multi-source sensor 21 and a scrambler 41. The multi-source sensor is used to receive image data and / or radar point cloud data of the perceived environment and target.
[0121] The multi-source sensor is the core data acquisition unit of the device described in this embodiment, specifically referring to a composite sensor system capable of simultaneously acquiring multiple types of sensor data. In practical applications, this sensor system typically includes the following components: Optical imaging components: such as visible light cameras and infrared thermal imagers, are used to collect image data of the environment and targets.
[0122] For example, multi-source sensors determine image data by receiving light signals such as visible light reflected from the perceived environment and target, and infrared light emitted from the target.
[0123] Radar detection components, such as millimeter-wave radar and lidar, are used to acquire three-dimensional point cloud information of targets.
[0124] For example, multi-source sensors determine radar point cloud data in at least one of the following ways: After the radar in the multi-source sensor actively transmits electromagnetic signals, the electromagnetic signals reflected by the sensed environment and target are received by the multi-source sensor. Alternatively, the radar passively receives electromagnetic signals transmitted by the perceived environment and targets.
[0125] The first sensing data is obtained by using only multi-source sensors to receive light or electromagnetic signals from the perceived environment and target.
[0126] The scrambler is the physical environment simulation unit of the device described in this embodiment. Its core function is to reproduce harsh environmental conditions through controllable physical means. Typical implementation methods include: Meteorological simulation subsystems include systems such as adjustable rainfall intensity sprinkler systems and programmable wind tunnels.
[0127] Optical interference subsystems: such as high-brightness floodlights, polarized light generators, etc.
[0128] Motion simulation platforms, such as six-degree-of-freedom vibration tables, can simulate the bumps and vibrations of a vehicle in motion.
[0129] These subsystems work in coordination through a central controller, accurately reproducing various complex environmental combinations, such as the combined harsh conditions of "heavy rain + strong backlight + bumpy road". By adding physical disturbances such as rainfall, wind force, and strong light to the multi-source sensors and / or the sensed environment and target through a scrambler, the data sensed by the multi-source sensors at this time becomes third-party sensing data.
[0130] It should be noted that the multi-source sensor in this embodiment and the multi-source sensor module described above can be the same device used in different method steps, or they can be different devices.
[0131] It should also be noted that the scrambler in this embodiment is not the same device as the scrambling module 22. The scrambling module uses artificial intelligence to add simulated scrambling effects to the already collected multi-source sensing data, while the scrambler in this embodiment applies real physical disturbances to the multi-source sensors and / or the collected environment and target in a real environment.
[0132] In one embodiment, by setting an active feed 42 at the sensed environment and target, different light intensity illumination and / or different intensity bistatic radar transmission signals are provided to the sensed environment and target.
[0133] The active feed source is an optional auxiliary unit of the device described in this embodiment, mainly used to enhance the controllability of the characteristics of the sensed environment and target. Its typical configuration includes: Lighting systems: such as LED arrays equipped with filters, which can simulate lighting conditions at different times of day (dawn / dusk / noon).
[0134] Radar signal enhancers: such as 60 GHz millimeter-wave signal generators, used to enhance the radar reflection characteristics of specific targets.
[0135] Infrared marking devices: such as modulated infrared beacons, used to maintain target detectability in low-visibility environments.
[0136] By programming and controlling the operating parameters of these devices, the influence of different environmental variables on the sensed data can be studied systematically.
[0137] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] Therefore, this application also proposes a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods described in any embodiment of this application.
[0139] Furthermore, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any embodiment of this application.
[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0143] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0144] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 shown is merely an example and should not impose any limitations on the function and scope of use of the embodiments of this application. It includes: one or more processors 520; and a storage device 510 for storing one or more programs, which, when executed by the one or more processors 520, cause the one or more processors 520 to implement the multi-source sensing data augmentation method provided in the embodiments of this application, the method including: Acquire first sensing data; the first sensing data is multi-source sensing data under normal conditions; The first sensing data is input into a pre-trained neural network model to generate a second sensing dataset; the second sensing dataset is a collection of multi-source sensing data with different perturbations generated by the neural network model performing spatiotemporally consistent perturbation simulation on the first sensing data.
[0145] The electronic device 500 also includes an input device 530 and an output device 540; the processor 520, storage device 510, input device 530 and output device 540 in the electronic device can be connected by a bus or other means, as shown in the figure, which is connected by a bus 550.
[0146] Storage device 510, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and module units, such as the program instructions corresponding to the multi-source sensing data augmentation method in this embodiment. Storage device 510 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on terminal usage. Furthermore, storage device 510 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, storage device 510 may further include memory remotely located relative to processor 520, and these remote memories can be connected via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0147] Input device 530 can be used to receive input digital, character, or voice information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 540 may include electronic devices such as a display screen and a speaker.
[0148] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0149] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be understood that when a device or component is “connected” to another device or component, it may be directly connected to the other device or component, or there may be an intermediary device or component. Furthermore, the term “connection” as used herein may include partially wireless connections as well as partially wired connections.
[0150] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0151] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for augmenting multi-source sensing data, characterized in that, Includes the following steps: Acquire first sensing data; the first sensing data is multi-source sensing data under normal conditions; The first sensing data is input into a pre-trained neural network model to generate a second sensing dataset; the second sensing dataset is a collection of multi-source sensing data with different perturbations generated by the neural network model performing spatiotemporally consistent perturbation simulation on the first sensing data.
2. The multi-source sensing data augmentation method according to claim 1, characterized in that, The pre-training process of the neural network model includes the following steps: Acquire first-sensory data and third-sensory data; the third-sensory data is environmental scrambling data collected synchronously with the first-sensory data, and the two are spatiotemporally aligned and have the same scene. The neural network model is trained using the first perception data as input and the third perception data as a supervision signal.
3. The multi-source sensing data augmentation method according to claim 1, characterized in that, The neural network model is a generative model.
4. The multi-source sensing data augmentation method according to claim 1, characterized in that, The disturbance simulation includes: Add optical perturbations to the image data; Add electromagnetic disturbances to radar point cloud data; Maintain the spatial correspondence between image data and point cloud data.
5. The multi-source sensing data augmentation method according to claim 1, characterized in that, Before the first perceptual data is input into the neural network model, the process further includes the following steps: Geometric transformations are performed on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
6. The multi-source sensing data augmentation method according to claim 2, characterized in that, The third sensing data is multi-source sensing data of the multi-source sensors being physically disturbed and / or the sensed environment and target being physically disturbed.
7. A multi-source sensing data augmentation system, used to implement the multi-source sensing data augmentation method according to any one of claims 1 to 6, characterized in that, include: Multi-source sensor module and scrambling module; The multi-source sensor module is used to acquire the first sensing data; The scrambling module is used to perform spatiotemporally consistent perturbation simulation on the first sensing data, generating a set of multi-source sensing data with different perturbations.
8. The multi-source sensing data augmentation system according to claim 7, characterized in that, It also includes a transformation module, which is used to perform geometric transformations on the image or radar point cloud data of the first sensing data to obtain the first sensing data after several geometric transformations.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-7.