Image data integration method based on multi-point image acquisition

By constructing an image data integration method for multi-point image acquisition and utilizing a reinforcement learning training framework and agent design, the problems of hardware adaptation and synchronization error in multi-point image acquisition are solved, achieving high-fidelity and stable image integration to meet the detail accuracy requirements of industrial inspection.

CN121600361BActive Publication Date: 2026-04-21CHINA RAILWAY XIAN GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA RAILWAY XIAN GRP CO LTD
Filing Date
2026-01-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing multi-point image acquisition technologies have shortcomings in hardware adaptation, synchronization errors, and noise interference, resulting in data distortion and multi-view image alignment deviations and detail loss during the integration process. This makes it difficult to meet the requirements of high-fidelity integration, and traditional integration strategies lack flexibility and stability.

Method used

A method for integrating image data from multi-point image acquisition is constructed. This method involves building an image integration reinforcement learning training framework adapted to multi-point image acquisition scenarios, utilizing a digital twin scenario model and a multi-point image integration agent, setting up state transition and reward feedback mechanisms, combining the NSGA-II algorithm to generate Pareto optimal integration strategies, designing sharpness, consistency, and integrity reward functions, and optimizing the image integration process.

Benefits of technology

It achieves reliable integration strategies under different multi-point acquisition scenarios, improves training stability and strategy generalization, ensures the clarity, consistency and integrity of image integration, solves the limitations of hardware adaptation and dynamic scene changes in traditional methods, and meets the detailed accuracy requirements of industrial inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600361B_ABST
    Figure CN121600361B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image management technology, specifically to an image data integration method based on multi-point image acquisition. It constructs an adaptive image integration reinforcement learning training framework, relying on a digital twin scene model to replicate the hardware characteristics and physical interference of the actual acquisition environment. Combined with the dynamic input / output design of the multi-point image integration agent, a priority experience playback mechanism, and stable training optimization logic, it improves training stability and policy generalization, avoiding fluctuations in integration effects caused by hardware parameter drift in real-world scenarios, and breaking the limitations of traditional fixed processes adapting to dynamic scenarios. Simultaneously, it designs an innovative reward system: a sharpness reward eliminates background interference and filters false gradients; a consistency reward focuses on overlapping areas and avoids extreme deviation interference; and a completeness reward distinguishes defect weights to ensure the retention of key defects. These three factors guide the agent strategy to meet industrial inspection requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image management technology, and more specifically, to a method for integrating image data based on multi-point image acquisition. Background Technology

[0002] With the increasing demand for image data accuracy and integrity in fields such as industrial inspection and 3D modeling, multi-point image acquisition technology has become a key means of acquiring comprehensive target information because it can compensate for the limitations of single-point acquisition through multi-view collaboration. However, existing multi-point image integration methods generally suffer from insufficient hardware adaptability: they fail to fully replicate the dynamic characteristics of hardware devices in the actual acquisition environment (such as parameter drift and aging during device operation), and they do not effectively reproduce the synchronization errors when multiple devices are linked (such as sampling time delay and data transmission jitter). At the same time, they are insufficient in simulating and compensating for physical interference such as noise and illumination fluctuations, which makes the acquired data prone to distortion. In the subsequent integration process, problems such as multi-view image alignment deviation and loss of details frequently occur, making it difficult to meet the requirements of high-fidelity integration.

[0003] At the level of constructing intelligent integration strategies, traditional methods often rely on fixed processes or simple rules (such as a single denoising algorithm or fixed fusion weights), lacking a flexible architecture that can adapt to different numbers of collection points and dynamic changes in the scenario. Existing training frameworks not only fail to establish a high-fidelity simulation environment that fits the actual scenario, thus failing to provide a real interactive scenario for policy learning, but also suffer from low efficiency in utilizing experience samples and poor training stability—either failing to distinguish sample priorities, resulting in high-value samples being ignored, or lacking an effective bias correction mechanism. The resulting integration strategy has weak generalization ability, making it difficult to cope with hardware differences and interference changes in different multi-point collection scenarios, and failing to output stable and reliable integration results.

[0004] Based on the above, this invention proposes an image data integration method based on multi-point image acquisition. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of this invention is to provide an image data integration method based on multi-point image acquisition.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] The image data integration method based on multi-point image acquisition is characterized by the following steps:

[0008] Step 1: Construct an image integration reinforcement learning training framework adapted to multi-point image acquisition scenarios;

[0009] A reinforcement learning training framework for image integration adapted to multi-point image acquisition scenarios is constructed as follows: First, a digital twin scenario model for multi-point acquisition is built, a multi-point image integration agent is constructed, a state transition and reward feedback mechanism is set, an experience playback mechanism is established, and training optimization and parameter update logic are defined.

[0010] The reward and feedback mechanism is as follows: Define clarity reward Consistency rewards Completeness reward Further define the multi-point image integration reward function ;in, For clarity weight, For consistency weights, Integrity weight; , , Dynamically changing with the current integration phase ;

[0011] Step 2: Define the state space and feature space for image integration;

[0012] Step 3: Generate a Pareto optimal integration strategy set based on the NSGA-II algorithm and execute the optimal integration strategy.

[0013] Furthermore, a digital twin scene model of the multi-point acquisition scenario is constructed, as follows: the actual scene of multi-point image acquisition is determined, a digital twin model of the image acquisition hardware is established, a standard image is constructed based on the physical properties of the scene targets in the actual scene, dynamic rules of scene physical interference are defined, and interference parameters are superimposed on the standard image according to the imaging logic of the real device to generate the original acquisition data.

[0014] Furthermore, a multi-point image integration agent is constructed, specifically as follows: the input dimensions and feature sources of the multi-point image integration agent are defined, a network structure adapted to the image data integration task is designed, an action space matching the integration sub-task is defined, and training parameters are set.

[0015] Furthermore, the input dimensions and feature sources of the multi-point image integration agent are clarified: the input includes two categories: image features and scene state. Image features include local matching features and global texture features; scene state includes the individual state of the collection point, the associated state of the collection point, and the state of the integration process.

[0016] Furthermore, clarity bonus ;in, The total number of valid pixels, i.e., the set of valid pixels. The number of pixels contained; For the effective pixel set; For regional weighting coefficients,

[0017] ; For pixels Edge gradient strength; This is an adaptive threshold for noise.

[0018] Furthermore, consistency rewards ; q represents a single pixel, It is the set of pixels in all overlapping regions of multiple views; Pixel weight coefficients are the weights assigned to the differences between different pixel regions, using a simplified binary weighting:

[0019] ; It represents the absolute value of the difference in grayscale values ​​of the same pixel q in two images from different viewpoints; This represents the total number of pixels in the globally overlapping region, i.e. .

[0020] Furthermore, integrity rewards M represents the total number of original defects; m represents the defect index. The importance weight of the defect, i.e., the importance coefficient of the m-th defect, is determined using a simplified binary weighting.

[0021] ; Indicates the crossover ratio of defect regions, and the original defect. Defects matched with the integrated image The regional overlap is calculated using a simplified formula as follows: ; This represents the original defect, specifically the m-th defect marked in the original scene. This represents the matching defects in the integrated image; in the image integrated by the agent, the defects are compared with the original defects. The corresponding defects.

[0022] Furthermore, the mathematical expression for the state space of image integration is: ; For image acquisition hardware status, For the target state of the scene, This indicates the status of the integration process.

[0023] Furthermore, the optimal integration strategy is executed, specifically as follows: The values ​​of sharpness weight, consistency weight, and integrity weight are determined based on the current integration stage; and the optimal integration strategy set multi-point image integration reward function is determined based on the values ​​of sharpness weight, consistency weight, and integrity weight. The solution with the largest value is marked as the optimal integration strategy, and that optimal integration strategy is executed.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] The method of this invention constructs an image integration reinforcement learning training framework adapted to multi-point image acquisition scenarios. Relying on a digital twin scene model, it accurately replicates the hardware characteristics (such as device dynamic drift and multi-device coordination errors) and physical interference (such as noise and illumination fluctuations) of the actual acquisition environment, providing a high-fidelity interactive scenario for the agent. At the same time, combined with the dynamic input and output design of the multi-point image integration agent (features and state dimensions adjusted according to the number of acquisition points), priority experience playback mechanism, and stable training optimization logic, the agent can fully learn the strategy of multi-acquisition point collaborative integration in the simulated scenario, and effectively improve the training stability and strategy generalization. It avoids the fluctuation of integration effect caused by hardware parameter drift and multi-device synchronization deviation in the actual scenario, and ensures that reliable integration strategies can be output in different multi-point acquisition scenarios, breaking the limitation of traditional fixed process integration methods that are difficult to adapt to dynamic scenarios.

[0026] The sharpness reward design breaks through the limitations of traditional global gradient measurement. It eliminates meaningless background interference through effective region filtering, focusing only on valuable target areas in the image. Combined with a dynamic noise threshold (adaptively adjusted based on local noise intensity), it filters false gradients, preventing noise from being misjudged as sharp edges. Simultaneously, it assigns higher weight to critical areas such as defects, accurately focusing on details crucial in industrial inspection. This design, while suppressing noise, maximizes the preservation of target edges and defect features, solving the problems of traditional sharpness measurement methods being susceptible to background noise and insufficient preservation of key details. It provides a high-quality image foundation for subsequent calibration and fusion processes, aligning with the core requirements of detail accuracy in industrial scenarios.

[0027] The consistency reward, by focusing on overlapping areas from multiple perspectives, introducing pixel weight coefficients (prioritizing defect areas), and an extreme difference truncation mechanism, accurately quantifies the pixel continuity of images from different perspectives. This avoids interference from non-overlapping areas or extreme deviations in consistency judgment, ensuring that key features are not misaligned after multiple images are aligned, thus solving the problem of traditional consistency measures ignoring differences in regional importance. The integrity reward, by distinguishing the importance weight of defects (prioritizing key defects) and combining the intersection-union ratio of defect areas, accurately measures the degree of retention of defect features during integration. This effectively avoids secondary defects interfering with core requirements and prevents key defects (such as defects affecting structural strength) from being lost during integration. Together, these three elements constitute a highly targeted reward system, guiding the agent's integration strategy to better meet the actual needs of industrial inspection and application. Attached Figure Description

[0028] Figure 1 This is a flowchart of an image data integration method based on multi-point image acquisition.

[0029] Figure 2 The flowchart shows the process of generating the optimal integration strategy set. Detailed Implementation

[0030] Reference Figures 1 to 2 An image data integration method based on multi-point image acquisition is described below:

[0031] Step 1: Construct an image integration reinforcement learning training framework adapted to multi-point image acquisition scenarios;

[0032] A reinforcement learning training framework for image integration adapted to multi-point image acquisition scenarios is constructed as follows: First, a digital twin scenario model for multi-point acquisition scenarios is constructed, a multi-point image integration agent is constructed, a state transition and reward feedback mechanism is set, an experience playback mechanism is built, and training optimization and parameter update logic is set.

[0033] The digital twin scene model of the multi-point acquisition scenario is constructed as follows: the actual scene of multi-point image acquisition is determined, a digital twin model of the image acquisition hardware is established, a standard image is constructed based on the physical properties of the scene targets in the actual scene, dynamic rules of scene physical interference are defined, and interference parameters are superimposed on the standard image according to the imaging logic of the real device to generate the original acquisition data.

[0034] Build a digital twin model of the image acquisition hardware, including: Digital mapping of hardware parameters: Digital mapping of internal parameters: Establish a parameter library according to the device type (industrial camera), and store the focal length (500 - 2000 pixels, uniformly sampled), principal point coordinates (image center ±50 pixels, sampled normally), and distortion coefficients (radial k1 = -0.3 to 0.3, tangential p1 = -0.001 to 0.001); Use JSON structure to associate "parameter range - probability distribution" to support random call; Digital mapping of external parameters: Convert the spatial attitude of the acquisition point into a mathematical matrix, rotation angle (±30° around the X / Y / Z axis, convert Euler angle to rotation matrix), translation vector (0 - 500mm, uniformly sampled in spherical coordinates), to achieve coordinate mapping from the physical space to the digital space; Imaging physical modeling: Basic projection model: Based on the principle of a pinhole camera, construct a projection link of "3D world point → camera coordinate system → 2D pixel", integrate internal parameters to calculate pixel coordinates, and superimpose a distortion model (radial / tangential distortion formula) to correct imaging deviation; Device-specific simulation: The industrial camera adds an association model between sensor noise (dark current, readout noise) and exposure time (1 - 100ms); Dynamic response simulation: Parameter time-varying simulation: Simulate the hardware operation drift, such as the focal length changing with temperature (temperature coefficient 1e-5 / °C), the distortion coefficient aging with usage duration (aging coefficient 1e-6 / h), and the rotation angle fluctuating due to vibration (±0.1°, period 1 - 5s). Multi-device collaborative simulation: Process the spatio-temporal synchronization errors of K acquisition points, including sampling time ±10ms delay, external parameter calibration ±0.5° / ±2mm deviation, and data transmission 10 - 100ms jitter, and restore the actual multi-device linkage scenario. Model calibration and verification: Calibration: Use a real device to capture a standard checkerboard, obtain the real internal / external parameters through the Zhang Zhengyou calibration method, and inversely correct the digital model parameters to make the error between the simulated parameters and the real values <5%; Verification: Use "reprojection error < 1.5 pixels", "correlation coefficient between simulated and real images > 0.9", and "dynamic drift RMSE < 1%" as indicators to ensure the model fidelity;

[0035] Standard images are constructed based on the physical properties of targets in real-world scenarios, including: Decomposition of physical properties of targets: Industrial scenarios (e.g., mechanical parts): decomposed into geometric parameters (size, aperture, tolerance), material properties, defect features, and physical properties → image feature mapping; Geometric property mapping: converting target size to image pixel scale according to the "actual size-pixel" conversion relationship of the hardware digital twin model (e.g., a φ20mm shaft diameter corresponds to a circle with a diameter of 200 pixels in the image); The geometric shape of the 3D target is generated through "3D model → 2D projection" (based on a pinhole camera model with hardware extrinsic parameters) to ensure viewpoint matching; Material / density property mapping: reflectivity / density corresponds to image grayscale / RGB values ​​(e.g., high reflectivity of metal → grayscale value 200-255); Structure / defect property mapping: defects correspond to image grayscale values. Gradient (crack defects → continuous areas with gray values ​​30-50 lower than surrounding areas), standard image generation and detail enhancement: Industrial scenarios: Export the target geometric contour from the CAD model, fill the grayscale according to the material reflectivity, and overlay the precise pixel area of ​​the defect (e.g., mark the crack with a rectangle, coordinates (x1,y1)-(x2,y2)); Standard image quality verification: Ensure the consistency between the standard image and physical properties through indicators: Geometric accuracy: The error between the target size in the image and the actual size is <1% (e.g., for an actual 20mm shaft diameter, the pixels in the image corresponding to a size deviation ≤0.2mm); Feature integrity: All key physical properties (e.g., part defects, organ structures) are clearly reflected in the image without omission; Consistency: Standard images of the same target from different perspectives have a geometric shape matching degree >95% (calculated through feature point matching);

[0036] Define dynamic rules for scene physical interference, and superimpose interference parameters onto the standard image according to the imaging logic of the real device to generate raw acquired data. Specifically, simulate dynamic factors affecting image quality in real scenes to generate a sequence of raw images with interference. (k is the data collection point, t is the time step); This represents the raw image data acquired by the k-th image acquisition hardware (such as a camera) at the t-th time step of multi-point image acquisition, which has been superimposed with real-world physical interference (noise, illumination deviation, etc.) and has not undergone subsequent processing; Illumination dynamic model: Illumination intensity ,in As the reference brightness, For fluctuation coefficient, , For a period of time, (Simulating periodic changes in illumination); local illumination deviations are modeled using a Gaussian distribution with a mean of 0 and a standard deviation of 5-20 for grayscale values. It is a time variable.

[0037] Equipment noise model: Mixed noise generation, including Gaussian noise (standard deviation) (Adjusted dynamically according to sensor temperature drift), salt-and-pepper noise (density 0.01~0.05, simulating random sensor failure).

[0038] Dynamic adjustment of viewing angle deviation: changes in extrinsic parameters between adjacent time steps (±0.5° rotation) (±1mm translation) simulates the viewpoint drift caused by equipment vibration or target micro-motion.

[0039] Construct a multi-point image integration agent as follows: clarify the input dimension and feature source of the multi-point image integration agent, design a network structure that adapts to the image data integration task, define the action space that matches the integration sub-task, and set the training parameters.

[0040] Clearly define the input dimensions and feature sources of the multi-point image integration agent: the input includes two categories: image features and scene state. The total dimension is dynamically adjusted according to the number of collection points K to ensure coverage of image content attributes and dynamic scene constraints.

[0041] 1. Image features (fixed dimension, independent of K):

[0042] Local matching features: Take the top-200 SIFT feature points in the image of each acquisition point, and cluster them into two central feature vectors (each 128-dimensional) using K-means, and concatenate them to form a 256-dimensional vector (dimensionality source: 2×128-dimensional).

[0043] Global texture features: Calculate the gray-level co-occurrence matrix (distance d=1,2,3,4; angle 0°,45°,90°,135°) for each image at each acquisition point, extract four types of statistics: contrast, energy, entropy, and correlation, and take the mean vector of K images, which has a total of 16 dimensions (dimensionality source: 4 distances × 4 angles → 4 statistics, the dimension remains unchanged after meaning).

[0044] Total dimensions of image features: 256 + 16 = 272 dimensions.

[0045] 2. Scene state (dimensions change dynamically with K):

[0046] Individual status of sampling points (K-dimensional): Noise intensity of each sampling point (calculated based on local variance, 1 dimension / point);

[0047] Acquisition point association status (K×(K-1) / 2 dimensions): the overlap of viewpoints between any two points (calculated by feature matching rate, 1 dimension / point pair);

[0048] Integration process status (3D): Structural similarity deviation between intermediate results and standard images (SSIM deviation, 1D), current integration stage marker (preprocessing=0 / calibration=1 / fusion=2, 1D), equipment average exposure time (1D).

[0049] Total dimensions of scene state: K + [K × (K-1) / 2] + 3.

[0050] Total input dimensions: When K=4 (typical 4-point acquisition scenario), the total dimensions = 272 + [4 + 6 + 3] = 272 + 13 = 285 dimensions; adjusting K can adapt to different multi-point acquisition scenarios, ensuring that the input information matches the quantity and correlation characteristics of the actual scenario.

[0051] Design a network structure adapted to image data integration tasks, including: Input layer: Image feature branch (receives 272-dimensional image features), Scene state branch (receives 13-dimensional scene state, K=4); Feature extraction layer: Image branch: 3-layer CNN (64→128→256 channels, 3×3 convolution, ReLU activation) + 1-layer global pooling (output 256 dimensions); Scene branch: 3-layer fully connected (128→64→32 neurons, ReLU activation), output 32 dimensions; Cross-attention fusion layer: Employs a multi-head attention mechanism (8 heads, 3 dimensions). 2) Calculate the mutual attention weights between image features and scene state (e.g., "high noise intensity → enhanced denoising feature weights"), outputting a 256-dimensional fused feature; Decision layer: Value stream: 2 fully connected layers (128→64→1 neuron), outputting state value V(s); Advantage stream: 2 fully connected layers (128→64→32 neurons), outputting action advantage A(s,a); Integrated output: Q(s,a)=V(s)+A(s,a)-mean(A(s,a)), where mean(A(s,a)) is the average value of the action advantage;

[0052] Define the action space matching the fusion subtasks: map each action to an integer index from 0 to 31, and decompose the subtasks into preprocessing subtasks, calibration subtasks, and fusion subtasks; the action type of the preprocessing subtasks: adaptive denoising parameter adjustment; the number of actions in the preprocessing subtasks: 10; the specific action definition of the preprocessing subtasks: Gaussian filter standard deviation. (Step size 0.5, corresponding index 0-9); Action type of calibration subtask: feature matching strategy + threshold combination; Number of actions in calibration subtask: 12; Specific action definition of calibration subtask: matching strategy (4 types: SIFT=0, ORB=1, SURF=2, AKAZE=3) × matching threshold (3 types: 0.7=0, 0.75=1, 0.8=2), index 10-21 (calculation formula: strategy × 3 + threshold + 10); Action type of fusion subtask: dynamic weight allocation; Number of actions in fusion subtask: 10; Specific action definition of fusion subtask: multi-image fusion weight coefficient α∈{0.1,0.2,...,1.0} (step size 0.1, corresponding index 22-31), α is the weight of the main acquisition point, and the remaining points are allocated according to (1-α) / (K-1); the main acquisition point refers to the acquisition point that is dynamically selected based on the current scene state and has the highest priority in contributing to the final integration result;

[0053] Training parameters are set as follows: Learning rate: initial η = 0.001, using exponential decay (decreasing by 10% every 1000 steps). Experience replay batch: 32 samples / batch. Discount factor γ = 0.95 (balancing immediate and long-term rewards). Target network update frequency: synchronize main network parameters every 1000 steps. Exploration rate. Initially 1.0 (fully random exploration), linearly decays to 0.1 (maintained after 100,000 steps), balancing exploration and exploitation.

[0054] The state transition mechanism is as follows:

[0055] Action parsing and execution: The action index (0-31) output by the agent is parsed into specific operation instructions by the environment module, and executed in three categories:

[0056] Preprocessing steps (0-9): Perform Gaussian filtering on the images of the K acquisition points ( =1.0-5.5), the filter strength is linearly related to the action index (index i corresponds to σ=1.0+0.5i).

[0057] Calibration action (10-21): Parse as a combination of "matching strategy + threshold" (e.g., index 10 = SIFT + 0.7), call the corresponding algorithm to calculate the transformation matrix (e.g., homography matrix H) between K images, and complete the viewpoint alignment;

[0058] Fusion action (22-31): Parse as fusion weight α (index 22 corresponds to α=0.1, each increase of index 1 α+0.1), calculate the weighted fused image according to the rule of "weight of main acquisition point α + equal distribution of other points (1-α)".

[0059] State update rules: New state It consists of three parts, and the update formula is as follows: ,in: Images of K acquisition points after motion processing (such as denoised images, aligned images after calibration, and fused intermediate results); Updated image features (re-extracting SIFT clustering features and gray-level co-occurrence matrix features); The updated scene status includes: Noise intensity: recalculated based on local variance after denoising (noise intensity decreases by 30%±5% if denoising is performed); Viewpoint overlap: recalculated based on feature matching rate after calibration (overlap increases by 20%±8% after alignment); Integration stage marker: automatically incremented (e.g., preprocessing → calibration → fusion, marked as 3 after fusion is completed to indicate termination).

[0060] Termination condition: When the integration stage is marked as 3 (fusion complete) or the number of iteration steps reaches the upper limit (T=50 steps), the state transition terminates and the final integration result is output.

[0061] The reward and feedback mechanism is as follows: Define clarity reward Consistency rewards Completeness reward (The regions of all three are [0,1]), further defining the multi-point image integration reward function. ;in, For clarity weight, For consistency weights, Integrity weight; , , Dynamically changing with the current integration phase ;

[0062] Preprocessing stage (marker=0): , , (Prioritize improving clarity); Calibration phase (marker=1): , , (Prioritize consistency); Fusion Phase (Mark=2): , , (Prioritize preserving integrity);

[0063] Clarity Bonus ;in, The total number of valid pixels, i.e., the set of valid pixels. The number of pixels contained (i.e. If the effective region is the 256×256 region at the center of the image, then ; The effective pixel set is the set of pixels consisting of "meaningful regions that are not part of the background" in the image. The region weight coefficient represents the weight assigned to the gradient values ​​of different pixel regions; it is a location-dependent function. ; For pixels The edge gradient strength is calculated using a simplified method: Among them, , These are the x and y gradients calculated using the 3×3 Sobel operator (reflecting the rate of grayscale change of a pixel in the horizontal / vertical direction). The noise adaptive threshold is a threshold that dynamically filters out "spurious gradients" caused by noise. It is calculated as follows: ; For pixels The standard deviation of grayscale in a 3×3 local area (a measure of noise intensity in that area).

[0064] Consistency rewards ;q represents a single pixel (simplified notation of coordinates (x,y)). It is "the set of pixels in all overlapping regions of multiple views" (for example, when K=4 sampling points, it includes the overlapping pixels of all adjacent point pairs such as (1-2), (2-3), (3-4), etc.). Pixel weight coefficients are the weights assigned to the differences between different pixel regions, using a simplified binary weighting:

[0065] ; This represents the absolute value of the difference in grayscale values ​​of the same pixel q in two images viewed from different perspectives. ,in, , These are the grayscale values ​​of pixel q in viewpoints a and b, respectively. Gray-scale difference threshold: In industrial scenarios, if the equipment (camera) has stable performance and small light fluctuations (such as in a closed workshop), the gray-scale difference range of image noise is fixed, and it can be set to 50, which is suitable for 8-bit grayscale images. This represents the total number of pixels in the globally overlapping region, i.e. ;

[0066] Completeness reward M represents the total number of original defects; m represents the defect index. The importance weight of the defect, i.e., the importance coefficient of the m-th defect, is determined using a simplified binary weighting. ; Indicates the crossover ratio of defect regions, and the original defect. Defects matched with the integrated image The regional overlap is calculated using a simplified formula as follows: (Value range [0,1], the larger the value, the more the two defect areas overlap and the higher the matching degree). Represents the original defect, the m-th defect marked in the original scene (or standard image); This represents the matching defects in the integrated image, which are in the image after agent integration (denoising, calibration, and fusion) compared to the original defects. Corresponding defects;

[0067] The experience replay mechanism is established as follows: A sample storage structure is defined, using a priority experience replay pool, and the sample storage format is as follows: ; For the current state, For actions, For single-step rewards (from the reward feedback mechanism) For the next state, This is a termination flag (1 = terminate, 0 = continue); This represents the number of sampling points corresponding to the current sample (used for subsequent generalization training). Samples are stored in categories based on the "number of sampling points K" (e.g., samples with K=4, K=6, etc.) to facilitate targeted sampling. Priority quantization: Sample priority is calculated based on TD error (time difference error). ; The basic priority of a sample directly determines the probability that the sample will be sampled in the experience replay. This represents the absolute value of the TD error; It is a very small positive number (usually taken as 1e−5); The discount factor can be set to 0.95; For the target network, the "next state" The estimation of the "optimal action value"; The target network (with the same structure as the main network, but with a lower parameter update frequency, used for stable training). Take the action with the largest Q value among all possible actions (i.e., the "optimal action value of the next state"). The main network's current state Execute action Value estimation; Sampling probability: The probability of a sample being selected is positively correlated with its priority. , , n is the priority influence coefficient (n controls the degree of priority influence; when n=0, it degenerates into uniform sampling); importance sampling weight: corrects the bias caused by priority sampling. , N represents the "total number of samples" in the experience replay pool. This is an adjustment parameter for "bias correction strength"; Capacity and elimination mechanism: Total capacity is 1 million samples, divided into K-value bins (maximum 200,000 samples per bin). When the capacity exceeds this, the sample with the smallest TD error is eliminated (high-value samples are retained). Batch sampling: 32 samples are drawn from the replay pool for each training iteration, with 80% from the current K-value bin (ensuring specificity) and 20% from other K-value bins (enhancing generalization). Regular updates: The TD error and priority of all samples are recalculated every 5000 steps to ensure that high-value samples (such as high-reward samples in the fusion phase) are learned first.

[0068] Define the training optimization and parameter update logic, including loss function design, parameter update process and stability assurance measures;

[0069] Loss function design: Weighted TD loss is adopted, combining priority weights and multi-objective rewards. The formula is as follows: ;in, The main network parameters are updated by minimizing the loss function through gradient descent; t is the "sample index" (from 1 to B) within the batch.

[0070] Parameter update process: Network parameter initialization: The initial parameters of the main network and the target network (with the same structure) are initialized using a He normal distribution to ensure a stable training starting point. Optimizer configuration: The Adam optimizer is used, with an initial learning rate η = 0.001;

[0071] Stability assurance measures:

[0072] Iterative process: The multi-point image integration agent interacts with the digital twin scene model of the multi-point acquisition scenario to generate samples and store them in the experience replay pool (each interaction takes ≤50ms); when the number of samples in the replay pool is ≥10,000, training begins: 32 samples are sampled from the pool and the loss L is calculated; the main network parameters are updated through backpropagation (gradient clipping threshold = 1.0 to prevent gradient explosion); the target network is synchronized every 1,000 steps, and a model snapshot is saved every 10,000 steps.

[0073] Convergence criteria: Reward stability: Average total reward fluctuation over 100 consecutive episodes ≤ 5%; Integration quality: Integration image SSIM ≥ 0.92 and defect retention rate ≥ 95% on the test set (K=4 / 6 / 8); Policy consistency: Under the same initial conditions, the overlap rate of action sequences output over 10 consecutive times ≥ 80%.

[0074] Step 2: Define the state space and feature space for image integration;

[0075] The mathematical expression for the state space of image integration is: ; For image acquisition hardware status, For the target state of the scene, For integrating process status; image acquisition hardware status The core variables include camera intrinsic parameters (3K dimensions (3 parameters per point)), camera extrinsic parameters (6K dimensions (6 parameters per point)), and hardware dynamic disturbances (2K dimensions (2 parameters per point)). The camera intrinsic parameters include focal length (500-2000 pixels), principal point coordinates (image center ± 50 pixels), and distortion coefficients (…). Camera extrinsic parameters include rotation angle (±30° around the X / Y / Z axes, expressed in Euler angles) and translation vector (0-500mm, 3D coordinates); hardware dynamic interference includes noise intensity at each acquisition point (local variance calculation, 0-20 grayscale value) and exposure time (1-100ms); scene target state. The core variables include target geometric attributes (2+2M dimensions (M is the number of defects, with 2 fixed key dimensions)) and target material and lighting (1+K dimensions (1 global reflectance, 1 lighting deviation per point)); target geometric attributes include key dimensions of mechanical parts (e.g., shaft diameter φ20±0.02mm → pixel scale 200±2 pixels) and defect locations (coordinates (x,y)); target material and lighting include material reflectance (0.3-0.8, distinguishing between metal and plastic) and local lighting deviation; integration process status. The core variables include integration stage markers (1D), intermediate result quality (2D), and collection point correlation (…). The integration stage markers include preprocessing=0, calibration=1, fusion=2, and termination=3 (discrete values); the intermediate result quality includes the SSIM deviation between the intermediate image and the standard image (0-0.1) and the defect retention rate (0.8-1.0); the acquisition point correlation includes the viewpoint overlap between any two points; Example (when K=4): total dimension of state space = (3×4+6×4+2×4)+(2+2×1)+(1+2+4×3 / 2)=44+4+9=57 dimensions;

[0076] The feature space of image integration includes an image feature subspace (fixed dimension, independent of K) and a scene state feature subspace (dynamically changes with K).

[0077] The image feature subspace includes local matching features and global texture features;

[0078] Local matching features: Extract the top-200 SIFT feature points (128 dimensions / point) from each image acquisition point, and cluster them into two center vectors (representing the two core features of "part edge" and "defect area") using K-means. After concatenation, the dimension is 2×128=256.

[0079] Global texture features: Calculate the gray-level co-occurrence matrix for each image at each acquisition point (distance d=1,2,3,4, angle 0°, 45°, 90°, 135°), extract four types of statistics: "contrast, energy, entropy, and correlation", and take the mean vector of K images (to eliminate single-point interference). Dimension = 16 dimensions (4 distances × 4 angles → 4 statistics).

[0080] Total dimensions of the image feature subspace: 256 + 16 = 272 dimensions.

[0081] The scene state feature subspace includes hardware interference features, inter-point correlation features, and process quality features;

[0082] Hardware interference characteristics: noise intensity at each acquisition point (local variance calculation, 1 dimension / point), average exposure time (1 dimension global), dimension = K+1 dimension;

[0083] Inter-point association features: the overlap of viewpoints between any two points (feature matching rate, 1 dimension / point pair), dimension = K×(K-1) / 2 dimensions;

[0084] Process quality characteristics: SSIM deviation of intermediate results (1-dimensional), integration stage markers (1-dimensional, one-hot encoding: [1,0,0] represents preprocessing), defect retention rate (1-dimensional), dimension = 3-dimensional;

[0085] The total dimension of the scene state feature subspace is: (K+1)+K×(K-1) / 2+3=K+K(K-1) / 2+4 dimensions.

[0086] Step 3: Generate a Pareto optimal integration strategy set based on the NSGA-II algorithm and execute the optimal integration strategy;

[0087] The Pareto optimal integration strategy set is generated based on the NSGA-II algorithm, as follows: Initialize the population: Generate a candidate integration strategy set; Population size: Set to 100 individuals; Continuous parameters: Uniformly randomized sampling within the value range; Discrete parameters: Randomly selected according to the document action space probability; Constraint verification: Ensure that the individual parameters meet the scenario constraints; Non-dominated sorting (hierarchical selection of high-quality strategies): For each individual in the population, its non-dominated relationship is determined according to three objective functions: If individual A is not inferior to individual B on all objectives, and is superior to B on at least one objective, then A dominates B; Divide the population into different "non-dominated levels" (Front): Front1 is the optimal level without any individual domination, Front2 is the level dominated only by Front1, and so on, to ensure that integration strategies with strong non-dominance are retained first. Crowding degree calculation (to ensure solution set diversity): For each individual within the Front, calculate its "crowding degree" in three target dimensions (measuring the distance between the individual and its neighbors to avoid solution set clustering): For each target dimension, sort the individuals within the Front according to the target value; calculate the distance between the individual and its front and rear neighbors in that dimension, and set the crowding degree of boundary individuals (maximum / minimum target value) to infinity (preferably retained); sum the distances in the three dimensions to obtain the individual crowding degree, and the larger the value, the stronger the "representativeness" of the individual in the solution set.

[0088] Genetic operations (offspring generation strategy, iterative optimization):

[0089] Selection: "Tournament selection" was adopted, and 3 individuals were randomly selected from the parent population, with priority given to individuals with higher front levels and greater crowding, for a total of 100 individuals as parents;

[0090] Continuous parameters: Use simulated binary crossover (SBX) with a crossover probability of 0.8 to ensure that the offspring parameters are within a reasonable range;

[0091] Discrete parameters (matching strategy, threshold): Use single-point crossover to exchange discrete parameter fragments of the parent individuals (e.g., parent 1 strategy = ORB, parent 2 strategy = SIFT, the offspring may be ORB).

[0092] Continuous parameters: Use polynomial mutation with a mutation probability of 0.1, and adjust the parameters slightly.

[0093] Discrete parameters: randomly replaced with other valid values;

[0094] Convergence condition: After 50 iterations, or after 5 consecutive generations, the individual target value fluctuation of Front1 is ≤3% (refer to the document "Reinforcement Learning Convergence Judgment Logic").

[0095] Optimal set selection: After the iteration, the Front1 (non-dominated level optimal) of the final population is taken as the "Pareto optimal integrated strategy set" to ensure that the strategies in the set do not dominate each other on the three objectives.

[0096] The optimal integration strategy is implemented as follows: Based on the current integration stage, the values ​​of sharpness weight, consistency weight, and integrity weight are determined. Based on these values, the optimal integration strategy set for multi-point image integration reward function is then determined. The solution with the largest value is marked as the optimal integration strategy, and that optimal integration strategy is executed.

[0097] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.

[0098] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0099] It should be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0100] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0103] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image data integration method based on multi-point image acquisition, characterized in that, The steps are as follows: Step 1: Construct an image integration reinforcement learning training framework adapted to multi-point image acquisition scenarios; A reinforcement learning training framework for image integration adapted to multi-point image acquisition scenarios is constructed as follows: First, a digital twin scenario model for multi-point acquisition scenarios is constructed, a multi-point image integration agent is constructed, a state transition and reward feedback mechanism is set, an experience playback mechanism is built, and training optimization and parameter update logic is set. Construct a multi-point image integration agent as follows: clarify the input dimension and feature source of the multi-point image integration agent, design a network structure that adapts to the image data integration task, define the action space that matches the integration sub-task, and set the training parameters. Clearly define the input dimensions and feature sources of the multi-point image integration agent: the input includes two categories: image features and scene state. The total dimension is dynamically adjusted according to the number of collection points K to ensure coverage of image content attributes and dynamic scene constraints. Design a network structure adapted to image data integration tasks, including: Input layer: image feature branch and scene state branch; Feature extraction layer: Image branch: 3 CNN layers + 1 global pooling layer; Scene branch: 3 fully connected layers, outputting 32 dimensions; Cross-attention fusion layer: Employs a multi-head attention mechanism to calculate the mutual attention weights between image features and scene states, outputting 256-dimensional fused features; Decision layer: Value stream: 2 fully connected layers, outputting state value V(s); Advantage stream: 2 fully connected layers, outputting action advantage A(s,a); Integrated output: Q(s,a)=V(s)+A(s,a)-mean(A(s,a)), where mean(A(s,a)) is the average value of the action advantage; Define the action space matching the fusion subtasks: map each action to an integer index from 0 to 31, and decompose the subtasks into preprocessing subtasks, calibration subtasks, and fusion subtasks; the action type of the preprocessing subtasks: adaptive denoising parameter adjustment; the number of actions in the preprocessing subtasks: 10; the specific action definition of the preprocessing subtasks: Gaussian filter standard deviation. The calibration subtask's action type is a combination of feature matching strategy and threshold. The number of actions in the calibration subtask is 12, and the specific action definition is: matching strategy × matching threshold, indexed 10-21. The fusion subtask's action type is dynamic weight allocation. The number of actions in the fusion subtask is 10. The specific action definition is: multi-image fusion weight coefficient α∈{0.1,0.2,...,1.0}, where α is the weight of the main acquisition point, and the remaining points are allocated according to (1-α) / (K-1). The main acquisition point refers to the acquisition point that is dynamically selected based on the current scene state and has the highest priority in contributing to the final integration result. Training parameters were set as follows: Learning rate: initial η = 0.001, using exponential decay; Experience replay batch: 32 samples / batch; Discount factor γ = 0.95; Target network update frequency: synchronize main network parameters every 1000 steps; Exploration rate... Initially 1.0, linearly decaying to 0.1, balancing exploration and utilization; The state transition mechanism is as follows: Action parsing and execution: Action indices 0-31 output by the agent are parsed into specific operation instructions by the environment module and executed in three categories: Preprocessing actions 0-9: Perform Gaussian filtering on the images of K acquisition points, with the filtering intensity linearly related to the action index; Calibration Action 10-21: Parse as a combination of "matching strategy + threshold", call the corresponding algorithm to calculate the transformation matrix between K images, and complete the viewpoint alignment; Fusion Actions 22-31: Parse as fusion weight α, calculate the weighted fused image according to the rule of weight α of the main acquisition points + equal distribution (1-α) of the remaining points; State update rules: New state It consists of three parts, and the update formula is as follows: ,in: Images of K acquisition points after motion processing; Updated image features; The updated scene state includes: noise intensity: recalculated based on local variance after denoising; viewpoint overlap: recalculated based on feature matching rate after calibration; integration stage marker: automatically incremented. The reward and feedback mechanism is as follows: Define clarity reward Consistency rewards Completeness reward Further define the multi-point image integration reward function ;in, For clarity weight, For consistency weights, Integrity weight; , , Dynamically changing with the current integration phase ; Clarity Bonus ;in, The total number of valid pixels, i.e., the set of valid pixels. The number of pixels contained; For the effective pixel set; For regional weighting coefficients, ; For pixels Edge gradient strength; An adaptive threshold for noise; Consistency rewards ; q represents a single pixel, It is the set of pixels in all overlapping regions of multiple views; Pixel weight coefficients are the weights assigned to the differences between different pixel regions, using a simplified binary weighting: ; It represents the absolute value of the difference in grayscale values ​​of the same pixel q in two images from different viewpoints; that is... ,in, , These are the grayscale values ​​of pixel q in viewpoints a and b, respectively. The grayscale difference threshold; This represents the total number of pixels in the globally overlapping region, i.e. ; Completeness reward M represents the total number of original defects; m represents the defect index. The importance weight of the defect, i.e., the importance coefficient of the m-th defect, is determined using a simplified binary weighting. ; Indicates the crossover ratio of defect regions, and the original defect. Defects matched with the integrated image The regional overlap is calculated using a simplified formula as follows: ; This represents the original defect, specifically the m-th defect marked in the original scene. This represents the matching defects in the integrated image; in the image integrated by the agent, the defects are compared with the original defects. Corresponding defects; Step 2: Define the state space and feature space for image integration; Step 3: Generate a Pareto optimal integration strategy set based on the NSGA-II algorithm and execute the optimal integration strategy.

2. The image data integration method based on multi-point image acquisition according to claim 1, characterized in that, The digital twin scene model for multi-point acquisition scenarios is constructed as follows: the actual scene of multi-point image acquisition is determined, a digital twin model of the image acquisition hardware is established, a standard image is constructed based on the physical properties of the scene targets in the actual scene, dynamic rules of scene physical interference are defined, and interference parameters are superimposed on the standard image according to the imaging logic of the real device to generate the original acquisition data.

3. The image data integration method based on multi-point image acquisition according to claim 1, characterized in that, The mathematical expression for the state space of image integration is: ; For image acquisition hardware status, For the target state of the scene, This indicates the status of the integration process.

4. The image data integration method based on multi-point image acquisition according to claim 1, characterized in that, The optimal integration strategy is implemented as follows: Based on the current integration stage, the values ​​of sharpness weight, consistency weight, and integrity weight are determined. Based on these values, the optimal integration strategy set for multi-point image integration reward function is then determined. The solution with the largest value is marked as the optimal integration strategy, and that optimal integration strategy is executed.

Citation Information

Patent Citations

  • Defect detection method and device, equipment and storage medium

    CN116823758A

  • Method of lightweight simultaneous localization and mapping performed on a real-time computing and battery operated wheeled device

    US20220187841A1