Multi-modal image recognition remote control method and system for high-precision operation of robot
By employing multimodal image recognition and a spiral layer-by-layer removal strategy, the problems of inaccurate localization of internal defects in composite materials and thermal damage during laser removal were solved, achieving high-precision defect repair.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are inaccurate in locating defects inside composite materials, and the laser removal process is prone to thermal damage, making it difficult to achieve high-precision repair.
Multimodal image recognition methods are used to acquire multi-physics field stereo imaging data. A three-dimensional defect localization model is generated by multimodal feature extraction and fusion. An adaptive laser removal is performed using a spiral layer-by-layer removal strategy and a thermal accumulation prediction model.
It enables precise location and high-precision removal of internal defects in composite materials, avoiding interlayer delamination damage and ensuring the structural integrity of the material.
Smart Images

Figure CN121805602A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of composite material testing and robot control technology, and more specifically, to a multimodal image recognition remote control method and system for high-precision robot operation. Background Technology
[0002] In the field of aerospace composite component manufacturing, carbon fiber / epoxy resin composites are widely used due to their excellent specific strength and specific stiffness. However, during the manufacturing process, tiny metal inclusions and conductive particles may become embedded within multilayer composite materials, severely affecting the material's electromagnetic compatibility and structural integrity.
[0003] Existing technologies typically employ single-modal imaging methods for defect detection, such as millimeter-wave imaging or ultrasonic testing, and use laser removal processes with fixed parameters for defect repair.
[0004] However, existing technologies have the following technical problems: complex interlayer interfaces and fiber orientation changes exist inside multilayer composite materials, and a single imaging mode is difficult to accurately distinguish between metal inclusions and carbon fiber structures, resulting in inaccurate defect localization; fixed-parameter laser removal processes cannot be adaptively adjusted according to material depth and thermal accumulation state, and are prone to causing interlayer delamination damage during the removal process, thus destroying the integrity of the material structure. Summary of the Invention
[0005] The present invention aims to provide a multimodal image recognition remote control method and system for high-precision robot operation, in order to solve the technical problems of difficulty in accurately locating internal defects in multilayer composite materials and the easy thermal damage caused by the removal process in related technologies.
[0006] In a first aspect, the present invention provides a multimodal image recognition remote control method for high-precision robot operation, comprising the following steps: Multimodal imaging data of multilayer composite materials are acquired to generate a multiphysics stereo imaging dataset. The multimodal imaging data includes millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals in multiple polarization directions, and multi-angle polarization scattering image sequences. Based on the multi-physics stereo imaging dataset, multimodal feature extraction and fusion are performed to output a three-dimensional defect localization model. The multimodal feature extraction includes tomographic reconstruction processing of millimeter-wave tomographic projection data to extract millimeter-wave feature maps, polarization decomposition of millimeter-wave transmission signals in multiple polarization directions to obtain polarization feature distribution maps, and Mueller matrix decomposition of multi-angle polarization scattering image sequences to obtain optical property maps. The fusion is achieved through a multimodal deep fusion network. The three-dimensional defect localization model includes a defect segmentation mask and a depth regression vector. Based on the three-dimensional defect localization model, a spiral layer-by-layer removal strategy is adopted to generate a precision removal trajectory sequence with depth adaptation. In the spiral layer-by-layer removal strategy, the process proceeds spirally from the defect edge to the center, and the removal depth of each layer is adaptively adjusted according to the material layer thickness and the defect depth. Based on the precise removal trajectory sequence, remote control commands for the robot are output. The control commands include the position coordinates, attitude angle, movement speed, laser power, pulse frequency, and spot diameter parameters of the robot's end effector. In the process of generating a depth-adaptive precision removal trajectory sequence, the temperature field distribution during the continuous removal process is predicted by a thermal accumulation prediction model based on the material's thermal conductivity, specific heat capacity, and interlayer thermal resistance parameters. When the predicted temperature exceeds a preset ratio of the material's glass transition temperature, the laser power and scanning speed are automatically adjusted or a cooling waiting time is inserted.
[0007] In a preferred embodiment, the process of acquiring the millimeter-wave tomographic projection data includes: Millimeter-wave tomographic projection data were acquired at angular intervals of 1° to 5° within the incident angle range of 0° to 180°. The phase of the millimeter-wave tomographic projection data at each incident angle is recorded synchronously. The spatial distribution of the relative permittivity inside the material is calculated based on the radial gradient of the phase. The calculation process is as follows: the mathematical relationship between the phase gradient and the relative permittivity is established based on the electromagnetic wave propagation theory. The phase gradient is related to the millimeter-wave wavelength and the square root of the relative permittivity. The spatial distribution of the relative permittivity is obtained through inversion calculation. The epoxy resin matrix, carbon fiber, and metal inclusions were distinguished based on the significant differences in the spatial distribution of their relative permittivity. The permittivity of the epoxy resin matrix ranged from 3 to 4, that of the carbon fiber ranged from 10 to 15, and that of the metal inclusions was greater than 100.
[0008] In a preferred embodiment, the step of polarization decomposition of the millimeter-wave transmission signal in multiple polarization directions to obtain a polarization feature distribution map includes: Construct a polarization scattering matrix that includes horizontal transmission and horizontal reception, vertical transmission and vertical reception, and cross-polarization scattering coefficients; The polarization scattering matrix is decomposed into eigenvalues to obtain eigenvalues and corresponding eigenvectors. By analyzing the magnitude relationship of eigenvalues and the direction of eigenvectors, the electromagnetic responses of different layers are separated. When the first eigenvalue is much larger than the second eigenvalue, it indicates that the layer is the dominant scattering layer. When the two eigenvalues are close, it indicates that the layer has multiple scattering. The electromagnetic response output of each depth layer is used as a polarization feature distribution map, where the main scattering component corresponds to the strong reflection feature of the metal inclusions, and the cross scattering component corresponds to the anisotropic scattering feature of the carbon fiber.
[0009] In a preferred embodiment, the step of performing Mueller matrix decomposition on the multi-angle polarization scattering image sequence to obtain the optical property map includes: A 4×4 Mueller matrix was established by measuring the light intensity changes under different combinations of incident and outgoing polarization states. The Mueller matrix is decomposed into the product of the depolarization matrix, the optical rotation matrix, and the birefringence matrix using the Lu-Chipman decomposition method; The depolarization degree parameter, linear birefringence, circular birefringence, and fiber orientation angle distribution are extracted from the decomposition results. The depolarization degree is obtained by calculation through the depolarization matrix, the linear birefringence and circular birefringence are obtained by Euler angle decomposition of the rotation matrix, and the fiber orientation angle is determined by the birefringence matrix.
[0010] In a preferred embodiment, the multimodal deep fusion network includes: Three parallel coding branches process millimeter-wave feature maps, polarization feature distribution maps, and optical property maps respectively to obtain three modal features. Each coding branch contains four convolutional blocks. The cross-modal attention mechanism achieves feature association by calculating the similarity matrix between features of different modalities. This includes mapping each modal feature to a query vector, a key vector, and a value vector, calculating attention weights based on the query vector and key vector, and fusing the value vector to obtain the fused feature. An adaptive weight allocation strategy is used to generate normalized fusion weight coefficients for each modality and depth layer through a learnable weight generation network. The three fusion features of each layer are weighted and summed using the fusion weight coefficients to obtain the final fusion features. The final fusion features of each depth layer are the three-dimensional defect localization model. The depth sensing module generates a confidence map for each depth layer by analyzing the signal attenuation patterns and scattering characteristics at different depths. The depth perception module receives intermediate features from each coding branch at different depth layers as input; The deep perception module includes a deep feature extraction subnetwork and a confidence generation subnetwork. The deep feature extraction subnetwork uses an LSTM structure to process cross-layer feature sequences, combining features from four deep layers. The LSTM outputs a sequence of hidden states, which are then fed into the LSTM. ; Confidence-based subnetworks generate the hidden states of each depth layer Two convolutional layers are applied, and finally a confidence map is generated using the sigmoid activation function. Its value ranges from 0 to 1, representing the reliability weight of the feature at that depth level; The output of the depth perception module is a sequence of confidence maps for four depth layers. The confidence map is multiplied element-wise with the fusion features of the corresponding depth layer to obtain the weighted fusion features. ,in This indicates an element-wise multiplication operation, and the weighted fusion features are input into the subsequent decoding branch for final defect localization and classification.
[0011] In a preferred embodiment, the spiral layer-by-layer removal strategy includes: Extract the defect boundary contour and obtain the outermost boundary contour point set through morphological erosion operation; The number of depth layers of the defect is determined based on the depth regression vector, and the defect depth range is divided equally. The thickness of each layer meets the constraint that it does not exceed the set proportion of the thickness of a single layer of material. For each layer, a spiral trajectory is generated from the boundary profile at a constant angular velocity and radial contraction velocity. The spiral trajectory parameters include the maximum radius, the minimum radius, and the spiral duration of a single layer. The polar coordinates of the spiral trajectory are converted to Cartesian coordinates to generate the trajectory point sequence of the current layer. For each trajectory point in the current layer trajectory point sequence, based on its spatial position and defect depth information, the thermal accumulation prediction model is called to calculate the process parameters including the optimal laser power and scanning speed. The Cartesian coordinates, depth and process parameters are combined to form the complete trajectory point parameters. By concatenating the trajectory point sequences of each layer in depth order, a depth-adaptive, precisely removed trajectory sequence is obtained.
[0012] In a preferred embodiment, the thermal accumulation prediction model is a multilayer perceptron structure, comprising: The input layer receives a feature vector containing process parameters at the current time and historical time, as well as material physical parameters. The historical time process parameters include the laser power sequence and scanning speed sequence of the most recent multiple time steps. Four hidden layers are used, employing the ReLU activation function and adding a Dropout layer to prevent overfitting; The output layer outputs the temperature field distribution of a three-dimensional spatial mesh; The heat accumulation prediction model is trained through supervised learning. The training dataset is generated by finite element thermo-coupling simulation software, and the loss function is a weighted combination of mean square error loss and relative error loss.
[0013] In a preferred embodiment, the multimodal image recognition remote control method for high-precision robot operation, as described in 1, further includes: During the removal operation, real-time multimodal imaging data is acquired simultaneously; The removal effect and material integrity are evaluated by comparing the changes in imaging features before and after removal. When an abnormal signal is detected, the laser power and scanning speed are automatically adjusted to generate a closed-loop control correction command.
[0014] In a preferred embodiment, the multimodal image recognition remote control method for high-precision robot operation further includes establishing a machine learning-based knowledge graph system: Collect historical removal case data, including defect type, material structure parameters, process parameters, and removal effect; A graph neural network structure is used to mine association patterns. The graph neural network includes three graph convolutional layers and a graph attention mechanism. The network is trained by optimizing the strategy through a combination of semi-supervised learning and reinforcement learning. The reward function comprehensively considers the removal quality score, processing time and the degree of thermal damage. It provides intelligent decision-making suggestions for new removal tasks, enabling adaptive optimization of the spiral layer-by-layer removal strategy.
[0015] Secondly, the present invention provides a multimodal image recognition remote control system for high-precision robot operation, for performing the aforementioned multimodal image recognition remote control method for high-precision robot operation, comprising: The multimodal imaging module is used to acquire multimodal imaging data of multilayer composite materials and generate a multiphysics stereo imaging dataset. The multimodal imaging data includes millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals in multiple polarization directions, and multi-angle polarization scattering image sequences. The feature fusion module is used to perform multimodal feature extraction and fusion based on the multiphysics stereo imaging dataset, and output a three-dimensional defect localization model. The multimodal feature extraction includes tomographic reconstruction processing of millimeter-wave tomographic projection data to extract millimeter-wave feature maps, polarization decomposition of millimeter-wave transmission signals in multiple polarization directions to obtain polarization feature distribution maps, and Mueller matrix decomposition of multi-angle polarization scattering image sequences to obtain optical property maps. The fusion is achieved through a multimodal deep fusion network. The three-dimensional defect localization model includes a defect segmentation mask and a depth regression vector. The trajectory planning module is used to generate a depth-adaptive precision removal trajectory sequence based on the three-dimensional defect location model using a spiral layer-by-layer removal strategy. In the spiral layer-by-layer removal strategy, the process proceeds spirally from the defect edge to the center, and the removal depth of each layer is adaptively adjusted according to the material layer thickness and the defect depth. The control execution module is used to output remote control commands for the robot based on the precise removal trajectory sequence. The control commands include the position coordinates, attitude angle, movement speed, laser power, pulse frequency, and spot diameter parameters of the robot end effector. The real-time monitoring module is used to acquire multimodal imaging data in real time during the removal process and generate closed-loop control correction commands; The knowledge optimization module includes a knowledge graph system based on graph neural networks, which is used to continuously optimize the spiral layer-by-layer removal strategy.
[0016] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: This invention utilizes a multimodal imaging data fusion mechanism, leveraging millimeter-wave tomography to provide depth resolution and polarization scattering imaging to capture the optical properties of materials. By associating different physical field information through a multimodal depth fusion network, it solves the technical problems of blurred imaging within composite materials and difficulty in identifying metal inclusions and carbon fiber structures using a single imaging modality. This achieves the technical effect of accurately distinguishing different material types and precisely locating the three-dimensional position of defects. Through a depth-adaptive spiral layer-by-layer removal strategy and a thermal accumulation prediction model, the laser power and scanning speed are dynamically adjusted based on the material layer thickness, defect depth, and temperature field distribution. This solves the technical problems of uneven energy distribution and thermal damage risk caused by interlayer interfaces during laser removal, achieving the technical effect of ensuring material structural integrity and avoiding interlayer delamination damage. Furthermore, through real-time monitoring feedback and a knowledge graph system based on graph neural networks, closed-loop control dynamic adjustment and continuous optimization of the removal strategy are achieved, solving the technical problem of poor adaptability of fixed-parameter removal methods and achieving high-precision adaptive removal of embedded defects in composite materials. Attached Figure Description
[0017] Figure 1 This is a flowchart of a multimodal image recognition remote control method for high-precision robot operation provided by an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0020] Example like Figure 1As shown in the figure, this embodiment of the invention provides a multimodal image recognition remote control method for high-precision robot operation, including the following steps: Step 100: Acquire multimodal imaging data of multilayer composite materials and generate a multiphysics stereo imaging dataset.
[0021] In this embodiment of the invention, the multimodal imaging data includes millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals with multiple polarization directions, and a multi-angle polarization scattering image sequence. Specifically, millimeter-wave tomographic projection data is acquired at angular intervals of 1° to 5° within an incident angle range of 0° to 180°; the multiple polarization directions include horizontal polarization, vertical polarization, left-handed circular polarization, and right-handed circular polarization; the acquisition angle range of the multi-angle polarization scattering image sequence is -60° to 60°, encompassing four main polarization directions: 0°, 45°, 90°, and 135°.
[0022] In this embodiment of the application, in order to improve the signal-to-noise ratio and depth resolution of the imaging data, the phase at each incident angle is also recorded simultaneously when acquiring millimeter-wave tomographic projection data. The spatial distribution of the relative permittivity inside the material is calculated using the phase gradient, and the contrast of different materials is enhanced by the difference in relative permittivity.
[0023] Furthermore, the method for calculating the spatial distribution of the relative permittivity within a material using phase gradient includes: for the phase in millimeter-wave tomographic projection data... First, calculate the phase gradient along the radial direction. The phase gradient is calculated using the finite difference method, specifically as follows: ,in The radial sampling interval is given; then, according to electromagnetic wave propagation theory, the relationship between the phase gradient and the relative permittivity of the material is: ,in For millimeter wave wavelength, Given the relative permittivity, the spatial distribution of the relative permittivity is obtained through inversion calculation. .
[0024] Furthermore, the millimeter wave wavelength The value is determined based on the operating frequency of the millimeter-wave imaging system. For the detection of internal defects in composite materials, the commonly used millimeter-wave frequency range is 30GHz to 300GHz, corresponding to a wavelength range of 1mm to 10mm.
[0025] Step 200: Based on the multi-physics stereo imaging dataset, perform multi-modal feature extraction and fusion to output a three-dimensional defect localization model.
[0026] In this embodiment of the invention, the multimodal feature extraction includes: performing tomographic reconstruction processing on millimeter-wave tomographic projection data to extract millimeter-wave feature maps; performing polarization decomposition on millimeter-wave transmission signals in multiple polarization directions to obtain polarization feature distribution maps; and performing Mueller matrix decomposition on multi-angle polarization scattering image sequences to obtain optical property maps. The tomographic reconstruction employs a filtered back-projection algorithm, combined with polarization decomposition technology to separate the electromagnetic responses of different layers, generating a depth-resolved tomographic cross-section sequence. Mueller matrix decomposition extracts depolarization degree, birefringence parameters, and polarization preservation features, and calculates the fiber orientation distribution.
[0027] Furthermore, the step of polarization decomposition of the millimeter-wave transmission signal in multiple polarization directions to obtain a polarization feature distribution map includes: for the millimeter-wave transmission signal in multiple polarization directions, including horizontal polarization... Vertical polarization Left-handed circular polarization and right-hand circular polarization First, construct the polarization scattering matrix. ,in The scattering coefficient is for horizontal transmission and horizontal reception. The scattering coefficient for vertical transmission and vertical reception. and The cross-polarization scattering coefficients are then used; then the polarization scattering matrix is... Perform eigenvalue decomposition ,in It is a unitary matrix. It is a diagonal matrix. and For eigenvalues, express The conjugate transpose of ; by analyzing the magnitude relationship of eigenvalues and the direction of eigenvectors, the electromagnetic responses of different layers are separated, when When this indicates that the layer is the dominant scattering layer, when The time indicates that there is multiple scattering in the layer; the polarization decomposition technique outputs the main scattering component and cross scattering component of each depth layer as a polarization feature distribution map. The main scattering component corresponds to the strong reflection characteristics of the metal inclusions, and the cross scattering component corresponds to the anisotropic scattering characteristics of the carbon fiber.
[0028] The aforementioned filtered back-projection algorithm takes millimeter-wave tomographic projection data at different incident angles as input, and the millimeter-wave tomographic projection data is the size... Two-dimensional matrix ,in For the number of projection angles, The value ranges from 0° to 180°, with angular intervals ranging from 1° to 5°. The number of radial sampling points at each projection angle. For radial coordinates.
[0029] Furthermore, the number of radial sampling points and radial sampling interval The method for determining the radial sampling interval is as follows: The radial sampling interval is set based on the maximum detection size of the material and the required spatial resolution. The Nyquist sampling theorem should be satisfied, and the wavelength should be set to one-quarter of the millimeter wave wavelength. Radial sampling interval For detection ranges with a radius The number of radial sampling points in the circular region radial coordinates The range of values is Discrete sampling points are ,in .
[0030] The filtering back projection algorithm first processes the two-dimensional matrix for each incident angle. Perform a one-dimensional Fourier transform to obtain the frequency domain representation. Then, a Ram-Lak filter or a Shepp-Logan filter is applied for high-pass filtering, and the filtered frequency domain data is... ,in Let be the frequency response function of the filter; then, perform an inverse Fourier transform on the filtered frequency domain data to obtain the filtered projection data. Finally, the filtered projection data from all angles are back-projected and accumulated radially to reconstruct the millimeter-wave feature map. The millimeter-wave feature map is of size. The tomographic reconstruction image represents the electromagnetic response distribution across the material's cross-section, where... and This refers to spatial resolution.
[0031] Furthermore, the spatial resolution of the tomographic reconstructed image and The determination method is as follows: based on the actual size of the material detection area and the pixel spacing of the reconstructed image, the pixel spacing of the reconstructed image is set to be equal to the radial sampling interval. They are the same, both being 0.8mm.
[0032] The input to the aforementioned Mueller matrix decomposition is a multi-angle polarization scattering image sequence, which contains scattered light intensity images acquired in four main polarization directions (0°, 45°, 90°, 135°), with an image size of [size missing] for each direction. Light intensity measurements in four polarization directions Constructing Stokes vectors ,in Indicates total light intensity. This represents the difference between the 0° and 90° linear polarization components. This represents the difference between the 45° and 135° linear polarization components; the process of obtaining an optical property map by Mueller matrix decomposition of the multi-angle polarization scattering image sequence includes: establishing a 4×4 Mueller matrix by measuring the light intensity changes under different combinations of incident and outgoing polarization states. The Mueller matrix describes the polarization response characteristics of a material; the elements of the Mueller matrix... The linear equations were obtained by solving a system of linear equations based on multiple sets of polarization measurements; the Mueller matrix was decomposed into a product of three fundamental matrices through extreme decomposition. ,in For the depolarization matrix, For optical rotation matrix, The birefringence matrix is used; the output of the Mueller matrix decomposition includes the depolarization parameter. Linear birefringence Circular birefringence The polarization retention coefficient and fiber orientation angle distribution diagram reflect the optical anisotropy characteristics of the material and the fiber arrangement direction.
[0033] Furthermore, the specific algorithm for the extreme decomposition of the Mueller matrix adopts the Lu-Chipman decomposition method, which decomposes the Mueller matrix... Decomposed into depolarization matrix Optical rotation matrix and birefringence matrix The product, specifically the steps are as follows: First, from the Mueller matrix... Extracting a 3×3 submatrix The 3×3 submatrix is composed of The second to fourth rows and the second to fourth columns constitute the submatrix; then the submatrix is processed. Perform symmetric polar decomposition , where the symmetric matrix The rotation matrix is obtained through the square root operation. Indicates the optical rotation properties of a material; degree of depolarization. By symmetric matrix The eigenvalues are calculated as follows: ,in for The three eigenvalues, The first element of the Mueller matrix; linear birefringence Circular birefringence By rotation matrix Obtained by Euler angle decomposition, the rotation matrix can be expressed as ,in The rotation angle corresponding to the x-axis represents linear birefringence. The rotation angle corresponding to the z-axis represents circular birefringence; the fiber orientation angle is determined by the fast axis direction angle calculated based on the birefringence matrix. ,in , , , Birefringence matrix The elements in the corresponding rows and columns, the fast axis direction angle reflects the main arrangement direction of carbon fibers in the material.
[0034] In this embodiment of the invention, multimodal fusion refers to inputting the extracted millimeter-wave feature map, polarization feature distribution and optical property map into a multimodal deep fusion network. The multimodal deep fusion network contains three parallel encoding branches that process different modal data respectively. It associates features at different depths through a cross-modal attention mechanism. The fusion layer adopts an adaptive weight allocation strategy to dynamically adjust the fusion weights according to the reliability of each modality at different depths.
[0035] The aforementioned multimodal deep fusion network's input layer includes three independent input ports, each receiving signals of size [size missing]. The millimeter-wave feature map, with a size of The polarization characteristic distribution map and size are The optical property spectrum, in which and For spatial resolution, , and These represent the number of feature channels for each modality; before inputting into the multimodal deep fusion network, the feature maps of the three modalities are normalized based on the range.
[0036] The three parallel encoding branches all employ a convolutional neural network structure, with each branch containing four convolutional blocks. Each convolutional block consists of two convolutional layers, a batch normalization layer, and a ReLU activation function. The cross-modal attention mechanism achieves feature association by calculating the similarity matrix between features of different modalities. Specifically, for the first... The three modal characteristics of the layer , , First through Convolution maps them to the same channel dimension , obtain the query vector Key vector Sum value vector ,in Then, attention weights are calculated based on the query vector and key vector. Finally, the fused feature is obtained by fusing the value vectors. The adaptive weight allocation strategy generates fused weight coefficients for each modality and depth layer through a learnable weight generation network. The weight generation network consists of two fully connected layers, with the input being the global average pooling vector of each modality feature and the embedding vector of the current depth index, and the output being three normalized fused weight coefficients. , , ,satisfy The final fusion feature is .
[0037] Furthermore, the method for generating the current depth index embedding vector is as follows: for the four depth layers of the encoding branch, the depth index... These correspond to the output layers of the four convolutional blocks in the network; the depth index embedding uses a sinusoidal positional encoding method for each depth index. Generation dimension is Embedded vector , where the embedded dimension ; the first embedding vector The formula for calculating each element is: when When it is even, ;when When it is an odd number, The sinusoidal positional encoding method ensures the uniqueness and distinguishability of embedding vectors at different depths, facilitating the weight generation network to adjust the fusion weights based on depth information. The embedding vector is concatenated with the global average pooling vector of each modality feature, and then used as the input to the weight generation network. The concatenated vector has a dimension of [missing information]. ,in The dimension of the global average pooling vector for each modality.
[0038] The output layer of the aforementioned multimodal deep fusion network contains two parallel decoding branches. The first decoding branch upsamples the fused features to the original resolution using four deconvolution blocks, with an output size of [missing information]. The defect segmentation mask, where Given the number of defect categories (including metal inclusions, carbon fiber, and background), a softmax activation function is used to generate the probability distribution of each pixel belonging to each category; the second decoding branch processes the global average pooling vector of the fused features through three fully connected layers, with an output size of [size missing]. The depth regression vector represents the continuous positional distribution of defects along the material depth direction, where... The depth is the number of discretization layers. The output values are normalized to between 0 and 1 using the sigmoid activation function.
[0039] The decoding process of the aforementioned multimodal deep fusion network output is as follows: For the defect segmentation mask, the class probability distribution at each pixel location is decoded through the argmax operation, and the class with the highest probability is selected as the final class label for that pixel, generating a size of... The discrete category label map; for the depth regression vector, the output values normalized to between 0 and 1 are mapped back to the actual depth values, and the mapping formula is: ,in For the first The normalized output value of the layer, and These are the minimum and maximum values of the material depth, respectively. This represents the actual depth value after decoding. By filtering layers in the depth regression vector whose values are greater than a threshold (set to 0.5), the depth layer location of the defect is extracted, forming the defect depth range. .
[0040] The aforementioned multimodal deep fusion network adopts a supervised learning training mode, and the training dataset contains multimodal imaging samples labeled with defect location, category, and depth information.
[0041] The loss function of the aforementioned multimodal deep fusion network adopts a multi-task learning framework, and the total loss function is: ,in For defect segmentation loss, a weighted combination of Dice loss and cross-entropy loss is used. Dice loss is defined as ,in To predict probabilities, This is a real label; For deep regression loss, smoothed L1 loss is used. ,in To predict depth values, This represents the actual depth value. To mitigate overfitting, L2 weight decay is used as the regularization loss; the weight coefficients are set to... , , .
[0042] In this embodiment of the application, in order to improve the accuracy of defect identification, a depth perception module is introduced into the multimodal deep fusion network. The depth perception module generates a confidence map for each depth layer by analyzing the signal attenuation law and scattering characteristic changes at different depths, guiding the network to focus on feature fusion in high confidence areas, thereby accurately distinguishing metal inclusions and carbon fiber structures.
[0043] The aforementioned deep perception module receives intermediate features from different depth layers of each coding branch as input, with the input feature size being... ,in The size is , This indicates different depth levels.
[0044] Furthermore, the number of channels for each depth layer feature Corresponding to the number of output channels of the convolutional blocks in the coding branch, according to the network structure definition, the number of channels of the four convolutional blocks are 64, 128, 256, and 512, respectively.
[0045] The deep perception module includes a deep feature extraction subnetwork and a confidence generation subnetwork. The deep feature extraction subnetwork uses an LSTM structure to process cross-layer feature sequences. Specifically, it processes the intermediate features of four deep layers. The input is sequentially fed into an LSTM, with the hidden state dimension set to 256. A timing gating mechanism is used to capture the signal attenuation pattern. The output hidden state sequence of the LSTM is as follows: ; Confidence-based subnetworks generate the hidden states of each depth layer Two convolutional operations are applied respectively, with the first convolutional kernel size being [size missing]. The number of channels is 128, and the size of the second convolutional kernel is [size missing]. The number of channels is 1, and finally a confidence map is generated using the sigmoid activation function. Its value ranges from 0 to 1, representing the reliability weight of the feature at that depth level.
[0046] The output of the aforementioned depth perception module is a sequence of confidence maps for four depth layers. These confidence maps are multiplied element-wise with the fusion features of the corresponding depth layers to obtain weighted fusion features. ,in This indicates an element-wise multiplication operation, and the weighted fusion features are input into the subsequent decoding branch for final defect localization and classification.
[0047] The aforementioned deep perception module employs an end-to-end joint training mode, training simultaneously with the multimodal deep fusion network. The hidden states of the LSTM are initialized to zero vectors, and the training process uses the backpropagation algorithm, with gradients propagated back to the LSTM and feature extraction layers via the confidence map. To improve the accuracy of confidence prediction, an auxiliary supervised loss is introduced. Where BCE is the binary cross-entropy loss, The ground truth confidence value is generated based on the true defect depth labels, set to 1 for depth layers containing defects and 0 for depth layers not containing defects; the auxiliary supervision loss is added to the total loss function with a weight of 0.2, forming... .
[0048] Furthermore, the true confidence level The generation process is as follows: First, obtain the true defect depth annotation information of the training samples. The true defect depth annotation information is in the form of depth range. The form is given, indicating the start and end positions of the defect in the material depth direction; then, based on the downsampling factor of the network coding branch, the first... The receptive field depth range corresponding to the layer features The first layer corresponds to the material surface to the depth. The area, the second layer corresponds to the depth. to The area, the third layer corresponds to the depth to The area, the 4th layer corresponds to the depth The depth boundary of each layer in the region down to the bottom layer of the material. , , The total thickness of the material is determined by dividing it according to the downsampling ratio.
[0049] Furthermore, the depth boundary points of each layer , , The specific calculation method is as follows: Let the total thickness of the material be... After four convolutional blocks, the network coding branch performs downsampling once after the second and fourth convolutional blocks, with cumulative downsampling ratios of [missing values]. (No downsampling in the first layer) (Second layer downsampling once) (No further downsampling was performed on the third layer) (Fourth layer downsampling twice); the material depth direction is divided according to the receptive field size, and the receptive field size is proportional to the downsampling factor. The receptive field depth of the layer accounts for the proportion of The formula for calculating the depth boundary point is: , , The depth range of the first layer of receptive field is... The second layer is The third layer is The 4th floor is The depth boundary point division method ensures that the receptive field coverage depth of each layer matches the spatial resolution of the features of each depth layer. Shallow features correspond to the surface information of the material, and deep features correspond to the deep information of the material.
[0050] Next, the intersection of the true defect depth range and the receptive field depth range of each layer is calculated. If the... The receptive field depth of the layer overlaps with the defect depth, i.e., the condition is met. If the layer contains defects, then the confidence level is true. Set as an all-1 matrix, with matrix size equal to the first... Spatial dimensions of layer features If there is no intersection, then the layer does not contain defects, and the confidence level is true. Set to an all-zero matrix; for boundary layers (i.e., when the defect depth range is close to the layer boundary point), if the proportion of the intersection depth to the depth range of the layer is less than the threshold (set to 0.2), then the layer is also considered to not contain defects, and the confidence ground value is set to 0 to avoid the interference of boundary ambiguity on training.
[0051] Furthermore, the boundary layer judgment threshold is set to 0.2 based on the following: the intersection depth ratio is defined as... , indicating that the defect is in the first The proportion within the layer depth range; through analysis of boundary layer samples in the training set, when the intersection depth ratio is less than 0.2, the signal response characteristics of this layer are mainly dominated by non-defect regions, and the defect signal is submerged in background noise, resulting in unstable confidence prediction; statistical experiments show that when At that time, the defect signal significantly contributes to the features of this layer, and the confidence prediction accuracy can reach over 85%. When the confidence prediction accuracy drops below 60%, 0.2 is used as the boundary layer judgment threshold, which ensures the reliability of the true confidence value and avoids the negative impact of weak signal layers on network training.
[0052] Step 300: Generate a depth-adaptive precision removal trajectory sequence based on the 3D defect localization model.
[0053] It should be noted that generating a precise removal trajectory sequence refers to calculating the optimal removal path using a path planning algorithm based on the three-dimensional spatial distribution of defects and the interlayer structural characteristics of the material. The path planning adopts a spiral layer-by-layer removal strategy, advancing spirally from the edge of the defect towards the center. The removal depth of each layer is adaptively adjusted according to the material layer thickness and defect depth to ensure that the removal depth in a single operation does not exceed a set proportion of the single layer thickness (preferably 30% in this embodiment of the invention), thus avoiding interlayer separation.
[0054] The aforementioned path planning algorithm employs a spiral layer-by-layer removal strategy, the input of which is the defect segmentation mask output by the 3D defect localization model. and deep regression vector and material interlayer structural parameters, including the thickness of each layer. and the depth and location of the defect; the output of the spiral layer-by-layer removal strategy is a precise removal trajectory sequence. Each trajectory point Includes spatial coordinates Laser power and scanning speed .
[0055] The aforementioned spiral layer-by-layer removal strategy includes the following steps: Step S1: Extract the defect boundary contour and perform defect segmentation masking. Apply morphological etching operation, and set the etching core size to... Pixels, to obtain the eroded mask Then calculate the difference between the original mask and the eroded mask. Extract the outermost boundary contour point set of the defect. ; Step S2: Based on the deep regression vector Determine the depth and number of layers of the defect , range of defect depth Divided into equal parts Layers, each layer thickness At the same time, check whether the thickness of each layer meets the constraints. ,in The thickness of a single layer of material where the defect is located; if this requirement is not met, the number of layers is increased. Until the constraints are satisfied.
[0056] Furthermore, the layer increase process employs an iterative calculation method, with the initial layer number set to... ,in This means rounding up to the nearest integer and then calculating the thickness of each layer. And verify the constraints. Does it satisfy the condition? If not, then let The calculation is then repeated until the constraints are met; to prevent excessive processing time due to too many layers, an upper limit constraint on the number of layers is set. ,in That is, the minimum layer thickness removed in a single step is not less than 0.05 mm. If during the iteration process... Exceed The constraints still cannot be met. Then take The defective areas are marked as deep defects that require segmented processing.
[0057] Step S3: For the first layer( Remove ) and initialize the current layer boundary contour as ,like but ,otherwise This is the final contour of the previous spiral trajectory; Step S4: Define the boundary profile of the current layer. Sort the points clockwise and select the starting point from the set of contour points. The starting point selection strategy is to select the contour point closest to the end point of the previous layer. If it is the first layer, the point with the smallest x-coordinate in the contour point set is selected. Step S5: Along the boundary contour With a constant angular velocity and radial contraction speed Generate a spiral trajectory; the spiral parametric equation is: , ,in For time parameters, The maximum radius of the current boundary profile. The minimum radius to shrink to the center. This refers to the duration of a single-layer spiral.
[0058] Furthermore, the maximum radius of the spiral trajectory parameter , minimum radius and single-layer spiral duration The method for determining the maximum radius is as follows: By calculating the current layer boundary contour point set The maximum distance to the centroid of the contour is obtained, and the centroid coordinates are obtained. Maximum radius ;Minimum radius Set to half the laser spot diameter to ensure that the helical trajectory completely covers the defect area when it contracts to the center. (Regarding the spot diameter...) laser, Single-layer spiral duration Based on the laser scanning speed and the total length of the spiral trajectory, the approximate length of the spiral trajectory is: ,in The radial spacing of the spiral trajectory is set to 80% of the laser spot diameter to ensure appropriate overlap between adjacent scanning trajectories. The total angle of the helix is given; to simplify calculations, the helix duration is approximated as... ,in The average radius, The typical value for laser scanning speed is 10 mm / s.
[0059] polar coordinates Convert to Cartesian coordinates Generate the trajectory point sequence of the current layer. ,in This represents the number of trajectory points in the current layer.
[0060] Furthermore, the number of trajectory points in the current layer The calculation method is as follows: discretization sampling is performed based on the helical parameter equation and the time step, and the time step is set. This ensures that the spatial distance between trajectory points sampled at adjacent time steps is approximately half the diameter of the laser spot, guaranteeing trajectory continuity and coverage integrity; the number of trajectory points ,in This indicates rounding up; adding 1 is to include the start and end points.
[0061] Step S6: For each point in the current layer trajectory point sequence Based on its spatial location and defect depth information, the optimal laser power for the trajectory point is calculated using a thermal accumulation prediction model. and scanning speed spatial coordinates With process parameters Combine them to form complete trajectory point parameters; The process for determining the optimal laser power and scanning speed is as follows: First, a set of candidate process parameters is constructed. Candidate laser power values are sampled at 10W intervals within the range of 50W to 500W, and candidate scanning speed values are sampled at 0.5mm / s intervals within the range of 1mm / s to 20mm / s. There are 10 candidate parameter combinations, among which Number of power candidates The number of candidate speeds; for each candidate parameter combination This is compared with historical process parameter sequences and current location coordinates. The material physical parameters are combined into a feature vector, which is then input into the thermal accumulation prediction model to obtain the predicted temperature field distribution under this parameter combination. .
[0062] Furthermore, the historical process parameter sequence refers to the record of process parameters for trajectory points executed before the current trajectory point, used to capture thermal accumulation effects; specifically, it consists of: a laser power sequence. Includes the laser power values of the 10 most recently executed trajectory points. If the current point is one of the top 10 points (i.e., ... If the value is insufficient, fill the missing part with zero; scan speed sequence It includes the scan speed values of the 10 most recently executed trajectory points, with any missing values filled with zeros; the historical process parameter sequence contains a total of 20 numerical parameters (10 power values and 10 speed values).
[0063] Perform constraint checks on the predicted temperature field and calculate the maximum temperature value in the temperature field. and the average temperature value within the defect removal area. The temperature constraint is that the maximum temperature does not exceed the glass transition temperature of the material. 80% of .
[0064] Among the candidate parameter combinations that satisfy the temperature constraint, the parameter combination with the highest removal efficiency is selected as the optimal parameter. To eliminate the influence of dimensions, the laser power, scanning speed, and average temperature are first normalized, and then the normalized removal efficiency index is calculated. The optimal parameter combination is determined by solving an optimization problem. Obtain, among which A set of candidate parameters to satisfy temperature constraints.
[0065] Furthermore, the removal efficiency index The time dimension is reflected in: laser power The scanning speed represents the energy input per unit time. The product of the two represents the distance scanned per unit time. This reflects the amount of material removed per unit time, which is proportional to the laser energy density and scanning speed; average temperature value. This reflects the material's thermal accumulation state at the current moment. Higher temperatures indicate more absorbed energy, requiring longer cooling times before further processing. Therefore, the temperature term in the denominator reflects the constraint of time cost. The physical meaning of the removal efficiency index is the normalized removal rate under unit temperature increase conditions. Maximizing this index is equivalent to minimizing the time required to remove a unit amount of material while ensuring no thermal damage occurs, achieving a time-optimal removal strategy. For a given defect area, the total removal time... ,in The dwell time at each trajectory point Maximizing the removal efficiency metric at each point, which is the sum of all cooling wait times, minimizes the total removal time.
[0066] If none of the candidate parameters satisfy the temperature constraint, then choose to make Minimize the parameter combination and insert a cooling wait time after that trajectory point. ,in is the cooling rate constant of the material.
[0067] Furthermore, the material cooling rate constant The determination method is as follows: through finite element thermal analysis simulation or experimental testing, after the laser stops, monitor the temperature decrease curve of the material surface over time, and analyze the temperature-time curve in the high-temperature range (temperature higher than...). A linear fit is performed on the interval ( ), and the slope of the fitted line is the cooling rate constant. The unit is ℃ / s; for carbon fiber / epoxy resin composites, under natural convection conditions Typical values range from 5°C / s to 15°C / s under forced air cooling conditions. The rate can be increased to 20℃ / s to 40℃ / s; the measurement location of the cooling rate constant is selected in a region 1mm to 3mm away from the center of the laser action point, and the temperature response characteristics of this region can represent the overall heat dissipation capacity of the material.
[0068] Step S7: Determine whether trajectory planning for all depth layers has been completed. Then let Return to step S3 to continue planning the next layer of trajectory; if Then, the trajectory point sequences of all layers are stitched together in depth order to obtain a complete and precise trajectory removal sequence. Output the trajectory sequence.
[0069] It should be noted that depth adaptation refers to dynamically adjusting the laser power parameters based on the depth of the defect and the properties of the surrounding materials. Finite element thermo-mechanical coupling simulation analysis is used to analyze the laser energy transmission path in a multi-layered structure, considering the reflection coefficients of the interlayer interfaces and the absorption coefficients of each layer, to calculate the required laser power at each depth location and generate a power-depth curve.
[0070] Furthermore, the method for obtaining the interlayer interface reflection coefficient and the absorption coefficient of each layer material is as follows: interlayer interface reflection coefficient The reflection coefficient is calculated using Fresnel's reflection formula for the case of perpendicular incidence. ,in and These represent the refractive indices of two adjacent layers of material with respect to the laser wavelength. For a laser with a wavelength of 1064 nm, the refractive index of the epoxy resin is... The refractive index of carbon fiber The reflectance coefficient of the epoxy resin-carbon fiber interlayer interface This indicates that approximately 2% of the laser energy is reflected at the interface; the absorption coefficient of each layer of material Measured using Beer-Lambert's law, the thickness of the laser irradiation was... Material samples, measuring transmitted light intensity With incident light intensity The ratio of absorption coefficient .
[0071] In this embodiment of the application, in order to minimize the heat-affected zone and protect the surrounding undamaged materials, a heat accumulation prediction model is introduced when generating the removal trajectory. This model predicts the temperature field distribution during the continuous removal process based on the material's thermal conductivity, specific heat capacity, and interlayer thermal resistance parameters. When the predicted temperature exceeds 80% of the material's glass transition temperature, a cooling waiting time is automatically inserted to ensure that the removal process does not cause thermal damage.
[0072] The aforementioned thermal accumulation prediction model employs a multilayer perceptron (MLP) structure, wherein the input layer receiver dimension of the MLP is... The feature vector contains process parameters and material physical parameters at the current and historical times, specifically including: laser power sequence (last 10 time steps), scanning speed sequence (last 10 time steps), spot diameter, material thermal conductivity, material specific heat capacity, material density, interlayer thermal resistance coefficient, current position coordinates, and removal time.
[0073] To eliminate the influence of differences in the dimensions and numerical ranges of different physical quantities, each parameter in the input feature vector is subjected to Z-score standardization; the output layer of the MLP is a fully connected layer with an output dimension of... , representing the temperature field distribution of a three-dimensional spatial grid, where , , These represent the number of discrete points of the temperature field grid in three spatial directions. The output layer does not use an activation function and directly outputs the predicted temperature value (unit: degrees Celsius).
[0074] Furthermore, the number of discrete points in the temperature field grid , , The method for determining the number of discrete points in the horizontal direction is as follows: The number of discrete points is set based on the spatial dimensions of the defect removal area and the required temperature field resolution. and Using 0.5 times the laser spot diameter as the grid spacing, for the spot diameter... The grid spacing is 0.05mm. If the horizontal dimensions of the defect area are 5mm × 5mm, then... Number of discrete points in the depth direction Using 0.2 times the thickness of a single layer of material as the grid spacing, for a typical single layer thickness... The depth grid spacing is 0.04mm. If the total material thickness is 10mm, then... For typical parameters , , Output dimension .
[0075] The aforementioned thermal accumulation prediction model adopts a supervised learning training mode. The training dataset is generated by finite element thermo-mechanical coupling simulation software and contains temperature field evolution data under different combinations of process parameters. Each training sample contains an input feature vector and the corresponding true value of the temperature field.
[0076] The loss function of the aforementioned heat accumulation prediction model uses a weighted combination of mean squared error loss and relative error loss. Among them, mean square error loss Measuring the absolute error of temperature prediction, To predict temperature values, This represents the true temperature value; relative error loss. The relative error of temperature prediction is measured, where To prevent division by zero by small constants; the weighting coefficient is set to , The optimization objective is to balance absolute error and relative error.
[0077] Step 400: Based on the precise removal of the trajectory sequence, output the robot remote control command.
[0078] It should be noted that the output control command refers to converting the removed trajectory sequence into motion control parameters and laser control parameters that the robot can execute, including parameters such as the position coordinates, attitude angle, motion speed, laser power, pulse frequency, and spot diameter of the robot end effector, which are sent to the robot control system through a remote communication interface.
[0079] In this embodiment of the application, in order to achieve real-time monitoring and dynamic adjustment of the removal process, during the removal operation, real-time multimodal imaging data (millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals in multiple polarization directions, and multi-angle polarization scattering image sequence images) change data are acquired simultaneously. By comparing the changes in imaging characteristics before and after removal, the removal effect and material integrity are evaluated. When an abnormal signal is detected, the laser power and scanning speed are automatically adjusted to generate closed-loop control correction commands.
[0080] In this embodiment of the application, in order to continuously optimize the removal strategy, a knowledge graph system based on machine learning is established. This system collects historical removal case data, including defect type, material structure parameters, process parameters and removal effect. Through deep learning algorithms, it mines the correlation patterns between parameters and provides intelligent decision suggestions for new removal tasks, thereby realizing adaptive optimization of the removal strategy.
[0081] The aforementioned knowledge graph system employs a graph neural network (GNN) structure for association pattern mining and decision recommendation. The input to the GNN is the node feature matrix and adjacency matrix of the knowledge graph. Node types include defect nodes, material nodes, process parameter nodes, and removal effect nodes. Defect node features include defect type encoding, defect size, defect depth, etc.; material node features include material category encoding, number of layers, thickness, etc.; process parameter node features include laser power, scanning speed, pulse frequency, etc.; and removal effect node features include removal integrity score, thermal damage degree, processing time, etc. For categorized features, defect type is converted into a numerical vector using one-hot encoding, and material category is also processed using one-hot encoding. For numerical features, including defect size, defect depth, number of layers, thickness, laser power, scanning speed, pulse frequency, removal integrity score, thermal damage degree, and processing time, range-based normalization is performed to normalize each feature to between 0 and 1, eliminating the influence of differences in the dimensions of different physical quantities on the learning of the graph neural network. The adjacency matrix is constructed based on the co-occurrence relationship of parameters in historical cases. If two nodes appear in the same case, the corresponding position is set to 1; otherwise, it is 0.
[0082] The aforementioned GNN output layer is divided into two branches based on the task type. The first branch is the process parameter regression branch, which receives the node features output by the graph convolution, aggregates all node information through global average pooling, and then passes through two fully connected layers, with an output dimension of... The process parameter vector includes recommended laser power, scanning speed, pulse frequency, and other parameter values. The hidden layer of the fully connected layer has a dimension of 256, and the output layer does not use an activation function. The second branch is the removal effect prediction branch, which also uses global average pooling and two fully connected layers to output a removal effect score with a dimension of 1, representing the predicted removal completeness. The output layer uses a sigmoid activation function to normalize the score to between 0 and 1.
[0083] The decoding process of the aforementioned GNN output is as follows: For the output of the process parameter regression branch, since the process parameters were normalized during training, inverse normalization is required during inference to obtain the actual executable process parameter values. The inverse normalization formula is as follows: ,in These are the normalized parameter values output by the network. and These are the minimum and maximum values of the process parameter in the training set, respectively. These are the actual parameter values after decoding; specifically, the inverse normalization range of laser power is 50W to 500W, the inverse normalization range of scanning speed is 1mm / s to 20mm / s, and the inverse normalization range of pulse frequency is determined according to the actual working range of the laser equipment; for the output of the removal effect prediction branch, since the output value has been normalized to between 0 and 1 by the sigmoid function, it can be directly used as a score for removal integrity. A score greater than 0.8 indicates that the expected removal effect is good, a score between 0.5 and 0.8 indicates that the removal effect is average and should be used with caution, and a score less than 0.5 indicates that this combination of process parameters is not recommended.
[0084] The aforementioned GNN employs a semi-supervised learning training mode combined with a reinforcement learning optimization strategy. The semi-supervised learning stage uses historical labeled case data; the training dataset contains 500 fully labeled removal cases, each including input parameters, corresponding process parameters, and the ground truth value of the removal effect. The reinforcement learning stage uses the GNN as the policy network, interacting with the simulation environment or the actual removal system to obtain delayed reward signals. The network parameters are updated using the proximal policy optimization (PPO) algorithm, and the reward function is defined as... ,in To normalize the removal quality score to a range of 0 to 1, To normalize the processing time to between 0 and 1, To normalize the thermal damage level to a value between 0 and 1, all three values are dimensionless and have the same dimensions, with weighting coefficients. , , .
[0085] Furthermore, the processing time The normalization method is as follows: First, calculate the actual processing time. This duration includes the laser application time and cooling waiting time for all trajectory points, and the calculation formula is as follows: ,in For the first Layer The dwell time of each trajectory point The sum of all cooling wait times is used; then the actual processing time is mapped to between 0 and 1 by range normalization.
[0086] Furthermore, the removal of quality scores The calculation method is as follows: The defect residue rate in the removed area is obtained through multimodal imaging detection after removal, and the calculation formula is as follows: ,in The area of the defect that remains after removal. The total area of defects before removal; the removal quality score is defined as... This indicates the completeness of defect removal, when When the defect is completely removed, This indicates that the defect has not been removed.
[0087] Furthermore, the degree of thermal damage The calculation method is as follows: by detecting the structural integrity of the material surrounding the removed area, ultrasonic C-scan or polarization imaging technology is used to detect the area and depth of interlayer delamination damage, and a thermal damage index is defined. ,in For the volume of delamination damage caused by thermal effects, The total volume for defect removal; the degree of thermal damage is mapped to between 0 and 1 through range normalization.
[0088] The Adam optimizer was used during the training process.
[0089] The loss function of the aforementioned GNN in the semi-supervised learning stage is a multi-task loss. ,in Mean square error loss for process parameter prediction , For the predicted process parameters, These are the actual process parameters; Binary cross-entropy loss for removal effect prediction ,in To ensure accurate label removal, To predict the removal effect; The loss is the graph structure regularization loss.
[0090] The following is a specific application example: During the production of carbon fiber / epoxy resin composite laminates for the main wing spars of a Boeing 787 passenger aircraft, an aerospace company's quality inspection department discovered a suspected metallic inclusion defect inside a composite material sample (sample number CF-20231029-A05) during non-destructive testing. The sample measured 200mm × 150mm × 10mm, used T800 carbon fiber reinforced epoxy resin matrix, and consisted of 16 layers of unidirectional prepreg, each layer being 0.625mm thick. The glass transition temperature was... Preliminary ultrasonic C-scan showed abnormal echo signals located in the central region of the sample within an area of approximately 30 mm × 25 mm, with an estimated depth between 5 mm and 7 mm.
[0091] Step 100 Implementation Example: Acquiring Multimodal Imaging Data of Multilayer Composite Materials A multimodal imaging system was used to perform a comprehensive scan of the sample. The millimeter-wave imaging system was set to a frequency of 94 GHz, corresponding to a wavelength of 3.19 mm. Tomographic projection data was acquired at 2° intervals from 0° to 180°, for a total of 91 projection angles. Transmission signals were also acquired in four directions: horizontal polarization, vertical polarization, left-handed circular polarization, and right-handed circular polarization. The polarization scattering imaging module acquired scattering images in four polarization directions: 0°, 45°, 90°, and 135°, with a resolution of 2048×2048 pixels for each direction.
[0092] Table 1. Sample millimeter-wave tomographic projection data (partial angles)
[0093] Table 2. Optical property parameters extracted by Mueller matrix decomposition
[0094] Step 200 Implementation Example: Based on the multi-physics stereo imaging dataset, perform multi-modal feature extraction and fusion to output a three-dimensional defect localization model; The acquired millimeter-wave tomographic projection data was reconstructed into a 125×125 pixel tomographic cross-sectional image using a filtered back-projection algorithm, with a radial sampling interval of... Polarization decomposition was performed on the millimeter-wave transmission signal in multiple polarization directions to obtain polarization feature distribution maps. Optical property maps were constructed from features such as depolarization degree and birefringence parameters obtained from Mueller matrix decomposition. The feature maps of the three modes were normalized by range and then input into a multimodal deep fusion network. After 200 epochs of training, the network achieved a Dice coefficient of 0.927 on the validation set, and a mean absolute error of 0.31 mm for depth regression.
[0095] The multimodal deep fusion network outputs a 3D defect localization model, which includes a defect segmentation mask. (3 categories: metal inclusions, carbon fiber, background) and depth regression vector (The total material thickness of 10mm is discretized into 20 layers at 0.5mm intervals.) By performing argmax decoding on the defect segmentation mask, the pixel positions of the metal inclusion categories were extracted, revealing a defect region containing 661 pixels. The defect center coordinates were (102.3mm, 78.6mm), and the defect planar dimensions were 28.5mm × 23.2mm. Applying a 0.5 threshold to the depth regression vector, the output values of layers 11 to 14 were detected to be 0.68, 0.92, 0.87, and 0.61, respectively, all exceeding the threshold, corresponding to a depth range of 5.0mm to 7.0mm. After precise mapping, the defect depth range was determined to be [5.2mm, 6.8mm], corresponding to layers 9 to 11 of the material.
[0096] Table 3. Decoding results of the three-dimensional defect localization model
[0097] Analysis of the confidence maps generated by the depth sensing module for the four depth layers showed that the average confidence level of the third layer (depth range 5mm to 6.67mm) was 0.91, the average confidence level of the fourth layer (depth range 6.67mm to 10mm) was 0.73, while the confidence levels of the first and second layers were both below 0.15, indicating that the defects are mainly distributed in the lower layer of the material.
[0098] Step 300 Implementation Example: Based on the three-dimensional defect localization model, a depth-adaptive precision removal trajectory sequence is generated using a spiral layer-by-layer removal strategy. Based on the defect depth range of 5.2mm to 6.8mm, calculate the defect span depth. Single layer thickness Based on the constraint that the depth of removal in a single pass should not exceed 30% of the single-layer thickness, the depth of removal in a single layer is limited to [specific value]. Initial layer calculation layer.
[0099] Table 4. Parameters for spiral-type layer-by-layer removal trajectory planning
[0100] The heat accumulation prediction model optimizes process parameters for typical trajectory points. Input features include power and velocity sequences from the past 10 steps, and material thermophysical parameters (thermal conductivity). Specific heat capacity ,density The model outputs temperature field predictions, along with the current location coordinates, to filter for those that meet the criteria. Candidate parameters for constraints.
[0101] Table 5. Optimal process parameters for typical trajectory points
[0102] At point P in the third layer of the trajectory 3,150 For location, the predicted maximum temperature for all candidate parameter combinations exceeds 132°C. The system selects the parameter combination corresponding to the minimum temperature (power 150W, speed 10mm / s, predicted temperature 136.8°C) and inserts a cooling wait time. .
[0103] Step 400 Implementation Example: Based on the precise removal trajectory sequence, output remote control commands for the robot. The generated 2653 trajectory points (the sum of 9 layers of trajectory points) are converted into robot control commands, including the joint angles of the six-axis robot, the pose of the end effector, the movement speed, and the power and pulse parameters of the laser. The control commands are sent to the remote robot control system via TCP / IP protocol (port 8080), with a communication frequency of 100Hz and a command latency of less than 10ms.
[0104] Table 6. Sample Robot Control Commands (Partial Trajectory Points)
[0105] During the process, the real-time millimeter-wave monitoring system collected imaging data of the removed area at a frequency of 5Hz. When the removal of the 5th layer was completed (progress 56%), the local temperature was detected to exceed the set threshold. The system automatically reduced the laser power of subsequent trajectory points by 15% (from the original 240W to 204W) and increased the scanning speed by 10% (from 6.8mm / s to 7.48mm / s), successfully avoiding thermal damage.
[0106] The entire removal process took approximately 358 seconds, including 312 seconds of laser action and 46 seconds of cooling time. Multimodal imaging re-inspection after removal showed that the metal inclusions were completely removed, with a defect residue rate of 0%. No interlayer delamination damage was found in the material surrounding the removed area, and the ratio of the heat-affected zone volume to the removed volume was 0.08, meeting the quality standards for aerospace composite materials.
[0107] Table 7, data after removing before and after comparisons.
[0108] The knowledge graph system inputs the parameters and results of this removal case into the database, and the graph neural network updates node features and edge weights based on the newly added cases. This is useful when subsequently processing similar defects (depth 5mm to 7mm, area 500mm²). 2 Up to 800mm 2 When dealing with metallic inclusions, the system recommends initial process parameters of 200W±20W laser power and 7.5mm / s±1mm / s scanning speed, with a predicted removal effect score of 0.92. The deviation from the actual execution parameters in this example is within 10%, verifying the effectiveness of the knowledge graph system.
[0109] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multimodal image recognition remote control method for high-precision robot operation, characterized in that, Includes the following steps: Multimodal imaging data of multilayer composite materials are acquired to generate a multiphysics stereo imaging dataset. The multimodal imaging data includes millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals in multiple polarization directions, and multi-angle polarization scattering image sequences. Based on the multi-physics stereo imaging dataset, multimodal feature extraction and fusion are performed to output a three-dimensional defect localization model. The multimodal feature extraction includes tomographic reconstruction processing of millimeter-wave tomographic projection data to extract millimeter-wave feature maps, polarization decomposition of millimeter-wave transmission signals in multiple polarization directions to obtain polarization feature distribution maps, and Mueller matrix decomposition of multi-angle polarization scattering image sequences to obtain optical property maps. The fusion is achieved through a multimodal deep fusion network. The three-dimensional defect localization model includes a defect segmentation mask and a depth regression vector. Based on the three-dimensional defect localization model, a spiral layer-by-layer removal strategy is adopted to generate a precision removal trajectory sequence with depth adaptation. In the spiral layer-by-layer removal strategy, the process proceeds spirally from the defect edge to the center, and the removal depth of each layer is adaptively adjusted according to the material layer thickness and the defect depth. Based on the precise removal trajectory sequence, remote control commands for the robot are output.
2. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The process of acquiring the millimeter-wave tomographic projection data includes: Millimeter-wave tomographic projection data were acquired at angular intervals of 1° to 5° within the incident angle range of 0° to 180°. The phase of the millimeter-wave tomographic projection data at each incident angle is recorded synchronously. The spatial distribution of the relative permittivity inside the material is calculated based on the radial gradient of the phase. The calculation process is as follows: the mathematical relationship between the phase gradient and the relative permittivity is established based on the electromagnetic wave propagation theory. The phase gradient is related to the millimeter-wave wavelength and the square root of the relative permittivity. The spatial distribution of the relative permittivity is obtained through inversion calculation. The epoxy resin matrix, carbon fiber, and metal inclusions were distinguished based on the significant differences in the spatial distribution of their relative permittivity. The permittivity of the epoxy resin matrix ranged from 3 to 4, that of the carbon fiber ranged from 10 to 15, and that of the metal inclusions was greater than 100.
3. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The polarization decomposition of the millimeter-wave transmission signal in multiple polarization directions to obtain the polarization feature distribution map includes: Construct a polarization scattering matrix that includes horizontal transmission and horizontal reception, vertical transmission and vertical reception, and cross-polarization scattering coefficients; The polarization scattering matrix is decomposed into eigenvalues to obtain eigenvalues and corresponding eigenvectors. By analyzing the magnitude relationship of eigenvalues and the direction of eigenvectors, the electromagnetic responses of different layers are separated. When the first eigenvalue is much larger than the second eigenvalue, it indicates that the layer is the dominant scattering layer. When the two eigenvalues are close, it indicates that the layer has multiple scattering. The electromagnetic response output of each depth layer is used as a polarization feature distribution map, where the main scattering component corresponds to the strong reflection feature of the metal inclusions, and the cross scattering component corresponds to the anisotropic scattering feature of the carbon fiber.
4. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The process of obtaining optical property maps by Mueller matrix decomposition of multi-angle polarization scattering image sequences includes: A 4×4 Mueller matrix was established by measuring the light intensity changes under different combinations of incident and outgoing polarization states. The Mueller matrix is decomposed into the product of the depolarization matrix, the optical rotation matrix, and the birefringence matrix using the Lu-Chipman decomposition method; The depolarization degree parameter, linear birefringence, circular birefringence, and fiber orientation angle distribution are extracted from the decomposition results. The depolarization degree is obtained by calculation through the depolarization matrix, the linear birefringence and circular birefringence are obtained by Euler angle decomposition of the rotation matrix, and the fiber orientation angle is determined by the birefringence matrix.
5. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The multimodal deep fusion network includes: Three parallel coding branches process millimeter-wave feature maps, polarization feature distribution maps, and optical property maps respectively to obtain three modal features. Each coding branch contains four convolutional blocks. The cross-modal attention mechanism achieves feature association by calculating the similarity matrix between features of different modalities. This includes mapping each modal feature to a query vector, a key vector, and a value vector, calculating attention weights based on the query vector and key vector, and fusing the value vector to obtain the fused feature. An adaptive weight allocation strategy is used to generate normalized fusion weight coefficients for each modality and depth layer through a learnable weight generation network. The three fusion features of each layer are weighted and summed using the fusion weight coefficients to obtain the final fusion features. The final fusion features of each depth layer are the three-dimensional defect localization model. The depth sensing module generates a confidence map for each depth layer by analyzing the signal attenuation patterns and scattering characteristics at different depths. The depth perception module receives intermediate features from each coding branch at different depth layers as input; The deep perception module includes a deep feature extraction subnetwork and a confidence generation subnetwork. The deep feature extraction subnetwork uses an LSTM structure to process cross-layer feature sequences, combining intermediate features from four deep layers. The LSTM outputs a sequence of hidden states, which are then fed into the LSTM. ; Confidence-based subnetworks generate the hidden states of each depth layer Two convolutional layers are applied, and finally a confidence map is generated using the sigmoid activation function. Its value ranges from 0 to 1, representing the reliability weight of the feature at that depth level; The output of the depth perception module is a sequence of confidence maps for four depth layers. The confidence map is multiplied element-wise with the fusion features of the corresponding depth layer to obtain the weighted fusion features. ,in This indicates an element-wise multiplication operation, and the weighted fusion features are input into the subsequent decoding branch for final defect localization and classification.
6. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The spiral layer-by-layer removal strategy includes: Extract the defect boundary contour and obtain the outermost boundary contour point set through morphological erosion operation; The number of depth layers of the defect is determined based on the depth regression vector, and the defect depth range is divided equally. The thickness of each layer meets the constraint that it does not exceed the set proportion of the thickness of a single layer of material. For each layer, a spiral trajectory is generated from the boundary profile at a constant angular velocity and radial contraction velocity. The spiral trajectory parameters include the maximum radius, the minimum radius, and the spiral duration of a single layer. The polar coordinates of the spiral trajectory are converted to Cartesian coordinates to generate the trajectory point sequence of the current layer. For each trajectory point in the current layer trajectory point sequence, based on its spatial position and defect depth information, the thermal accumulation prediction model is called to calculate the process parameters including the optimal laser power and scanning speed. The Cartesian coordinates, depth and process parameters are combined to form the complete trajectory point parameters. By concatenating the trajectory point sequences of each layer in depth order, a depth-adaptive, precisely removed trajectory sequence is obtained.
7. The multimodal image recognition remote control method for high-precision robot operation according to claim 6, characterized in that, The heat accumulation prediction model is a multilayer perceptron structure, including: The input layer receives a feature vector containing process parameters at the current time and historical time, as well as material physical parameters. The historical time process parameters include the laser power sequence and scanning speed sequence of the most recent multiple time steps. Four hidden layers are used, employing the ReLU activation function and adding a Dropout layer to prevent overfitting; The output layer outputs the temperature field distribution of a three-dimensional spatial mesh; The heat accumulation prediction model is trained through supervised learning. The training dataset is generated by finite element thermo-coupling simulation software, and the loss function is a weighted combination of mean square error loss and relative error loss.
8. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The method further includes: During the removal operation, real-time multimodal imaging data is acquired simultaneously; The removal effect and material integrity are evaluated by comparing the changes in imaging features before and after removal. When an abnormal signal is detected, the laser power and scanning speed are automatically adjusted to generate a closed-loop control correction command.
9. The multimodal image recognition remote control method for high-precision robot operation according to claim 1, characterized in that, The method also includes establishing a machine learning-based knowledge graph system: Collect historical removal case data, including defect type, material structure parameters, process parameters, and removal effect; A graph neural network structure is used to mine association patterns. The graph neural network includes three graph convolutional layers and a graph attention mechanism. The network is trained by optimizing the strategy through a combination of semi-supervised learning and reinforcement learning. The reward function comprehensively considers the removal quality score, processing time and the degree of thermal damage. It provides intelligent decision-making suggestions for new removal tasks, enabling adaptive optimization of the spiral layer-by-layer removal strategy.
10. A multimodal image recognition remote control system for high-precision robot operation, used to execute the multimodal image recognition remote control method for high-precision robot operation as described in any one of claims 1-9, characterized in that, include: The multimodal imaging module is used to acquire multimodal imaging data of multilayer composite materials and generate a multiphysics stereo imaging dataset. The multimodal imaging data includes millimeter-wave tomographic projection data at different incident angles, millimeter-wave transmission signals in multiple polarization directions, and multi-angle polarization scattering image sequences. The feature fusion module is used to perform multimodal feature extraction and fusion based on the multiphysics stereo imaging dataset, and output a three-dimensional defect localization model. The multimodal feature extraction includes tomographic reconstruction processing of millimeter-wave tomographic projection data to extract millimeter-wave feature maps, polarization decomposition of millimeter-wave transmission signals in multiple polarization directions to obtain polarization feature distribution maps, and Mueller matrix decomposition of multi-angle polarization scattering image sequences to obtain optical property maps. The fusion is achieved through a multimodal deep fusion network. The three-dimensional defect localization model includes a defect segmentation mask and a depth regression vector. The trajectory planning module is used to generate a depth-adaptive precision removal trajectory sequence based on the three-dimensional defect location model using a spiral layer-by-layer removal strategy. In the spiral layer-by-layer removal strategy, the process proceeds spirally from the defect edge to the center, and the removal depth of each layer is adaptively adjusted according to the material layer thickness and the defect depth. The control execution module is used to output remote control commands for the robot based on the precise removal trajectory sequence; The real-time monitoring module is used to acquire multimodal imaging data in real time during the removal process and generate closed-loop control correction commands; The knowledge optimization module includes a knowledge graph system based on graph neural networks, which is used to continuously optimize the spiral layer-by-layer removal strategy.