Workpiece surface roughness detection method, device, equipment and medium
Through the explicit dual-pooling cross-modal alignment network model, combined with RGB and laser speckle feature maps for complementary feature correction and fusion, the problem of low surface roughness measurement accuracy of machined workpieces is solved, and efficient and accurate surface roughness detection is achieved.
Patent Information
- Application Number
- CN202510785791.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
AI Technical Summary
In the existing technology, the measurement methods of the surface roughness of machined workpieces have problems such as low precision, easy damage to soft materials and slow detection speed. In particular, non-contact measurement methods are affected by light interference and laser speckle noise, resulting in low prediction accuracy.
An explicit dual-pooling cross-modal alignment network model is adopted. Through the bidirectional cross-modal attention feature correction module and the parallel dual-pooling cross-modal fusion module, the RGB feature map and the laser speckle feature map are combined to perform complementary feature correction and fusion processing. The fully connected layer is used for classification to improve the prediction accuracy of surface roughness.
It improves the prediction accuracy of the surface roughness of machined workpieces, realizes efficient and accurate surface roughness detection, adapts to different materials, reduces information loss, and improves detection speed.
Smart Images

Figure CN120689310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent detection of workpiece surfaces, and in particular to a method, device, equipment and medium for detecting the surface roughness of a workpiece. Background Art
[0002] Currently, surface roughness measurement methods for machined workpieces are primarily divided into contact and non-contact methods. Contact methods primarily involve stylus measurement, which offers high accuracy but can miss tiny indentations and easily damage soft materials. Furthermore, their slow speed makes them difficult to meet the demands of real-time industrial testing.
[0003] Non-contact measurement prediction methods primarily use machine learning algorithms to process RGB images captured by industrial cameras or laser speckle patterns. However, RGB images are susceptible to environmental interference, such as lighting, and the high reflectivity of machined workpieces can easily cause flare in the images, making it difficult to capture fine surface textures and other detailed features, thus reducing prediction accuracy. The inherent noise and micro-vibrations in laser speckle patterns can degrade image quality, thereby affecting the accuracy of surface roughness prediction. Summary of the Invention
[0004] The object of the present invention is to provide a workpiece surface roughness detection method, device, equipment and medium for improving the prediction accuracy of the surface roughness of a machined workpiece.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides a workpiece surface roughness detection method, which is applied to a trained explicit dual-pooling cross-modal alignment network model, wherein the explicit dual-pooling cross-modal alignment network model includes a bidirectional cross-modal attention feature correction module, a parallel dual-pooling cross-modal fusion module, and a fully connected layer. The method includes:
[0007] Obtaining the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected;
[0008] Determining a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correcting the laser speckle feature map based on the first complementary feature to obtain a first corrected feature;
[0009] determining a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature;
[0010] Using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature;
[0011] The fully connected layer is used to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0012] Optionally, the using the bidirectional cross-modal attention feature correction module to determine a first complementary feature related to the laser speckle feature map from the RGB feature map, and correcting the laser speckle feature map based on the first complementary feature to obtain the first corrected feature includes:
[0013] Using an average pooling function to reduce the height and width of the RGB feature map according to a preset ratio to obtain a downsampled RGB feature map;
[0014] Using the formula:
[0015]
[0016] determining a first complementary feature;
[0017] Among them, Att RGB is the first complementary characteristic, Q RGB is the query vector obtained from the downsampled RGB feature map, is the transpose of the bond vector obtained from the laser speckle feature map, V Speckle is the value vector obtained from the laser speckle feature map, h is the header number;
[0018] Restoring the first complementary feature by upsampling and reshaping operations, and fusing the restored first complementary feature with the RGB feature map to obtain a fused feature;
[0019] The multi-layer perceptron is used to process the fusion features and obtain the multi-layer perceptron result;
[0020] The multi-layer perceptron result is fused with the fusion feature to obtain a first correction feature.
[0021] Optionally, the using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature includes:
[0022] Using a first average pooling layer and a first maximum pooling layer in parallel to process the first correction feature, respectively, to obtain a first average pooling result and a first maximum pooling result;
[0023] Concatenate the first average pooling result and the first maximum pooling result to obtain a first double-pooling feature;
[0024] Using a second average pooling layer and a second maximum pooling layer in parallel to process the second correction feature, respectively, to obtain a second average pooling result and a second maximum pooling result;
[0025] Concatenate the second average pooling result and the second maximum pooling result to obtain a second double pooling feature;
[0026] The first dual-pooling feature and the second dual-pooling feature are concatenated to obtain a cross-modal fusion feature.
[0027] Optionally, the method further includes determining a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correcting the laser speckle feature map based on the first complementary feature to obtain a first corrected feature.
[0028] Obtaining training samples; the training samples include RGB feature map training data, laser speckle feature map training data and corresponding surface roughness results;
[0029] determining a first training complementary feature related to the laser speckle feature map training data from the RGB feature map training data, and correcting the laser speckle feature map training data based on the first training complementary feature to obtain a first training correction feature;
[0030] Determining a second training complementary feature related to the RGB feature map training data from the laser speckle feature map training data, and correcting the RGB feature map training data based on the second training complementary feature to obtain a second training corrected feature;
[0031] Performing double-pooling processing on the first training correction feature and the second training correction feature respectively to obtain a first training double-pooling feature and a second training double-pooling feature;
[0032] The first training double-pooling feature and the second training double-pooling feature are fused to obtain the cross-modal fusion feature of the training sample;
[0033] Based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature and the cross-modal fusion feature of the training sample, the initial model of the explicit dual-pooling cross-modal alignment network is trained to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0034] Optionally, the training of the explicit dual-pooling cross-modal alignment network initial model based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature, and the cross-modal fusion feature of the training sample to obtain the trained explicit dual-pooling cross-modal alignment network model includes:
[0035] Calculating a first auxiliary loss function based on the first training dual-pooling feature and the surface roughness result;
[0036] Calculating a second auxiliary loss function based on the second training dual-pooling feature and the surface roughness result;
[0037] Calculating a main loss function based on the cross-modal fusion features of the training sample and the surface roughness result;
[0038] The first auxiliary loss function, the second auxiliary loss function and the main loss function are summed to obtain the target loss function;
[0039] Based on the target loss function, the parameters of the initial model of the explicit dual-pooling cross-modal alignment network are adjusted to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0040] Optionally, the parallel dual-pooling cross-modal fusion module using the explicit dual-pooling cross-modal alignment network model fuses the first correction feature and the second correction feature, and classifies the fused features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected, and then further includes:
[0041] Acquiring experimental data of the surface roughness of the workpiece to be tested;
[0042] The prediction results that are the same as the experimental data are confirmed as correctly classified samples, and the prediction results other than the correctly classified samples are confirmed as incorrectly classified samples; the correctly classified samples include correctly classified positive samples and correctly classified negative samples, and the incorrectly classified samples include incorrectly classified positive samples and incorrectly classified negative samples;
[0043] Using the formula:
[0044]
[0045] Determine the accuracy of the forecast;
[0046] Using the formula:
[0047]
[0048] Determine the recall of the predictions;
[0049] Calculating the harmonic mean of the recall and the precision;
[0050] Using the formula:
[0051]
[0052] Calculate the accuracy of the prediction;
[0053] Among them, Precision is the accuracy, Recall is the recall rate, Accuracy is the accuracy rate, TP is the number of correctly classified positive samples, TN is the number of correctly classified negative samples, FN is the number of incorrectly classified positive samples, and FP is the number of incorrectly classified negative samples.
[0054] Optionally, obtaining an RGB feature map and a laser speckle feature map of the surface of the workpiece to be inspected includes:
[0055] Acquire RGB images and laser speckle images of the surface of the workpiece to be inspected;
[0056] Using the MobileNet-V3 Large model to perform feature extraction on the RGB image and the laser speckle image, respectively, to obtain a first extracted feature and a second extracted feature;
[0057] The first extracted features are flattened and transposed to obtain an RGB feature map, and the second extracted features are flattened and transposed to obtain a laser speckle feature map.
[0058] Compared with the prior art, the present invention provides a workpiece surface roughness detection method, which includes: obtaining an RGB feature map and a laser speckle feature map of the surface of a workpiece to be detected; using a bidirectional cross-modal attention feature correction module to determine a first complementary feature related to the laser speckle feature map from the RGB feature map, and correcting the laser speckle feature map based on the first complementary feature to obtain a first corrected feature; determining a second complementary feature related to the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature; using a parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first corrected feature and the second corrected feature to obtain a cross-modal fusion feature; and using a fully connected layer to classify the cross-modal fusion feature to obtain a prediction result of the roughness of the surface of the workpiece to be detected. Since the intrinsic modal heterogeneity between the RGB feature map and the laser speckle feature map shows representation differences in encoding, it is impossible for them to fully learn from each other. However, this application adopts a bidirectional cross-modal attention feature correction module to allow the RGB feature map and the laser speckle feature map to retrieve each other and learn each other's complementary features from the other feature map, thereby deeply integrating the advantageous features of the two feature maps and improving the surface roughness prediction accuracy. In addition, the parallel dual-pooling cross-modal fusion module uses a dual-pooling method to help the model better capture local and global information with a smaller amount of computation to eliminate information loss, further improving the prediction accuracy of the surface roughness of the processed workpiece.
[0059] In a second aspect, the present invention provides a workpiece surface roughness detection device, which is applied to the workpiece surface roughness detection method, and the device includes:
[0060] RGB feature map and laser speckle feature map acquisition module, used to obtain the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected;
[0061] a first correction module, configured to determine a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correct the laser speckle feature map based on the first complementary feature to obtain a first correction feature;
[0062] a second correction module, configured to determine a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correct the RGB feature map based on the second complementary feature to obtain a second correction feature;
[0063] a cross-modal fusion module, configured to perform double-pooling and fusion processing on the first correction feature and the second correction feature using the parallel double-pooling cross-modal fusion module to obtain a cross-modal fusion feature;
[0064] A classification module is used to classify the cross-modal fusion features using the fully connected layer to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0065] In a third aspect, the present invention provides a workpiece surface roughness detection device, which is applied to the workpiece surface roughness detection method, and the device includes:
[0066] Communication unit / communication interface, used to obtain RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected;
[0067] a processing unit / processor, configured to determine, using the bidirectional cross-modal attention feature correction module, a first complementary feature associated with the laser speckle feature map from the RGB feature map, and correct the laser speckle feature map based on the first complementary feature to obtain a first corrected feature;
[0068] determining a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature;
[0069] Using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature;
[0070] The fully connected layer is used to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0071] In a fourth aspect, the present invention provides a computer-readable storage medium having instructions stored therein. When the instructions are executed by a processor, a method for detecting the surface roughness of a workpiece is implemented.
[0072] Compared with the prior art, the beneficial effects of the second aspect device type solution, the third aspect equipment type solution, and the fourth aspect computer-readable storage medium type solution provided by the present invention are the same as the beneficial effects of a workpiece surface roughness detection method described in the above technical solution, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0074] Figure 1 Schematic diagram of the explicit dual-pooling cross-modal alignment network model provided by the present invention;
[0075] Figure 2A flow chart of a workpiece surface roughness detection method provided by the present invention;
[0076] Figure 3 Schematic diagram of the bidirectional cross-modal attention feature correction module provided by the present invention;
[0077] Figure 4 The parallel dual-pooling cross-modal fusion module and classification principle diagram provided by the present invention;
[0078] Figure 5 Sample data of the RGB image and laser speckle image of the machined workpiece provided by the present invention;
[0079] Figure 6 A radar chart comparing the results of the present invention with other multimodal detection methods;
[0080] Figure 7 A schematic structural diagram of a workpiece surface roughness detection device provided by the present invention;
[0081] Figure 8 This is a structural schematic diagram of a workpiece surface roughness detection device provided by the present invention. DETAILED DESCRIPTION
[0082] To facilitate a clear description of the technical solutions of the embodiments of the present invention, the words "first" and "second" are used in the embodiments of the present invention to distinguish between identical or similar items with substantially the same functions and effects. For example, the first threshold and the second threshold are merely used to distinguish between different thresholds and do not limit their order. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.
[0083] It should be noted that, in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present invention should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0084] In the present invention, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or multiple.
[0085] Before introducing the embodiments of the present invention, the following definitions are given for the relevant terms involved in the embodiments of the present invention:
[0086] Main modality: refers to the image type that plays a dominant role in the task, and the model mainly relies on its information for discrimination or prediction.
[0087] Auxiliary modality: provides information complementary to the main modality to enhance the robustness and discriminative ability of the model.
[0088] The learnable weight matrix is a core component in deep learning models. Its essence is a matrix that automatically adjusts parameters in a data-driven manner. It is used to perform linear transformations on input data and combine nonlinear activation functions to capture complex patterns.
[0089] The surface roughness of metal parts can significantly impact their mechanical and functional properties, such as contact stiffness, wear resistance, and fatigue performance. Furthermore, tool wear can significantly impact surface roughness, potentially compromising part integrity and accelerating mechanical failure. However, real-time roughness prediction enables tool condition monitoring, ensuring workpiece quality and facilitating timely tool replacement to prevent potential mechanical failures. Therefore, efficient and accurate surface roughness measurement is essential to ensure the surface quality and service life of machined workpieces.
[0090] The statistical characteristics of laser speckle have a special correlation with surface roughness parameters. At the same time, it contains rich geometric and texture information and is highly sensitive to surface micromorphology. Therefore, a large number of studies are currently exploring surface roughness prediction methods based on laser speckle. At the same time, due to its fast detection speed and strong adaptability to different materials, it has been widely used in industrial scenarios. However, the inherent noise and micro-vibration of laser speckle images will reduce the image quality of laser speckle, thereby affecting the final surface roughness prediction accuracy. It is difficult to obtain fine surface texture from RGB images. Therefore, the existing surface roughness prediction methods for machined workpieces based on RGB images or laser speckle images have low accuracy.
[0091] In order to solve the above problems, the present invention provides a workpiece surface roughness detection method, device, equipment and medium, which are described below with reference to the accompanying drawings.
[0092] See also Figure 1 , a workpiece surface roughness detection method provided by the present invention is applied to a trained explicit dual-pooling cross-modal alignment network model, the explicit dual-pooling cross-modal alignment network model includes a bidirectional cross-modal attention feature correction module, a parallel dual-pooling cross-modal fusion module PPCFM and a classification module, the classification module includes a three-branch fully connected layer, through the three-branch fully connected layer, three prediction results of class1, calss2 and class3 can be obtained, wherein the prediction results corresponding to the fully connected layers of two branches are used to assist in the training of the model. In practical applications, the bidirectional cross-modal attention feature correction module includes BCACM-RGB and BCACM-LS. Among them, BCACM-RGB performs correction processing on the RGB feature map, and BCACM-LS performs correction processing on the laser speckle feature map.
[0093] like Figure 2 As shown, the present invention provides a method for detecting the surface roughness of a workpiece, comprising the following steps:
[0094] Step 201: Acquire an RGB feature map and a laser speckle feature map of the surface of a workpiece to be inspected;
[0095] As an optional method, the steps for obtaining the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected are as follows:
[0096] Step 2010: Acquire an RGB image and a laser speckle image of the surface of the workpiece to be inspected;
[0097] The RGB image is obtained by photographing the surface of the workpiece to be inspected with a camera, and the laser speckle image is collected when the workpiece to be inspected is activated by a semiconductor laser. The RGB image and the laser speckle image are collected at the same position of the workpiece to be inspected.
[0098] Step 211: using the backbone network MobileNet-V3 Large model to perform feature extraction on the RGB image and the laser speckle image respectively to obtain a first extracted feature and a second extracted feature;
[0099] The MobileNet-V3 Large model is a large version of the third-generation MobileNet model. Designed specifically for mobile and edge computing devices, it is a lightweight convolutional neural network designed to optimize model efficiency and performance. The MobileNet-V3 Large model can be trained using the ImageNet dataset, a large-scale visual dataset of paramount importance in the field of computer vision. The MobileNet-V3 Large model uses a common neural network model, so the feature extraction process is not detailed here. Alternatively, other neural network models that can extract features from RGB images and laser speckle images can be used.
[0100] Step 2012: Flatten and transpose the first extracted features to obtain an RGB feature map, and flatten and transpose the second extracted features to obtain a laser speckle feature map.
[0101] like Figure 3 As shown, the first extracted feature F RGB and the second extracted feature F Speckle The size of is H×W×C, and the first extracted feature F can be RGB and the second extracted feature F Speckle Rearrange to N×C.
[0102] Step 202: Determine a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correct the laser speckle feature map based on the first complementary feature to obtain a first corrected feature.
[0103] When correcting the RGB feature map, the RGB feature map is used as the main mode, the laser speckle feature map is used as the auxiliary mode, and the RGB feature map is used as the query. Figure 3 Step 202 is described below. Step 202 can be implemented by using the following steps:
[0104] Step 2021: Use the average pooling function to reduce the height and width of the RGB feature map according to a preset ratio to obtain a downsampled RGB feature map, as shown in formula (1):
[0105]
[0106] Wherein, is the downsampled RGB feature map, N = H × W, r is the preset downsampling ratio, C is the number of feature map channels, H is the height of the feature map, and W is the width of the feature map. Indicates that the feature map is transformed from Convert to AvgPool(·) is the average pooling function. Step 2021 corresponds to Figure 3 The downsampling and transposition operations in .
[0107] Step 2022: Use formula (2):
[0108]
[0109] determining a first complementary feature;
[0110] Among them, Att RGB is the first complementary characteristic, Q RGB is the query vector obtained from the downsampled RGB feature map, is the transpose of the bond vector obtained from the laser speckle feature map, V Speckle is the value vector obtained from the laser speckle feature map; h is the number of headers, which can be set to 8 in the present invention. Figure 3 in is the matrix multiplication operation, is the addition operation.
[0111] The query vector is obtained by multiplying the downsampled RGB feature map with the learnable weight matrix of the query vector; the key vector is obtained by multiplying the laser speckle feature map with the learnable weight matrix of the key vector; and the value vector is obtained by multiplying the laser speckle feature map with the learnable weight matrix of the value vector.
[0112] Step 2023: restore the first complementary feature by upsampling and reshaping operations, and fuse the restored first complementary feature with the RGB feature map, i.e., add them together to obtain a fused feature; the calculation is as shown in formula (3):
[0113]
[0114] Among them, F′ RGB is the fusion feature, F RGB is the RGB feature map.
[0115] Step 2024: Process the fusion features using a multi-layer perceptron to obtain a multi-layer perceptron result;
[0116] Step 2025: Fuse the multi-layer perceptron result with the fusion feature to obtain the first correction feature. The multi-layer perceptron and fusion calculation are shown in formula (4):
[0117]
[0118] in, is the first correction feature, and MLP(·) is a multi-layer perceptron.
[0119] Step 203: determining a second complementary feature related to the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature;
[0120] When calibrating the laser speckle feature map, the laser speckle feature map is used as the primary modality, the RGB feature map is used as the auxiliary modality, and the laser speckle feature map is used as the query. Specifically, step 203 is to calibrate the laser speckle feature map as the primary modality. The principle is the same as that of step 202, and the process will not be repeated here.
[0121] Existing implicit cross-modal feature interaction is usually a passive supplementary strategy, which allows the auxiliary modality to generate queries to retrieve basic information of the main modality. This implicit interaction will lead to semantic dislocation, that is, the auxiliary modality dominates the interaction process, resulting in the semantic information of the corrected modality being weakened during the feature aggregation process. Therefore, steps 202 and 203 of the present application adopt an explicit cross-modal complementary retrieval mechanism, introduce a semantic-driven reverse mining paradigm, and use the main modality as a query to deeply explore the complementary features of the auxiliary modality while retaining the core semantics of the main modality, thereby providing a more robust multimodal representation for subsequent cross-modal feature fusion.
[0122] Step 204: using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature;
[0123] Among them, two parallel dual-pooling cross-modal fusion modules can be provided to process the first correction feature and the second correction feature respectively.
[0124] Specifically, such as Figure 4 As shown, AP represents an average pooling operation, MP represents a maximum pooling operation, and C represents a concatenation operation. Step 204 can be implemented based on the following steps:
[0125] Using a first average pooling layer and a first maximum pooling layer in parallel to process the first correction feature, respectively, to obtain a first average pooling result and a first maximum pooling result;
[0126] Concatenate the first average pooling result and the first maximum pooling result to obtain a first double-pooling feature;
[0127] Using a second average pooling layer and a second maximum pooling layer in parallel to process the second correction feature, respectively, to obtain a second average pooling result and a second maximum pooling result;
[0128] Concatenate the second average pooling result and the second maximum pooling result to obtain a second double pooling feature;
[0129] The first dual-pooling feature and the second dual-pooling feature are concatenated to obtain a cross-modal fusion feature.
[0130] The calculation formula for the above steps is shown in (5):
[0131]
[0132] Among them, F fused To fuse features across modalities, MaxPool(·) represents the maximum pooling operation, and Concat(·) represents the feature concatenation operation.
[0133] Step 205: using the fully connected layer to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0134] The above method uses a bidirectional cross-modal attention feature correction module to allow the RGB feature map and the laser speckle feature map to retrieve each other and learn each other's complementary features from the other feature map, thereby deeply integrating the advantageous features of the two feature maps and improving the surface roughness prediction accuracy. In addition, the parallel dual-pooling cross-modal fusion module uses a dual-pooling method to help the model capture local information and global information first and then fuse them with a smaller amount of computation, which can eliminate information loss, efficiently improve the final fusion effect, and further improve the prediction accuracy of the surface roughness of the processed workpiece.
[0135] As an optional method, before step 202, the explicit dual-pooling cross-modal alignment network model is also trained. The training steps are as follows:
[0136] Step 101: Obtain training samples; the training samples include RGB feature map training data, laser speckle feature map training data, and corresponding surface roughness results;
[0137] Among them, the training samples are obtained by collecting RGB images and laser speckle images and processing the two. Specifically, a multimodal image acquisition platform is first built according to the specifications of the workpiece to be processed, the hardware selection of the camera and laser, and sandpaper with different mesh sizes, such as 46 mesh, 80 mesh, 100 mesh and 120 mesh, is used to grind the workpiece made of Q235 material. At the same time, a wire drawing machine is used to process the workpiece of the same material. Then, a probe surface roughness tester with the model of MitutoyoSJ-410 is used to measure the roughness Ra of the machined workpiece surface, and the average value of the measurement of the same processed surface is taken as the overall roughness of the surface. The measured roughness range is 0-1.6μm, and then it is divided into three categories according to its roughness value, category 1: 0-0.4μm, category 2: 0.4-0.8μm and category 3: 0.8-1.6μm. Figure 5 As shown, Figure (a) and Figure (b) are category 1, Figure (c) and Figure (d) are category 2, and Figure (e) is category 3. The processing parameters and the corresponding roughness ranges and categories are shown in Table 1:
[0138] Table 1: Different processing parameters and corresponding categories
[0139] Processing parameters Surface roughness (μm) category Grinding-46 mesh 0.4-0.8 2 Grinding-80 mesh 0.4-0.8 2 Grinding-100 mesh 0-0.4 1 Grinding-120 mesh 0-0.4 1 wire drawing machine 0.8-1.6 3
[0140] The workpiece was then placed on a multimodal image acquisition platform for image data acquisition. During natural light imaging, the semiconductor laser source was deactivated to prevent spectral interference, allowing RGB images of the surface texture to be acquired under natural light illumination. Conversely, laser speckle images were acquired with the semiconductor laser activated. To ensure consistency, a fixed-mounted industrial camera captured paired images at the same workpiece position under both lighting conditions. The RGB and laser speckle image pairs were combined to create a dataset. To ensure the effectiveness of neural network training, data augmentation methods such as random horizontal flipping were used to increase the diversity and size of the dataset.
[0141] Before training, the images were scaled to a resolution of 512×512. The dataset consisted of 5120 pairs of laser speckle images and RGB images at the same location. 2049 of these pairs were used for training, and the remaining 511 pairs were used for validation. The training data was processed using the methods in steps 2011 and 2012 to generate training samples. The validation data was processed using the same method as the training data.
[0142] Step 102: determining a first training complementary feature related to the laser speckle feature map training data from the RGB feature map training data, and correcting the laser speckle feature map training data based on the first training complementary feature to obtain a first training correction feature;
[0143] Step 103: determining a second training complementary feature related to the RGB feature map training data from the laser speckle feature map training data, and correcting the RGB feature map training data based on the second training complementary feature to obtain a second training corrected feature;
[0144] Step 104: performing double-pooling processing on the first training correction feature and the second training correction feature respectively to obtain a first training double-pooling feature and a second training double-pooling feature;
[0145] Step 105: Fusing the first training dual-pooling feature and the second training dual-pooling feature to obtain a cross-modal fusion feature of the training sample;
[0146] Step 106: Based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature, and the cross-modal fusion feature of the training sample, the initial model of the explicit dual-pooling cross-modal alignment network is trained to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0147] Specifically, step 106 can be implemented based on the following steps:
[0148] Step 1061: Calculate a first auxiliary loss function based on the first training dual-pooling feature and the surface roughness result;
[0149] Specifically, the first training double-pooled feature is processed using a first fully connected layer to obtain a first processing result, and the first processing result and the surface roughness result are substituted into the auxiliary loss function calculation formula to calculate the first auxiliary loss function;
[0150] Step 1062: Calculate a second auxiliary loss function based on the second training dual-pooling feature and the surface roughness result;
[0151] Specifically, the second training double-pooled feature is processed using a second fully connected layer to obtain a second processing result, and the second processing result and the surface roughness result are substituted into the auxiliary loss function calculation formula to calculate the second auxiliary loss function;
[0152] Step 1063: Calculate a main loss function based on the cross-modal fusion features of the training sample and the surface roughness result;
[0153] Specifically, the cross-modal fusion features of the training samples and the surface roughness results are substituted into the main loss function calculation formula to obtain the main loss function. It should be noted that the auxiliary loss function and the main loss function can select appropriate loss function calculation formulas as needed. Since the loss function calculation formula is a commonly used formula, it will not be listed again.
[0154] Step 1064: summing the first auxiliary loss function, the second auxiliary loss function, and the main loss function to obtain a target loss function;
[0155] Step 1065: Adjust the parameters of the initial model of the explicit dual-pooling cross-modal alignment network based on the target loss function to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0156] Specifically, the parameters can be adjusted by using a reverse transfer method.
[0157] It should be noted that the first training dual-pooling feature, the second training dual-pooling feature, and the cross-modal fusion feature of the training sample are not used to classify and predict a certain category separately, but are the result of the joint action of these three.
[0158] The above steps use parallel auxiliary heads to introduce auxiliary supervision, which can enhance cross-modal fusion while retaining the features of specific modalities. In order to achieve cross-modal interaction, two consecutive fully connected layers will process the merged dual-pooled features to generate complementary joint representations. At the same time, the specific modality auxiliary heads reconstruct the original features from their respective pooled inputs as auxiliary reconstruction losses to prevent overfitting and maintain the uniqueness of single-modal features. This design ensures the robustness of the features. The splicing of dual-pooled features reduces information loss, while the auxiliary loss retains the uniqueness of the specific modality. In addition, the lightweight architecture also takes into account cross-modal integration and specific modality protection. Finally, considering the data imbalance of these modalities, the present invention adopts two additional supervisions for these modalities to improve the final prediction.
[0159] As an optional method, the parallel dual-pooling cross-modal fusion module of the explicit dual-pooling cross-modal alignment network model is used to fuse the first correction feature and the second correction feature, and classify the fused features to obtain the prediction result of the roughness of the surface of the workpiece to be detected, and then the overall detection performance of the model is reflected by calculating the evaluation index. In order to comprehensively evaluate the computational efficiency and real-time performance, the performance of the model is quantified by three main indicators: inference speed FPS, parameter quantity Parameters and floating-point operation FLOPs. In addition, the five indicators of Accuracy, Precision, Recall and F1-score are used to compare the model proposed by the present invention with other models to measure the accuracy of model prediction, and the present invention uses the average accuracy as the main evaluation comparison standard. The specific steps are as follows:
[0160] Acquiring experimental data of the surface roughness of the workpiece to be tested;
[0161] The prediction results that are the same as the experimental data are confirmed as correctly classified samples, and the prediction results other than the correctly classified samples are confirmed as incorrectly classified samples; the correctly classified samples include correctly classified positive samples and correctly classified negative samples, and the incorrectly classified samples include incorrectly classified positive samples and incorrectly classified negative samples;
[0162] Using formula (6):
[0163]
[0164] Determine the accuracy of the prediction; accuracy is the ratio of the number of correctly predicted positive samples to the number of samples predicted as positive samples, reflecting the reliability of the model in making positive sample predictions.
[0165] Using formula (7):
[0166]
[0167] Determine the recall rate of the prediction; the recall rate represents the ratio of correctly predicted positive samples to the number of positive samples, reflecting the ability of the model to identify positive samples.
[0168] Calculate the harmonic mean of the recall rate and the precision, i.e., the F1-score, to comprehensively evaluate the accuracy and coverage of the model on positive samples, as shown in formula (8);
[0169]
[0170] Using formula (9):
[0171]
[0172] Calculate the accuracy of the prediction, which is used to measure the proportion of correct predictions made by the model as a whole;
[0173] Among them, Precision is the accuracy, Recall is the recall rate, F1-score is the F1-score, Accuracy is the accuracy, TP is the number of correctly classified positive samples, TN is the number of correctly classified negative samples, FN is the number of incorrectly classified positive samples, and FP is the number of incorrectly classified negative samples.
[0174] like Figure 6As shown in the figure, in order to verify the effectiveness of the method for collaborative prediction of workpiece surface roughness by multimodal data in the present invention, under the same experimental parameter settings, the prediction results of the unimodal model are first compared with the surface roughness results of the multimodal prediction to verify the effectiveness of the method proposed in the present invention; finally, the bidirectional cross-modal attention feature correction module based on the explicit cross-modal complementary retrieval mechanism proposed in the present invention is combined with the parallel dual-pooling cross-modal fusion module and the performance is compared with the existing multimodal classification models: CRDNet, VAFNet, CCR-Net, MGMNet, Co-CNN, MDL-RS, GDNet and other models. It can be seen that the accuracy of the detection of machined workpiece surface roughness proposed in the present invention is higher.
[0175] As an optional method, after step 206, the effectiveness of the surface roughness prediction method of the present invention may be verified as follows:
[0176] Using MobileNetV3-large pre-trained on ImageNet as the backbone network, the performance of two different single-modal models and multi-modal models of image data are compared respectively. In addition, the model performance of the explicit multi-modal information interaction proposed in the present invention and the ordinary feature superposition are compared. Table 2 shows the corresponding comparison results. It can be seen that the accuracy and F1-score of the explicit multi-modal feature interaction proposed in the present invention are much better than the results of single-modal learning. In addition, when simply superposing multi-modal features, not only can the overall performance not be improved, but the prediction accuracy will be far lower than that of single-modal learning. At the same time, the prediction speed of the method of the present invention can still maintain 432.08FPS, indicating that it can meet the requirements of real-time detection in industrial applications.
[0177] Table 2: Comparison of single-modal detection method and multi-modal detection method proposed in this invention
[0178]
[0179] The embodiments of the present invention can be divided into functional modules according to the above-mentioned method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into a single processing module. The above-mentioned integrated modules can be implemented in the form of hardware or software functional modules. It should be noted that the division of modules in the embodiments of the present invention is illustrative and is only a logical functional division. In actual implementation, other division methods may be used.
[0180] In the case of dividing each functional module into corresponding functional modules, Figure 7 The schematic diagram of the structure of a workpiece surface roughness detection device provided by the present invention is shown. The device is applied to a workpiece surface roughness detection method, such as Figure 7 As shown, the device includes:
[0181] RGB feature map and laser speckle feature map acquisition module 701, used to acquire the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected;
[0182] a first correction module 702, configured to determine a first complementary feature associated with the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correct the laser speckle feature map based on the first complementary feature to obtain a first correction feature;
[0183] a second correction module 703, configured to determine a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correct the RGB feature map based on the second complementary feature to obtain a second correction feature;
[0184] A cross-modal fusion module 704 is configured to perform a double-pooling and fusion process on the first correction feature and the second correction feature using the parallel double-pooling cross-modal fusion module to obtain a cross-modal fusion feature;
[0185] The classification module 705 is configured to classify the cross-modal fusion features using the fully connected layer to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0186] Optionally, the first correction module 702 may specifically include:
[0187] A downsampling and average pooling unit, configured to reduce the height and width of the RGB feature map by a preset ratio using an average pooling function to obtain a downsampled RGB feature map;
[0188] The first complementary feature calculation unit is configured to use the formula:
[0189]
[0190] determining a first complementary feature;
[0191] Among them, Att RGB is the first complementary characteristic, Q RGB is the query vector obtained from the downsampled RGB feature map, is the transpose of the bond vector obtained from the laser speckle feature map, V Speckle is the value vector obtained from the laser speckle feature map, h is the header number;
[0192] a fusion unit, configured to restore the first complementary feature by upsampling and reshaping operations, and fuse the restored first complementary feature with the RGB feature map to obtain a fused feature;
[0193] A multi-layer perception mechanism unit is used to process the fusion features using a multi-layer perception machine to obtain a multi-layer perception machine result;
[0194] The first correction unit is used to fuse the multi-layer perceptron result with the fusion feature to obtain a first correction feature.
[0195] Optionally, the cross-modal fusion module 704 may specifically include:
[0196] A first dual pooling unit is used to process the first correction feature using a parallel first average pooling layer and a first maximum pooling layer, respectively, to obtain a first average pooling result and a first maximum pooling result;
[0197] A first splicing unit, configured to splice the first average pooling result and the first maximum pooling result to obtain a first double-pooling feature;
[0198] A second dual pooling unit is used to process the second correction feature using a parallel second average pooling layer and a second maximum pooling layer, respectively, to obtain a second average pooling result and a second maximum pooling result;
[0199] A second splicing unit, configured to splice the second average pooling result and the second maximum pooling result to obtain a second double-pooling feature;
[0200] A cross-modal fusion unit is used to concatenate the first dual-pooling feature and the second dual-pooling feature to obtain a cross-modal fusion feature.
[0201] Optionally, the device further includes a training module, which may specifically include:
[0202] A training sample acquisition unit is used to acquire training samples; the training samples include RGB feature map training data, laser speckle feature map training data and corresponding surface roughness results;
[0203] a first training correction unit, configured to determine a first training complementary feature related to the laser speckle feature map training data from the RGB feature map training data, and correct the laser speckle feature map training data based on the first training complementary feature to obtain a first training correction feature;
[0204] a second training correction unit, configured to determine, from the laser speckle feature map training data, a second training complementary feature associated with the RGB feature map training data, and correct the RGB feature map training data based on the second training complementary feature to obtain a second training correction feature;
[0205] A training dual-pooling unit is used to perform dual-pooling processing on the first training correction feature and the second training correction feature, respectively, to obtain a first training dual-pooling feature and a second training dual-pooling feature;
[0206] A training cross-modal fusion unit is used to fuse the first training double-pooling feature and the second training double-pooling feature to obtain a cross-modal fusion feature of the training sample;
[0207] A model adjustment unit is used to train the initial model of the explicit dual-pooling cross-modal alignment network based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature, and the cross-modal fusion feature of the training sample to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0208] Optionally, the model adjustment unit may be specifically configured to:
[0209] Calculating a first auxiliary loss function based on the first training dual-pooling feature and the surface roughness result;
[0210] Calculating a second auxiliary loss function based on the second training dual-pooling feature and the surface roughness result;
[0211] Calculating a main loss function based on the cross-modal fusion features of the training sample and the surface roughness result;
[0212] The first auxiliary loss function, the second auxiliary loss function and the main loss function are summed to obtain the target loss function;
[0213] Based on the target loss function, the parameters of the initial model of the explicit dual-pooling cross-modal alignment network are adjusted to obtain a trained explicit dual-pooling cross-modal alignment network model.
[0214] Optionally, the device further includes a model evaluation module, which may specifically include:
[0215] An experimental data acquisition unit, used to acquire experimental data of the surface roughness of the workpiece to be detected;
[0216] The negative sample and positive sample classification unit is used to confirm the prediction results that are the same as the experimental data as the correctly classified samples, and confirm the prediction results other than the correctly classified samples as the incorrectly classified samples; the correctly classified samples include correctly classified positive samples and correctly classified negative samples, and the incorrectly classified samples include incorrectly classified positive samples and incorrectly classified negative samples;
[0217] Precision calculation unit, used to use the formula:
[0218]
[0219] Determine the accuracy of the forecast;
[0220] Recall calculation unit, used to use the formula:
[0221]
[0222] Determine the recall of the predictions;
[0223] a harmonic mean calculation unit, configured to calculate a harmonic mean of the recall rate and the precision;
[0224] Accuracy calculation unit, used to use the formula:
[0225]
[0226] Calculate the accuracy of the prediction;
[0227] Among them, Precision is the accuracy, Recall is the recall rate, Accuracy is the accuracy rate, TP is the number of correctly classified positive samples, TN is the number of correctly classified negative samples, FN is the number of incorrectly classified positive samples, and FP is the number of incorrectly classified negative samples.
[0228] Optionally, the RGB feature map and laser speckle feature map acquisition module 701 may specifically include:
[0229] RGB image and laser speckle image acquisition unit, used to acquire RGB image and laser speckle image of the surface of the workpiece to be inspected;
[0230] a feature extraction unit, configured to perform feature extraction on the RGB image and the laser speckle image respectively using a MobileNet-V3 Large model to obtain a first extracted feature and a second extracted feature;
[0231] A feature map processing unit is configured to flatten and transpose the first extracted features to obtain an RGB feature map, and flatten and transpose the second extracted features to obtain a laser speckle feature map.
[0232] The above mainly introduces the solution provided by the embodiment of the present invention from the perspective of the interaction between the various modules. It can be understood that in order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0233] In the case of using the corresponding integrated unit, Figure 8 The schematic diagram of the structure of a workpiece surface roughness detection device provided by the present invention is shown. The device is applied to a workpiece surface roughness detection method, such as Figure 8 As shown, the device includes:
[0234] Communication unit / communication interface, used to obtain RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected;
[0235] a processing unit / processor, configured to determine, using the bidirectional cross-modal attention feature correction module, a first complementary feature associated with the laser speckle feature map from the RGB feature map, and correct the laser speckle feature map based on the first complementary feature to obtain a first corrected feature;
[0236] determining a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature;
[0237] Using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature;
[0238] The fully connected layer is used to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
[0239] Among them, the processing unit can be a processor or controller, for example, a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute the various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. The processor can also be a combination that implements computing functions, for example, a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication module can be a transceiver, a transceiver circuit or a communication interface, and the like. The storage module can be a memory.
[0240] like Figure 8 As shown, the processor can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. There can be one or more communication interfaces. The communication interface can use any device such as a transceiver for communicating with other devices or a communication network.
[0241] like Figure 8 As shown, the terminal device may further include a communication line. The communication line may include a path for transmitting information between the components.
[0242] Optional, such as Figure 8 As shown, the terminal device may further include a memory. The memory is used to store computer-executable instructions for executing the solution of the present invention, and the execution is controlled by the processor. The processor is used to execute the computer-executable instructions stored in the memory, thereby implementing the method provided by the embodiment of the present invention.
[0243] like Figure 8As shown, the memory can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a communication line. The memory can also be integrated with the processor.
[0244] Optionally, the computer-executable instructions in the embodiment of the present invention may also be referred to as application program codes, which is not specifically limited in the embodiment of the present invention.
[0245] In a specific implementation, as an embodiment, Figure 8 As shown, the processor may include one or more CPUs, such as Figure 8 CPU0 and CPU1 in.
[0246] In a specific implementation, as an embodiment, Figure 8 As shown, the terminal device may include multiple processors, such as Figure 8 Each of these processors can be a single-core processor or a multi-core processor.
[0247] On the one hand, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are executed by a processor, the above-mentioned workpiece surface roughness detection method is implemented.
[0248] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video disc (DVD); or a semiconductor medium, such as a solid-state drive (SSD).
[0249] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0250] Although the present invention has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the invention. It will be apparent that various modifications and variations may be made to the present invention by those skilled in the art without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such modifications and variations as fall within the scope of the claims of the present invention and their equivalents.
Claims
1. A method for detecting the surface roughness of a workpiece, characterized in that: The method is applied to a trained explicit dual-pooling cross-modal alignment network model, wherein the explicit dual-pooling cross-modal alignment network model includes a bidirectional cross-modal attention feature correction module, a parallel dual-pooling cross-modal fusion module, and a fully connected layer. Obtaining the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected; Determining a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correcting the laser speckle feature map based on the first complementary feature to obtain a first corrected feature; determining a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature; Using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature; The fully connected layer is used to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
2. The method for detecting workpiece surface roughness according to claim 1, wherein: The using the bidirectional cross-modal attention feature correction module to determine a first complementary feature related to the laser speckle feature map from the RGB feature map, and correcting the laser speckle feature map based on the first complementary feature to obtain a first correction feature includes: Using an average pooling function to reduce the height and width of the RGB feature map according to a preset ratio to obtain a downsampled RGB feature map; Using the formula: determining a first complementary feature; Among them, Att RGB is the first complementary characteristic, Q RGB is the query vector obtained from the downsampled RGB feature map, is the transpose of the bond vector obtained from the laser speckle feature map, V Speckle is the value vector obtained from the laser speckle feature map, h is the header number; Restoring the first complementary feature by upsampling and reshaping operations, and fusing the restored first complementary feature with the RGB feature map to obtain a fused feature; The multi-layer perceptron is used to process the fusion features and obtain the multi-layer perceptron result; The multi-layer perceptron result is fused with the fusion feature to obtain a first correction feature.
3. The method for detecting workpiece surface roughness according to claim 1, wherein: The parallel dual-pooling cross-modal fusion module is used to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature, including: Using a first average pooling layer and a first maximum pooling layer in parallel to process the first correction feature, respectively, to obtain a first average pooling result and a first maximum pooling result; Concatenate the first average pooling result and the first maximum pooling result to obtain a first double-pooling feature; Using a second average pooling layer and a second maximum pooling layer in parallel to process the second correction feature, respectively, to obtain a second average pooling result and a second maximum pooling result; Concatenate the second average pooling result and the second maximum pooling result to obtain a second double pooling feature; The first dual-pooling feature and the second dual-pooling feature are concatenated to obtain a cross-modal fusion feature.
4. The method for detecting workpiece surface roughness according to claim 1, wherein: The method further includes determining a first complementary feature related to the laser speckle feature map from the RGB feature map by using the bidirectional cross-modal attention feature correction module, and correcting the laser speckle feature map based on the first complementary feature to obtain a first correction feature. Obtaining training samples; the training samples include RGB feature map training data, laser speckle feature map training data and corresponding surface roughness results; determining a first training complementary feature related to the laser speckle feature map training data from the RGB feature map training data, and correcting the laser speckle feature map training data based on the first training complementary feature to obtain a first training correction feature; Determining a second training complementary feature related to the RGB feature map training data from the laser speckle feature map training data, and correcting the RGB feature map training data based on the second training complementary feature to obtain a second training corrected feature; Performing double-pooling processing on the first training correction feature and the second training correction feature respectively to obtain a first training double-pooling feature and a second training double-pooling feature; The first training double-pooling feature and the second training double-pooling feature are fused to obtain the cross-modal fusion feature of the training sample; Based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature and the cross-modal fusion feature of the training sample, the initial model of the explicit dual-pooling cross-modal alignment network is trained to obtain a trained explicit dual-pooling cross-modal alignment network model.
5. The method for detecting workpiece surface roughness according to claim 4, wherein: The step of training the explicit dual-pooling cross-modal alignment network initial model based on the surface roughness result, the first training dual-pooling feature, the second training dual-pooling feature, and the cross-modal fusion feature of the training sample to obtain the trained explicit dual-pooling cross-modal alignment network model includes: Calculating a first auxiliary loss function based on the first training dual-pooling feature and the surface roughness result; Calculating a second auxiliary loss function based on the second training dual-pooling feature and the surface roughness result; Calculating a main loss function based on the cross-modal fusion features of the training sample and the surface roughness result; The first auxiliary loss function, the second auxiliary loss function and the main loss function are summed to obtain the target loss function; Based on the target loss function, the parameters of the initial model of the explicit dual-pooling cross-modal alignment network are adjusted to obtain a trained explicit dual-pooling cross-modal alignment network model.
6. The method for detecting workpiece surface roughness according to claim 1, wherein: The cross-modal fusion features are classified by the fully connected layer to obtain a prediction result of the roughness of the surface of the workpiece to be detected, and then the method further includes: Acquiring experimental data of the surface roughness of the workpiece to be tested; The prediction results identical to the experimental data are confirmed as correctly classified samples, and the prediction results other than the correctly classified samples are confirmed as incorrectly classified samples; the correctly classified samples include correctly classified positive samples and correctly classified negative samples, and the incorrectly classified samples include incorrectly classified positive samples and incorrectly classified negative samples; Using the formula: Determine the accuracy of the forecast; Using the formula: Determine the recall of the predictions; Calculating the harmonic mean of the recall and the precision; Using the formula: Calculate the accuracy of the prediction; Among them, Precision is the accuracy, Recall is the recall rate, Accuracy is the accuracy rate, TP is the number of correctly classified positive samples, TN is the number of correctly classified negative samples, FN is the number of incorrectly classified positive samples, and FP is the number of incorrectly classified negative samples.
7. The method for detecting workpiece surface roughness according to claim 1, wherein: The obtaining of the RGB feature map and the laser speckle feature map of the surface of the workpiece to be inspected comprises: Acquire RGB images and laser speckle images of the surface of the workpiece to be inspected; Using the MobileNet-V3 Large model to perform feature extraction on the RGB image and the laser speckle image, respectively, to obtain a first extracted feature and a second extracted feature; The first extracted features are flattened and transposed to obtain an RGB feature map, and the second extracted features are flattened and transposed to obtain a laser speckle feature map.
8. A workpiece surface roughness detection device, characterized in that: The method for detecting the surface roughness of a workpiece according to any one of claims 1 to 7, wherein the device comprises: RGB feature map and laser speckle feature map acquisition module, used to obtain the RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected; a first correction module, configured to determine a first complementary feature related to the laser speckle feature map from the RGB feature map using the bidirectional cross-modal attention feature correction module, and correct the laser speckle feature map based on the first complementary feature to obtain a first correction feature; a second correction module, configured to determine a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correct the RGB feature map based on the second complementary feature to obtain a second correction feature; a cross-modal fusion module, configured to perform double-pooling and fusion processing on the first correction feature and the second correction feature using the parallel double-pooling cross-modal fusion module to obtain a cross-modal fusion feature; A classification module is used to classify the cross-modal fusion features using the fully connected layer to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
9. A workpiece surface roughness detection device, characterized in that: The workpiece surface roughness detection method according to any one of claims 1 to 7, wherein the device comprises: Communication unit / communication interface, used to obtain RGB feature map and laser speckle feature map of the surface of the workpiece to be inspected; a processing unit / processor, configured to determine, using the bidirectional cross-modal attention feature correction module, a first complementary feature associated with the laser speckle feature map from the RGB feature map, and correct the laser speckle feature map based on the first complementary feature to obtain a first corrected feature; determining a second complementary feature associated with the RGB feature map from the laser speckle feature map, and correcting the RGB feature map based on the second complementary feature to obtain a second corrected feature; Using the parallel dual-pooling cross-modal fusion module to perform dual-pooling and fusion processing on the first correction feature and the second correction feature to obtain a cross-modal fusion feature; The fully connected layer is used to classify the cross-modal fusion features to obtain a prediction result of the roughness of the surface of the workpiece to be inspected.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by the processor, the workpiece surface roughness detection method described in any one of claims 1 to 7 is implemented.