Clothing cultural relic non-contact measurement method based on deep learning and multi-modal large model
Through multimodal data acquisition and deep learning technology, combined with high-resolution cameras, laser scanners and infrared sensors, the problem of damage to clothing cultural relics caused by traditional measurement methods has been solved, and high-precision non-contact measurement and data consistency reconstruction have been achieved.
Patent Information
- Application Number
- CN202510848914.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional measurement methods cause physical damage to clothing artifacts and are difficult to accurately capture complex textures and deformed structures. The inconsistency problem of multi-angle data fusion has not been effectively solved.
It adopts a multimodal data acquisition module, a deep learning feature extraction module, a multimodal data fusion and verification module, and a 3D reconstruction and platemaking data generation module, combined with a high-resolution camera, a laser scanner, and an infrared sensor, and uses deep learning and a multimodal large model for contactless measurement. It ensures data consistency through timestamp synchronization, self-attention mechanism, and anomaly detection to generate high-precision platemaking data.
It achieves non-contact high-precision measurement, avoids damage to cultural relics, improves the accuracy of capturing complex textures and deformed structures, and ensures the consistency and reconstruction accuracy of multi-angle data.
Smart Images

Figure CN120707777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of measurement and digital reconstruction of clothing artifacts, and in particular to a contactless measurement method for clothing artifacts based on deep learning and a multimodal large model. Background Art
[0002] The measurement and digital reconstruction of clothing artifacts are of vital importance in cultural heritage protection. It helps to preserve artifact information for a long time, conduct research, and carry out artifact restoration.
[0003] However, traditional measurement methods have many limitations:
[0004] 1. Contact measurement, such as using a tape measure or caliper, involves direct contact with the artifact. This can cause physical damage to fragile, complex garments. For example, the fibers of ancient silk garments are extremely fragile, and friction from the tape measure can break the threads, damaging the integrity of the artifact.
[0005] Conventional non-contact measurement technologies, such as laser scanning and photogrammetry, often lack the required accuracy when processing complex textures and deformable structures. For example, for cultural relics with exquisite embroidery and intricate folds, laser scanning may not accurately capture the subtle textures of the embroidery, while photogrammetry may lose key details during reconstruction due to lighting and angle issues.
[0006] 3. Furthermore, the fusion and consistency verification of multi-angle measurement data presents a technical challenge. Existing methods are ineffective in handling data inconsistencies. For example, in images taken at different angles, the color and shape of the same location may deviate due to differences in lighting and shooting position. Existing fusion methods may not be able to effectively address these issues, resulting in distorted reconstruction results.
[0007] In recent years, deep learning and large multimodal models have made significant progress in fields such as image processing and 3D reconstruction. Deep learning technology can train models using large amounts of data, automatically extract features, and make high-precision predictions. For example, in the field of facial recognition, deep learning models can learn various features from massive amounts of facial images, achieving high-precision recognition. Large multimodal models can integrate data from different sensors to improve the consistency and accuracy of measurements. For example, in the field of autonomous driving, large multimodal models can fuse data from multiple sensors such as cameras and radar to more accurately perceive the surrounding environment. However, applying these technologies to the contactless measurement of clothing artifacts still faces challenges, especially in dealing with complex textures, deformable structures, and multi-angle data fusion, which existing technologies have not yet fully addressed. Summary of the Invention
[0008] The present invention aims to solve many problems of traditional measurement methods in the measurement of clothing artifacts, and provide a contactless measurement method for clothing artifacts based on deep learning and multimodal large models.
[0009] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0010] A contactless measurement method for clothing artifacts based on deep learning and multimodal large models, including a multimodal data acquisition module, a deep learning feature extraction module, a multimodal data fusion and verification module, a 3D reconstruction and plate-making data generation module, and an automated measurement process control module;
[0011] The multimodal data acquisition module uses a multi-sensor system of high-resolution cameras, laser scanners, and infrared sensors to collect multi-angle images and depth data of clothing artifacts. It uses timestamp synchronization technology and is equipped with an automatic calibration function to ensure data synchronization and consistency.
[0012] The deep learning feature extraction module uses a pre-trained convolutional neural network combined with an attention mechanism to automatically identify and extract the complex textures, edges, and deformation structures of clothing and cultural relics, accurately capturing important features;
[0013] The multimodal data fusion and verification module uses a large multimodal model with a Transformer architecture. It automatically aligns and integrates multi-angle measurement data through a self-attention mechanism, and introduces anomaly detection algorithms to eliminate inconsistent data points to ensure data consistency.
[0014] The 3D reconstruction and plate-making data generation module uses Poisson reconstruction and mesh generation algorithms to build a 3D model based on the fused multimodal data. It also uses deep learning optimization to accurately restore the details of the cultural relics and generate high-precision plate-making data.
[0015] The automated measurement process control module uses intelligent algorithms to achieve automated control of the measurement process, dynamically adjusts measurement parameters based on the characteristics of the cultural relics, and provides a user interface to support real-time monitoring and adjustment to ensure measurement accuracy.
[0016] In the multimodal data acquisition module, high-resolution cameras are used to capture the surface texture and color information of cultural relics; laser scanners are used to obtain the geometric shape and depth information of cultural relics; and infrared sensors are used to record the material properties of cultural relics, including thermal conductivity and reflectivity.
[0017] In the multimodal data acquisition module, timestamp synchronization technology is used to ensure data synchronization. Each sensor will be accurately timestamped when collecting data, ensuring that data collected by different sensors at the same time point can be accurately aligned. It is equipped with an automatic calibration function, which automatically adjusts the position and angle of the sensor through a preset calibration algorithm to ensure data consistency between different sensors.
[0018] The specific implementation steps of the deep learning feature extraction module are as follows:
[0019] S1. Data preprocessing: First, the collected multimodal data is preprocessed, including image denoising, normalization, and data enhancement. Image denoising uses the Gaussian filter algorithm, and the formula is:
[0020]
[0021] in,
[0022] (x,y) is the coordinate of the pixel relative to the center of the kernel;
[0023] G(x,y) is the Gaussian kernel function;
[0024] σ is the standard deviation, which is used to control the smoothness of the filter;
[0025] Normalization scales the image pixel values to the range [0,1], and the formula is:
[0026]
[0027] in,
[0028] I norm (x,y) is the normalized image pixel value;
[0029] I(x,y) is the original image pixel value;
[0030] I min and I max are the minimum and maximum pixel values of the image, respectively;
[0031] Data augmentation includes random rotation, flipping, and scaling to increase the diversity of training data;
[0032] S2. Convolutional neural network feature extraction: Use the pre-trained convolutional neural network model for feature extraction. The convolutional neural network gradually extracts low-level to high-level features of the image through multi-layer convolution and pooling operations. The convolution operation formula is:
[0033]
[0034] in,
[0035] C(x,y) is the convolution output, which represents the value at position (x,y) in the output feature map;
[0036] k is the size of the convolution kernel;
[0037] W(i,j) is the weight of the convolution kernel at position (i,j);
[0038] I(x+i,y+j) is the local area of the input image;
[0039] b is the bias term;
[0040] The pooling operation uses maximum pooling, and the formula is:
[0041]
[0042] in,
[0043] P(x,y) is the pooling output;
[0044] R is the pooling window area;
[0045] S3. Attention mechanism: The attention mechanism calculates the attention weight of the feature map. The formula is:
[0046]
[0047] in,
[0048] A(x,y) is the attention weight;
[0049] S(x,y) is the score of the feature map at position (x,y);
[0050] S(i,j) is the similarity score between the i-th element and the j-th element in the input sequence, which is the similarity between the query vector and the key vector;
[0051] The final feature map is weighted summed by attention weights, and the formula is:
[0052] F att (x,y)=A(x,y)·F(x,y)
[0053] in,
[0054] F att (x,y) is the weighted feature map;
[0055] F(x,y) is the original feature map;
[0056] S4. Feature fusion: Fuse features from different modalities to generate a comprehensive feature representation. Feature fusion uses the weighted average method, and the formula is:
[0057]
[0058] in,
[0059] F fusion (x,y) is the fused feature;
[0060] m is the index of the modality, indicating different data sources or types;
[0061] M is the number of modalities, indicating the total number of modalities involved in fusion;
[0062] w m is the weight of the mth mode;
[0063] F m (x,y) is the feature of the mth mode.
[0064] The specific implementation steps of the multimodal data fusion and verification module are as follows:
[0065] P1. Data alignment and preprocessing: First, align and preprocess the multi-angle measurement data from different sensors. The alignment operation matches the timestamp and spatial position information to ensure that the data collected by different sensors at the same time point can accurately correspond. Preprocessing includes data normalization and denoising. The normalization formula is:
[0066]
[0067] in,
[0068] D norm (x,y) is the normalized data value;
[0069] D(x,y) is the original data value;
[0070] D min and D max are the minimum and maximum values of the data respectively;
[0071] Denoising uses the wavelet transform algorithm, the formula is:
[0072]
[0073] in,
[0074] W(a,b) is the wavelet coefficient;
[0075] a is the scale parameter;
[0076] b is the translation parameter;
[0077] D(t) is the original signal that needs to be wavelet transformed;
[0078] ψ(t) is the wavelet basis function;
[0079] P2. Transformer-based multimodal model: A Transformer-based multimodal model is used for data fusion. The Transformer model automatically learns the relationship between different modal data through the self-attention mechanism. The calculation formula of the self-attention mechanism is:
[0080]
[0081] in,
[0082] Q is the query matrix;
[0083] K is the bond matrix;
[0084] V is the value matrix;
[0085] d k is the dimension of the key vector;
[0086] Through the multi-head attention mechanism, the model learns richer feature representations from different subspaces. The formula is:
[0087] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O
[0088] in,
[0089] head i =Attention(QW i Q ,KW i K ,VW i V ), W i Q 、W i K 、W i V and W O is a learnable weight matrix;
[0090] h is the number of attention heads;
[0091] P3. Anomaly Detection Algorithm: Introduce an anomaly detection algorithm based on isolation forest to automatically identify and remove inconsistent or abnormal data points. Isolation forest constructs a tree structure by randomly selecting features and split values. Anomalies are usually isolated at shallow nodes. The anomaly detection formula is:
[0092]
[0093] in,
[0094] h(x) is the path length of data point x in the tree;
[0095] E(h(x)) is the expected value of the path length;
[0096] c(n) is the average path length of the tree;
[0097] n is the sample size;
[0098] By setting a threshold, outliers are removed from the data set;
[0099] P4. Data consistency check: The self-attention mechanism of the multimodal large model is used to check the consistency of the fused data. The consistency check is performed by calculating the similarity between data of different modalities. The formula is:
[0100]
[0101] in,
[0102] M1 and M2 are data representations of different modalities;
[0103] Represents the dot product;
[0104] ||·|| represents the norm;
[0105] By setting the similarity threshold, the consistency of different modal data in key areas is ensured.
[0106] The specific steps of the 3D reconstruction and plate making data generation module are as follows:
[0107] H1. Poisson reconstruction algorithm: The Poisson reconstruction algorithm generates the surface of a 3D model by solving the Poisson equation. First, the point cloud data is converted into a signed distance field, and then the Poisson equation is solved to reconstruct the surface. The form of the Poisson equation is:
[0108]
[0109] in,
[0110] is a scalar field;
[0111] v is a vector field;
[0112] △ is the Laplace operator;
[0113] ▽· is the divergence operator;
[0114] By solving the equation, the surface of the three-dimensional model can be obtained. In order to improve the reconstruction accuracy, the algorithm introduces a multi-scale reconstruction strategy, which solves the Poisson equation at different scales and then fuses the results. The formula is:
[0115]
[0116] in,
[0117] is the indicator function of the final reconstructed surface;
[0118] is the reconstruction result of the i-th scale;
[0119] w i is the weight coefficient;
[0120] n is the number of scales;
[0121] H2. Mesh generation algorithm: Based on Poisson reconstruction, a mesh generation algorithm is used to convert the reconstructed surface into a triangular mesh. The mesh generation algorithm generates a high-quality triangular mesh through Delaunay triangulation and surface smoothing technology. The formula for Delaunay triangulation is:
[0122]
[0123] in,
[0124] T i is the i-th triangle;
[0125] Area(T i ) is the area of the triangle;
[0126] The surface smoothing technique reduces the noise on the mesh surface by using the Laplace smoothing algorithm. The formula is:
[0127]
[0128] in,
[0129] v i new is the new coordinate position after smoothing;
[0130] v i is the position of vertex i;
[0131] N(i) is the neighborhood of vertex i;
[0132] λ is the smoothing coefficient;
[0133] H3. Deep learning optimization: Generative adversarial networks are used for detail enhancement. The generator network generates a high-resolution 3D model through multi-layer convolution and deconvolution operations, while the discriminator network judges the authenticity of the generated model through multi-layer convolution operations. The loss function of GAN is:
[0134]
[0135] in,
[0136] G is the generator;
[0137] D is the discriminator;
[0138] x is the real data;
[0139] z is the noise vector;
[0140] Through adversarial training, the generator can generate high-precision 3D models;
[0141] H4. Plate making data generation: Based on the optimization of the 3D model, high-precision plate making data is generated. The plate making data includes size, shape and texture information. The calculation formula for size data is:
[0142] Length=max(x i )-min(x i )
[0143] Width=max(y i )-min(y i )
[0144] Height=max(z i )-min(z i )
[0145] in,
[0146] x i ,y i ,z i are the vertex coordinates of the three-dimensional model;
[0147] Texture information is generated through texture mapping technology, which maps a two-dimensional texture image onto the surface of a three-dimensional model.
[0148] The specific steps for implementing the automated measurement process control module are as follows:
[0149] K1. Intelligent Algorithm Control: This system automatically controls the entire measurement process by combining rule-based intelligent algorithms with machine learning models. First, the system uses a preset rule library to make preliminary judgments on the basic characteristics of artifacts, including material, complexity, and size range. Second, the system uses machine learning models to perform more detailed classification and prediction of artifact characteristics. The machine learning model's inputs include the artifact's material, texture complexity, and geometric shape, and its output is the optimal measurement parameters, including scanning accuracy, shooting angle, and sensor type. The machine learning model's training data is derived from historical measurement records, and prediction accuracy is improved through continuous optimization of model parameters.
[0150] K2. Dynamic parameter adjustment: During the measurement process, the system dynamically adjusts the measurement parameters based on the real-time collected data. The dynamic adjustment algorithm is based on feedback control theory. The specific formula is:
[0151]
[0152] in,
[0153] u(t) is the adjusted measurement parameter;
[0154] e(t) is the error between the current measurement data and the target value;
[0155] K p ,K i ,K d are the proportional, integral and differential coefficients respectively;
[0156] The system can adjust measurement parameters in real time to ensure measurement accuracy;
[0157] K3. User interface support: The module provides a user-friendly graphical interface that allows operators to monitor the measurement process in real time and make adjustments and interventions. The interface also supports data visualization, showing measurement results through charts and 3D models, helping operators understand the data more intuitively.
[0158] K4. Anomaly detection and processing: The module integrates an anomaly detection algorithm that can automatically identify and process anomalies that occur during the measurement process. The anomaly detection algorithm is based on statistical methods and machine learning models. It identifies abnormal data points by analyzing the distribution characteristics of the measurement data.
[0159] In step K3, the content displayed on the interface includes the current measurement progress, real-time collected data, and adjustment suggestions for measurement parameters.
[0160] In step K4, the strategies for handling the exception include re-collecting data, adjusting measurement parameters, or switching sensor types.
[0161] The beneficial effects of the present invention are: the present invention adopts deep learning combined with a multimodal large model for verification, so that accurate plate-making data can be obtained without touching the cultural relics. It is specifically aimed at the special needs of clothing cultural relics, and can effectively solve the problems of insufficient accuracy and data inconsistency of traditional measurement methods in the measurement of clothing cultural relics, providing more efficient and accurate technical support for the protection of cultural heritage. BRIEF DESCRIPTION OF THE DRAWINGS
[0162] Figure 1 Schematic diagram of the multimodal data acquisition module of the present invention;
[0163] Figure 2 Schematic diagram of the deep learning feature extraction module in the present invention;
[0164] Figure 3 Schematic diagram of the multimodal data fusion and verification module in the present invention;
[0165] Figure 4 Schematic diagram of a 3D reconstruction and plate-making data generation module in the present invention;
[0166] Figure 5 Schematic diagram of the automated measurement process control module of the present invention;
[0167] The following is a detailed description of the embodiments of the present invention with reference to the accompanying drawings. DETAILED DESCRIPTION
[0168] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention. The following paragraphs describe the present invention in more detail by way of example with reference to the accompanying drawings. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly illustrate the purpose of the embodiments of the present invention.
[0169] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0170] The present invention will be further described below with reference to the accompanying drawings and examples:
[0171] The contactless measurement method for clothing artifacts based on deep learning and multimodal large models includes a multimodal data acquisition module, a deep learning feature extraction module, a multimodal data fusion and verification module, a three-dimensional reconstruction and plate-making data generation module, and an automated measurement process control module.
[0172] 1. Multimodal data acquisition module:
[0173] like Figure 1As shown, the multimodal data acquisition module uses a multi-sensor system consisting of a high-resolution camera, a laser scanner, and an infrared sensor to collect multi-angle images and depth data of clothing artifacts. It utilizes timestamp synchronization technology and is equipped with an automatic calibration function to ensure data synchronization and consistency. Specifically, the high-resolution camera captures the artifact's surface texture and color information; the laser scanner acquires its geometry and depth information; and the infrared sensor records its material properties, such as thermal conductivity and reflectivity. To ensure data synchronization, timestamp synchronization technology is used. Each sensor accurately timestamps its data when it is collected, ensuring accurate alignment of data collected by different sensors at the same time. Furthermore, an automatic calibration function automatically adjusts the sensor position and angle using a pre-set calibration algorithm to ensure data consistency across sensors. For example, during the measurement process, the camera and laser scanner capture and scan at the same angular interval, while the infrared sensor records the artifact's material information in real time. The module automatically calibrates each sensor to ensure consistent shooting angles and timing.
[0174] Specific implementation method: In actual operation, the clothing artifact is first placed on the measuring platform to ensure its fixed position. Then, the multimodal data acquisition module is activated. A high-resolution camera captures multi-angle images of the artifact at regular intervals. A laser scanner scans at the same angle to obtain the artifact's geometric shape and depth information. The infrared sensor records the artifact's material properties in real time. Each sensor is accurately timestamped when collecting data, ensuring that data collected by different sensors at the same time is accurately aligned. The module automatically calibrates each sensor to ensure consistency in shooting angle and time. For example, when measuring a clothing artifact with exquisite embroidery, the camera can accurately capture the color and texture of the embroidery pattern, the laser scanner can accurately obtain the embroidery's geometric shape and depth information, and the infrared sensor records the characteristics of the embroidery material. Through timestamp synchronization technology and automatic calibration functions, the data collected by different sensors are accurately aligned, generating high-quality multimodal data, providing a reliable foundation for subsequent feature extraction and 3D reconstruction.
[0175] 2. Deep learning feature extraction module:
[0176] like Figure 2 As shown in the figure, the deep learning feature extraction module uses a pre-trained convolutional neural network (CNN) combined with an attention mechanism to automatically identify and extract the complex textures, edges, and deformation structures of clothing artifacts, accurately capturing important features.
[0177] The specific implementation steps of the deep learning feature extraction module are as follows:
[0178] S1. Data preprocessing: First, the collected multimodal data is preprocessed, including image denoising, normalization, and data enhancement. Image denoising uses the Gaussian filter algorithm, and the formula is:
[0179]
[0180] in,
[0181] (x,y) is the coordinate of the pixel relative to the center of the kernel;
[0182] G(x,y) is the Gaussian kernel function;
[0183] σ is the standard deviation, which is used to control the smoothness of the filter;
[0184] Normalization scales the image pixel values to the range [0,1], and the formula is:
[0185]
[0186] in,
[0187] I norm (x,y) is the normalized image pixel value;
[0188] I(x,y) is the original image pixel value;
[0189] I min and I max are the minimum and maximum pixel values of the image, respectively;
[0190] Data augmentation includes random rotation, flipping, and scaling to increase the diversity of training data;
[0191] S2. Convolutional neural network feature extraction: Use the pre-trained convolutional neural network model for feature extraction. The convolutional neural network gradually extracts low-level to high-level features of the image through multi-layer convolution and pooling operations. The convolution operation formula is:
[0192]
[0193] in,
[0194] C(x,y) is the convolution output, which represents the value at position (x,y) in the output feature map;
[0195] k is the size of the convolution kernel;
[0196] W(i,j) is the weight of the convolution kernel at position (i,j);
[0197] I(x+i,y+j) is the local area of the input image;
[0198] b is the bias term;
[0199] The pooling operation uses maximum pooling, and the formula is:
[0200]
[0201] in,
[0202] P(x,y) is the pooling output;
[0203] R is the pooling window area;
[0204] S3. Attention mechanism: The attention mechanism calculates the attention weight of the feature map. The formula is:
[0205]
[0206] in,
[0207] A(x,y) is the attention weight;
[0208] S(x,y) is the score of the feature map at position (x,y);
[0209] S(i,j) is the similarity score between the i-th element and the j-th element in the input sequence, which is the similarity between the query vector and the key vector;
[0210] The final feature map is weighted summed by attention weights, and the formula is:
[0211] F att (x,y)=A(x,y)·F(x,y)
[0212] in,
[0213] F att (x,y) is the weighted feature map;
[0214] F(x,y) is the original feature map;
[0215] S4. Feature fusion: Fuse features from different modalities to generate a comprehensive feature representation. Feature fusion uses the weighted average method, and the formula is:
[0216]
[0217] in,
[0218] F fusion (x,y) is the fused feature;
[0219] m is the index of the modality, indicating different data sources or types;
[0220] M is the number of modalities, indicating the total number of modalities involved in fusion;
[0221] wm is the weight of the mth mode;
[0222] F m (x,y) is the feature of the mth mode.
[0223] Specific implementation method: In actual operation, the collected multimodal data is first preprocessed. For example, an image of a clothing artifact with exquisite embroidery is subjected to Gaussian filtering and denoising to remove random noise in the image. Then, the image pixel values are normalized to the range of [0,1], and data enhancement operations such as random rotation, flipping, and scaling are performed. Next, a pre-trained ResNet-50 model is used for feature extraction. Through multi-layer convolution and pooling operations, low-level to high-level features of the image are gradually extracted. The introduction of an attention mechanism enables the model to focus on embroidery patterns, folds, and stitching details, ignoring surrounding background interference. Finally, features from different modalities are fused to generate a comprehensive feature representation, providing a reliable foundation for subsequent multimodal data fusion and verification. For example, when processing a clothing artifact with complex folds, the model can accurately extract the geometric shape and texture features of the folds, while ignoring irrelevant background information and generating high-quality feature representations.
[0224] 3. Multimodal data fusion and verification module:
[0225] like Figure 3 As shown in the figure, the multimodal data fusion and verification module adopts a large multimodal model based on the Transformer architecture. It automatically aligns and integrates multi-angle measurement data from different sensors through the self-attention mechanism, and introduces anomaly detection algorithm to eliminate inconsistent or abnormal data points to ensure data consistency and accuracy.
[0226] The specific implementation steps of the multimodal data fusion and verification module are as follows:
[0227] P1. Data alignment and preprocessing: First, align and preprocess the multi-angle measurement data from different sensors. The alignment operation matches the timestamp and spatial position information to ensure that the data collected by different sensors at the same time point can accurately correspond. Preprocessing includes data normalization and denoising. The normalization formula is:
[0228]
[0229] in,
[0230] D norm (x,y) is the normalized data value;
[0231] D(x,y) is the original data value;
[0232] D min and Dmax are the minimum and maximum values of the data respectively;
[0233] Denoising uses the wavelet transform algorithm, the formula is:
[0234]
[0235] in,
[0236] W(a,b) is the wavelet coefficient;
[0237] a is the scale parameter;
[0238] b is the translation parameter;
[0239] D(t) is the original signal that needs to be wavelet transformed;
[0240] ψ(t) is the wavelet basis function;
[0241] P2. Transformer-based multimodal model: A Transformer-based multimodal model is used for data fusion. The Transformer model automatically learns the relationship between different modal data through the self-attention mechanism. The calculation formula of the self-attention mechanism is:
[0242]
[0243] in,
[0244] Q is the query matrix;
[0245] K is the bond matrix;
[0246] V is the value matrix;
[0247] d k is the dimension of the key vector;
[0248] Through the multi-head attention mechanism, the model learns richer feature representations from different subspaces. The formula is:
[0249] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O in,
[0250] head i =Attention(QW i Q ,KW i K ,VW i V ), W i Q 、Wi K 、W i V and W O is a learnable weight matrix;
[0251] h is the number of attention heads;
[0252] P3. Anomaly Detection Algorithm: Introduce an anomaly detection algorithm based on isolation forest to automatically identify and remove inconsistent or abnormal data points. Isolation forest constructs a tree structure by randomly selecting features and split values. Anomalies are usually isolated at shallow nodes. The anomaly detection formula is:
[0253]
[0254] in,
[0255] h(x) is the path length of data point x in the tree;
[0256] E(h(x)) is the expected value of the path length;
[0257] c(n) is the average path length of the tree;
[0258] n is the sample size;
[0259] By setting a threshold, outliers are removed from the data set;
[0260] P4. Data consistency check: The self-attention mechanism of the multimodal large model is used to check the consistency of the fused data. The consistency check is performed by calculating the similarity between data of different modalities. The formula is:
[0261]
[0262] in,
[0263] M1 and M2 are data representations of different modalities;
[0264] Represents the dot product;
[0265] ||·|| represents the norm;
[0266] By setting the similarity threshold, the consistency of different modal data in key areas is ensured.
[0267] Specific implementation: In practice, multi-angle measurement data from different sensors is first aligned and preprocessed. For example, the timestamps and spatial positions of camera images and depth information acquired by a laser scanner are matched to ensure that the data correspond at the same time and spatial location. The data is then normalized and denoised to remove noise interference. Next, a large multimodal model based on the Transformer architecture is used for data fusion, automatically learning the relationships between the different modal data through a self-attention mechanism.
[0268] 4. 3D reconstruction and plate making data generation module:
[0269] like Figure 4 As shown, the 3D reconstruction and plate-making data generation module builds a 3D model based on the fused multimodal data using Poisson reconstruction and mesh generation algorithms, and accurately restores the details of the cultural relics through deep learning optimization to generate high-precision plate-making data.
[0270] The specific steps of the 3D reconstruction and plate making data generation module are as follows:
[0271] H1. Poisson reconstruction algorithm: The Poisson reconstruction algorithm generates the surface of a 3D model by solving the Poisson equation. First, the point cloud data is converted into a signed distance field, and then the Poisson equation is solved to reconstruct the surface. The form of the Poisson equation is:
[0272]
[0273] in,
[0274] is a scalar field;
[0275] v is a vector field;
[0276] △ is the Laplace operator;
[0277] ▽· is the divergence operator;
[0278] By solving the equation, the surface of the three-dimensional model can be obtained. In order to improve the reconstruction accuracy, the algorithm introduces a multi-scale reconstruction strategy, which solves the Poisson equation at different scales and then fuses the results. The formula is:
[0279]
[0280] in,
[0281] is the indicator function of the final reconstructed surface;
[0282] is the reconstruction result of the i-th scale;
[0283] wi is the weight coefficient;
[0284] n is the number of scales;
[0285] H2. Mesh generation algorithm: Based on Poisson reconstruction, a mesh generation algorithm is used to convert the reconstructed surface into a triangular mesh. The mesh generation algorithm generates a high-quality triangular mesh through Delaunay triangulation and surface smoothing technology. The formula for Delaunay triangulation is:
[0286]
[0287] in,
[0288] T i is the i-th triangle;
[0289] Area(T i ) is the area of the triangle;
[0290] The surface smoothing technique reduces the noise on the mesh surface by using the Laplace smoothing algorithm. The formula is:
[0291]
[0292] in,
[0293] v i new is the new coordinate position after smoothing;
[0294] v i is the position of vertex i;
[0295] N(i) is the neighborhood of vertex i;
[0296] λ is the smoothing coefficient;
[0297] H3. Deep learning optimization: Generative adversarial networks are used for detail enhancement. The generator network generates a high-resolution 3D model through multi-layer convolution and deconvolution operations, while the discriminator network judges the authenticity of the generated model through multi-layer convolution operations. The loss function of GAN is:
[0298]
[0299] in,
[0300] G is the generator;
[0301] D is the discriminator;
[0302] x is the real data;
[0303] z is the noise vector;
[0304] Through adversarial training, the generator can generate high-precision 3D models;
[0305] H4. Plate making data generation: Based on the optimization of the 3D model, high-precision plate making data is generated. The plate making data includes size, shape and texture information. The calculation formula for size data is:
[0306] Length=max(x i )-min(x i )
[0307] Width=max(y i )-min(y i )
[0308] Height=max(z i )-min(z i )
[0309] in,
[0310] x i ,y i ,z i are the vertex coordinates of the three-dimensional model;
[0311] Texture information is generated through texture mapping technology, which maps a two-dimensional texture image onto the surface of a three-dimensional model.
[0312] Specific implementation: In practice, the fused multimodal data is first converted into the surface of a 3D model using a Poisson reconstruction algorithm. For example, for a garment artifact with complex folds, the Poisson reconstruction algorithm can accurately reconstruct its surface shape. A mesh generation algorithm is then used to convert the reconstructed surface into a triangular mesh, and surface smoothing techniques are used to reduce noise on the mesh surface. Next, a generative adversarial network is used to optimize the 3D model, accurately recreating the artifact's details. Finally, highly accurate plate-making data is generated, including size, shape, and texture information.
[0313] 5. Automated measurement process control module:
[0314] like Figure 5 As shown, the automated measurement process control module realizes automated control of the measurement process through intelligent algorithms, dynamically adjusts measurement parameters according to the characteristics of cultural relics, and provides a user interface to support real-time monitoring and adjustment to ensure measurement accuracy.
[0315] The specific steps for implementing the automated measurement process control module are as follows:
[0316] K1. Intelligent algorithm control: The entire measurement process is automatically controlled by combining rule-based intelligent algorithms and machine learning models. First, the system makes a preliminary judgment on the basic characteristics of the cultural relics (such as material, complexity, size range, etc.) through a preset rule library. For example, for clothing cultural relics made of silk, the system will automatically reduce the scanning intensity of the laser scanner to avoid damage to fragile fibers. Secondly, the system uses machine learning models (such as support vector machines (SVM)) to perform more detailed classification and prediction of the characteristics of cultural relics. The input of the machine learning model includes the material, texture complexity, geometric shape and other characteristics of the cultural relics, and the output is the optimal measurement parameters, such as scanning accuracy, shooting angle, sensor type, etc. The training data of the machine learning model comes from historical measurement records, and the accuracy of the prediction is improved by continuously optimizing the model parameters.
[0317] K2. Dynamic parameter adjustment: During the measurement process, the system dynamically adjusts the measurement parameters based on the real-time collected data. For example, when the system detects that the texture of a certain part is too complex, it will automatically increase the camera resolution or increase the scanning density of the laser scanner. The dynamic adjustment algorithm is based on feedback control theory, and the specific formula is:
[0318]
[0319] in,
[0320] u(t) is the adjusted measurement parameter;
[0321] e(t) is the error between the current measurement data and the target value;
[0322] K p ,K i ,K d are the proportional, integral and differential coefficients respectively;
[0323] The system can adjust measurement parameters in real time to ensure measurement accuracy;
[0324] K3. User interface support: The module provides a user-friendly graphical interface that allows operators to monitor the measurement process in real time and make necessary adjustments and interventions. The interface displays content including the current measurement progress, real-time collected data, and measurement parameter adjustment suggestions. For example, the operator can view the scanning results of the laser scanner on the interface and manually adjust the scanning angle or intensity according to the actual situation. The interface also supports data visualization, showing the measurement results through charts and 3D models, helping operators to understand the data more intuitively.
[0325] K4. Anomaly Detection and Handling: The module also integrates an anomaly detection algorithm that automatically identifies and handles anomalies that occur during the measurement process. Based on statistical methods and machine learning models, the anomaly detection algorithm analyzes the distribution characteristics of the measured data to identify anomalous data points. For example, if the system detects that a portion of the measured data deviates significantly from the expected value, it automatically marks the data point and prompts the operator to review or remeasure. Strategies for handling anomalies include recollecting data, adjusting measurement parameters, or switching sensor types.
[0326] Specific implementation: In actual operation, the system first uses a rule base and machine learning model to make a preliminary assessment and prediction of the artifact's characteristics and determine the optimal measurement parameters. For example, for a garment with intricate embroidery, the system automatically increases the camera's resolution and the laser scanner's scanning density. During the measurement process, the system dynamically adjusts measurement parameters based on real-time data to ensure accuracy. For example, if the system detects an overly complex texture on a particular part, it automatically increases the laser scanner's scanning density. The operator can monitor the measurement process in real time through the user interface and make manual adjustments as needed. For example, the operator can view the laser scanner's scan results on the interface and manually adjust the scanning angle or intensity based on the actual situation. Furthermore, the system automatically detects and handles anomalies during the measurement process. For example, if the system detects a significant deviation from the expected value for a particular part, it automatically marks the data point and prompts the operator to review or remeasure. Through this automated process, the system efficiently and accurately completes the measurement of garment artifacts.
[0327] The purpose of this invention is to provide a non-contact, high-precision measurement method for clothing artifacts based on deep learning and a multimodal large model, addressing the many issues with traditional measurement methods for clothing artifacts. Traditional contact measurement can cause physical damage to artifacts, which is unacceptable for precious clothing artifacts.
[0328] The present invention uses non-contact technology to avoid direct contact with the cultural relics, ensuring their integrity and preventing any damage during the measurement process. Traditional non-contact measurement technology has shortcomings in terms of accuracy and cannot accurately capture the complex textures and deformed structures of clothing cultural relics.
[0329] This invention leverages deep learning technology, extensive training data, and advanced algorithms to accurately identify and extract various features of cultural relics, thereby improving measurement accuracy. Traditional methods for processing multi-angle measurement data are ineffective in fusion and consistency verification, resulting in distorted reconstruction results.
[0330] This invention uses a large multimodal model to integrate measurement data from different sensors, automatically aligning and integrating multi-angle measurement data to ensure data consistency and accuracy. Traditional measurement methods often require extensive manual intervention and are inefficient. This invention automates and intelligentizes the measurement process, reducing manual labor and improving measurement efficiency.
[0331] The core innovation of this invention lies in the use of deep learning combined with a multimodal large-scale model for verification, enabling accurate plate-making data to be obtained without touching the artifacts. This approach specifically addresses the unique needs of clothing artifacts. This method effectively addresses the limitations of traditional measurement methods for clothing artifacts, such as insufficient precision and inconsistent data, and provides more efficient and accurate technical support for cultural heritage protection.
[0332] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various improvements are made using the method concept and technical solution of the present invention, or they are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.
Claims
1. A contactless measurement method for clothing artifacts based on deep learning and multimodal large models, characterized by: It includes multimodal data acquisition module, deep learning feature extraction module, multimodal data fusion and verification module, 3D reconstruction and plate making data generation module and automated measurement process control module; The multimodal data acquisition module uses a multi-sensor system of high-resolution cameras, laser scanners, and infrared sensors to collect multi-angle images and depth data of clothing artifacts. It uses timestamp synchronization technology and is equipped with an automatic calibration function to ensure data synchronization and consistency. The deep learning feature extraction module uses a pre-trained convolutional neural network combined with an attention mechanism to automatically identify and extract the complex textures, edges, and deformation structures of clothing and cultural relics, accurately capturing important features; The multimodal data fusion and verification module uses a large multimodal model with a Transformer architecture. It automatically aligns and integrates multi-angle measurement data through a self-attention mechanism, and introduces anomaly detection algorithms to eliminate inconsistent data points to ensure data consistency. The 3D reconstruction and plate-making data generation module uses Poisson reconstruction and mesh generation algorithms to build a 3D model based on the fused multimodal data. It also uses deep learning optimization to accurately restore the details of the cultural relics and generate high-precision plate-making data. The automated measurement process control module uses intelligent algorithms to achieve automated control of the measurement process, dynamically adjusts measurement parameters based on the characteristics of the cultural relics, and provides a user interface to support real-time monitoring and adjustment to ensure measurement accuracy.
2. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 1 is characterized in that: In the multimodal data acquisition module, high-resolution cameras are used to capture the surface texture and color information of cultural relics; laser scanners are used to obtain the geometric shape and depth information of cultural relics; and infrared sensors are used to record the material properties of cultural relics, including thermal conductivity and reflectivity.
3. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 2 is characterized in that: In the multimodal data acquisition module, timestamp synchronization technology is used to ensure data synchronization. Each sensor will be accurately timestamped when collecting data, ensuring that data collected by different sensors at the same time point can be accurately aligned. It is equipped with an automatic calibration function, which automatically adjusts the position and angle of the sensor through a preset calibration algorithm to ensure data consistency between different sensors.
4. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 1 is characterized in that: The specific implementation steps of the deep learning feature extraction module are as follows: S1. Data preprocessing: First, the collected multimodal data is preprocessed, including image denoising, normalization, and data enhancement. Image denoising uses the Gaussian filter algorithm, and the formula is: in, (x,y) is the coordinate of the pixel relative to the center of the kernel; G(x,y) is the Gaussian kernel function; σ is the standard deviation, which is used to control the smoothness of the filter; Normalization scales the image pixel values to the range [0,1], and the formula is: in, I norm (x,y) is the normalized image pixel value; I(x,y) is the original image pixel value; I min and I max are the minimum and maximum pixel values of the image, respectively; Data augmentation includes random rotation, flipping, and scaling to increase the diversity of training data; S2. Convolutional neural network feature extraction: Use the pre-trained convolutional neural network model for feature extraction. The convolutional neural network gradually extracts low-level to high-level features of the image through multi-layer convolution and pooling operations. The convolution operation formula is: in, C(x,y) is the convolution output, which represents the value at position (x,y) in the output feature map; k is the size of the convolution kernel; W(i,j) is the weight of the convolution kernel at position (i,j); I(x+i,y+j) is the local area of the input image; b is the bias term; The pooling operation uses maximum pooling, and the formula is: in, P(x,y) is the pooling output; R is the pooling window area; S3. Attention mechanism: The attention mechanism calculates the attention weight of the feature map. The formula is: in, A(x,y) is the attention weight; S(x,y) is the score of the feature map at position (x,y); S(i,j) is the similarity score between the i-th element and the j-th element in the input sequence, which is the similarity between the query vector and the key vector; The final feature map is weighted summed by attention weights, and the formula is: F att (x,y)=A(x,y)·F(x,y) in, F att (x,y) is the weighted feature map; F(x,y) is the original feature map; S4. Feature fusion: Fuse features from different modalities to generate a comprehensive feature representation. Feature fusion uses the weighted average method, and the formula is: in, F fusion (x,y) is the fused feature; m is the index of the modality, indicating different data sources or types; M is the number of modalities, indicating the total number of modalities involved in fusion; w m is the weight of the mth mode; F m (x,y) is the feature of the mth mode.
5. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 1 is characterized in that: The specific implementation steps of the multimodal data fusion and verification module are as follows: P1. Data alignment and preprocessing: First, align and preprocess the multi-angle measurement data from different sensors. The alignment operation matches the timestamp and spatial position information to ensure that the data collected by different sensors at the same time point can accurately correspond. Preprocessing includes data normalization and denoising. The normalization formula is: in, D norm (x,y) is the normalized data value; D(x,y) is the original data value; D min and D max are the minimum and maximum values of the data respectively; Denoising uses the wavelet transform algorithm, the formula is: in, W(a,b) is the wavelet coefficient; a is the scale parameter; b is the translation parameter; D(t) is the original signal that needs to be wavelet transformed; ψ(t) is the wavelet basis function; P2. Transformer-based multimodal model: A Transformer-based multimodal model is used for data fusion. The Transformer model automatically learns the relationship between different modal data through the self-attention mechanism. The calculation formula of the self-attention mechanism is: in, Q is the query matrix; K is the bond matrix; V is the value matrix; d k is the dimension of the key vector; Through the multi-head attention mechanism, the model learns richer feature representations from different subspaces. The formula is: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O in, head i =Attention(QW i Q ,KW i K ,VW i V ), W i Q 、W i K 、W i V and W O is a learnable weight matrix; h is the number of attention heads; P3. Anomaly Detection Algorithm: Introduce an anomaly detection algorithm based on isolation forest to automatically identify and remove inconsistent or abnormal data points. Isolation forest constructs a tree structure by randomly selecting features and split values. Anomalies are usually isolated at shallow nodes. The anomaly detection formula is: in, h(x) is the path length of data point x in the tree; E(h(x)) is the expected value of the path length; c(n) is the average path length of the tree; n is the sample size; By setting a threshold, outliers are removed from the data set; P4. Data consistency check: The self-attention mechanism of the multimodal large model is used to perform consistency check on the fused data. The consistency check is performed by calculating the similarity between data of different modalities. The formula is: in, M1 and M2 are data representations of different modalities; Represents the dot product; ||·|| represents the norm; By setting the similarity threshold, the consistency of different modal data in key areas is ensured.
6. The contactless measurement method for clothing and cultural relics based on deep learning and multimodal large models according to claim 1 is characterized in that: The specific steps of the 3D reconstruction and plate making data generation module are as follows: H1. Poisson reconstruction algorithm: The Poisson reconstruction algorithm generates the surface of a 3D model by solving the Poisson equation. First, the point cloud data is converted into a signed distance field, and then the Poisson equation is solved to reconstruct the surface. The form of the Poisson equation is: in, is a scalar field; v is a vector field; △ is the Laplace operator; ▽· is the divergence operator; By solving the equation, the surface of the three-dimensional model can be obtained. In order to improve the reconstruction accuracy, the algorithm introduces a multi-scale reconstruction strategy, which solves the Poisson equation at different scales and then fuses the results. The formula is: in, is the indicator function of the final reconstructed surface; is the reconstruction result of the i-th scale; w i is the weight coefficient; n is the number of scales; H2. Mesh generation algorithm: Based on Poisson reconstruction, a mesh generation algorithm is used to convert the reconstructed surface into a triangular mesh. The mesh generation algorithm generates a high-quality triangular mesh through Delaunay triangulation and surface smoothing technology. The formula for Delaunay triangulation is: in, T i is the i-th triangle; Area(T i ) is the area of the triangle; The surface smoothing technique reduces the noise on the mesh surface by using the Laplace smoothing algorithm. The formula is: in, v i new is the new coordinate position after smoothing; v i is the position of vertex i; N(i) is the neighborhood of vertex i; λ is the smoothing coefficient; H3. Deep learning optimization: Generative adversarial networks are used for detail enhancement. The generator network generates a high-resolution 3D model through multi-layer convolution and deconvolution operations, while the discriminator network judges the authenticity of the generated model through multi-layer convolution operations. The loss function of GAN is: in, G is the generator; D is the discriminator; x is the real data; z is the noise vector; Through adversarial training, the generator can generate high-precision 3D models; H4. Plate making data generation: Based on the optimization of the 3D model, high-precision plate making data is generated. The plate making data includes size, shape and texture information. The calculation formula for size data is: Length=max(x i )-min(x i ) Width=max(y i )-min(y i ) Height=max(z i )-min(z i ) in, x i ,y i ,z i are the vertex coordinates of the three-dimensional model; Texture information is generated through texture mapping technology, which maps a two-dimensional texture image onto the surface of a three-dimensional model.
7. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 6 is characterized in that: The specific steps for implementing the automated measurement process control module are as follows: K1. Intelligent Algorithm Control: This system automatically controls the entire measurement process by combining rule-based intelligent algorithms with machine learning models. First, the system uses a pre-set rule library to make a preliminary assessment of the basic characteristics of the artifact, including material, complexity, and size range. Second, the system utilizes machine learning models to perform more detailed classification and prediction of the artifact's characteristics. The machine learning model's input includes the artifact's material, texture complexity, and geometric features, and its output is the optimal measurement parameters, including scanning accuracy, shooting angle, and sensor type. The machine learning model's training data comes from historical measurement records, and the accuracy of predictions is improved by continuously optimizing the model parameters. K2. Dynamic parameter adjustment: During the measurement process, the system dynamically adjusts the measurement parameters based on the real-time collected data. The dynamic adjustment algorithm is based on feedback control theory. The specific formula is: in, u(t) is the adjusted measurement parameter; e(t) is the error between the current measurement data and the target value; K p ,K i ,K d are the proportional, integral and differential coefficients respectively; The system can adjust measurement parameters in real time to ensure measurement accuracy; K3. User interface support: The module provides a user-friendly graphical interface that allows operators to monitor the measurement process in real time and make adjustments and interventions. The interface also supports data visualization, showing measurement results through charts and 3D models, helping operators understand the data more intuitively. K4. Anomaly detection and processing: The module integrates an anomaly detection algorithm that can automatically identify and process anomalies that occur during the measurement process. The anomaly detection algorithm is based on statistical methods and machine learning models. It identifies abnormal data points by analyzing the distribution characteristics of the measurement data.
8. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 7 is characterized in that: In step K3, the content displayed on the interface includes the current measurement progress, real-time collected data, and adjustment suggestions for measurement parameters.
9. The contactless measurement method for clothing artifacts based on deep learning and multimodal large models according to claim 8 is characterized in that: In step K4, the strategies for handling the exception include re-collecting data, adjusting measurement parameters, or switching sensor types.