A roughness prediction method and system based on long sequence missing perception of multi-modal data, a terminal and a storage medium
By constructing a roughness prediction method that senses the long sequence of multimodal data missingness, the problem of prediction performance degradation caused by mixed missingness of long sequence of multimodal time-series data in CNC machining is solved, and high-precision part surface roughness prediction is achieved, improving the robustness and prediction accuracy of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies suffer from reduced prediction performance due to long-sequence mixed missing data in multimodal time series data during CNC machining. This makes it difficult to balance missing data modeling, multimodal fusion, and feature optimization, resulting in reduced accuracy in predicting the surface roughness of parts.
A roughness prediction method based on long sequence missing perception of multimodal data is adopted. By acquiring long sequence signals of multimodal data, missing structure is parsed, missing semantic representation is constructed, and long sequence missing perception mask and learnable token are used for weighted processing. Feature extraction and optimization are combined with shared encoder, context enhancement and conditional diffusion model. Finally, the surface roughness of the part is predicted through a dual-head prediction module.
Achieving high-precision prediction of part surface roughness under long sequence complex missing conditions improves the robustness and prediction accuracy of the model, suppresses modal bias problems, and enhances cross-sample discrimination ability and cross-modal distribution consistency.
Smart Images

Figure CN122490469A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of CNC machining and artificial intelligence, and in particular to a roughness prediction method, system, terminal, and computer-readable storage medium based on long sequence missing perception of multimodal data. Background Technology
[0002] Surface roughness prediction of parts (e.g., cutting tools) is an important research area in CNC machining quality monitoring. Its core lies in establishing a mapping relationship between multimodal sensor signals acquired during machining and the workpiece surface roughness. Compared to mapping based on a single signal, multimodal data provides more information, leading to more accurate predictions. In real industrial environments, multimodal sensor signals typically exhibit strong nonlinearity, long-term dependence, and complex coupling characteristics. Furthermore, due to the complex acquisition environment, sensor data often suffers from varying degrees of data loss, with the loss exhibiting a clear structured pattern, primarily manifested as long-sequence channel-level continuous loss and whole-modal loss.
[0003] To address the issue of missing data, existing methods primarily handle it by deleting missing samples or imputing missing data. Deletion methods directly remove samples with missing data, which can lead to data distribution imbalance when the missing proportion is high, making it unsuitable for real-world processing scenarios. Data imputation methods estimate missing values using existing data, with common strategies including mean imputation and regression models. However, interpolation methods rely on assumptions about data distribution and struggle to handle complex nonlinear dynamic relationships, especially under long-term missing sequence conditions, leading to a continuous accumulation of reconstruction errors. In addition, deep learning methods have been introduced into missing data processing in recent years, including time series modeling methods based on recurrent neural networks and self-attention mechanisms, as well as data reconstruction methods based on diffusion models. These methods have improved reconstruction accuracy and feature representation capabilities to some extent, but still have the following shortcomings: First, they lack explicit modeling of the missing structure, making it difficult to distinguish the semantic information of different missing patterns; second, the diffusion model and feature learning process are independent of each other, failing to form an effective collaborative optimization mechanism; and third, the multimodal fusion process is easily dominated by the strong mode, leading to modality bias. Therefore, achieving high-precision prediction of part surface roughness under complex missing mechanisms remains a significant challenge.
[0004] Existing technologies have significant limitations in handling the problem of long sequence multimodal missing data. Traditional missing data processing methods cannot handle complex long sequence missing data structures and have low prediction accuracy. In addition, although diffusion models have strong modeling capabilities, the diffusion model and the feature learning process are independent of each other, making it difficult to achieve effective information co-optimization.
[0005] In other words, in actual machining environments, due to factors such as sensor failure, network anomalies, channel malfunctions, and other unexpected failures, the acquired vibration signals, cutting force signals, and acoustic emission signals often exhibit prolonged continuous loss or complete mode loss. Furthermore, significant differences in signal quality between different modes lead to modal biases during model training, resulting in reduced accuracy in predicting part surface roughness. Existing methods struggle to simultaneously address missing modeling, multimodal fusion, and feature optimization under long sequence conditions.
[0006] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0007] The main objective of this invention is to provide a roughness prediction method, system, terminal, and computer-readable storage medium based on the perception of long sequence missing data in multimodal data, aiming to solve the problem of decreased prediction performance caused by long sequence mixed missing data in multimodal time series data during CNC machining in the prior art.
[0008] To achieve the above objectives, this invention provides a roughness prediction method based on long-sequence missing data perception of multimodal data. The roughness prediction method based on long-sequence missing data perception of multimodal data includes the following steps: A multimodal long sequence signal is acquired, and the missing structure of the multimodal long sequence signal is parsed. For channel-level missing and modality-level missing, a missing semantic representation is constructed. The missing regions are weighted and feature replaced by long sequence missing perception mask and long sequence perception learnable token, respectively. The feature extraction network extracts long sequence features by channel based on the weighted and replaced multimodal long sequence signal. The long sequence features of each channel are input into a shared encoder, which extracts the wear features of the part surface and embeds them into modal shared features. The modal shared features are then fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level contrastive learning and multimodal alignment to obtain the fused features. The fused features are progressively denoised, optimized, and reconstructed using a conditional diffusion model. An intermediate feedback mechanism is constructed during the diffusion process, and a feedback loss is generated based on the denoising results. The feedback loss is then applied inversely to the shared encoder and the context enhancement process. The surface wear features of the part are embedded and the diffusion-optimized features are input into the dual-head prediction module. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the diffusion-optimized features and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
[0009] Optionally, the roughness prediction method based on long sequence missing awareness of multimodal data, wherein the steps of acquiring multimodal long sequence signals, performing missing structure analysis on the multimodal long sequence signals, constructing missing semantic representations for channel-level and modality-level missing data, and using long sequence missing awareness masks and long sequence awareness learnable tokens to perform weighted processing and feature substitution on the missing regions respectively, and the feature extraction network extracting long sequence features by channel based on the weighted and substituted multimodal long sequence signals, specifically includes: Acquire multimodal long sequence signals from each channel The inputs include vibration signals from four channels, cutting force signals from three channels, and acoustic emission signals from two channels. For batch size, 9 represents the length of the long sequence window and the total number of channels. Missing labels based on the type of missing data It automatically identifies two missing modes: whole mode missing and channel local missing. Based on missing labels Construct a globally weighted mask Different weights are assigned to valid segments, missing whole modes, and locally missing segments in the channel to achieve soft mask weighting; The global semantic embedding is obtained by concatenating three parts: segmented position embedding, missing length embedding, and channel representation embedding. : ; in, For segmented position embedding, For missing length embedding, The channel is embedded; Channel-level and sample-level aggregation is performed on the global missing semantics to obtain sample-level missing semantics. and channel efficiency ; Design two types of learnable tokens: complete modal tokens and segmented missing tokens; Full Modal Token Used to maintain modal integrity: ; Missing token segment Used to preserve the location of localized channel defects: ; in, This indicates a feature concatenation operation. Indicates a basic learnable token. Represents a modally learnable token. This indicates the location where the entire mode is missing. Indicates the length of the missing integer mode. This indicates that the channel can learn tokens. Indicates the position of the starting point of the missing segment. Indicates the length of the missing segment; Calculate missing-aware imputation features: ; in, This represents a weighted multimodal long sequence signal; Represents a multimodal long sequence signal; ; in, This represents the weighted projection characteristics of the original signal. Represents a learnable linear transformation; ; when When, it indicates a valid information segment, when When, it indicates that the entire mode is missing, when When this occurs, it indicates that a single channel is partially missing; in, This indicates missing perceptual filling features. This represents the learnable fusion coefficients, used to balance the weights of the original signal and the segmented tokens; Fill in features with missing perception The data is fed into a feature extraction network based on the Conformer architecture. This network performs joint encoding of multiple channels within a modality while preserving single-channel details, outputting a 9-channel long sequence feature. .
[0010] Optionally, the roughness prediction method based on long sequence missing perception of multimodal data, wherein the long sequence features of each channel are input into a shared encoder, the shared encoder extracts surface wear features of the part and embeds them with modal shared features, the modal shared features are fused into the long sequence features of each channel through context enhancement, and the fused long sequence features of each channel are subjected to instance-level contrastive learning and multimodal alignment to obtain fused features, specifically including: Long sequence features of 9 channels Modal grouping is performed, and the three types of modalities are aggregated to obtain modal-level features. : ; in, Indicates modal category, This indicates the operation of average pooling based on modality grouping; Modal-level features With 9-channel long sequence characteristics Modal attention features are obtained by performing modal attention encoding and channel attention encoding respectively. Channel attention features : ; ; in, This represents modal attention encoding. Indicates channel attention encoding; Modality attention features Broadcast to channel dimension, and channel attention features and 9-channel long sequence features Fusion to generate cross-modal shared features : ; in, Indicates expansion or broadcasting; Cross-modal feature sharing Global pooling is performed to obtain the embedded surface wear features of the part. Then embed the wear characteristics of the part surface. Perform global pooling to obtain modality-shared features. ; The dimensions of the modality-shared features are projected onto the same dimensions as the original encoded features to obtain the projected shared features. ; Channel efficiency With sample-level missing semantics The embedded features are transformed into gating conditions to obtain channel-efficient embeddings. and missing semantic embedding : ; ; in, This represents the sigmoid activation function. Indicates a linear layer. This indicates the number of dimensions added to the tensor; Long sequence features based on 9 channels Shared features after projection Missing semantic embeddings Generate dynamic fusion gating : ; Shared features after gated weighted fusion projection With efficient embedding of channels Generate context-enhanced features : ; Context-enhanced features and 9-channel long sequence features As input, the final contextual features are generated through a feedforward network and layer normalization. : ; in, Representation layer normalization, Indicates a feedforward network. Indicates residual connection; For final context features Two random augmented views are constructed. Differentiated perturbations are applied to the effective channels, while the masks are preserved for the missing channels. Masked temporal pooling is then performed on the effective channels in the two random augmented views to obtain the instance-level features of the two random augmented views. and ; Instance-level features using two randomly augmented views and Flattened instance features and Perform projection and calculate instance-level contrast loss. : ; in, , representing the total number of instances. Represents cosine similarity. Indicates the temperature coefficient. Indicates instance weight, Indicating the second randomized enhanced view, the first Feature vectors of each instance; Based on the comparison results, linear layers and normalization operations were used to refine the final context features. Enhancement is performed to obtain contrast enhancement features. : ; ; in, Represents characteristic residuals, Indicates the activation function; For final context features Contrast Enhancement Features Modal pooling is performed separately to obtain modal-level context features. Contrast enhancement features at the modal level ; Modal-level context features Contrast enhancement features at the modal level Using the key value, modal attention alignment is performed to obtain modal alignment features. Modality alignment features Broadcast to 9-channel dimensions, with contrast enhancement features Perform channel attention alignment to generate fused features .
[0011] Optionally, the roughness prediction method based on long sequence missing perception of multimodal data, wherein the long sequence features of each channel are input into a shared encoder, the shared encoder extracts surface wear features of the part and embeds them with modal shared features, the modal shared features are fused into the long sequence features of each channel through context enhancement, the fused long sequence features of each channel are subjected to instance-level contrastive learning and multimodal alignment to obtain fused features, and then further includes: Calculate modal alignment loss Constrained contrast enhancement features With fusion features Consistency: ; ; in, This indicates modality alignment features. Features after broadcasting to 9-channel dimensions Indicates target alignment features. Represents the numerically stable term. Indicates the sample index; Indicates the channel index; Indicates the time step index.
[0012] Optionally, the roughness prediction method based on long sequence missing information of multimodal data, wherein the conditional diffusion model is used to progressively denoise, optimize, and reconstruct the fused features, and an intermediate feedback mechanism is constructed during the diffusion process, a feedback loss is generated based on the denoising result, and the feedback loss is applied inversely to the shared encoder and the context enhancement process, specifically including: Fusion features Latent features are obtained by temporal pooling based on modality grouping. ; latent features Input conditional diffusion model, at time step Internal potential characteristics Add noise: ; in, express Modal latent features after adding noise Represents the adaptive noise figure. This represents noise that follows a standard normal distribution. Represents the identity matrix; Denoising latent features are estimated through a reverse process: ; in, express Potential features after denoising step This represents a conditional noise prediction network. This represents the average completeness of the entire mode. State semantics representing modality; The training objective of the conditional diffusion model is to minimize the mean square error between the predicted noise and the actual noise. ; in, Indicates the loss of diffused noise. Represents the modal loss weights. Indicates predicted noise; Use the intermediate feedback loss to backfeed the shared encoder and context enhancement: ; in, Indicates intermediate feedback loss. Represents the expectation operator. Indicates the weight of each time step. Indicates a feedback scalar; Total loss of the conditional diffusion model for: ; in, Indicates the weight of the spread noise loss. This represents the intermediate feedback loss weight.
[0013] Optionally, the roughness prediction method based on long-sequence missing information of multimodal data, wherein the embedded wear features of the part surface and the diffusion-optimized features are jointly input into a dual-head prediction module, the first prediction head in the dual-head prediction module is used to predict the wear of the part surface to obtain the predicted wear value of the part surface, and the second prediction head combines the diffusion-optimized features and the predicted wear value of the part surface to perform regression prediction on the surface roughness of the part surface to obtain the predicted surface roughness value of the part surface, specifically including: Embedding part surface wear features output by shared encoder As input, the predicted wear value of the part surface is obtained through a linear layer. : ; in, express layer, express Activation function; Roughness prediction includes a basic prediction branch and a diffusion enhancement branch. The two branches have different inputs. The roughness values predicted by the two branches are then dynamically fused. and Let represent the inputs to the basic prediction branch and the diffusion enhancement branch, respectively, and then we obtain the outputs of the basic prediction branch. and The output of the basic prediction branch and To merge; Among them, the roughness prediction structure is: ; in, Indicates the input for different branches, Indicates the output of different branches; Input to the basic prediction branch: ; ; in, Indicates the potential features after flattening. This indicates the flattening operation, used to flatten a high-dimensional feature vector into a one-dimensional vector. Input to the basic prediction branch The input is fed into the basic prediction branch structure to obtain the output of the basic prediction branch. ; Input to the diffusion enhancement branch: ; ; in, The diffusion fusion feature represents the modal features after diffusion repair. and potential characteristics Enhanced features obtained through fusion It is the diffusion and fusion characteristic after flattening; Input to the diffusion enhancement branch The input is fed into the diffusion enhancement branch structure to obtain the diffusion enhancement branch output. ; The output of the basic prediction branch and Perform dynamic fusion: ; in, This represents the final predicted value of the part's surface roughness. This represents the diffusion gating coefficient.
[0014] Optionally, the roughness prediction method based on long-sequence missing perception of multimodal data, wherein the embedded wear features of the part surface and the diffusion-optimized features are jointly input into a dual-head prediction module, the first prediction head in the dual-head prediction module is used to predict the wear of the part surface to obtain the predicted wear value of the part surface, and the second prediction head combines the diffusion-optimized features and the predicted wear value of the part surface to perform regression prediction on the surface roughness of the part surface to obtain the predicted surface roughness value of the part surface, and then further includes: Calculate the regression loss: ; in, Indicates regression loss, Indicates the first Predicted surface wear values for a sample of parts. Indicates the first The true value of surface wear of a sample of parts. Show the first Predicted surface roughness values for a sample of parts. Indicates the first The true surface roughness values of the parts in each sample; Total design loss: ; in, This represents the total target loss of the entire framework. This indicates the instance-level contrastive loss weights. Indicates the modal alignment loss weights. This represents the total loss weight of the conditional diffusion model.
[0015] Furthermore, to achieve the above objectives, the present invention also provides a roughness prediction system based on long-sequence missing data perception of multimodal data, wherein the roughness prediction system based on long-sequence missing data perception of multimodal data includes: The missing information-aware encoding module is used to acquire multimodal long sequence signals, perform missing structure parsing on the multimodal long sequence signals, construct missing semantic representations for channel-level and modality-level missing information, and use long sequence missing information-aware masks and long sequence-aware learnable tokens to perform weighted processing and feature substitution on the missing regions respectively. The feature extraction network extracts long sequence features by channel based on the weighted and substituted multimodal long sequence signals. A multimodal dual alignment module is used to input the long sequence features of each channel into a shared encoder. The shared encoder extracts the wear features of the part surface and embeds them with modal shared features. The modal shared features are fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level comparative learning and multimodal alignment to obtain fused features. The diffusion module is used to perform stepwise denoising optimization and reconstruction of the fused features using a conditional diffusion model, and to build an intermediate feedback mechanism during the diffusion process. It generates a feedback loss based on the denoising result and applies the feedback loss inversely to the shared encoder and context enhancement process. The multi-task prediction module is used to input the surface wear features of the part into the dual-head prediction module together with the features after diffusion optimization. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the features after diffusion optimization and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
[0016] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a roughness prediction program based on long-sequence missing perception of multimodal data stored in the memory and executable on the processor, wherein when the roughness prediction program based on long-sequence missing perception of multimodal data is executed by the processor, it implements the steps of the roughness prediction method based on long-sequence missing perception of multimodal data as described above.
[0017] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a roughness prediction program based on long sequence missing perception of multimodal data, and when the roughness prediction program based on long sequence missing perception of multimodal data is executed by a processor, it implements the steps of the roughness prediction method based on long sequence missing perception of multimodal data as described above.
[0018] In this invention, a multimodal long sequence signal is acquired, and a missing structure analysis is performed on the multimodal long sequence signal. For channel-level and modal-level missing data, a missing semantic representation is constructed. Long sequence missing-aware masks and long sequence-aware learnable tokens are used to perform weighted processing and feature substitution on the missing regions, respectively. A feature extraction network extracts long sequence features by channel based on the weighted and substituted multimodal long sequence signal. The long sequence features of each channel are input into a shared encoder. The shared encoder extracts surface wear features of the part and modal shared features. Context enhancement is used to fuse the modal shared features into the long sequence features of each channel. The fused long sequence features of each channel are then processed... Instance-level contrastive learning and multimodal alignment are used to obtain fused features. A conditional diffusion model is employed to progressively denoise, optimize, and reconstruct these fused features, with an intermediate feedback mechanism built during the diffusion process. A feedback loss is generated based on the denoising results and applied inversely to the shared encoder and context enhancement process. The surface wear features of the part are embedded and the diffusion-optimized features are input into a dual-head prediction module. The first prediction head in the dual-head prediction module predicts the surface wear of the part, obtaining a predicted surface wear value. The second prediction head combines the diffusion-optimized features and the predicted surface wear value to perform regression prediction on the surface roughness of the part, obtaining a predicted surface roughness value. This invention introduces long-sequence missing information masks and long-sequence learningable tokens to achieve data augmentation that accommodates different missing modes. A multimodal dual alignment mechanism consisting of instance-level contrastive learning and modal-level feature alignment is introduced to enhance cross-sample discrimination capability and constrain cross-modal distribution consistency. Simultaneously, a diffusion model feedback closed-loop optimization mechanism is used to adaptively optimize feature representation. Through the synergistic effect of these mechanisms, high-precision prediction of part surface roughness can be achieved under conditions of continuous long-sequence channel-level missing data and mixed whole-modal missing data. Attached Figure Description
[0019] Figure 1 This is a flowchart of a preferred embodiment of the roughness prediction method based on long sequence missing perception of multimodal data according to the present invention; Figure 2 This is a schematic diagram illustrating the entire process of predicting the surface roughness of a part in a preferred embodiment of the roughness prediction method based on long sequence missing perception of multimodal data according to the present invention. Figure 3 This is a structural diagram of a preferred embodiment of the roughness prediction system based on long sequence missing information of multimodal data according to the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0021] Existing technologies have significant limitations in handling long-sequence multimodal missing data. Traditional missing data processing methods cannot handle complex long-sequence missing structures, resulting in low prediction accuracy. Furthermore, while diffusion models possess strong modeling capabilities, they are independent of the feature learning process, making it difficult to achieve effective collaborative optimization of information. The purpose of this invention is to construct a roughness prediction method based on long-sequence missing data perception in multimodal data. This method achieves collaborative feature modeling and information compensation under complex missing conditions in long sequences, as well as adaptive optimization of feature representation, effectively suppressing modality bias and improving the robustness and prediction accuracy of the model under complex missing conditions in long sequences.
[0022] The roughness prediction method based on long sequence missing information in multimodal data according to a preferred embodiment of the present invention includes a biomedical visual language model comprising an adaptive prototype visual locator and a co-causal decoupling module; as shown in the following example. Figure 1 and Figure 2 As shown, the roughness prediction method based on long sequence missing information of multimodal data includes the following steps: Step S10: Obtain multimodal long sequence signals, perform missing structure analysis on the multimodal long sequence signals, construct missing semantic representations for channel-level and modal-level missing signals, and use long sequence missing perception masks and long sequence perception learnable tokens to perform weighted processing and feature substitution on the missing regions respectively. The feature extraction network extracts long sequence features by channel based on the weighted and substituted multimodal long sequence signals.
[0023] Specifically, this step is mainly responsible for performing missing information processing on the original multi-channel signal with missing data. Its core functions are to accurately identify missing patterns, generate missing semantic representations, and fill in missing location information to avoid feature distortion or information distortion caused by missing data, thus providing high-quality input for subsequent feature extraction.
[0024] Acquire multimodal long sequence signals from each channel The inputs include vibration signals from four channels, cutting force signals from three channels, and acoustic emission signals from two channels. For batch size, The length of the long sequence window. 9 represents the total number of channels.
[0025] Missing labels based on the type of missing data It automatically identifies two missing modes: whole mode missing and channel local missing. Missing labels based on the type of missing data It can automatically parse out two types of missing modes: whole-mode missing and channel-local missing. When When the missing markers in all channels of a certain mode are all true, it indicates that this segment is a valid information segment; when the missing markers in all channels of a certain mode are all true, it indicates that this segment is a valid information segment. When, it indicates that this mode is missing an integer mode; when When , it indicates that a single channel is partially missing.
[0026] Based on missing labels Construct a globally weighted mask Different weights are assigned to valid segments, segments with missing integer modes, and segments with locally missing channels to achieve soft mask weighting; among them, the weights assigned to valid segments are... Set to 1; missing weight for integer modes ; partial missing section of the channel Set it to 0.2.
[0027] Simultaneously, a global semantic embedding is obtained by concatenating three parts: segmented position embedding, missing length embedding, and channel representation embedding. : ; in, For segmented position embedding, For missing length embedding, The channel represents the embedding.
[0028] Then, channel-level and sample-level aggregation is performed on the global missing semantics to obtain sample-level missing semantics. and channel efficiency It acts on subsequent characterization enhancement.
[0029] When data is missing, learnable tokens are used to inject structured information such as modality, channel, position, and length to replace invalid values in the missing data. This avoids feature bias caused by directly filling in 0s or the mean, while preserving the correlation and temporal consistency between modalities. To adapt to both whole modality missing and partial channel missing modes, this invention designs two types of learnable tokens: whole modality tokens and segmented missing tokens.
[0030] Among them, complete modal tokens Used to maintain modal integrity: ; Among them, the segmented missing token Used to preserve the location of localized channel defects: ; in, This indicates a feature concatenation operation. Indicates a basic learnable token. This represents modal learnable tokens, with different modal learnable tokens for different channels of each modality. They are consistent. This indicates the location where the entire mode is missing. Indicates the length of the missing integer mode. This represents a channel-learnable token, and the channel-learnable token within each channel. They are consistent. Indicates the position of the starting point of the missing segment. Indicates the length of the missing segment.
[0031] Calculate missing-aware imputation features to provide stable data for subsequent shared encoders: ; in, This represents a weighted multimodal long sequence signal; It represents a multimodal long sequence signal.
[0032] ; in, This represents the weighted projection characteristics of the original signal. It represents a learnable linear transformation.
[0033] ; when When, it indicates a valid information segment, when When, it indicates that the entire mode is missing, when When , it indicates that a single channel is partially missing.
[0034] in, This indicates missing perceptual filling features. This represents the learnable fusion coefficient, used to balance the weights of the original signal and the segmented tokens, avoiding information distortion caused by fixed weighting.
[0035] Fill in features with missing perception The data is fed into a feature extraction network based on the Conformer architecture. This network performs joint encoding of multiple channels within a modality while preserving single-channel details, outputting a 9-channel long sequence feature. .
[0036] Step S20: Input the long sequence features of each channel into the shared encoder. The shared encoder extracts the wear features of the part surface and embeds them into the modal shared features. The modal shared features are fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level comparative learning and multimodal alignment to obtain the fused features.
[0037] Specifically, the input for this step is a long sequence of 9 channels. By completing four parts—shared encoder, context enhancement, instance-level contrastive learning, and multimodal alignment—the problem of feature quality degradation caused by modality bias and missing features is solved.
[0038] (1) Shared encoder.
[0039] Efficiency through the channel Weighted pooling is used to reduce the contribution of missing channels and avoid being misled by missing data.
[0040] First, the long sequence features of the 9 channels were analyzed. Modal grouping is performed, and the three types of modes are aggregated separately to obtain modal-level features. : ; in, Indicates modal category, This indicates the operation of average pooling by modality grouping.
[0041] Then, for modal-level features... With 9-channel long sequence characteristics Modal attention encoding and channel attention encoding are performed separately to preserve channel specificity and prevent the loss of modal shared information, thus obtaining modal attention features. Channel attention features : ; ; in, This represents modal attention encoding. This indicates channel attention encoding.
[0042] Modality attention features Broadcast to channel dimension, and channel attention features and 9-channel long sequence features Fusion to generate cross-modal shared features : ; in, It indicates expansion or broadcasting.
[0043] Cross-modal feature sharing Global pooling is performed to obtain the embedded surface wear features of the part. Then embed the wear characteristics of the part surface. Perform global pooling to obtain modality-shared features. This is used for subsequent prediction of part surface wear and roughness.
[0044] (2) Contextual enhancement.
[0045] To enhance the contextual relevance of features, this invention designs a context enhancement operation that dynamically fuses shared features and original encoded features through a gating mechanism.
[0046] First, project the dimensions of the modality-shared features onto the same dimensions as the original encoded features to obtain the projected shared features. At the same time, the channel efficiency will be improved. With sample-level missing semantics The embedded features are transformed into gating conditions to obtain channel-efficient embeddings. and missing semantic embedding : ; ; in, This represents the sigmoid activation function. Indicates a linear layer. This indicates the number of dimensions added to the tensor.
[0047] Then, based on the long sequence features of 9 channels Shared features after projection Missing semantic embeddings Generate dynamic fusion gating : ; Shared features after gated weighted fusion projection With efficient embedding of channels Generate context-enhanced features : ; Context-enhanced features and 9-channel long sequence features As input, the final contextual features are generated through a feedforward network and layer normalization. : ; in, Representation layer normalization, Indicates a feedforward network. This indicates a residual connection.
[0048] (3) Instance-level comparative learning.
[0049] To improve the feature stability of the model under missing perturbations and avoid feature drift caused by missing data, this invention designs an instance-level contrastive learning part, which constrains the feature consistency of the same instance by constructing dual random views and using weighted NT-Xent loss.
[0050] First, the final context features Two random augmented views are constructed. Differentiated perturbations are applied to the effective channels, while the masks are preserved for the missing channels. Masked temporal pooling is then performed on the effective channels in the two random augmented views to obtain the instance-level features of the two random augmented views. and .
[0051] Instance-level features using two randomly augmented views and Flattened instance features and Perform projection and calculate instance-level contrast loss (weighted NT-Xent contrast loss). : ; in, , representing the total number of instances. Represents cosine similarity. Indicates the temperature coefficient. , Indicates instance weight, , , Indicates the sample index; Indicates the channel index. Indicating the second randomized enhanced view, the first The feature vector of each instance.
[0052] Based on the comparison results, linear layers and normalization operations were used to refine the final context features. Enhancement is performed to obtain contrast enhancement features. : ; ; in, Represents the feature residuals, which are the features of the final context. The increment, This represents the activation function.
[0053] (4) Multimodal alignment.
[0054] The multimodal alignment design of this invention aims to solve the problems of modal imbalance and strong modality dominance. It utilizes two-layer attention alignment of modality and channel to achieve balanced representation of strong and weak modalities.
[0055] For final context features Contrast Enhancement Features Modal pooling is performed separately to obtain modal-level context features. Contrast enhancement features at the modal level .
[0056] Modal-level context features Contrast enhancement features at the modal level Using the key value, modal attention alignment is performed to obtain modal alignment features. Modality alignment features Broadcast to 9-channel dimensions, with contrast enhancement features Perform channel attention alignment to generate fused features .
[0057] Simultaneously calculate the modal alignment loss. Constrained contrast enhancement features With fusion features Consistency: ; ; in, This indicates modality alignment features. Features after broadcasting to 9-channel dimensions Indicates target alignment features. This represents a numerically stable term, used to avoid denominators of zero and ensure the stability of numerical calculations. Indicates the sample index; Indicates the channel index; Indicates the time step index.
[0058] Step S30: The fused features are progressively denoised, optimized, and reconstructed using a conditional diffusion model. An intermediate feedback mechanism is constructed during the diffusion process. A feedback loss is generated based on the denoising results, and the feedback loss is applied inversely to the shared encoder and context enhancement process.
[0059] Specifically, the fusion features Latent features are obtained by temporal pooling based on modality grouping. It is used as input to the diffusion model, and the diffusion-repaired modal features are obtained through the diffusion module. .
[0060] Noise diffusion and reverse denoising process: removing latent features Input conditional diffusion model, at time step Internal potential characteristics Add noise: ; in, express Modal latent features after adding noise Represents the adaptive noise figure. This represents noise that follows a standard normal distribution. Represents the identity matrix.
[0061] Denoising latent features are estimated through a reverse process: ; in, express Potential features after denoising step This represents a conditional noise prediction network. This represents the average completeness of the entire mode. State semantics representing modalities, using high-dimensional latent features By pooling, linear projection, and sigmoid activation, and compressing it into a scalar value in the interval [0, 1], this scalar value represents the modal state semantics. .
[0062] Loss function for diffusion model: The training objective of the conditional diffusion model is to minimize the mean square error between the predicted noise and the actual noise. ; in, Indicates the loss of diffused noise. This represents the modal loss weight, which is set to 0.25 in this invention. This represents the predicted noise.
[0063] Use the intermediate feedback loss to backfeed the shared encoder and context enhancement: ; in, Indicates intermediate feedback loss. Represents the expectation operator. This represents the weight of each time step, which changes as time progresses. The weight is higher in later denoising steps, allowing the model to focus on approximating the latent features. It represents the feedback scalar, a semantic supervision signal generated by the denoising features, which inversely optimizes the representation quality of modal semantics.
[0064] Total loss of the conditional diffusion model for: ; in, Indicates the weight of the spread noise loss. This represents the intermediate feedback loss weight.
[0065] Step S40: The surface wear features of the part are embedded and the diffusion-optimized features are input into the dual-head prediction module. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the diffusion-optimized features and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
[0066] Specifically, part surface wear prediction: embedding the part surface wear features output by the shared encoder. As input, the predicted wear value of the part surface is obtained through a linear layer. : ; in, express The layer randomly sets the output of some neurons to 0 with a fixed probability to prevent overfitting and enhance generalization ability. express Activation function.
[0067] Roughness prediction: Roughness prediction includes a basic prediction branch and a diffusion enhancement branch. The two branches have different inputs, and the roughness values predicted by each branch are then dynamically fused. The structures of the two branches are identical. and Let represent the inputs to the basic prediction branch and the diffusion enhancement branch, respectively, and then we obtain the outputs of the basic prediction branch. and The output of the basic prediction branch and To integrate.
[0068] Among them, the roughness prediction structure is: ; in, Indicates the input for different branches, This indicates the output of different branches.
[0069] (1) Input to the basic prediction branch: ; ; in, Indicates the potential features after flattening. This indicates a flattening operation, used to flatten a high-dimensional feature vector into a one-dimensional vector.
[0070] Input to the basic prediction branch The input is fed into the basic prediction branch structure to obtain the output of the basic prediction branch. .
[0071] (2) Input of the diffusion enhancement branch: ; ; in, The diffusion fusion feature represents the modal features after diffusion repair. and potential characteristics Enhanced features obtained through fusion It is the diffusion and fusion characteristic after flattening.
[0072] Input to the diffusion enhancement branch The input is fed into the diffusion enhancement branch structure to obtain the diffusion enhancement branch output. .
[0073] The output of the basic prediction branch and Perform dynamic fusion: ; in, This represents the final predicted value of the part's surface roughness. This represents the diffusion gating coefficient. It takes values [0, 1] as a learnable gating value used to dynamically balance the sharing between the two branches.
[0074] Calculate the regression loss: ; in, Indicates regression loss, Indicates the first Predicted surface wear values for a sample of parts. Indicates the first The true value of surface wear of a sample of parts. Show the first Predicted surface roughness values for a sample of parts. Indicates the first The true surface roughness value of a sample part.
[0075] Total design loss: ; in, This represents the total target loss of the entire framework. This indicates the instance-level contrastive loss weights. Indicates the modal alignment loss weights. This represents the total loss weight of the conditional diffusion model.
[0076] The innovative aspects of this invention: (1) Long sequence missing information perception mechanism: By introducing a long sequence missing information perception mask, the missing regions are explicitly marked and weighted. At the same time, long sequence perception learnable tokens are used to semantically replace the missing regions, thereby achieving compensation and reconstruction of missing information at the feature level. This mechanism can not only distinguish different missing patterns, but also maintain the continuity and stability of feature expression under long-term continuous missing conditions, thus improving the model's ability to model complex missing structures.
[0077] (2) Multimodal Dual-Alignment Feature Learning Mechanism: A dual-alignment mechanism consisting of instance-level contrastive learning and modality-level feature alignment is constructed to enhance cross-sample discrimination capability and constrain cross-modal distribution consistency. At the sample level, contrastive learning brings the feature distance between samples in the same processing state closer and widens the feature distribution of samples in different states, thereby improving feature discrimination capability. At the modality level, alignment constraints are applied to the feature distributions of different modalities to ensure that each modality maintains a consistent expression in a unified semantic space. This mechanism effectively alleviates the modality bias problem caused by information imbalance during multimodal fusion and improves the robustness of fused features.
[0078] (3) Diffusion Feedback Closed-Loop Optimization Mechanism: A conditional diffusion model is introduced in the feature optimization stage. The fused features are gradually reconstructed through a multi-step noise addition and reverse denoising process. An intermediate feedback path is constructed during the diffusion process to convert the denoising error into a feedback loss that is then passed back to the feature extraction and fusion stages. This mechanism forms a closed-loop structure of feature extraction, diffusion optimization, and feedback correction, which couples the feature learning process with the diffusion reconstruction process, achieving dynamic optimization of feature representation and gradual approximation of the true distribution, thereby improving the quality of feature representation.
[0079] The beneficial effects of this invention are: (1) Information preservation capability under long sequence mixed missing structure: In actual processing, long-term continuous missing and whole modality missing will lead to serious loss of key information. By constructing a missingness-aware mask and a learnable token mechanism, explicit modeling and semantic compensation are performed in the missing region, so that the model can still maintain stable feature expression under long sequence missing conditions, effectively reducing the risk of performance degradation caused by missing.
[0080] (2) Diffusion-driven adaptive optimization capability of features: In traditional methods, feature learning and reconstruction processes are independent of each other and it is difficult to form an effective synergy. This invention introduces a diffusion feedback closed-loop mechanism to apply the feedback loss in reverse to the feature extraction stage, so as to dynamically adjust the feature representation according to the reconstruction results, so that the features are gradually optimized and the distribution is approximated, thereby significantly improving the model's expressive ability under complex data conditions.
[0081] (3) Suppressing modal bias and improving prediction robustness: Due to the differences in the quality of different modal signals, traditional methods tend to rely on strong modes, leading to modal bias problems. This invention achieves implicit balancing of the contributions of each mode at the feature level through the synergistic effect of a multimodal alignment mechanism and diffusion optimization process, enabling the model to maintain stable prediction performance even when some modes are missing or their quality degrades.
[0082] This invention verifies its feasibility and effectiveness through experiments. The experiments used approximately 1.6 billion samples from a real dataset, which includes vibration signals from four channels, cutting force signals from three channels, and acoustic emission signals from two channels. In the experimental design, the case of long sequences with high missing values was first simulated. The simulated missing value conditions included a mixture of three types of missing values: continuous channel-level missing values, segmented channel-level missing values, and whole-mode missing values, with a missing value trigger probability set to 70%. Two coefficients of determination (CQDs) for wear prediction and roughness prediction were used in the experiment. The coefficient of determination (COD) is used as an evaluation metric to measure the goodness of fit between the predicted and actual values. The closer the COD is to 1, the higher the consistency between the model's predicted structure and the actual roughness, and the better the model's performance. Experimental results show that using only the missing perceptual coding module for wear value prediction... The roughness is 0.9869. The wear value is 0.9773; the wear value prediction is improved after adding the multimodal dual alignment module. The roughness is 0.9889. The value is 0.9851; while the wear value is predicted after adding the three-step diffusion. The roughness reached 0.9900. The value is 0.9850. This result clearly demonstrates that the long-sequence deletion-sensing diffusion method proposed in this invention can effectively suppress modal bias and improve the prediction accuracy of the model under conditions of mixed deletions in long sequences.
[0083] Furthermore, such as Figure 3 As shown, based on the above-mentioned roughness prediction method based on long-sequence missing data perception of multimodal data, the present invention also provides a roughness prediction system based on long-sequence missing data perception of multimodal data, wherein the roughness prediction system based on long-sequence missing data perception of multimodal data includes: The missing information-aware encoding module 51 is used to acquire multimodal long sequence signals, perform missing structure analysis on the multimodal long sequence signals, construct missing semantic representations for channel-level and modal-level missing information, and use long sequence missing information-aware mask and long sequence-aware learnable token to perform weighted processing and feature substitution on the missing regions respectively. The feature extraction network extracts long sequence features by channel based on the weighted and substituted multimodal long sequence signals. The multimodal dual alignment module 52 is used to input the long sequence features of each channel into the shared encoder. The shared encoder extracts the wear features of the part surface and embeds them with the modal shared features. The modal shared features are fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level comparative learning and multimodal alignment to obtain the fused features. The diffusion module 53 is used to perform stepwise denoising optimization and reconstruction of the fused features using a conditional diffusion model, and to build an intermediate feedback mechanism during the diffusion process. It generates a feedback loss based on the denoising result and applies the feedback loss in reverse to the shared encoder and context enhancement process. The multi-task prediction module 54 is used to input the surface wear features of the part and the diffusion-optimized features into the dual-head prediction module. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the diffusion-optimized features and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
[0084] Furthermore, such as Figure 4 As shown, based on the above-mentioned roughness prediction method and system based on the perception of long sequence missing data of multimodal data, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0085] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a roughness prediction program 40 based on multimodal data long sequence missing perception. This roughness prediction program 40 can be executed by the processor 10 to implement the roughness prediction method based on multimodal data long sequence missing perception in this application.
[0086] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the roughness prediction method based on long sequence missing perception of multimodal data.
[0087] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminal's processor 10, memory 20, and display 30 communicate with each other via a system bus.
[0088] In one embodiment, when the processor 10 executes the roughness prediction program 40 based on long sequence missing perception of multimodal data in the memory 20, it implements the steps of the roughness prediction method based on long sequence missing perception of multimodal data as described above.
[0089] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a roughness prediction program based on long sequence missing perception of multimodal data, and the roughness prediction program based on long sequence missing perception of multimodal data, when executed by a processor, implements the steps of the roughness prediction method based on long sequence missing perception of multimodal data as described above.
[0090] In summary, this invention provides a roughness prediction method, system, terminal, and computer-readable storage medium based on long-sequence missing perception of multimodal data. The method includes: acquiring multimodal long-sequence signals; parsing the missing structure of the multimodal long-sequence signals; constructing missing semantic representations for channel-level and modal-level missing data; and using long-sequence missing perception masks and long-sequence perception learnable tokens to perform weighted processing and feature substitution on the missing regions, respectively; a feature extraction network extracting long-sequence features by channel based on the weighted and substituted multimodal long-sequence signals; inputting the long-sequence features of each channel into a shared encoder; the shared encoder extracting surface wear features of the part and modal shared features; and fusing the modal shared features into each channel through context enhancement. In the long-sequence features of each channel, the fused long-sequence features of each channel are subjected to instance-level contrastive learning and multimodal alignment to obtain fused features. A conditional diffusion model is used to progressively denoise, optimize, and reconstruct the fused features, and an intermediate feedback mechanism is constructed during the diffusion process. A feedback loss is generated based on the denoising results, and this feedback loss is used inversely to the shared encoder and context enhancement process. The surface wear features of the part are embedded and the diffusion-optimized features are input into a dual-head prediction module. The first prediction head in the dual-head prediction module predicts the surface wear of the part, obtaining a predicted value for the surface wear. The second prediction head combines the diffusion-optimized features and the predicted surface wear value to perform regression prediction on the surface roughness of the part, obtaining a predicted value for the surface roughness. This invention achieves data augmentation that takes into account different missing modes by introducing a long-sequence missing-aware mask and a long-sequence missing-aware learnable token. It introduces a multimodal dual alignment mechanism composed of instance-level contrastive learning and modal-level feature alignment to achieve enhanced cross-sample discrimination capability and cross-modal distribution consistency constraints. Simultaneously, a diffusion model feedback closed-loop optimization mechanism is used to adaptively optimize the feature representation. Through the synergistic effect of the above mechanisms, high-precision prediction of part surface roughness can be achieved even in the case of continuous missing long-sequence channel levels and mixed missing whole modes.
[0091] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0092] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0093] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A roughness prediction method based on long sequence missing information in multimodal data, characterized in that, The roughness prediction method based on long sequence missing information of multimodal data includes: A multimodal long sequence signal is acquired, and the missing structure of the multimodal long sequence signal is parsed. For channel-level missing and modality-level missing, a missing semantic representation is constructed. The missing regions are weighted and feature replaced by long sequence missing perception mask and long sequence perception learnable token, respectively. The feature extraction network extracts long sequence features by channel based on the weighted and replaced multimodal long sequence signal. The long sequence features of each channel are input into a shared encoder, which extracts the wear features of the part surface and embeds them into modal shared features. The modal shared features are then fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level contrastive learning and multimodal alignment to obtain the fused features. The fused features are progressively denoised, optimized, and reconstructed using a conditional diffusion model. An intermediate feedback mechanism is constructed during the diffusion process, and a feedback loss is generated based on the denoising results. The feedback loss is then applied inversely to the shared encoder and the context enhancement process. The surface wear features of the part are embedded and the diffusion-optimized features are input into the dual-head prediction module. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the diffusion-optimized features and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
2. The roughness prediction method based on long sequence missing information of multimodal data according to claim 1, characterized in that, The process involves acquiring multimodal long-sequence signals, performing missing structure analysis on the signals, constructing semantic representations for both channel-level and modal-level missing data, and then using long-sequence missing-aware masks and long-sequence-aware learnable tokens to perform weighted processing and feature substitution on the missing regions, followed by feature extraction. The network extracts long sequence features by channel based on the weighted substitution multimodal long sequence signal, specifically including: Acquire multimodal long sequence signals from each channel As input, the signals include vibration signals from four channels, cutting force signals from three channels, and acoustic emission signals from two channels. For batch size, 9 represents the length of the long sequence window and the total number of channels. Missing labels based on the type of missing data It automatically identifies two missing modes: whole mode missing and channel local missing. Based on missing labels Construct a globally weighted mask Different weights are assigned to valid segments, missing whole modes, and locally missing segments in the channel to achieve soft mask weighting; The global semantic embedding is obtained by concatenating three parts: segmented position embedding, missing length embedding, and channel representation embedding. : ; in, For segmented position embedding, For missing length embedding, The channel is embedded; Channel-level and sample-level aggregation is performed on the global missing semantics to obtain sample-level missing semantics. and channel efficiency ; Design two types of learnable tokens: complete modal tokens and segmented missing tokens; Full Modal Token Used to maintain modal integrity: ; Missing token segment Used to preserve the location of localized channel defects: ; in, This indicates a feature concatenation operation. Indicates a basic learnable token. Represents a modally learnable token. This indicates the location where the entire mode is missing. Indicates the length of the missing integer mode. This indicates that the channel can learn tokens. Indicates the position of the starting point of the missing segment. Indicates the length of the missing segment; Calculate missing-aware imputation features: ; in, This represents a weighted multimodal long sequence signal; Represents a multimodal long sequence signal; ; in, This represents the weighted projection characteristics of the original signal. Represents a learnable linear transformation; ; when When, it indicates a valid information segment, when When, it indicates that the entire mode is missing, when When this occurs, it indicates that a single channel is partially missing; in, This indicates missing perceptual filling features. This represents the learnable fusion coefficients, used to balance the weights of the original signal and the segmented tokens; Fill in features with missing perception The data is fed into a feature extraction network based on the Conformer architecture. This network performs joint encoding of multiple channels within the modality while preserving single-channel details, outputting a long sequence of 9 channels. .
3. The roughness prediction method based on long sequence missing information of multimodal data according to claim 2, characterized in that, The long-sequence features of each channel are input into a shared encoder. The shared encoder extracts surface wear features of the part and embeds them with modal shared features. Context enhancement is used to fuse the modal shared features into the long-sequence features of each channel. The fused long-sequence features of each channel are then subjected to instance-level contrastive learning and multimodal alignment to obtain fused features. Specifically, this includes: Long sequence features of 9 channels Modal grouping is performed, and the three types of modalities are aggregated to obtain modal-level features. : ; in, Indicates modal category, This indicates the operation of average pooling based on modality grouping; Modal-level features With 9-channel long sequence characteristics Modal attention features are obtained by performing modal attention encoding and channel attention encoding respectively. and channel attention features : ; ; in, This represents modal attention encoding. Indicates channel attention encoding; Modality attention features Broadcast to channel dimension, and channel attention features and 9-channel long sequence features Fusion to generate cross-modal shared features : ; in, Indicates expansion or broadcasting; Cross-modal feature sharing Global pooling is performed to obtain the embedded surface wear features of the part. Then embed the wear characteristics of the part surface. Perform global pooling to obtain modality-shared features. ; The dimensions of the modality-shared features are projected onto the same dimensions as the original encoded features to obtain the projected shared features. ; Channel efficiency With sample-level missing semantics The embedded features are transformed into gating conditions to obtain channel-efficient embeddings. and missing semantic embedding : ; ; in, This represents the sigmoid activation function. Indicates a linear layer. This indicates the number of dimensions added to the tensor; Long sequence features based on 9 channels Shared features after projection Missing semantic embeddings Generate dynamic fusion gating : ; Shared features after gated weighted fusion projection With efficient embedding of channels Generate context-enhanced features : ; Context-enhanced features and 9-channel long sequence features As input, the final contextual features are generated through a feedforward network and layer normalization. : ; in, Representation layer normalization, Indicates a feedforward network. Indicates residual connection; For final context features Two random augmented views are constructed. Differentiated perturbations are applied to the effective channels, while the missing channels are masked. Masked temporal pooling is then performed on the effective channels in the two random augmented views to obtain instance-level features of the two random augmented views. and ; Instance-level features using two randomly augmented views and Flattened instance features and Perform projection and calculate instance-level contrast loss. : ; in, , representing the total number of instances. Represents cosine similarity. Indicates the temperature coefficient. Indicates instance weight, Indicating the second randomized enhanced view, the first Feature vectors of each instance; Based on the comparison results, linear layers and normalization operations were used to refine the final context features. Enhancement is performed to obtain contrast enhancement features. : ; ; in, Represents characteristic residuals, Indicates the activation function; For final context features Contrast Enhancement Features Modal pooling is performed separately to obtain modal-level context features. Contrast enhancement features at the modal level ; Modal-level context features Contrast enhancement features at the modal level Using the key value, modal attention alignment is performed to obtain modal alignment features. Align features with modality Broadcast to 9-channel dimensions, with contrast enhancement features Perform channel attention alignment to generate fused features .
4. The roughness prediction method based on long sequence missing information of multimodal data according to claim 3, characterized in that, The process involves inputting long-sequence features from each channel into a shared encoder, which extracts surface wear features and modal shared features from the part. Context enhancement is then used to fuse the modal shared features into the long-sequence features of each channel. The fused long-sequence features of each channel are then subjected to instance-level contrastive learning and multimodal alignment to obtain fused features. The process further includes: Calculate modal alignment loss Constrained contrast enhancement features With fusion features Consistency: ; ; in, This indicates modality alignment features. Features after broadcasting to 9-channel dimensions Indicates target alignment features. Represents the numerically stable term. Indicates the sample index; Indicates the channel index; Indicates the time step index.
5. The roughness prediction method based on long sequence missing information of multimodal data according to claim 4, characterized in that, The conditional diffusion model is used to progressively denoise, optimize, and reconstruct the fused features. An intermediate feedback mechanism is constructed during the diffusion process, and a feedback loss is generated based on the denoising results. This feedback loss is then applied inversely to the shared encoder and the context enhancement process. Specifically, this includes: Fusion features Latent features are obtained by temporal pooling based on modality grouping. ; latent features Input conditional diffusion model, at time step Internal potential characteristics Add noise: ; in, express Modal latent features after adding noise Represents the adaptive noise figure. This represents noise that follows a standard normal distribution. Represents the identity matrix; Denoising latent features are estimated through a reverse process: ; in, express Potential features after denoising This represents a conditional noise prediction network. This represents the average completeness of the entire mode. State semantics representing modality; The training objective of the conditional diffusion model is to minimize the mean square error between the predicted noise and the actual noise. ; in, Indicates the loss of diffused noise. Represents the modal loss weights. Indicates predicted noise; Use the intermediate feedback loss to backfeed the shared encoder and context enhancement: ; in, Indicates intermediate feedback loss. Represents the expectation operator. Indicates the weight of each time step. Indicates a feedback scalar; Total loss of the conditional diffusion model for: ; in, Indicates the weight of the spread noise loss. This represents the intermediate feedback loss weight.
6. The roughness prediction method based on long sequence missing information of multimodal data according to claim 5, characterized in that, The embedded and diffusion-optimized features of the part surface wear characteristics are input into a dual-head prediction module. The first prediction head in the dual-head prediction module predicts the part surface wear, obtaining a predicted part surface wear value. The second prediction head combines the diffusion-optimized features and the predicted part surface wear value to perform regression prediction on the part surface roughness, obtaining a predicted part surface roughness value. Specifically, this includes: Embedding part surface wear features output by shared encoder As input, the predicted wear value of the part surface is obtained through a linear layer. : ; in, express layer, express Activation function; Roughness prediction includes a basic prediction branch and a diffusion enhancement branch. The two branches have different inputs. The roughness values predicted by the two branches are then dynamically fused. and Let represent the inputs to the basic prediction branch and the diffusion enhancement branch, respectively, and then we obtain the outputs of the basic prediction branch. and The output of the basic prediction branch and To merge; Among them, the roughness prediction structure is: ; in, Indicates the input for different branches, Indicates the output of different branches; Input to the basic prediction branch: ; ; in, Indicates the potential features after flattening. This indicates the flattening operation, used to flatten a high-dimensional feature vector into a one-dimensional vector. Input to the basic prediction branch The input is fed into the basic prediction branch structure to obtain the output of the basic prediction branch. ; Input to the diffusion enhancement branch: ; ; in, The diffusion fusion feature represents the modal features after diffusion repair. and potential characteristics Enhanced features obtained through fusion It is the diffusion and fusion characteristic after flattening; Input to the diffusion enhancement branch The input is fed into the diffusion enhancement branch structure to obtain the diffusion enhancement branch output. ; The output of the basic prediction branch and Perform dynamic fusion: ; in, This represents the final predicted surface roughness value of the part. This represents the diffusion gating coefficient.
7. The roughness prediction method based on long sequence missing information of multimodal data according to claim 8, characterized in that, The embedded and diffusion-optimized features of the part surface wear characteristics are input into a dual-head prediction module. The first prediction head in the dual-head prediction module predicts the part surface wear, obtaining a predicted part surface wear value. The second prediction head combines the diffusion-optimized features and the predicted part surface wear value to perform regression prediction on the part surface roughness, obtaining a predicted part surface roughness value. This is followed by: Calculate the regression loss: ; in, Indicates regression loss, Indicates the first Predicted surface wear values for a sample of parts. Indicates the first The true value of surface wear of a sample of parts. Show the first Predicted surface roughness values for a sample of parts. Indicates the first The true surface roughness values of the parts in each sample; Total design loss: ; in, This represents the total target loss of the entire framework. This indicates the instance-level contrastive loss weights. Indicates the modal alignment loss weights. This represents the total loss weight of the conditional diffusion model.
8. A roughness prediction system based on long sequence missing information of multimodal data, characterized in that, The roughness prediction system based on long sequence missing information of multimodal data includes: The missing information-aware encoding module is used to acquire multimodal long sequence signals, perform missing structure parsing on the multimodal long sequence signals, construct missing semantic representations for channel-level and modality-level missing information, and use long sequence missing information-aware masks and long sequence-aware learnable tokens to perform weighted processing and feature substitution on the missing regions respectively. The feature extraction network extracts long sequence features by channel based on the weighted and substituted multimodal long sequence signals. A multimodal dual alignment module is used to input the long sequence features of each channel into a shared encoder. The shared encoder extracts the wear features of the part surface and embeds them with modal shared features. The modal shared features are fused into the long sequence features of each channel through context enhancement. The fused long sequence features of each channel are then subjected to instance-level comparative learning and multimodal alignment to obtain fused features. The diffusion module is used to perform stepwise denoising optimization and reconstruction of the fused features using a conditional diffusion model, and to build an intermediate feedback mechanism during the diffusion process. It generates a feedback loss based on the denoising result and applies the feedback loss inversely to the shared encoder and context enhancement process. The multi-task prediction module is used to input the surface wear features of the part into the dual-head prediction module together with the features after diffusion optimization. The first prediction head in the dual-head prediction module is used to predict the surface wear of the part and obtain the predicted value of the surface wear of the part. The second prediction head combines the features after diffusion optimization and the predicted value of the surface wear of the part to perform regression prediction on the surface roughness of the part and obtain the predicted value of the surface roughness of the part.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a roughness prediction program based on long-sequence missing data perception stored in the memory and executable on the processor. When the roughness prediction program based on long-sequence missing data perception is executed by the processor, it implements the steps of the roughness prediction method based on long-sequence missing data perception as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a roughness prediction program based on long-sequence missing detection of multimodal data, which, when executed by a processor, implements the steps of the roughness prediction method based on long-sequence missing detection of multimodal data as described in any one of claims 1-7.