Model training and device prediction method based on time sequence visual language industrial large model
Patent Information
- Application Number
- CN202611042680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-14
AI Technical Summary
[0003]然而,现有方案对时序动态、频域结构及语义知识的联合建模能力不足,异构模态间表征空间差异较大,导致特征融合困难;同时训练中易出现模态贡献失衡,使模型在复杂工况、少样本及跨域场景下的精度和泛化能力受限
[0041]本申请提供的一种基于时序视觉语言工业大模型的模型训练与设备预测方法,通过获取工业设备的多传感器原始运行时序信号和工业知识文本,并基于时序视觉语言工业大模型分别形成时序嵌入向量、视觉嵌入向量和文本嵌入向量,再以时序嵌入向量为基准对各模态嵌入进行梯度对齐与特征融合处理,能够有效缩小异构模态间的表征差异,增强时序动态、频域结构及语义知识之间的协同表征能力,进而在复杂工况、少样本及跨域场景下实现了提高预测准确性的技术效果。
Smart Images

Figure CN122548663B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial equipment prediction and health management technology, and in particular to a model training and equipment prediction method based on a temporal visual language industrial large model. Background Technology
[0002] In the field of industrial equipment prediction and health management, condition assessment and life prediction are typically carried out based on sensor time-series data, combined with frequency domain analysis or textual knowledge, using single-modal models or dual-modal fusion methods.
[0003] However, existing solutions lack the ability to jointly model temporal dynamics, frequency domain structure, and semantic knowledge. The large differences in representation spaces between heterogeneous modalities make feature fusion difficult. At the same time, modal contribution imbalances are prone to occur during training, which limits the accuracy and generalization ability of the model in complex working conditions, with few samples, and in cross-domain scenarios.
[0004] Therefore, how to achieve effective collaborative representation and stable fusion of three-modal information of industrial equipment in order to improve the accuracy of prediction has become an urgent technical problem to be solved. Summary of the Invention
[0005] This application provides a model training and equipment prediction method based on a temporal visual language industrial large model to solve the above-mentioned technical problems, thereby achieving the technical effect of improving prediction accuracy.
[0006] Firstly, this application provides a method for predicting the state of industrial equipment based on a temporal visual language industrial large-scale model, including:
[0007] Acquire raw runtime sequence signals from multiple sensors of industrial equipment, as well as industrial knowledge text of the industrial equipment;
[0008] The following operations are performed using a temporal visual language industrial large-scale model:
[0009] The original runtime sequence signal is processed by time-series feature encoding to obtain a time-series embedding vector;
[0010] The original runtime sequence signal is subjected to frequency domain transformation and visual feature encoding to obtain a visual embedding vector;
[0011] Based on temporal embedding vectors, visual embedding vectors, and industry knowledge text, cross-modal fusion prompt data is generated.
[0012] Semantic encoding is performed on the cross-modal fusion prompt data to obtain text embedding vectors;
[0013] Based on the temporal embedding vector, gradient alignment and feature fusion processing are performed on the temporal embedding vector, visual embedding vector and text embedding vector to obtain the fused embedding vector.
[0014] Based on the fused embedding vector, prediction processing is performed to generate prediction results of the operating status of industrial equipment.
[0015] Secondly, this application provides a training method for a large-scale industrial model of temporal visual language, including:
[0016] Acquire historical multi-sensor runtime sequence signals of industrial equipment and historical industrial knowledge text of industrial equipment;
[0017] Based on historical multi-sensor runtime sequence signals, determine the corresponding frequency domain image samples;
[0018] A three-modal training dataset is constructed based on historical multi-sensor runtime sequence signals, historical industrial knowledge texts, and frequency domain image samples.
[0019] An initial trimodal model is constructed, which includes a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer, and a prediction layer. The temporal coding layer includes multiple temporal encoders and a routing mechanism for assigning corresponding temporal encoders to temporal features.
[0020] Based on the trimodal training dataset, a pre-defined two-stage training strategy is used to train the initial trimodal model, resulting in a trained temporal visual language industrial large model. The temporal visual language industrial large model is used to execute the method provided in the first aspect above.
[0021] Thirdly, this application provides an industrial equipment state prediction device based on a temporal visual language industrial large model, comprising:
[0022] The first acquisition module is used to acquire the raw runtime sequence signals collected by the multiple sensors of the industrial equipment and the industrial knowledge text of the industrial equipment.
[0023] The execution module is used to perform the following operations using a temporal visual language industrial large model:
[0024] The original runtime sequence signal is processed by time-series feature encoding to obtain a time-series embedding vector;
[0025] The original runtime sequence signal is subjected to frequency domain transformation and visual feature encoding to obtain a visual embedding vector;
[0026] Based on temporal embedding vectors, visual embedding vectors, and industry knowledge text, cross-modal fusion prompt data is generated.
[0027] Semantic encoding is performed on the cross-modal fusion prompt data to obtain text embedding vectors;
[0028] Based on the temporal embedding vector, gradient alignment and feature fusion processing are performed on the temporal embedding vector, visual embedding vector and text embedding vector to obtain the fused embedding vector.
[0029] Based on the fused embedding vector, prediction processing is performed to generate prediction results of the operating status of industrial equipment.
[0030] Fourthly, this application provides a training device for a temporal visual language industrial large-scale model, comprising:
[0031] The second acquisition module is used to acquire historical multi-sensor runtime sequence signals of industrial equipment and historical industrial knowledge text of industrial equipment.
[0032] The determination module is used to determine the corresponding frequency domain image samples based on the historical multi-sensor runtime sequence signals;
[0033] The first building module is used to construct a three-modal training dataset based on historical multi-sensor runtime sequence signals, historical industrial knowledge text, and frequency domain image samples.
[0034] The second building module is used to build an initial trimodal model, which includes a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer, and a prediction layer. The temporal coding layer includes multiple temporal encoders and a routing mechanism for assigning corresponding temporal encoders to temporal features.
[0035] The training module is used to train the initial trimodal model using a preset two-stage training strategy based on the trimodal training dataset, to obtain the trained temporal visual language industrial large model; the temporal visual language industrial large model is used to execute the method provided in the first aspect.
[0036] Fifthly, this application provides an electronic device, including: a memory and a processor;
[0037] The memory stores the instructions that the computer executes;
[0038] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods provided in the first and / or second aspects above.
[0039] In a sixth aspect, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the first and / or second aspects above.
[0040] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the first and / or second aspects above.
[0041] This application provides a model training and equipment prediction method based on a temporal visual language industrial big model. By acquiring the original operating sequence signals of industrial equipment from multiple sensors and industrial knowledge text, and forming temporal embedding vectors, visual embedding vectors, and text embedding vectors based on the temporal visual language industrial big model, gradient alignment and feature fusion processing are performed on each modality embedding based on the temporal embedding vector. This can effectively reduce the representation differences between heterogeneous modalities, enhance the collaborative representation ability between temporal dynamics, frequency domain structure, and semantic knowledge, and thus achieve the technical effect of improving prediction accuracy in complex working conditions, low sample size, and cross-domain scenarios. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] Figure 1 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 1 ;
[0044] Figure 2 A flowchart illustrating the training method for the temporal visual language industrial large model provided in this application embodiment;
[0045] Figure 3 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 2 ;
[0046] Figure 4 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 3 ;
[0047] Figure 5 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 4 ;
[0048] Figure 6 A schematic diagram of the structure of the industrial equipment state prediction device based on a temporal visual language industrial large model provided in this application embodiment;
[0049] Figure 7 A schematic diagram of the structure of the training device for the temporal visual language industrial large model provided in the embodiments of this application;
[0050] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0051] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0052] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0053] Industrial equipment operation status prediction and health management technology is widely used in scenarios such as aero engines, wind turbine generators, rail transit traction systems, energy storage batteries, motor bearings, compressors, and key units in process industries. Its core objective is to assess the health status, failure trends, and life changes of equipment based on multi-source data during continuous operation, thereby providing a basis for maintenance scheduling, risk warning, and operation optimization.
[0054] In practical deployments, raw operating sequence signals such as vibration, temperature, pressure, current, voltage, and speed are typically collected by field sensors and uploaded to a monitoring platform via edge acquisition terminals or industrial gateways. The backend analysis system then performs status identification, anomaly detection, or lifespan prediction. In addition to sensor data, many industrial scenarios also rely on industrial knowledge texts such as maintenance records, operating procedures, fault case libraries, design specifications, expert experience descriptions, and equipment manuals. These texts contain fault mechanisms, degradation modes, and handling suggestions, providing significant auxiliary value for understanding the operating status under complex conditions. Especially when equipment operates under varying loads, varying speeds, multiple operating conditions, or even with limited sample sizes, a single data source is often insufficient to accurately depict the true health status of the equipment. Therefore, industrial sites typically require a comprehensive analysis architecture capable of simultaneously processing raw time-series signals, frequency domain characteristics, and industrial knowledge information to meet the demands of high reliability, high continuity, and strong real-time operation and maintenance.
[0055] In existing technologies, the prediction of the operating status of industrial equipment is mostly based on time series data modeling. The typical approach is to directly preprocess, segment, extract features and infer models from the raw operating time series signals collected by multiple sensors, and use recurrent neural networks, convolutional networks or attention-based time series models to output fault categories, health scores or remaining life results.
[0056] Another approach converts the original runtime sequence signal into a two-dimensional representation such as a spectrogram or time-frequency graph using Fourier transform, wavelet transform, or other frequency domain analysis methods. Then, it uses a visual model to extract structured features, improving the ability to identify frequency domain patterns such as impulses, harmonics, and energy distribution. Some approaches also attempt to incorporate industrial knowledge text, such as summarizing equipment operating status in textual descriptions, or inputting maintenance knowledge and fault mode texts into a language model to assist in state reasoning.
[0057] While these solutions have achieved some success in their respective applicable scenarios, they still generally have significant limitations:
[0058] When relying solely on time-series modeling, the model is more likely to focus on local numerical fluctuations and is less sensitive to changes in the frequency domain structure of non-stationary signals. This can easily lead to missed detections or misjudgments when there is significant noise, weak initial fault characteristics, or frequent switching of operating conditions.
[0059] While relying solely on frequency domain images can effectively describe the distribution of frequency band energy and texture, it is insufficient in depicting the temporal evolution during degradation and makes it difficult to accurately distinguish between state patterns that appear similar but have different evolutionary paths. Relying solely on textual knowledge often relies on artificially abstracted descriptions of the signal, easily resulting in the loss of high-frequency details and information on changes in physical quantities from the original signal.
[0060] Even existing dual-modal fusion schemes often face problems such as large differences in heterogeneous feature spaces, inconsistent fusion interfaces, and one modality dominating training while the contributions of other modalities are suppressed. This makes it difficult for multi-source information to coordinate stably. In particular, under complex working conditions, cross-device migration, few-sample learning, and insufficiently labeled data environments, the model accuracy and generalization ability decline more significantly.
[0061] Therefore, how to achieve effective collaborative representation and stable fusion of three-modal information of industrial equipment in order to improve the accuracy of prediction has become an urgent technical problem to be solved.
[0062] To address the aforementioned technical issues, this application proposes a model training and equipment prediction method based on a temporal visual language industrial big data model. During the continuous operation monitoring of industrial equipment, the method first acquires raw runtime timing signals collected by multiple sensors and corresponding industrial knowledge text. The temporal visual language industrial big data model is then used to perform temporal feature encoding on the raw runtime timing signals to obtain temporal embedding vectors. Simultaneously, frequency domain transformation and visual feature encoding are performed on the raw runtime timing signals to obtain visual embedding vectors. Subsequently, cross-modal fusion prompt data is generated based on the temporal embedding vectors, visual embedding vectors, and industrial knowledge text. Semantic encoding is then performed on this cross-modal fusion prompt data to obtain text embedding vectors. Further, using the temporal embedding vectors as a benchmark, gradient alignment and feature fusion processing are performed on the temporal embedding vectors, visual embedding vectors, and text embedding vectors to obtain fused embedding vectors. Finally, prediction processing is performed based on the fused embedding vectors to generate the predicted operating status of the industrial equipment.
[0063] By employing the methods described above, the synergistic utilization of temporal information, visual information, and textual knowledge can be achieved within a unified prediction process, thereby enhancing the technical effectiveness of prediction accuracy.
[0064] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0065] Figure 1 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 1 ,like Figure 1 As shown in this embodiment, the industrial equipment state prediction method based on a temporal visual language industrial large model includes:
[0066] S101. Acquire the raw operating sequence signals collected by multiple sensors of the industrial equipment and the industrial knowledge text of the industrial equipment.
[0067] In this embodiment, the original runtime sequence signal is used to characterize the continuous monitoring data formed by multiple sensors during the operation of industrial equipment. It can reflect the changes in vibration, temperature, pressure, current, voltage, speed and other operating condition-related physical quantities of the equipment at different time points. The industrial knowledge text is used to characterize the explanatory knowledge information related to the industrial equipment, which can include maintenance manuals, inspection records, fault mode descriptions, degradation mechanism descriptions, design specifications, operating procedures, expert experience texts and historical case texts.
[0068] Based on the above analysis, the core of this step lies in building the heterogeneous input foundation required for subsequent trimodal prediction, so that the original numerical time series, the visual information derived from the time series, and the knowledge semantic information have a unified data source in the same processing flow.
[0069] Optionally, the executing entity can be a predictive processing device deployed in an industrial field edge server, industrial gateway, monitoring platform server, or cloud analysis node.
[0070] In practice, the acquisition of the original runtime timing signals can be achieved by periodically sampling through acquisition units such as accelerometers, temperature sensors, pressure sensors, current transformers, voltage sampling modules, and speed encoders connected to industrial equipment, forming a multi-channel timing data stream.
[0071] The sampling frequency of different sensors can be set according to the equipment type and fault-sensitive frequency band. For example, vibration signals use a higher sampling frequency, while temperature and pressure signals use a relatively lower sampling frequency. Then, the edge acquisition terminal completes the timestamp alignment, channel identifier binding, and buffer encapsulation, and then transmits the data to the analysis end via industrial Ethernet, fieldbus, wireless private network, or message queue.
[0072] To ensure consistency in subsequent modeling, the acquired multi-channel signals can be further processed by missing value imputation, outlier suppression, detrending, normalization, and unified time window segmentation. Missing value imputation can be accomplished using nearest neighbor interpolation, moving average, or model estimation, while outlier suppression can be accomplished using median filtering, threshold truncation, or robust smoothing.
[0073] Optionally, industrial knowledge texts can be obtained from enterprise equipment management systems, computerized maintenance management systems, knowledge base platforms, electronic manual databases, or manual input terminals. They can directly extract texts corresponding to the target equipment model, or filter relevant knowledge fragments based on the equipment's current operating conditions, historical fault categories, and component levels.
[0074] For unstructured text, word segmentation, terminology standardization, stop fragment cleaning, professional dictionary mapping, and sentence segmentation can be used to form a text set suitable for subsequent processing. For structured records, such as fields for maintenance time, fault location, handling measures, and degradation level, they can be converted into natural language fragments or key-value text expressions for organization together with subsequent prompt data.
[0075] Based on the above processing, this step outputs the original runtime sequence signals and industrial knowledge text that are interconnected within the same monitoring period, providing a unified input foundation for subsequent time-series coding, frequency domain conversion, semantic modeling, and fusion prediction. By simultaneously introducing data-driven and knowledge-driven information, the incomplete representation problem caused by relying solely on single sensor time series can be avoided. This allows the fault mechanisms and degradation patterns implicit in the knowledge text under complex operating conditions to participate in subsequent predictions, thereby providing input layer support for improving prediction stability and generalization ability.
[0076] It should be understood that the above examples are merely illustrative and not limiting, and do not affect the scope of protection of this application.
[0077] S102. Using the temporal visual language industrial large model, perform the following operations: perform temporal feature encoding processing on the original runtime timing signal to obtain the temporal embedding vector.
[0078] In this embodiment, the temporal visual language industrial big model is used to jointly model the original runtime sequence signals of industrial equipment and industrial knowledge text, and to complete embedding generation, gradient alignment, fusion, and prediction output. The temporal embedding vector is used to characterize the dynamic change features in the original runtime sequence signals and serves as the benchmark for subsequent cross-modal fusion and gradient alignment.
[0079] For example, the formation process of the timing embedding vector is not simply to directly encode the entire timing sequence. Instead, the continuous timing signal is first divided into multiple overlapping local timing segments. Then, each segment is encoded, weighted, summed, and globally averaged through multiple timing encoders and routing mechanisms in the timing coding layer, thereby taking into account local fault fluctuations, periodic changes, and global degradation trends.
[0080] S103. Using the temporal visual language industrial large model, perform the following operations: perform frequency domain transformation and visual feature encoding on the original runtime temporal signal to obtain the visual embedding vector.
[0081] In this embodiment, the visual embedding vector is used to characterize the visual features obtained by frequency domain transformation of the original runtime sequence signal and participates in trimodal fusion. These visual features are not directly derived from traditional camera images, but rather are obtained by converting the industrial equipment runtime sequence signal into a two-dimensional image capable of representing frequency band energy, periodic structure, nonlinear topology, and time-frequency local texture through recursive graph transformation, Fourier transform, and wavelet transform. The structured representation is then extracted and projected using a visual coding layer.
[0082] For example, the purpose of this step is to compensate for the lack of perception of frequency domain structure and image texture patterns by pure temporal coding, so that the model can identify impact textures, harmonic distributions, energy accumulation regions and multi-scale frequency changes from the two-dimensional representation.
[0083] In one possible implementation, the temporal visual language industrial large-scale model includes a visual coding layer; correspondingly, the original runtime temporal signal undergoes frequency domain transformation and visual feature encoding processing to obtain a visual embedding vector, including:
[0084] The original runtime sequence signal is processed by recursive graph transformation, Fourier transform, and wavelet transform through a visual coding layer to generate corresponding recursive feature maps, Fourier amplitude maps, and wavelet coefficient maps. The recursive feature maps, Fourier amplitude maps, and wavelet coefficient maps are then fused through the visual coding layer to obtain a frequency domain pseudo-color image. Visual features are extracted from the frequency domain pseudo-color image through the visual coding layer to obtain corresponding visual features, which are then projected to obtain a visual embedding vector.
[0085] In this embodiment, recursive graph transform is used to characterize the nonlinear topological structure and self-similarity of the time-series signal, Fourier transform is used to extract global frequency components and their amplitude distribution, and wavelet transform is used to characterize multi-scale time-frequency localization features. These three transforms collectively reflect the dynamic changes of the device at different degradation stages. The recursive feature map, Fourier amplitude map, and wavelet coefficient map are fused to form a pseudo-color image in the frequency domain, ensuring that different frequency domain information maintains identifiable structural features within the same image carrier, facilitating unified processing by the visual coding network. Visual feature extraction can be achieved using convolutional networks, visual transformers, or combinations thereof. Projection processing is used to map the extracted high-dimensional visual features to a preset dimensional space to obtain a visual embedding vector consistent with other modal dimensions. The specific formulas for recursive graph transform, Fourier transform, and wavelet transform processing of the original runtime time-series signal are shown below:
[0086]
[0087] in, Let (i,j) be the value of the recurrence graph of the b-th sample at position (i,j). It is a step function. For recursive threshold, Let b be the sensor vector of the b-th sample at time step i. Let f(x) be the Fourier amplitude of the k-th frequency component in the c-th channel of the b-th sample. This represents the value of the b-th sample, the c-th channel, and the n-th time step. For the wavelet coefficients of the c-th channel of the b-th sample, The scaling parameter of the wavelet transform. For the translation index of the wavelet transform, This is the conjugate wavelet function after scaling and translation.
[0088] For example, the original runtime sequence signal can first undergo length normalization, denoising, and segmentation, and then be input into the recursive graph generation module, Fourier transform module, and wavelet transform module respectively to obtain three types of two-dimensional images (recursive feature map, Fourier amplitude map, and wavelet coefficient map). These three types of images can be fused using channel stitching, weight stacking, or color mapping. Color mapping can encode different frequency domain information into different color levels, enhancing the visibility of image texture differences and energy distribution differences. Subsequently, the frequency domain pseudo-color image enters the visual coding network, outputting a visual feature vector, which is then compressed to a predetermined embedding dimension through a linear projection layer or multilayer perceptron to form a visual embedding vector for subsequent cross-modal alignment and fusion. The visual embedding vector is obtained using the following formula:
[0089]
[0090] in, The original visual features of the b-th sample are... For encoding functions, For the frequency domain pseudo-color image of the b-th sample, Let be the visual embedding vector of the b-th sample. This is a mapping function.
[0091] It should be noted that in practical applications, the visual coding layer can also choose other equivalent visual feature extraction structures, and this application embodiment does not limit this.
[0092] Through the above structure, the temporal fluctuations, frequency domain energy distribution, and multi-scale local changes in the original runtime sequence signal can be synchronously transformed into learnable image features, thereby preserving key information such as fault impacts, periodic disturbances, and degraded textures. The resulting visual embedding vector possesses both structural expressiveness and semantic transferability, effectively supplementing the shortcomings of simple temporal coding in perceiving frequency domain patterns, and improving the accuracy and generalization ability of runtime state prediction under complex operating conditions.
[0093] S104. Using the temporal visual language industrial big model, perform the following operations: Generate cross-modal fusion prompt data based on the temporal embedding vector, visual embedding vector, and industrial knowledge text.
[0094] In this embodiment, cross-modal fusion prompt data is used to concatenate temporal features, visual features, and industrial knowledge text into prompt content that can be semantically encoded, thereby mapping heterogeneous modalities to a unified semantic expression.
[0095] For example, the quality of text embedding vectors depends on whether the cue content adequately expresses the key information of both the numerical and visual modalities. Therefore, this step is not simply concatenating the original text, but rather performing feature interpretation and semantic summarization on the temporal and visual embedding vectors first, and then combining them with industry knowledge text to form structured cue content and cross-modal fusion cue data.
[0096] Optionally, the construction of cross-modal fusion prompt data may include extracting baseline state features, fault-sensitive fluctuation features and life-related trend features based on temporal embedding vectors, extracting color distribution features, texture features, spatial layout features and density pattern features based on visual embedding vectors, and finally concatenating them with industrial knowledge text to form a unified prompt.
[0097] S105. Using the temporal visual language industrial large model, perform the following operations: perform semantic encoding processing on the cross-modal fusion prompt data to obtain the text embedding vector.
[0098] In this embodiment, the text embedding vector is used to represent the semantic information encoded from the cross-modal fusion cue data, and participates in the fusion together with the temporal embedding vector and the visual embedding vector. The semantic encoding process transforms the semantically reconstructed trimodal cue content into a continuous vector representation, enabling the fault mechanisms, operational contexts, and maintenance experiences in industrial knowledge texts to be associated with the state features extracted from the signals in a unified semantic space. The text embedding vector is obtained using the following formula:
[0099]
[0100] in, The original semantic features of the b-th sample are... For encoding functions, For the cross-modal fusion cue data of the b-th sample, Let b be the text embedding vector of the b-th sample. This is a mapping function.
[0101] For example, this step does not encode the original knowledge text separately, but encodes the cue data that already contains temporal summaries, visual summaries and knowledge context as a whole. Therefore, the output text embedding vector naturally carries cross-modal integrated semantics.
[0102] S106. Using the temporal visual language industrial large model, perform the following operations: Based on the temporal embedding vector, perform gradient alignment and feature fusion processing on the temporal embedding vector, visual embedding vector and text embedding vector to obtain the fused embedding vector.
[0103] In this embodiment, the temporal embedding vector serves as a benchmark, indicating that the fusion process centers on the physical dynamic features of the original running signal, aligning and coordinating the visual and textual auxiliary modalities. Gradient alignment is used to coordinate the update direction and intensity of different modal layers during training, preventing one modality from dominating training due to excessively strong gradients or being marginalized due to insufficient learning. Feature fusion is used to converge the three modalities into a unified fused embedding vector at the representation layer. According to the technical analysis, this process may include calculating modal attention weights based on the temporal embedding vector, performing weighted fusion of the three embedding vectors, and then using the temporal embedding vector as an anchor point, aggregating the weighted fused embedding vector through a cross-modal attention mechanism, while simultaneously combining adaptive gradient enhancement and cross-layer consistency constraints to achieve gradient-level coordination.
[0104] S107. Using the temporal visual language industrial large model, perform the following operations: perform prediction processing based on the fused embedding vector to generate the prediction results of the operating status of industrial equipment.
[0105] In this embodiment, the fused embedding vector is a unified representation that integrates temporal dynamic information, frequency domain visual information, and industrial knowledge semantic information. The prediction processing maps this unified representation to specific industrial equipment operating status results. The operating status prediction results can be presented as classification results, regression results, or ranking results according to actual business objectives, such as fault type, health status score, remaining useful life prediction value, risk level, degradation trend index, or a determination result of whether an alarm is required.
[0106] Based on the above analysis, this step transforms the advantages of the aforementioned three-modal collaborative representation into an output that can be directly used by the operation and maintenance system, thereby completing the closed loop from raw monitoring data to decision results.
[0107] In one possible implementation, the temporal visual language industrial big data model includes a prediction layer; accordingly, based on the fused embedding vectors, a prediction result of the operating status of industrial equipment is generated, including:
[0108] The prediction layer classifies the fused embedding vectors to determine the classification results and corresponding confidence levels of the industrial equipment's operating status. The prediction layer then performs regression processing on the fused embedding vectors to obtain the remaining service life of the industrial equipment and the corresponding degradation trend curve. Finally, the operating status classification results, confidence levels, remaining service life values, and degradation trend curves are integrated to obtain the predicted operating status results of the industrial equipment.
[0109] In this embodiment, the fused embedding vector is a comprehensive representation of temporal and frequency domain visual data and industrial knowledge text after alignment and fusion, which can simultaneously retain the current state of the equipment, degradation mode, and knowledge constraint information. The operating status classification result is used to quantify the fault type and severity of industrial equipment, the confidence score is used to characterize the credibility of the classification result, the remaining service life value is used to indicate the length of time the equipment can continue to operate safely, and the degradation trend curve is used to reflect the decay trajectory of the service life over time.
[0110] For example, the prediction layer may include a shared feature input unit, a classification output unit, and a regression output unit. The shared feature input unit fuses the embedded vectors into the subsequent network structure. The classification output unit can use a fully connected mapping combined with a normalized exponential function (Softmax) to normalize the probability distribution of each fault category, and take the category with the highest probability as the operational status classification result; this highest probability is the confidence level. The regression output unit can use a linear mapping or a multilayer perceptron to output the remaining service life value, and form a degradation trend curve based on the regression results within continuous time intervals or a sliding window. The classification task and the regression task can share the front-end representation and output it in the final layer to reduce the mutual interference of modal information in the output stage.
[0111] It should be noted that in practical applications, the prediction layer can also be selected from other neural network structures, and this application does not limit this.
[0112] Specifically, the fused embedding vector first enters the classification layer of the prediction layer to output whether the device is in a normal, warning, mildly degraded, severely degraded, or specific fault state, along with its corresponding probability. Then, the same fused embedding vector enters the regression layer to output the remaining service life and its temporal evolution trend. Finally, the classification results, confidence scores, remaining service life values, and degradation trend curves are uniformly encapsulated into an operational status prediction result for subsequent alarm, maintenance decision-making, or lifespan management. This approach allows the model to simultaneously address fault identification and lifespan prediction, improving the completeness, reliability, and interpretability of prediction results under complex operating conditions, and enhancing the ability to represent the entire process of equipment degradation.
[0113] In one possible implementation, after integrating the operating status classification results, confidence level, remaining useful life, and degradation trend curve to obtain the operating status prediction results of the industrial equipment, the method further includes:
[0114] Based on the classification results of operating status, the severity level of industrial equipment failure is determined; when the severity level of failure is higher than the preset level threshold, and / or the remaining service life is lower than the preset service life threshold, corresponding early warning information is generated.
[0115] In this embodiment, the fault severity level is a classification result used to characterize the current fault risk level of industrial equipment, which can be mapped from the operating status classification result. The operating status classification result is used to identify the normal, slightly abnormal, severely abnormal, or faulty state of the equipment, and can further correspond to different levels of fault severity. The preset level threshold is used to limit the minimum severity level for triggering an early warning, and the remaining service life value is used to reflect the time interval during which the equipment can continue to operate safely under the current operating conditions. The preset service life threshold is used to limit the minimum acceptable service life boundary. The early warning information is used to output risk prompts to the monitoring platform, operation and maintenance terminal, or alarm module so as to arrange maintenance, load reduction, or shutdown in a timely manner.
[0116] In practice, after integrating the operational status prediction results, the background analysis unit can read the operational status classification results and determine the severity level of the fault based on a pre-established level mapping table. The mapping table can be jointly determined by historical fault samples, expert rules, and equipment maintenance specifications. For example, minor wear, local anomalies, and critical failures can be assigned to low, medium, and high severity levels, respectively.
[0117] The process of determining the severity level of a fault can be carried out by using a fixed rule matching method, or by combining confidence level with a credible correction of the classification results. When the classification confidence level is low, the conservative judgment weight of the severity level is increased, thereby reducing the risk of missed detection.
[0118] The remaining service life value can be directly used as the output value of the prediction module and compared with the preset service life threshold. When any judgment condition is met, the alarm module generates an early warning message containing the equipment identifier, fault level, service life information and suggested handling measures, and outputs it through audible and visual alarms, message push or control system linkage.
[0119] Through the above processing, the system can further classify fault risks after providing operational status prediction results, and implement dual-condition triggering early warning based on severity level and lifespan boundary, thus making the early warning output closer to the actual health status of the equipment. This approach can improve the timeliness of handling after fault identification, reduce the probability of minor anomalies evolving into major faults, and enhance the safety and continuity of industrial equipment operation and maintenance management.
[0120] This application provides a method for predicting the state of industrial equipment based on a temporal visual language industrial big data model. By acquiring the original operating sequence signals of the industrial equipment from multiple sensors and industrial knowledge text, and forming temporal embedding vectors, visual embedding vectors, and text embedding vectors based on the temporal visual language industrial big data model, gradient alignment and feature fusion processing are performed on each modality embedding based on the temporal embedding vector. This method can effectively reduce the representation differences between heterogeneous modalities and enhance the collaborative representation ability between temporal dynamics, frequency domain structure, and semantic knowledge. Thus, it achieves the technical effect of improving prediction accuracy in complex working conditions, low sample size, and cross-domain scenarios.
[0121] Figure 2 A flowchart illustrating the training method for the temporal visual language industrial large-scale model provided in this application embodiment is shown below. Figure 2 As shown in this embodiment, the training method for the temporal visual language industrial large-scale model is provided. The temporal visual language industrial large-scale model is used to execute the industrial equipment state prediction method based on the temporal visual language industrial large-scale model provided in the above embodiment. The training method includes:
[0122] S201. Obtain historical multi-sensor runtime sequence signals of industrial equipment and historical industrial knowledge text of industrial equipment.
[0123] In this embodiment, historical multi-sensor runtime sequence signals of the industrial equipment serve as the raw time-series data source for training input, characterizing the dynamic state of the industrial equipment during past operation. Specifically, this data can be multi-channel time-series data continuously collected by multiple sensors installed on the equipment body, transmission components, load side, lubrication system, or environmental monitoring locations. Data types include vibration, temperature, pressure, current, voltage, speed, flow rate, acoustic emission, or displacement. Historical industrial knowledge text serves as the textual knowledge source for training input, providing domain semantic information for the equipment status. Specifically, this text can originate from unstructured or semi-structured text materials such as maintenance manuals, fault mode descriptions, inspection records, maintenance work orders, operating procedures, maintenance logs, design specifications, fault case libraries, and expert experience documents.
[0124] Based on the above analysis, the core of this step is to build a dual-source input foundation that can cover the physical state of equipment operation and the semantics of domain knowledge, so that subsequent training no longer depends on a single numerical signal, but establishes the correlation between time-series information and industrial knowledge in the initial stage of training.
[0125] In one possible embodiment, the executing entity can be an industrial internet platform server, a training server cluster deployed in an enterprise data center, a cloud training module in an edge computing node and cloud collaborative system, or a dedicated modeling device with graphics processor and storage resources.
[0126] In practice, the system can first receive multi-sensor runtime sequence signals generated by industrial equipment during its historical operating cycle through programmable logic controllers, data acquisition cards, edge gateways, or distributed control systems in the industrial field, and then write the collected results into a time-series database, data lake, or distributed file system through a unified communication interface.
[0127] Optionally, for sensor signals with different sampling frequencies, timestamp alignment, resampling, missing value compensation, abnormal peak removal, and dimension normalization can be performed first to form a standardized historical multi-sensor runtime sequence signal that can be used for subsequent modeling.
[0128] Optionally, to ensure the traceability of training samples, metadata such as device identifier, component identifier, operating condition label, sampling start and end time, sampling frequency, and operation batch number can be added to each historical time-series signal.
[0129] Correspondingly, historical industrial knowledge texts can be acquired through methods such as importing from enterprise knowledge bases, document parsing, manual annotation and entry, extraction from maintenance management systems, conversion of paper documents using optical character recognition, and generation of descriptive statements from structured tables.
[0130] It is important to note that for issues such as inconsistent terminology, mixed abbreviations, non-standard formatting, and noisy characters that may exist in the original text, text cleaning, terminology standardization, sentence segmentation, stop character filtering, professional dictionary mapping, and entity extraction processing can be performed to organize the knowledge content, including fault mechanisms, degradation patterns, diagnostic criteria, treatment plans, and operational limitations, into a historical industrial knowledge text set that can be used for semantic encoding.
[0131] In terms of specific organization methods, knowledge texts can be indexed according to equipment type, component level, fault type, operating condition category, or maintenance stage, and a preliminary association can be established based on equipment codes or sample time windows and time series data.
[0132] Based on the above processing flow, two types of complementary input resources can be formed before training. The historical multi-sensor runtime sequence signals reflect the continuous change trajectory of the equipment status, while the historical industrial knowledge text reflects the fault mechanism and handling semantics related to the change trajectory.
[0133] By acquiring both types of data simultaneously, the model can learn the implicit correspondence between numerical fluctuations, frequency structure, and semantic knowledge in subsequent training, thus laying the foundation for solving the problem of incomplete state characterization caused by relying on only a single data source under complex working conditions.
[0134] It should be understood that the above examples are merely illustrative and not limiting, and do not affect the scope of protection of this application.
[0135] S202. Determine the corresponding frequency domain image samples based on the historical multi-sensor runtime sequence signals.
[0136] In this embodiment, the frequency domain image samples are obtained by converting historical multi-sensor runtime sequence signals to provide visual modality training data. Their function is to map the dynamic signals that originally existed in one-dimensional time sequence form into two-dimensional image representations so that the visual coding layer can extract structured frequency domain features, texture distribution features, and local energy accumulation features.
[0137] Based on the above analysis, it can be seen that in the early stage of failure, during load changes or speed switching, the differences in some key states of industrial equipment are not always directly reflected in the amplitude of the original time domain waveform, but may be reflected in energy shifts in certain frequency bands, harmonic enhancement, sideband structure changes or transient impact textures. Therefore, further converting the historical multi-sensor runtime sequence signals into frequency domain image samples can help improve the model's ability to identify complex non-stationary patterns.
[0138] In one possible embodiment, the standardized historical multi-sensor runtime sequence signal can first be sliced according to a preset window length, window step size, and sample labeling strategy to form multiple time window samples. The window length can be determined based on the equipment rotation cycle, process cycle time, or duration of degradation characteristics. The window step size can be a fixed step size or an overlapping sliding window method to ensure that critical fault segments are not lost due to segmentation boundaries. Subsequently, frequency domain transformation is performed on each time window sample. The frequency domain transformation can specifically employ a Fast Fourier Transform (FFT) to map the one-dimensional time-series signal into a spectral amplitude map; alternatively, a Short-Time Fourier Transform (SFT) can be used to generate a time-spectrum map to simultaneously preserve temporal locality and frequency distribution information; or a continuous wavelet transform can be used to generate a scaled spectrum to enhance the analytical capability for transient impacts and non-stationary signals.
[0139] In another possible embodiment, a recursive graph can be used to visualize the repeating trajectory of the time-series signal in phase space, thereby characterizing the changes in the system's dynamic state. For multi-channel sensor data, different channels can be converted into separate images and then stacked, or the frequency domain results from multiple sensors can be stitched together to form a multi-channel frequency domain image sample.
[0140] In practice, to ensure consistency between the frequency domain image samples and the input requirements of the visual coding layer, the conversion results can be further processed with amplitude normalization, logarithmic compression, color mapping, pixel size unification, boundary cropping, and noise suppression. For example, for the spectral amplitude obtained from the Fast Fourier Transform, the difference between strong and weak frequency components can be reduced by taking the logarithm of the amplitude; for the time-frequency matrix formed by the Short-Time Fourier Transform, bilinear interpolation can be performed based on the preset image size; for the wavelet transform results, the scale axis can be mapped to the standard frequency axis to ensure comparability of visual distributions between different samples. The generated frequency domain image samples maintain a one-to-one correspondence with the original time window samples, and are supplemented with device number, sampling time period, sensor source, and label information, thereby enabling traceable pairing in the subsequent dataset construction and training stages.
[0141] Based on the above analysis, this step transforms the frequency features and local variation patterns in historical multi-sensor runtime sequential signals into a more visually extractable structured representation, giving the model a second representational perspective in addition to temporal coding. This not only compensates for the lack of perception of frequency domain details in single temporal modeling, but also establishes a corresponding foundation for temporal and visual modalities in subsequent cross-modal fusion, thereby improving the stability of trimodal collaborative learning.
[0142] It should be understood that the above examples are merely illustrative and not limiting, and do not affect the scope of protection of this application.
[0143] S203. Construct a three-modal training dataset based on historical multi-sensor runtime sequence signals, historical industrial knowledge texts, and frequency domain image samples.
[0144] In this embodiment, the trimodal training dataset is used to train the initial trimodal model. Essentially, it unifies the temporal, visual, and textual modalities into a training sample set that can be jointly input at the sample level. Historical multi-sensor runtime temporal signals correspond to temporal modal input, frequency domain image samples correspond to visual modal input, and historical industrial knowledge text corresponds to semantic modal input.
[0145] Based on the above analysis, it can be seen that the data of each modality alone is not enough to achieve stable training. It is also necessary to establish cross-modal mapping relationships under semantically consistent conditions such as the same device, the same time window, the same operating state, or the same fault category, so that the model can learn the correspondence, complementary relationship and collaborative representation mechanism between different modalities during the training process.
[0146] In one possible implementation, when constructing the trimodal training dataset, the time window of the historical multi-sensor runtime sequence signal can be used as the primary index of the sample.
[0147] For each time-series window sample, first bind the corresponding frequency domain image sample obtained by frequency domain transformation of the time-series window, and then retrieve the most relevant text fragments to the time-series window from the historical industrial knowledge text set according to the equipment number, component identification, fault type, operating condition description, maintenance stage, sampling time neighborhood or label rules.
[0148] If the knowledge text is a long-term stable document, such as an equipment manual or fault manual, it can be associated with the equipment model and the fault mode.
[0149] If the knowledge text consists of maintenance work orders and inspection records, they can be associated according to the principles of time proximity, component consistency, and status label consistency.
[0150] Optionally, for cases where a time window may correspond to multiple text segments, one or more text segments can be selected as training input based on relevance score, keyword matching strength, semantic vector similarity, or manual rules.
[0151] In terms of data organization, each training sample can contain time-series data tensors, visual image tensors, text sequences, and supervision information. Supervision information can include health status categories, fault categories, remaining lifespan labels, health scores, anomaly scores, or degradation stage identifiers. For different task types, classification labels, regression labels, or ranking labels can be selected.
[0152] Optionally, to improve training quality, class balancing, sample deduplication, label correction, and outlier removal can be performed during the dataset construction phase. For example, when faulty samples are significantly fewer than normal samples, the probability of minority class samples appearing in training can be increased through temporal sliding window enhancement, frequency domain image enhancement, text fragment recombination, or weighted sampling mechanisms; when there are redundant descriptions in the text source, more compact knowledge text input can be generated through keyword filtering, similar sentence aggregation, and semantic compression.
[0153] In another possible embodiment, to facilitate subsequent phased training, the training set, validation set, and test set can be pre-divided in the three-modal training dataset, and highly correlated samples within the same device's operating cycle can be prevented from leaking across sets.
[0154] For cross-device generalization scenarios, training and validation sets can be further divided by device ID to test the model's transferability. The dataset can also store the length information, effective mask, sample confidence score, and text association confidence score for each modality, thus supporting mask calculation, loss weighting, and anomaly robustness handling during subsequent model training.
[0155] In summary, the trimodal training dataset is no longer a simple data stack, but rather forms a cross-modal pairing structure with a unified sample semantic center. This structure enables the model to simultaneously receive temporal variation information, frequency domain visual information, and industrial knowledge semantic information during training, and learn how these three together point to the device state through supervised signals. This effectively alleviates the problems of inconsistent heterogeneous modal interfaces, weak correlations, and difficulties in collaborative learning in existing technologies, laying a data foundation for subsequent multi-layer encoding, fusion, and optimization.
[0156] It should be understood that the above examples are merely illustrative and not limiting, and do not affect the scope of protection of this application.
[0157] S204. Construct an initial trimodal model, which includes a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer, and a prediction layer. The temporal coding layer includes multiple temporal encoders and a routing mechanism for assigning corresponding temporal encoders to temporal features.
[0158] In this embodiment, the initial trimodal model is used to carry the complete processing chain of trimodal feature extraction, fusion and prediction.
[0159] The temporal coding layer is used to encode the historical multi-sensor runtime timing signals, and the visual coding layer is used to encode the frequency domain image samples. The multiple temporal encoders in the temporal coding layer are used to adapt to the temporal features of different complexities, different operating conditions, or different local segments. The routing mechanism is used to assign the temporal features to one or more temporal encoders for processing based on the input feature status, thereby forming a hybrid expert-style temporal modeling architecture.
[0160] The semantic coding layer is used to encode historical industrial knowledge text or prompt text constructed based on temporal and visual information;
[0161] Gradient alignment fusion layer is used to coordinate and fuse feature representations from different modalities;
[0162] The prediction layer is used to output the prediction results corresponding to the training task.
[0163] Based on the above analysis, this step is the core of the entire training method, which determines the flow, representation, and joint optimization of the three-modal data within the model.
[0164] In one possible embodiment, the temporal coding layer may include multiple temporal encoders configured in parallel. Each temporal encoder may be implemented using a one-dimensional convolutional network, a temporal transformer, a gated recurrent unit network, a long short-term memory network, or a hybrid structure combining convolution and attention. The input historical multi-sensor runtime temporal signal is first processed through linear projection, position encoding, and channel mapping before being fed into multiple temporal encoder candidate layers. The routing mechanism may be implemented using a separate lightweight routing network, whose inputs are the statistical features of the temporal segment, the initial embedding representation, or the local complexity score, and whose outputs are the routing weights corresponding to each temporal encoder.
[0165] For example, the routing mechanism can employ a top-k selection strategy, activating only one or more temporal encoders with the highest scores to reduce computational overhead and enhance the ability to model by division of labor. The outputs of each temporal encoder can be weighted and summed or concatenated according to the routing weights to form a temporal embedding vector. This setup allows samples with different operating modes, noise levels, and degradation stages to be represented by more suitable encoders, thereby improving the accuracy of temporal modeling.
[0166] The visual coding layer can employ convolutional neural networks, visual Transformers, or a fusion structure of convolution and attention to perform block segmentation, feature extraction, and global aggregation processing on frequency domain image samples, and output visual embedding vectors.
[0167] The semantic encoding layer can use a pre-trained language model to perform word segmentation, embedding, context encoding, and pooling on historical industrial knowledge texts to obtain text embedding vectors.
[0168] In one possible embodiment, the input to the semantic coding layer can be not only the original industrial knowledge text, but also cross-modal fusion prompt data generated based on time-series statistical features, frequency domain image descriptions, and equipment operation tags. For example, the mean, variance, dominant frequency, frequency band energy, and operating condition description can be filled into the template statement to enhance the semantic inheritance capability of the text modality to other modalities.
[0169] The gradient alignment fusion layer is used to receive temporal embedding vectors, visual embedding vectors, and text embedding vectors, and to perform feature alignment and fusion.
[0170] Specifically, different modalities can be embedded and projected onto a unified dimensional space through linear mapping, and then a fused embedding vector can be formed through cross-attention, multilayer perceptron gating fusion, tensor fusion, or weighted residual fusion.
[0171] During training, the gradient alignment fusion layer also participates in calculating the alignment loss and gradient coordination term, ensuring that each modality layer maintains a relatively consistent optimization trend during parameter updates. For example, based on the angle relationship between the gradient vectors of each layer, conflicting gradients can be projected or scaled to reduce the problem of one modality dominating training and inhibiting the learning of other modalities.
[0172] The prediction layer can use a fully connected layer, a classifier head, or a regression head to map the fused embedding vector to the fault category, health score, or remaining life prediction value.
[0173] Based on the above model structure, the initial trimodal model possesses a complete path from trimodal inputs to predicted outputs before training begins. Furthermore, the introduction of a temporal encoder and routing mechanism enables adaptive modeling of complex temporal patterns. The introduction of a gradient alignment fusion layer achieves balanced optimization of the trimodal layers during training. This structure directly addresses the needs of multi-condition, heterogeneous data, and knowledge-assisted analysis in industrial equipment, providing a structural solution to the problems of imbalanced modal contributions and inconsistent interfaces in existing fusion schemes.
[0174] It should be understood that the above examples are merely illustrative and not limiting, and do not affect the scope of protection of this application.
[0175] S205. Based on the trimodal training dataset, a pre-defined two-stage training strategy is adopted to train the initial trimodal model, thereby obtaining the trained temporal visual language industrial large model.
[0176] In this embodiment, a two-stage training strategy is used to optimize the initial three-modal model in stages, so that different modalities can gradually reach a cooperative state during the training process.
[0177] Based on the above analysis, this step is a key step in achieving training stability and fusion effectiveness. Its purpose is to avoid optimization oscillations, modal competition, and dominance imbalance that occur when the three-modal models are directly trained from random initialization states.
[0178] In one possible implementation, the two-stage training strategy includes:
[0179] In the first training phase, the parameters of the visual coding layer and the semantic coding layer are frozen, and the temporal coding layer is trained based on the trimodal training dataset.
[0180] In the second training phase, the network parameters of the temporal encoder in the temporal coding layer are frozen; the parameters of the visual coding layer and the semantic coding layer are fine-tuned based on the trimodal training dataset; the feature alignment loss of the temporal coding layer, the visual coding layer, and the semantic coding layer is determined based on the gradient alignment fusion layer; within each training batch, the second total routing weight and the second sample load of each temporal encoder are determined based on the routing weights assigned to each temporal encoder by the samples in the current batch; the second coefficient of variation is calculated based on the second total routing weight and the second sample load; the second balance regularization term of the temporal coding layer is generated based on the second coefficient of variation; and the parameters of the initial trimodal model are updated based on the feature alignment loss and the second balance regularization term.
[0181] In this embodiment, the second total routing weight is used to characterize the cumulative selection strength of each time encoder in the current batch, the second sample load is used to characterize the number of samples it carries, the second coefficient of variation is used to reflect the degree of dispersion of route allocation, and the second balance regularization term is used to suppress route collapse and maintain the participation of each time encoder in balance.
[0182] In practice, the first training phase only updates the parameters of the time-series coding layer, allowing it to learn the degradation patterns, impulse characteristics, and operating condition changes in the original time-series signal, thereby forming a stable basic representation of the time series.
[0183] The second training phase keeps the parameters of the temporal encoders in the temporal encoding layer unchanged to avoid drift of existing temporal representations during fusion. Simultaneously, the visual encoding and semantic encoding layers are fine-tuned to better adapt to the visual features and textual semantics constructed from the same training data. Subsequently, the gradient alignment fusion layer outputs the feature alignment loss between each modality. Combined with the routing weights of each sample in the current batch assigned to different temporal encoders, a second total routing weight and a second sample load are calculated. Based on this, a second coefficient of variation is calculated, and a second balanced regularization term is generated. Finally, both types of losses are used to drive parameter updates, improving training stability while maintaining modal coordination.
[0184] For example, feature alignment loss is used to address modality imbalance and convergence conflicts during the training process of the three modalities, ensuring the core dominance of the temporal layer while improving the learning performance of the visual and text layers. Its computation includes two core mechanisms: adaptive gradient enhancement and cross-layer consistency constraints.
[0185] The adaptive gradient enhancement mechanism dynamically adjusts the gradient update magnitude of each modality layer, assigning larger gradient coefficients to modalities with lower reliability to address the underfitting problem of auxiliary modalities caused by modality imbalance. In its implementation, it first calculates the confidence scores of each modality layer and the fusion layer based on the task loss. Then, it calculates the gradient enhancement coefficients for each modality based on the confidence differences. Finally, it normalizes the gradients of the visual and text layers using the temporal layer gradient as a benchmark, obtaining the temporal center-normalized gradient for model parameter updates. The specific calculation formula is as follows:
[0186]
[0187] in, The confidence score for the m-th modality is... The task loss for the m-th modal layer. The confidence score for the fusion layer. This represents the task loss of the fusion layer.
[0188]
[0189] in, Let be the gradient enhancement coefficient for the m-th mode. It is a positive part function. The gradient enhancement intensity coefficient, is the numerical stability constant.
[0190]
[0191] in, For the temporal center-normalized gradient of the m-th modal layer, The original gradient of the m-th modal layer. Let T be the L2 norm, and let V, K be the temporal, visual, and textual modalities, respectively.
[0192] Cross-layer consistency constraints are used to mitigate modality bias caused by skewed sample distribution, forcing the prediction logic of each modality layer to remain consistent. In practice, the reliability weights of modality pairs are first calculated based on the confidence levels of each modality. Then, a difference metric function is used to calculate the difference in prediction results between each modality pair. Finally, a weighted sum is taken to obtain the total consistency loss, as shown in the following formula:
[0193]
[0194] in, For the reliability weights of the mode pair (m, n), It is a set of modal pairs.
[0195]
[0196] in, For cross-layer consistency loss, For difference measurement function, Let B be the prediction result of the m-th modal layer for the b-th sample, and B be the training batch.
[0197] The final feature alignment loss is composed of gradient constraints generated by the adaptive gradient enhancement mechanism and cross-layer consistency loss.
[0198] Optionally, the parameters of the initial trimodal model can be updated based on the feature alignment loss and the second balance regularization term. This can be achieved by first constructing a total model loss function based on the feature alignment loss and the second balance regularization term, and then using the total model loss function to update the parameters of the initial trimodal model.
[0199] For example, the total training loss of the model is a weighted sum of the task loss of the fusion layer, the supervision loss of each modality layer, the balance regularization loss, and the feature alignment loss. By jointly optimizing the above loss terms, collaborative training and depth alignment of the three modalities are achieved. The specific calculation formula is as follows:
[0200]
[0201] in, Let be the total loss function of the model. The weighting coefficients for layer supervision loss. To balance the weight coefficients of the regularization term, To balance the regularization term, The cross-layer consistency loss weighting coefficient is used. For modal sets, This is a loss for the fusion layer tasks.
[0202] The training method described above enables the temporal coding layer to first obtain a reliable backbone representation, and then, under the condition of freezing the backbone, the visual and semantic layers are collaboratively adapted, thereby reducing optimization conflicts between heterogeneous modalities. By introducing gradient alignment constraints, the fusion bias caused by inconsistent orientations of different layers can be reduced; by using route equalization regularization, local overfitting caused by a few temporal encoders occupying too many samples for a long time can be avoided. This can improve the convergence stability, feature consistency, and generalization ability of the temporal vision-language industrial large-scale model under complex working conditions, with few samples, and in cross-device scenarios.
[0203] This application provides a training method for a temporal visual language industrial large-scale model, including acquiring historical multi-sensor runtime sequence signals of industrial equipment and historical industrial knowledge text of industrial equipment; determining corresponding frequency domain image samples based on the historical multi-sensor runtime sequence signals; constructing a three-modal training dataset based on the historical multi-sensor runtime sequence signals, historical industrial knowledge text, and frequency domain image samples; constructing an initial three-modal model including a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer, and a prediction layer; and training the initial three-modal model using a preset two-stage training strategy based on the three-modal training dataset to obtain the trained temporal visual language industrial large-scale model. In this embodiment, the original time-series signals, frequency domain images, and industrial knowledge text are organized in a unified training process. A time-series representation foundation adapted to complex working conditions is established through a time-series coding layer including multiple time-series encoders and routing mechanisms. Then, stable collaborative learning of three-modal features is achieved by using a gradient alignment fusion layer and a two-stage training strategy. This effectively alleviates problems such as large differences in heterogeneous modal spaces, unstable joint optimization, and imbalance of modal contributions during the training phase. As a result, the obtained time-series visual language industrial big model takes into account time-domain dynamic features, frequency domain structural features, and industrial knowledge semantic features, thereby improving the accuracy of predicting the operating status of industrial equipment.
[0204] Figure 3 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 2 ,like Figure 3 As shown, this embodiment, based on the above embodiments, provides a detailed description of the generation process of cross-modal fusion prompt data. This process is performed in the semantic encoding layer of the temporal visual language industrial large model, and includes:
[0205] S301. Based on the time-series embedding vector, extract the corresponding baseline state features, fault-sensitive fluctuation features, and life-related trend features.
[0206] In this embodiment, the baseline state characteristics are obtained by calculating the mean, variance, kurtosis, and baseline drift of the sliding window sequence of the time-series embedding vector, which is used to accurately characterize the basic health benchmark of the equipment in the fault-free stable operating range; the fault-sensitive fluctuation characteristics are obtained by extracting the local mutation amplitude, short-time energy change rate, and abnormal impulse response intensity from the time-series embedding vector, which can keenly capture non-stationary signal disturbances related to early faults such as bearing wear and gear pitting; the lifetime-related trend characteristics are obtained by performing linear regression fitting on the global sequence of the time-series embedding vector to calculate the degradation slope, cumulative change, and remaining lifetime correlation coefficient, which is used to quantitatively characterize the direction and rate of equipment performance degradation over time.
[0207] S302. Input the baseline state features, fault-sensitive fluctuation features and life-related trend features into the semantic encoding layer to generate time-series statistical description text.
[0208] For example, the semantic coding layer has a built-in predefined industrial equipment feature-text template library, which can automatically map the extracted numerical features into structured short sentences or descriptive fragments. For example, when the baseline drift is less than a preset threshold, it generates "the equipment baseline is stable"; when the short-term energy change rate exceeds the threshold, it generates "the equipment vibration fluctuation is significantly enhanced"; and when the degradation slope is greater than a preset value, it generates "the equipment performance degradation is accelerated", thus forming a complete time-series statistical description text.
[0209] S303. Based on the visual embedding vector, extract the corresponding color distribution features, texture features, spatial layout features, and density pattern features.
[0210] For example, color distribution features are obtained by calculating gray-level histograms and color moments for the red, green, and blue channels of the frequency domain pseudo-color image, respectively, to characterize the energy distribution of different frequency bands; texture features are obtained by extracting four statistics—contrast, correlation, energy, and entropy—from the gray-level co-occurrence matrix, to characterize edge, stripe, and fine-grained texture features in the frequency domain image; spatial layout features are obtained by calculating the centroid coordinates and distribution range of the spatial attention heatmap of the visual embedding vector, to characterize the positional relationship of high-energy frequency bands in the frequency domain space; and density pattern features are obtained by statistically analyzing the proportion and clustering degree of high-response pixels in the frequency domain pseudo-color image, to characterize the density distribution of fault-related frequency bands.
[0211] S304. Input the color distribution features, texture features, spatial layout features and density pattern features into the semantic encoding layer to generate visual semantic description text.
[0212] For example, the semantic coding layer generates corresponding descriptive fragments based on the physical meaning of different visual features. For instance, when the energy proportion of low-frequency channels increases significantly, it generates "low-frequency energy accumulation"; when the texture contrast exceeds the threshold, it generates "spectral texture intensification"; when the center of gravity of high-energy regions shifts to high frequencies, it generates "high-frequency components increase"; and when the proportion of high-response pixels exceeds the preset value, it generates "abnormal frequency band density increases," ultimately forming a complete visual semantic descriptive text.
[0213] S305. The time-series statistical description text, visual semantic description text, and industrial knowledge text are concatenated to generate cross-modal fusion prompt data.
[0214] For example, a fixed order of "industrial knowledge text + time-series statistical description text + visual semantic description text" is used for concatenation. This order conforms to the principle of "prior knowledge priority" in prompting engineering, which enables the model to learn domain knowledge first and then process it in combination with real-time observation data.
[0215] Through the above processing, temporal features can explicitly express their steady-state, fluctuating, and degradation meanings in text form, visual features can express spectral structure and texture changes in text form, and industrial knowledge text provides constraints on fault mechanisms and degradation patterns, thereby enabling semantic alignment of information from different sources within a unified cue space. This approach effectively reduces the expression differences between heterogeneous modalities, enhances the model's ability to identify early, weak faults and complex variable operating conditions, and significantly improves the stability and generalization ability of operational state prediction results.
[0216] The industrial equipment condition prediction method based on a temporal visual language industrial big data model provided in this application accurately quantifies the basic condition, early fault disturbances, and performance degradation trends of equipment by extracting baseline state features, fault-sensitive fluctuation features, and lifespan-related trend features. These features are then converted into temporal statistical descriptive text to achieve semantic expression of temporal information. Color distribution features, texture features, spatial layout features, and density pattern features are extracted to comprehensively characterize the energy distribution, texture, and structural information of the frequency domain image. These visual features are then converted into visual semantic descriptive text to achieve semantic expression of frequency domain visual information. The temporal statistical descriptive text, visual semantic descriptive text, and industrial knowledge text are concatenated to generate cross-modal fusion prompt data, reducing heterogeneous modal differences and improving the accuracy and generalization of model predictions.
[0217] Figure 4 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 3 ,like Figure 4 As shown, this embodiment, based on the above embodiments, provides a detailed explanation of the specific process for obtaining the fused embedding vector. This process is carried out in the gradient alignment fusion layer of the temporal visual language industrial large model, and includes:
[0218] S401. Determine the corresponding modal attention scores based on the temporal embedding vector, visual embedding vector, and text embedding vector.
[0219] In this embodiment, each embedding vector corresponds to a modal attention score. The modal attention score is used to characterize the strength of each modal embedding vector's contribution to the device state prediction task.
[0220] For example, the attention calculation module built into the gradient alignment fusion layer can generate modal attention scores based on the cosine similarity between each modal embedding vector and the temporal embedding vector, or based on the relevance score output by the learnable linear mapping layer. The modal attention score is obtained by the following formula:
[0221]
[0222] in, The modal attention score for the m-th modality of the b-th sample. This is the similarity calculation function (using cosine similarity). Let b be the temporal embedding vector of the b-th sample. The embedding vector for the m-th modality of the b-th sample (m=T,V,K correspond to the temporal, visual, and textual modalities, respectively).
[0223] S402. Normalize the attention score for each modality to obtain the attention weights corresponding to the temporal embedding vector, visual embedding vector, and text embedding vector, respectively.
[0224] In this embodiment, attention weights are used to map the strength of each modality's contribution to a normalized coefficient that can participate in the fusion operation.
[0225] For example, the normalization module built into the gradient alignment fusion layer uses the Softmax function to process the attention scores of all modalities, mapping each score to a weight value that sums to 1. This avoids a single modality from dominating the fusion process and ensures balanced fusion of the three modalities. The calculation process of the attention weights is as follows:
[0226]
[0227] in, Let m be the attention weight for the b-th sample and m-th modality. , It is a modal set.
[0228] S403. Based on each attention weight, the temporal embedding vector, visual embedding vector, and text embedding vector are weighted to obtain the corresponding weighted fusion embedding vector.
[0229] In this embodiment, the weighted fusion embedding vector is used to represent the information of each modality after weight adjustment, and each embedding vector corresponds to a weighted fusion embedding vector.
[0230] For example, the weighting module built into the gradient alignment fusion layer performs element-wise weighted multiplication on the three types of embedding vectors (temporal embedding vector, visual embedding vector, and text embedding vector). It multiplies the embedding vector of each modality with its corresponding attention weight to obtain three weighted feature vectors of the same dimension. Then, it adds these three weighted feature vectors element by element to obtain the final weighted fused embedding vector. The specific calculation formula is as follows:
[0231]
[0232] in, Let be the weighted fusion embedding vector of the b-th sample.
[0233] S404. Through the gradient alignment fusion layer, the weighted fusion embedding vector is subjected to cross-modal attention aggregation processing based on the temporal embedding vector to obtain the temporal center cross-modal representation.
[0234] In this embodiment, the time-series center cross-modal representation is used to retain the time-series mainline while introducing supplementary information from other modalities.
[0235] For example, the gradient alignment fusion layer has a built-in cross-modal aggregation module. This module uses the temporal embedding vector as the query vector and the weighted fusion embedding vector as the key vector and value vector, and performs multi-head cross-modal attention operation to selectively converge visual and textual information to form a temporal-centric cross-modal representation that emphasizes the evolution of device runtime. The specific calculation formula is as follows:
[0236]
[0237] in, For the temporal center cross-modal representation of the b-th sample, This is the time-centric cross-modal attention function.
[0238] S405. The weighted fusion embedding vector and the temporal center cross-modal representation are spliced and projected to obtain the fusion embedding vector.
[0239] For example, the gradient alignment fusion layer's built-in projection mapping module first concatenates the weighted fusion embedding vector with the temporal center cross-modal representation along the feature dimension to obtain a high-dimensional concatenated feature. Then, a fully connected layer compresses the high-dimensional concatenated feature to a preset fixed dimension, thereby obtaining a fusion embedding vector suitable for subsequent runtime classification and remaining lifetime regression tasks. The specific calculation formula for the fusion embedding vector is as follows:
[0240]
[0241] in, Let b be the fused embedding vector of the b-th sample. For mapping functions, This involves concatenating the weighted fusion embedding vector with the temporal center cross-modal representation along the feature dimension.
[0242] This fusion method can reduce the impact of intermodal distribution differences on training stability, enabling the fusion results to simultaneously possess dynamic evolution information and knowledge constraint information, improving the consistency of state recognition under complex working conditions, enhancing the separability of weak features in the early stage of faults, and improving the generalization ability under cross-device and few-sample conditions.
[0243] The industrial equipment state prediction method based on a temporal visual language industrial large-scale model provided in this application determines the modal attention score to clarify the contribution strength of each modal embedding vector to the equipment state prediction task. The modal attention scores are normalized to obtain attention weights, avoiding excessive dominance of a single modality during the fusion process. Weighted processing of each embedding vector yields a weighted fusion embedding vector, which inherits the information representation of each modality after weight adjustment. Cross-modal attention aggregation is performed based on the temporal embedding vector to obtain a temporal-centered cross-modal representation. The weighted fusion embedding vector and the temporal-centered cross-modal representation are concatenated and projected to obtain a fusion embedding vector, reducing the distribution differences between modalities and improving the stability and generalization ability of the model prediction.
[0244] Figure 5 A flowchart illustrating the industrial equipment state prediction method based on a temporal visual language industrial large model provided in this application embodiment. Figure 4 ,like Figure 5 As shown, this embodiment, based on the above embodiments, provides a detailed explanation of the specific process for obtaining the temporal embedding vector. This process is performed in the temporal coding layer of the temporal visual language industrial large model, and includes:
[0245] S501. Divide the original runtime timing signal into multiple overlapping local timing segments.
[0246] For example, when segmenting local time-series segments, the segment length and overlap step can be configured according to the frequency of equipment rotation speed fluctuations and the fault response cycle. For example, for rotating equipment, the segment length is set to cover at least one complete rotation cycle, and the overlap step is set to 50% of the segment length to ensure that sufficient continuity and contextual association are maintained between adjacent segments and to avoid local abnormal information being truncated.
[0247] S502. Through the temporal coding layer, linear projection and positional coding are performed on each local temporal segment to obtain multiple corresponding initial temporal features.
[0248] For example, linear projection uses a one-dimensional fully connected mapping layer to uniformly map local time segments of different lengths or with different numbers of channels to a preset fixed feature dimension, eliminating differences in different input formats. Position encoding uses a sine-cosine position encoding method, superimposing an encoding vector corresponding to the position of each local time segment in the original signal, preserving the sequential information of the segments, and enabling the model to perceive the temporal evolution of the time-series signal.
[0249] S503. According to the routing mechanism, assign a corresponding timing encoder to each initial timing feature to obtain the routing weight of each initial timing feature.
[0250] In this embodiment, the temporal coding layer includes multiple lightweight temporal encoders and a routing mechanism. The temporal encoders are used to extract local dynamic patterns for different local temporal segments. Internally, they may include temporal convolutional units, gating units, or self-attention units to enhance the characterization of shock fluctuations, periodic changes, and degradation trends. The routing mechanism calculates routing weights based on the cosine similarity between the initial temporal features and the predefined prototype vectors of each temporal encoder. After normalizing the weights, one or more best-matching temporal encoders are assigned to each initial temporal feature, achieving load sharing among multiple temporal encoders, reducing overload on a single encoder, and improving feature specialization capabilities. The routing weights can be obtained using the following formula:
[0251]
[0252] in, The routing weights assigned to the e-th temporal encoder for the b-th initial temporal feature. Let be the correlation score between the b-th initial temporal feature and the e-th temporal encoder. The correlation score can be obtained by the cosine similarity between the initial temporal feature and the vector. The set of time encoders selected for the b-th initial time feature, with routing weights satisfying... E represents the total number of timing encoders.
[0253] In one possible implementation, after assigning a corresponding time encoder to each initial time-series feature according to the routing mechanism and obtaining the routing weight for each initial time-series feature, the method further includes:
[0254] Based on each routing weight, determine the first total routing weight and the first sample load of each time encoder in the time coding layer; based on the first total routing weight and the first sample load, determine the corresponding first coefficient of variation; based on the first coefficient of variation, determine the corresponding first balance regularization term; based on the first balance regularization term, update the parameters of the routing mechanism in the time coding layer to obtain the updated routing mechanism.
[0255] In this embodiment, the first total routing weight characterizes the cumulative allocation strength of each temporal encoder across all initial temporal features, and the first sample load characterizes the actual number of samples carried by each temporal encoder. Together, they reflect whether the routing allocation is concentrated among a few encoders. The first coefficient of variation characterizes the degree of load imbalance among each temporal encoder; an increase in this coefficient indicates a more unbalanced routing allocation. The first balance regularization term transforms this imbalance into training constraints to suppress routing collapse and promote participation of each encoder in encoding.
[0256] For example, the temporal coding layer can be configured as a routing network containing multiple parallel temporal encoders. Each temporal encoder can be composed of a lightweight multilayer perceptron, a convolutional coding unit, or a gated attention unit, sharing or partially sharing parameters to adapt to the expression requirements of different initial temporal features. The routing mechanism calculates the corresponding routing weight based on the feature response value, similarity score, or gating probability of each initial temporal feature, and summarizes the routing weights to obtain the first total routing weight of each temporal encoder. At the same time, the number of selected initial temporal features is counted to form the first sample load. Subsequently, the coefficient of variation can be calculated based on the first total routing weight and the first sample load, and mapped to a regularization term and added to the total loss function. The routing network parameters are updated through backpropagation. The specific calculation formulas for the first total routing weight, the first sample load, the first coefficient of variation, and the first balanced regularization term are as follows:
[0257]
[0258] in, The first total routing weight of the e-th time encoder is... For the first sample load of the e-th time encoder, To select the sample set of the e-th timing encoder, The cardinality function of a set. Let be the square of the first coefficient of variation of variable q. Let be the variance of variable q. Let be the square of the mean of the variable q. This is a numerical stability constant to prevent the denominator from being zero. To balance the regularization term, For training batches.
[0259] This approach uses statistics and constraints on the routing allocation results to enable the temporal coding layer to dynamically adjust the usage frequency of each encoder during training, avoiding representation imbalance caused by a few encoders overloading features. By introducing a first balancing regularization term, the routing mechanism can maintain a relatively balanced load distribution while ensuring feature discrimination capability, thereby improving the stability and generalization ability of the temporal embedding vector.
[0260] S504. Encode each initial temporal feature using the assigned temporal encoder to obtain multiple corresponding encoded temporal features.
[0261] Each timing encoder independently encodes its assigned initial timing features, extracting the unique local dynamic features of that segment. Different timing encoders specialize in learning different types of timing patterns; for example, one encoder focuses on extracting steady-state operation features, another on extracting impact fault features, and a third on extracting progressive degradation features, thus comprehensively covering various state modes during the operation of industrial equipment. The specific calculation formula for the encoded timing features is as follows:
[0262]
[0263] in, Let be the encoded temporal features of the b-th sample at time step t. Let e be the encoding function of the e-th sequential encoder. Let be the initial temporal features of the b-th sample at the t-th time step.
[0264] S505. Based on each routing weight, perform weighted summation and global averaging on each encoded temporal feature to obtain the temporal embedding vector.
[0265] The weighted summation process multiplies each encoded temporal feature by its corresponding routing weight, and then sums all the weighted encoded temporal features to obtain a comprehensive feature that integrates the coding results of various experts. Global averaging compresses the dimensionality of the comprehensive features of all local segments, eliminating local differences between segments while preserving global temporal evolution information. The final temporal embedding vector can simultaneously retain local transient anomaly information and global degradation evolution information, providing a unified representation basis for the subsequent alignment and fusion of visual embedding vectors and text embedding vectors. The formula for calculating the temporal embedding vector is shown below:
[0266]
[0267] in, Let b be the temporal embedding vector of the b-th sample. Let L be the mapping function, and L be the time step length of the original runtime sequence signal.
[0268] It should be noted that the specific structural models of the aforementioned timing encoders can be replaced according to the deployment computing power, real-time requirements and equipment type, and this application embodiment does not limit this.
[0269] The industrial equipment status prediction method based on the temporal visual language industrial big model provided in this application adopts a combination of overlapping segmentation, location encoding, routing allocation and weighted convergence. The model can maintain good stability under varying operating conditions and with few samples, improve the representation ability of the original operating sequence signal, enhance the ability to simultaneously perceive fine-grained anomalies and long-term trends, and further improve the accuracy and generalization ability of the temporal visual language industrial big model in predicting the operating status of industrial equipment.
[0270] Figure 6 This is a schematic diagram of the industrial equipment state prediction device based on a temporal visual language industrial large-scale model provided in an embodiment of this application. The device in this embodiment can be in software and / or hardware form. Figure 6 As shown in the embodiment of this application, the industrial equipment state prediction device 600 based on a temporal visual language industrial large model includes: a first acquisition module 601 and an execution module 602.
[0271] The first acquisition module 601 is used to acquire the raw runtime sequence signals collected by the multiple sensors of the industrial equipment and the industrial knowledge text of the industrial equipment.
[0272] Execution module 602 is used to perform the following operations using a temporal visual language industrial large model:
[0273] The original runtime sequence signal is processed by time-series feature encoding to obtain a time-series embedding vector;
[0274] The original runtime sequence signal is subjected to frequency domain transformation and visual feature encoding to obtain a visual embedding vector;
[0275] Based on temporal embedding vectors, visual embedding vectors, and industry knowledge text, cross-modal fusion prompt data is generated.
[0276] Semantic encoding is performed on the cross-modal fusion prompt data to obtain text embedding vectors;
[0277] Based on the temporal embedding vector, gradient alignment and feature fusion processing are performed on the temporal embedding vector, visual embedding vector and text embedding vector to obtain the fused embedding vector.
[0278] Based on the fused embedding vector, prediction processing is performed to generate prediction results of the operating status of industrial equipment.
[0279] In one possible implementation, the temporal visual language industrial big model includes a prediction layer;
[0280] Accordingly, the execution module 602 is also used for:
[0281] The fused embedding vectors are classified through a prediction layer to determine the classification results of the operating status of industrial equipment and the corresponding confidence level. The classification results of the operating status are used to quantify the fault type and severity of industrial equipment.
[0282] By performing regression processing on the fused embedding vector through the prediction layer, the remaining service life of the industrial equipment and the corresponding degradation trend curve are obtained.
[0283] By integrating the operational status classification results, confidence level, remaining service life value, and degradation trend curve, the operational status prediction results of industrial equipment are obtained.
[0284] In one possible implementation, the execution module 602 is further configured to:
[0285] Based on the classification results of operating status, the severity level of industrial equipment failure is determined;
[0286] When the severity level of the fault is higher than the preset level threshold, and / or the remaining service life is lower than the preset service life threshold, a corresponding warning message is generated.
[0287] In one possible implementation, the temporal visual language industrial big model also includes a semantic encoding layer;
[0288] Accordingly, the execution module 602 is also used for:
[0289] Based on the temporal embedding vector, extract the corresponding baseline state features, fault-sensitive fluctuation features, and lifetime-related trend features;
[0290] Baseline state features, fault-sensitive fluctuation features, and lifetime-related trend features are input into the semantic encoding layer to generate time-series statistical description text.
[0291] Based on the visual embedding vector, extract the corresponding color distribution features, texture features, spatial layout features, and density pattern features;
[0292] Color distribution features, texture features, spatial layout features, and density pattern features are input into the semantic encoding layer to generate visual semantic descriptive text.
[0293] The time-series statistical description text, visual semantic description text, and industrial knowledge text are concatenated to generate cross-modal fusion prompt data.
[0294] In one possible implementation, the temporal visual language industrial big model also includes a gradient alignment fusion layer;
[0295] Accordingly, the execution module 602 is also used for:
[0296] Temporal embedding vectors, visual embedding vectors, and text embedding vectors are respectively input into the gradient alignment fusion layer to determine the corresponding modal attention scores. Each embedding vector corresponds to a modal attention score.
[0297] By using a gradient alignment fusion layer, the attention scores of each modality are normalized to obtain the attention weights corresponding to the temporal embedding vector, visual embedding vector, and text embedding vector, respectively.
[0298] Based on each attention weight, the temporal embedding vector, visual embedding vector, and text embedding vector are weighted through a gradient alignment fusion layer to obtain the corresponding weighted fused embedding vector.
[0299] By using a gradient alignment fusion layer, the weighted fusion embedding vector is subjected to cross-modal attention aggregation processing based on the temporal embedding vector to obtain the temporal center cross-modal representation.
[0300] By using a gradient-aligned fusion layer, the weighted fusion embedding vector and the temporal center cross-modal representation are spliced and projected to obtain the fusion embedding vector.
[0301] In one possible implementation, the temporal visual language industrial big model further includes a temporal coding layer, wherein the temporal coding layer includes multiple preset temporal encoders and a preset routing mechanism.
[0302] Accordingly, the execution module 602 is also used for:
[0303] The original runtime timing signal is divided into multiple overlapping local timing segments;
[0304] Through the temporal coding layer, linear projection and positional coding are performed on each local temporal segment to obtain multiple corresponding initial temporal features;
[0305] Based on the routing mechanism, a corresponding time encoder is assigned to each initial time series feature to obtain the routing weight of each initial time series feature;
[0306] Each initial temporal feature is encoded using the assigned temporal encoder to obtain multiple corresponding encoded temporal features;
[0307] Based on each routing weight, each encoded temporal feature is weighted and summed and then averaged globally to obtain a temporal embedding vector.
[0308] In one possible implementation, the execution module 602 is further configured to:
[0309] Based on each routing weight, determine the first total routing weight and the first sample load of each time encoder in the time coding layer;
[0310] Based on the first total routing weight and the first sample load, a corresponding first coefficient of variation is determined, wherein the first coefficient of variation is used to characterize the degree of load imbalance between each time encoder;
[0311] Based on the first coefficient of variation, determine the corresponding first balance regularization term;
[0312] Based on the first balanced regularization term, the parameters of the routing mechanism in the temporal coding layer are updated to obtain the updated routing mechanism.
[0313] In one possible implementation, the temporal visual language industrial big model also includes a visual coding layer;
[0314] Accordingly, the execution module 602 is also used for:
[0315] Through the visual coding layer, the original runtime sequence signal is processed by recursive graph transformation, Fourier transform and wavelet transform respectively to generate the corresponding recursive feature map, Fourier amplitude map and wavelet coefficient map.
[0316] By using a visual coding layer, the recursive feature map, Fourier amplitude map, and wavelet coefficient map are fused together to obtain a frequency domain pseudo-color image.
[0317] Visual features are extracted from the frequency domain pseudo-color image through a visual coding layer to obtain the corresponding visual features. The visual features are then projected to obtain the visual embedding vector.
[0318] The industrial equipment state prediction device based on the temporal visual language industrial large model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0319] Figure 7 This is a schematic diagram of the training device for a temporal visual language industrial large-scale model provided in an embodiment of this application. The device in this embodiment can be in software and / or hardware form. Figure 7 As shown in the embodiment of this application, the training device 700 for a temporal visual language industrial large model includes: a second acquisition module 701, a determination module 702, a first construction module 703, a second construction module 704, and a training module 705.
[0320] The second acquisition module 701 is used to acquire historical multi-sensor operation sequence signals of industrial equipment and historical industrial knowledge text of industrial equipment.
[0321] The determination module 702 is used to determine the corresponding frequency domain image samples based on the historical multi-sensor runtime sequence signals;
[0322] The first construction module 703 is used to construct a three-modal training dataset based on historical multi-sensor runtime sequence signals, historical industrial knowledge text, and frequency domain image samples.
[0323] The second building module 704 is used to build an initial trimodal model, wherein the initial trimodal model includes a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer and a prediction layer. The temporal coding layer includes multiple temporal encoders and a routing mechanism for assigning corresponding temporal encoders to temporal features.
[0324] Training module 705 is used to train the initial trimodal model based on the trimodal training dataset using a preset two-stage training strategy to obtain the trained temporal visual language industrial large model; the temporal visual language industrial large model is used to execute the method provided in the first aspect.
[0325] In one possible implementation, the training module 705 is also used for a two-stage training strategy, including:
[0326] In the first training phase, the parameters of the visual coding layer and the semantic coding layer are frozen, and the temporal coding layer is trained based on the trimodal training dataset.
[0327] In the second training phase, the network parameters of the temporal encoder in the temporal coding layer are frozen;
[0328] Fine-tuning the parameters of the visual encoding layer and the semantic encoding layer based on the trimodal training dataset;
[0329] Based on the gradient alignment fusion layer, determine the feature alignment loss of the temporal coding layer, visual coding layer, and semantic coding layer;
[0330] Within each training batch, the second total routing weight and second sample load of each time encoder are determined based on the routing weights assigned to each time encoder by the samples in the current batch.
[0331] The second coefficient of variation is calculated based on the second total route weight and the second sample load.
[0332] Based on the second coefficient of variation, a second balanced regularization term for the temporal coding layer is generated;
[0333] The parameters of the initial trimodal model are updated based on the feature alignment loss and the second balanced regularization term.
[0334] The training device for the temporal visual language industrial large model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0335] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 800 provided in this embodiment includes at least one processor 801 and a memory 802. Optionally, the device 800 further includes a communication component 803. The processor 801, memory 802, and communication component 803 are connected via a bus.
[0336] In a specific implementation, at least one processor 801 executes computer execution instructions stored in memory 802, causing at least one processor 801 to perform the above-described method.
[0337] The specific implementation process of processor 801 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0338] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0339] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0340] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0341] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0342] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0343] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0344] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0345] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0346] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0347] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0348] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0349] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0350] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for predicting the state of industrial equipment based on a temporal visual language industrial large-scale model, characterized in that, include: Acquire raw runtime sequence signals collected by multiple sensors of industrial equipment, as well as industrial knowledge text of the industrial equipment; The following operations are performed using a temporal visual language industrial large-scale model: The original runtime sequence signal is subjected to time-series feature encoding processing to obtain a time-series embedding vector; The original runtime sequence signal is subjected to frequency domain transformation and visual feature encoding to obtain a visual embedding vector; Based on the time-series embedding vector, extract the corresponding baseline state features, fault-sensitive fluctuation features, and lifetime-related trend features; The baseline status characteristics are used to characterize the basic health benchmark of the equipment within the fault-free and stable operating range. The fault-sensitive fluctuation feature is used to capture non-stationary signal disturbances related to early faults. The lifespan-related trend features are used to quantitatively characterize the direction and rate of equipment performance degradation over time. The baseline state features, the fault-sensitive fluctuation features, and the lifetime-related trend features are input into the semantic encoding layer to generate time-series statistical description text. Based on the visual embedding vector, the corresponding color distribution features, texture features, spatial layout features, and density pattern features are extracted; the color distribution features are used to characterize the energy distribution of different frequency bands; the texture features are used to characterize the edges, stripes, and fine-grained texture features in the frequency domain image; the spatial layout features are used to characterize the positional relationship of high-energy frequency bands in the frequency domain space; and the density pattern features are used to characterize the density distribution of fault-related frequency bands. The color distribution features, texture features, spatial layout features, and density pattern features are input into the semantic encoding layer to generate visual semantic description text; The time-series statistical description text, the visual semantic description text, and the industrial knowledge text are concatenated to generate cross-modal fusion prompt data; The cross-modal fusion prompt data is semantically encoded to obtain a text embedding vector; The temporal embedding vector, the visual embedding vector, and the text embedding vector are respectively input into the gradient alignment fusion layer of the temporal visual language industrial big data model to determine the corresponding modal attention score. Each embedding vector corresponds to a modal attention score. The gradient alignment fusion layer normalizes the attention score of each modality to obtain the attention weights corresponding to the temporal embedding vector, the visual embedding vector, and the text embedding vector, respectively. Based on each attention weight, the temporal embedding vector, the visual embedding vector, and the text embedding vector are weighted by the gradient alignment fusion layer to obtain the corresponding weighted fusion embedding vector. Through the gradient alignment fusion layer, using the temporal embedding vector as a reference, cross-modal attention aggregation processing is performed on the weighted fusion embedding vector to obtain a temporal-centered cross-modal representation. The temporal-centered cross-modal representation is used to retain the temporal main line while introducing supplementary information from other modalities. The gradient alignment fusion layer has a built-in cross-modal aggregation module. The cross-modal aggregation module uses the temporal embedding vector as the query vector and the weighted fusion embedding vector as the key vector and value vector, and performs multi-head cross-modal attention operation to selectively aggregate visual and textual information to form a temporal-centered cross-modal representation that emphasizes the evolution of device runtime. The gradient alignment fusion layer is used to concatenate and project the weighted fusion embedding vector and the temporal center cross-modal representation to obtain the fusion embedding vector. Based on the fused embedding vector, a prediction result for the operating status of the industrial equipment is generated.
2. The method according to claim 1, characterized in that, The temporal visual language industrial big model includes a prediction layer; Accordingly, generating the predicted operating status of the industrial equipment based on the fused embedding vector includes: The prediction layer classifies the fused embedding vector to determine the classification result of the operating status of the industrial equipment and the corresponding confidence level. The classification result of the operating status is used to quantify the fault type and severity of the industrial equipment. By performing regression processing on the fused embedding vector through the prediction layer, the remaining service life value of the industrial equipment and the corresponding degradation trend curve are obtained. The operating status classification results, the confidence level, the remaining service life value, and the degradation trend curve are integrated to obtain the operating status prediction results of the industrial equipment.
3. The method according to claim 2, characterized in that, After integrating the operating status classification results, the confidence level, the remaining service life, and the degradation trend curve to obtain the operating status prediction result of the industrial equipment, the method further includes: Based on the operational status classification results, the severity level of the industrial equipment fault is determined; When the severity level of the fault is higher than a preset level threshold, and / or the remaining service life is lower than a preset service life threshold, a corresponding warning message is generated.
4. The method according to claim 1, characterized in that, The temporal visual language industrial big model also includes a temporal coding layer, wherein the temporal coding layer includes multiple preset temporal encoders and a preset routing mechanism; Accordingly, the step of performing time-series feature encoding on the original runtime sequence signal to obtain a time-series embedding vector includes: The original runtime timing signal is divided into multiple overlapping local timing segments; Through the temporal coding layer, each local temporal segment is subjected to linear projection and positional coding to obtain multiple corresponding initial temporal features; According to the routing mechanism, a corresponding time encoder is assigned to each initial time series feature to obtain the routing weight of each initial time series feature; Each initial temporal feature is encoded using the assigned temporal encoder to obtain multiple corresponding encoded temporal features; Based on each routing weight, each encoded temporal feature is weighted and summed and then globally averaged to obtain the temporal embedding vector.
5. The method according to claim 4, characterized in that, After assigning a corresponding time encoder to each initial time-series feature according to the routing mechanism to obtain the routing weight of each initial time-series feature, the method further includes: Based on each of the routing weights, determine the first total routing weight and the first sample load of each of the time encoders in the time coding layer; Based on the first total routing weight and the first sample load, a corresponding first coefficient of variation is determined, wherein the first coefficient of variation is used to characterize the degree of load imbalance between each of the time encoders; Based on the first coefficient of variation, determine the corresponding first balance regularization term; Based on the first balanced regularization term, the parameters of the routing mechanism in the temporal coding layer are updated to obtain the updated routing mechanism.
6. The method according to claim 1, characterized in that, The temporal visual language industrial big model also includes a visual encoding layer; Accordingly, the step of performing frequency domain transformation and visual feature encoding on the original runtime sequence signal to obtain a visual embedding vector includes: The visual coding layer performs recursive graph transformation, Fourier transform, and wavelet transform on the original runtime sequence signal to generate corresponding recursive feature maps, Fourier amplitude maps, and wavelet coefficient maps. The recursive feature map, the Fourier amplitude map, and the wavelet coefficient map are fused through the visual coding layer to obtain a frequency domain pseudo-color image. The visual coding layer extracts visual features from the frequency domain pseudo-color image to obtain corresponding visual features, and then projects these visual features to obtain the visual embedding vector.
7. A training method for a temporal visual language industrial large-scale model, characterized in that, include: Acquire historical multi-sensor runtime sequence signals of industrial equipment and historical industrial knowledge text of the industrial equipment; Based on the historical multi-sensor runtime sequence signals, the corresponding frequency domain image samples are determined; A three-modal training dataset is constructed based on the historical multi-sensor runtime sequence signals, the historical industrial knowledge text, and the frequency domain image samples. An initial trimodal model is constructed, wherein the initial trimodal model includes a temporal coding layer, a visual coding layer, a semantic coding layer, a gradient alignment fusion layer, and a prediction layer. The temporal coding layer includes multiple temporal encoders and a routing mechanism for assigning corresponding temporal encoders to temporal features. Based on the trimodal training dataset, the initial trimodal model is trained using a preset two-stage training strategy to obtain a trained temporal visual language industrial large model; the temporal visual language industrial large model is used to perform the method described in any one of claims 1-6.
8. The method according to claim 7, characterized in that, The two-stage training strategy includes: In the first training phase, the parameters of the visual coding layer and the semantic coding layer are frozen, and the temporal coding layer is trained based on the trimodal training dataset. In the second training phase, the network parameters of the temporal encoder in the temporal coding layer are frozen; Fine-tune the parameters of the visual encoding layer and the semantic encoding layer based on the trimodal training dataset; Based on the gradient alignment fusion layer, the feature alignment loss of the temporal coding layer, visual coding layer, and semantic coding layer is determined; Within each training batch, based on the routing weights assigned to each of the time encoders by the samples in the current batch, the second total routing weight and the second sample load of each time encoder are determined; The second coefficient of variation is calculated based on the second total routing weight and the second sample load; Based on the second coefficient of variation, a second balanced regularization term is generated for the temporal coding layer; The parameters of the initial trimodal model are updated based on the feature alignment loss and the second balanced regularization term.
Citation Information
Patent Citations
Optical fiber sound wave event identification method, equipment and medium
CN120724138A
Bilingual harmful model factor identification method and bilingual harmful model factor identification system based on multiple modes and multiple views
CN121456389A
Multi-modal monitoring method and device for helicopter engine fault diagnosis
CN121705590A
Structural health monitoring abnormal data diagnosis method based on multi-task hybrid expert vision Transform
CN122023893A