Industrial data fusion and interoperation processing method and device, equipment and storage medium

By using multi-scale time axis reconstruction, cross-modal coding, and graph neural network processing to process industrial data, the problem of low efficiency in the fusion and interoperability of multi-source heterogeneous data has been solved, and efficient unified governance and intelligent analysis of industrial data have been achieved.

CN121997258APending Publication Date: 2026-05-08CHINA ACADEMY OF INFORMATION & COMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ACADEMY OF INFORMATION & COMM
Filing Date
2026-01-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies are inefficient and time-consuming when processing industrial data, making it difficult to achieve unified governance, intelligent analysis, and cross-system collaboration of multi-source heterogeneous data. In particular, they face difficulties in multimodal data fusion and protocol interoperability, resulting in high processing costs, low efficiency, and a high susceptibility to errors.

Method used

A multi-scale time axis reconstruction mechanism is used to preprocess multi-source industrial data. Features are extracted and mapped to a unified semantic space through a cross-modal coding framework. Automatic annotation is performed using a student model. An industrial association graph is constructed by combining a graph neural network. Protocol parsing is performed through a semantic coding model to achieve data fusion and interoperability.

Benefits of technology

It significantly improves the semantic understanding and fusion depth of industrial data, establishes dynamic interoperability across protocols and systems, solves the problems of low efficiency and long cycle when processing industrial data in existing technologies, and realizes seamless collaboration and system interconnection of multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997258A_ABST
    Figure CN121997258A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial data fusion and interoperation processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring multi-source industrial data from an industrial field environment; preprocessing the multi-source industrial data in combination with a multi-scale time axis reconstruction mechanism to obtain a standardized input data set; performing feature extraction and feature mapping on the standardized input data set through a cross-modal coding framework to obtain a multi-modal sample data set; labeling samples in the multi-modal sample data set based on a student model to obtain a multi-modal labeled data set; constructing an industrial association graph based on a graph neural network in combination with the multi-modal annotation data set, and processing the industrial association graph to obtain a unified state representation after fusion; semantic analysis is carried out on the industrial protocol through the semantic coding model, and cross-protocol mapping is carried out in combination with the industrial association diagram to obtain an interoperation result; and providing the fused unified state representation and interoperation result to an upper layer application through the micro-service architecture. The method improves the data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of industrial internet technology, and in particular to an industrial data fusion and interoperability processing method, apparatus, equipment and storage medium. Background Technology

[0002] With the continuous improvement of digitalization and intelligence in the manufacturing industry, the scale of data generated by industrial enterprises in various business processes such as production operation, process control, quality management, equipment maintenance, and supply chain collaboration continues to rise. The data types have also expanded from traditional structured field measurement data to multimodal forms such as images, videos, time-series signals, alarm logs, text records, and protocol messages. Industrial data exhibits typical characteristics in terms of quantity, speed, diversity, and value density, and its role in process optimization, condition diagnosis, predictive maintenance, quality control, and even production decision-making is becoming increasingly prominent. However, unlike the relatively standardized and format-consistent data in the internet field, industrial data is deeply coupled with equipment, processes, workshops, scenarios, and the enterprise's own logic. Its inherent heterogeneity and high specialization bring significant challenges to integrated processing and interoperability.

[0003] In real-world applications, industrial data is often diverse in quantity and type. Therefore, existing technologies often require a significant amount of manpower to reorganize and interpret industrial data, resulting in low processing efficiency and long processing cycles. Summary of the Invention

[0004] This invention provides an industrial data fusion and interoperability processing method, apparatus, equipment, and storage medium to solve the problems of low efficiency and long cycle time in the prior art when processing industrial data.

[0005] According to one aspect of the present invention, an industrial data fusion and interoperability processing method is provided, the method comprising: Acquire multi-source industrial data from industrial field environments; The multi-source industrial data is preprocessed using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; By extracting features from the standardized input dataset using a cross-modal coding framework, features from different modalities are mapped to a unified semantic space, resulting in a multimodal sample dataset with a unified representation. The samples in the multimodal sample dataset are automatically labeled based on the trained student model to obtain a multimodal labeled dataset; the student model is a model that performs knowledge transfer based on the teacher model; An industrial association graph is constructed based on a graph neural network combined with the multimodal labeled dataset. The industrial association graph is then processed to obtain a fused unified state representation. The semantic encoding model is used to perform semantic parsing of industrial protocols, and cross-protocol mapping is performed in combination with the industrial association graph to obtain interoperability results; The unified state representation and interoperability results are provided to upper-layer applications through a microservice architecture.

[0006] According to another aspect of the present invention, an industrial data fusion and interoperability processing apparatus is provided, the apparatus comprising: The acquisition module is used to acquire multi-source industrial data from the industrial field environment; The processing module is used to preprocess the multi-source industrial data by combining a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; The extraction module is used to extract features from the standardized input dataset through a cross-modal coding framework, mapping features from different modalities to a unified semantic space to obtain a multimodal sample dataset with a unified representation. The annotation module is used to automatically annotate the samples in the multimodal sample dataset based on the trained student model, so as to obtain a multimodal annotated dataset; the student model is a model that performs knowledge transfer based on the teacher model; The construction module is used to construct an industrial association graph based on a graph neural network combined with the multimodal labeled dataset, and to process the industrial association graph to obtain a fused unified state representation; The parsing module is used to perform semantic parsing of industrial protocols through a semantic coding model, and to perform cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results. A module is provided to provide the integrated unified state representation and the interoperability results to upper-layer applications through a microservice architecture.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the industrial data fusion and interoperability processing method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the industrial data fusion and interoperability processing method according to any embodiment of the present invention.

[0009] This invention discloses an industrial data fusion and interoperability processing method, apparatus, device, and storage medium. The method includes: acquiring multi-source industrial data from an industrial environment; preprocessing the multi-source industrial data using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; extracting features from the standardized input dataset using a cross-modal coding framework, mapping features of different modalities to a unified semantic space to obtain a unified representation of a multimodal sample dataset; automatically labeling samples in the multimodal sample dataset based on a trained student model to obtain a multimodal labeled dataset; wherein the student model is a model that performs knowledge transfer based on a teacher model; constructing an industrial association graph based on a graph neural network and the multimodal labeled dataset, processing the industrial association graph to obtain a fused unified state representation; performing semantic parsing of industrial protocols using a semantic coding model, performing cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results; and providing the fused unified state representation and the interoperability results to upper-layer applications using a microservice architecture. This method, taking into account the characteristics of "multi-source, multi-modal, strongly correlated, and highly specialized" industrial sites, constructs a unified technical framework that can operate in a closed loop, integrating data access, semantic extraction, knowledge transfer, intelligent annotation, graph structure fusion, protocol semantic alignment, and reinforcement learning-driven collaborative optimization. This not only significantly improves the semantic understanding and fusion depth of industrial data, but also establishes dynamic interoperability capabilities across protocols and systems, solving the problems of low efficiency and long cycle time in processing industrial data in existing technologies.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating an industrial data fusion and interoperability processing method provided in Embodiment 1 of the present invention; Figure 2 A flowchart illustrating an industrial data fusion and interoperability processing method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an industrial data fusion and interoperability processing device provided in Embodiment 2 of the present invention; Figure 4This is a schematic diagram of the electronic device used in the industrial data fusion and interoperability processing method according to an embodiment of the present invention. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. It should be understood that the various steps described in the method embodiments of the present invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, any variations of the terms "comprising" and "having," etc., are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0018] Over long-term practice, industrial enterprises have developed numerous localized data models based on systems such as Programmable Logic Controllers (PLCs), Distributed Control Systems (DCS), Supervisory Control and Data Acquisition (SCADA), and Manufacturing Execution Systems (MES). Between different equipment manufacturers, control architectures, and production lines, there are significant differences not only in variable naming, data formats, sampling periods, and semantic definitions, but even within the same production line, due to historical evolution and upgrades, multiple communication protocols, data standards, and acquisition methods may coexist. This state of "multi-source heterogeneity and semantic fragmentation" often requires enterprises to invest significant manpower in reorganizing, aligning, and interpreting data when conducting unified data governance, intelligent analysis, or cross-system collaboration. This process is time-consuming, costly, and prone to errors. While existing technologies can achieve interoperability between different systems through pre-configured rule bases or mapping tables manually maintained by engineers, this approach lacks a unified semantic foundation, cannot withstand frequently changing operating conditions, and is difficult to adapt to the dynamic generation and real-time processing of massive amounts of data in complex industrial scenarios.

[0019] In terms of data annotation and quality processing, the difficulties are even more pronounced in industrial scenarios. Taking visual inspection as an example, the types of defects generated in different process stages vary significantly in morphology, texture, and scale. Most defects are highly domain-specific and lack universally recognizable features. Manual work is not only costly and inefficient, but also makes it difficult to guarantee the consistency of annotation. Time-series data and log data rely more heavily on the experience and judgment of professional engineers, and the annotation process is often time-consuming and significantly affected by subjective factors. Although many enterprises have accumulated massive amounts of data, the lack of annotation systems and the excessively high cost of annotation mean that the data cannot be directly used for model training or intelligent algorithms in most cases. The lack of high-quality industrial data has also become one of the bottlenecks restricting the development of intelligent manufacturing.

[0020] Meanwhile, there are significant technological gaps in the fusion and analysis of multimodal industrial data. Traditional methods primarily focus on single-modal processing, typically training models separately for image, time-series, and text data, and then achieving coarse-grained fusion through simple post-processing. These methods fail to truly understand the relationships between different modes and struggle to capture the dynamic mapping between different physical quantities in actual production processes. For example, in complex manufacturing scenarios, changes in vibration signals are often potentially correlated with image defects, process parameter fluctuations, or equipment log events. Only by modeling within a highly unified semantic space can reliable causal analysis, predictive inference, and collaborative decision-making be achieved. However, current technologies lack a universal mechanism for the unified expression and effective fusion of multimodal characteristics in industry.

[0021] The situation is equally complex regarding interoperability of industrial protocols. Enterprises commonly employ various industrial communication protocols, such as Open Platform Communications Unified Architecture (OPC UA), Modbus, Ethernet Industrial Protocol (EtherNet / IP), and Process Field Network (Profinet). These protocols differ fundamentally in semantic structure, encoding methods, and namespace organization, making direct data exchange and sharing difficult. Existing interoperability methods primarily rely on manually constructed mapping rules, requiring engineers to parse, organize, and configure each protocol description item by item—a cumbersome process heavily dependent on personal experience. Once equipment is upgraded, processes are adjusted, or protocols change, these mapping rules must be re-maintained, resulting in poor system scalability, high maintenance costs, and insufficient flexibility, making it difficult to support rapidly iterating smart manufacturing scenarios. Due to the lack of technological means to automatically learn protocol semantics and intelligently establish cross-protocol mapping relationships, data exchange and collaboration between industrial systems still face significant obstacles.

[0022] In recent years, large-scale modeling technology has achieved significant breakthroughs in natural language processing, visual understanding, and cross-modal representation, bringing new possibilities to industrial data processing. Large-scale models possess powerful language understanding, knowledge abstraction, and semantic reasoning capabilities, and can extend general semantic capabilities to specialized industrial domains through knowledge transfer and domain adaptation. However, the direct implementation of large-scale models in industrial scenarios still faces both technical and engineering challenges. On the one hand, data samples in the industrial field are scarce and knowledge is proprietary, making simple scaling insufficient for satisfactory results. On the other hand, industrial environments demand extremely high real-time performance and stability; the complex structure of large-scale models involves large computational demands during the inference phase, making low-latency deployment at edge computing or production lines difficult. Furthermore, while large-scale models excel at text and visual processing, they lack natural advantages in handling time-series signals, protocol structures, control variables, and industrial process logic, requiring in-depth modeling tailored to specific real-world scenarios.

[0023] As intelligent manufacturing levels continue to improve, enterprises are increasingly demanding cross-modal data fusion, unified semantic expression, automatic protocol mapping, and efficient system interoperability. Especially in typical applications such as state prediction, fault diagnosis, autonomous optimization, and flexible scheduling, only when different types of data can seamlessly collaborate and different devices and systems can automatically communicate can future-oriented adaptive industrial intelligent systems be built. Currently prevalent methods relying on manual configuration, rule-driven approaches, or lightweight model-driven approaches can no longer meet the real-time, accuracy, and scalability requirements of large-scale industrial scenarios. There is an urgent need for a new technological system capable of uniformly processing multi-source data, automatically performing semantic reasoning, adapting to different operating conditions, and possessing sustainable learning capabilities.

[0024] To address the challenges of complex data types, inconsistent semantics, diverse protocols, high annotation costs, difficulties in multimodal fusion, and low interoperability efficiency in the industrial sector, this invention provides a method that leverages the knowledge transfer capabilities of large models to construct a unified semantic space. This enables intelligent data annotation, semantic understanding, multi-source fusion, and automatic protocol interoperability. The method is characterized by its generalizability, scalability, engineering deployability, and sustainable optimization. It can meet the real-time intelligent processing needs of intelligent manufacturing environments, driving the industrial system from data fragmentation to data collaboration, from isolated equipment to system interconnection, and from local optimization to global intelligence.

[0025] Example 1 Figure 1This is a flowchart illustrating an industrial data fusion and interoperability processing method provided in Embodiment 1 of the present invention. This method is applicable to processing industrial data acquired from an industrial field environment. The method can be executed by an industrial data fusion and interoperability processing device, which can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes, but is not limited to, devices such as computers.

[0026] like Figure 1 As shown in Embodiment 1 of the present invention, an industrial data fusion and interoperability processing method includes the following steps: S110. Acquire multi-source industrial data from the industrial field environment.

[0027] Industrial site environment refers to the site where industrial operations are carried out. Multi-source industrial data refers to various types of data acquired from different data collection objects, different data collection devices, and different data types within the industrial site environment. Multi-source industrial data can include image data, time-series data, text and log data, and protocol message data collected from production lines, equipment control systems, sensor networks, vision inspection systems, and business systems.

[0028] In this embodiment, multi-source industrial data can be acquired from the industrial site environment. For example, considering the high heterogeneity of multi-source data in industrial sites, a unified data acquisition adaptation layer can be constructed. This layer supports synchronous or asynchronous data retrieval from MES, SCADA, PLC controllers, DCS systems, industrial cameras, vibration sensors, temperature, pressure and flow multi-physical quantity sensors, work order systems, log servers, and device protocol endpoints. When necessary, an edge gateway can be introduced to achieve lightweight data normalization.

[0029] S120. The multi-source industrial data is preprocessed using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset.

[0030] Multi-scale timeline reconstruction mechanisms can be techniques for aligning various types of data into a standardized timeline system with a unified, multi-granularity time scale. Preprocessing can include steps such as format conversion, time alignment, missing data filling, noise suppression, and metadata binding.

[0031] In this embodiment, a multi-scale time axis reconstruction mechanism can be used to preprocess multi-source industrial data to form a standardized input dataset that can be processed by the model.

[0032] In one embodiment, the multi-source industrial data includes at least high-frequency data, low-frequency data, event-type data, and image-type data. Correspondingly, the preprocessing of the multi-source industrial data using a multi-scale timeline reconstruction mechanism to obtain a standardized input dataset includes: performing window aggregation and low-pass filtering on the high-frequency data to obtain multi-scale time slices; recovering missing segments in the low-frequency data through interpolation, migration interpolation, and morphological smoothing to obtain recovered low-frequency data; aligning the low-frequency data with the multi-scale time slices to obtain aligned low-frequency data; performing timestamp normalization and context binding on the event-type data, inserting the event-type data into the multi-scale time slices to obtain embedded event-type data; generating association indexes for the image-type data corresponding to the operating conditions and variable states at the corresponding moments in the multi-scale time slices to obtain image-type data with association indexes; and using the multi-scale time slices, aligned low-frequency data, embedded event-type data, and image-type data with association indexes as data in the standardized input dataset.

[0033] High-frequency data refers to time-series data that is continuously collected at high frequencies, including the operating status of equipment and physical quantities of core processes in industrial settings. The collection frequency can be greater than 1 Hz. Low-frequency data refers to data that is collected or recorded at low frequencies, including process parameters, batch information, and equipment ledgers in industrial production. The collection / recording frequency is typically less than once per minute. Event-based data refers to discretely triggered, non-continuous data with clear timestamps and business semantics during industrial production. Image-based data refers to visual data collected intermittently or triggered from workpieces, equipment, and workstation environments in industrial settings using industrial cameras, vision sensors, and other devices. Window aggregation refers to dividing continuous high-frequency time-series data into preset time windows and performing statistical aggregation on the raw data within each window. Low-pass filtering is a technique that allows low-frequency components in a signal to pass through while attenuating or blocking high-frequency components, thereby filtering out high-frequency noise and preserving the true trend of the data. Multi-scale time slices can be preset multi-level time-granularity data blocks. Timestamp normalization is the process of converting timestamps from different sources and formats into a standardized form. Context binding can be a processing method that binds the context information such as the business scenario, environment, and associated identifiers of the data to the target data, so that the data carries complete background information.

[0034] In this embodiment, different processing methods can be used for high-frequency data, low-frequency data, event-type data, and image-type data. For high-frequency data, window aggregation and low-pass filtering are performed to obtain multi-scale time slices. For low-frequency data, time alignment can be performed between the low-frequency data and the multi-scale time slices to obtain aligned low-frequency data. For event-type data, timestamp normalization and context binding can be performed to insert the event-type data into the multi-scale time slices to obtain embedded event-type data. For image-type data, an association index can be generated with the working conditions and variable states at the corresponding time in the multi-scale time slices to obtain image-type data with association indexes. Finally, the multi-scale time slices, aligned low-frequency data, embedded event-type data, and image-type data with association indexes are used as data in the standardized input dataset.

[0035] For example, by introducing a multi-scale time axis reconstruction mechanism, time slice fragmentation caused by differences in sampling frequencies between different systems can be eliminated. For high-frequency data (such as vibration signals), window aggregation and low-pass filtering can be used to obtain stable multi-scale time slices. For low-frequency data (such as process parameters and machine status), interpolation, migration interpolation, morphological smoothing and other methods can be used to recover missing segments. For event-type data (alarms, work orders), timestamp normalization and context binding can be performed to accurately insert them into the time series. For image-type data, an association index with the corresponding working conditions and variable status at the time can be generated. Finally, a standardized multimodal input dataset with temporal consistency, modal homogeneity and context integrity is formed, thus laying a unified data foundation for subsequent cross-modal feature modeling.

[0036] S130. Features are extracted from the standardized input dataset using a cross-modal coding framework, and features from different modalities are mapped to a unified semantic space to obtain a multimodal sample dataset with a unified representation.

[0037] The cross-modal coding framework can be a hierarchical feature extraction and semantic mapping system for heterogeneous modal data such as text, images, and time series data in industrial scenarios. The semantic space can be a high-dimensional vector space constructed based on industry-specific semantic rules.

[0038] In this embodiment, a cross-modal coding framework can be used to extract features from a standardized input dataset and map features from different modalities to a unified semantic space, resulting in a multimodal sample dataset with unified representation. For example, a cross-modal coding framework can be constructed based on pre-trained large-scale language and visual models. For text data, an improved Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model can be used for encoding. For image data, visual features can be extracted using Residual Neural Network (ResNet) or Vision Transformer (ViT) and other visual networks. For time-series data, Bidirectional Long Short-Term Memory (BiLSTM) or Transformer encoding can be used. Subsequently, a self-attention mechanism and a feature alignment network can be used to map features from different modalities to a unified high-dimensional semantic space.

[0039] In one embodiment, the cross-modal coding framework includes a text coding sub-model, an image coding sub-model, a temporal coding sub-model, and a cross-modal alignment network, wherein the data types in the standardized input dataset include text type, image type, and temporal type.

[0040] The text encoding sub-model can be a semantic feature extraction module designed for text data specific to the industrial field. The image encoding sub-model can be a spatial feature extraction module for industrial visual data. The temporal encoding sub-model can be a temporal feature extraction module designed for industrial time-series data. The cross-modal alignment network can be the core linkage module of the cross-modal encoding framework.

[0041] In this embodiment, the cross-modal coding framework can be composed of a text coding sub-model, an image coding sub-model, a temporal coding sub-model, and a cross-modal alignment network. The data types in the standardized input dataset can include text, image, and temporal types.

[0042] Furthermore, the step of extracting features from the standardized input dataset using a cross-modal coding framework, mapping features from different modalities to a unified semantic space to obtain a multimodal sample dataset with a unified representation, includes: extracting semantic features from the text type data using the text coding sub-model to obtain text features; extracting spatial features from the image type data using the image coding sub-model; obtaining temporal features from the temporal type data using the temporal coding sub-model; and mapping the text features, spatial features, and temporal features to a unified semantic space using the cross-modal alignment network to obtain a multimodal sample dataset with a unified representation. The text coding sub-model is a large language model trained twice on industrial corpora; the image coding sub-model uses a fusion structure of deep convolutional networks and visual Transformers; and the temporal coding sub-model uses a combination structure of bidirectional long short-term memory networks, temporal Transformers, and frequency domain convolutions.

[0043] In this embodiment, text features can be obtained by semantic extraction of text-type data through a text encoding sub-model, spatial features of image-type data can be extracted through an image encoding sub-model, and temporal features of temporal-type data can be obtained through a temporal encoding sub-model. Finally, text features, spatial features, and temporal features are mapped to a unified semantic space through a cross-modal alignment network to obtain a unified representation of a multimodal sample dataset. This enables a unified representation of industrial multimodal data and provides a solid semantic foundation for subsequent knowledge transfer, semantic reasoning, graph structure fusion, and interoperability mapping.

[0044] For example, in the feature extraction stage, a multimodal representation system is constructed, consisting of a text encoding sub-model, an image encoding sub-model, a temporal encoding sub-model, and a cross-modal alignment network. The text encoding sub-model can employ a Large Language Model (LLM) or BERT, pre-trained on industrial corpora (such as equipment instructions, maintenance manuals, alarm dictionaries, and protocol documents), to semantically extract variable meanings, process terms, and protocol field descriptions, obtaining embedded representations with industrial semantics as text features. The image encoding sub-model can use a fusion structure of a deep convolutional network (ResNet) and a visual Transformer (ViT) to extract spatial features such as workpiece appearance, equipment movement, and defect morphology, and enhance alignment with process parameters and equipment status through an attention mechanism. The temporal encoding sub-model can use a combination structure of bidirectional LSTM, temporal Transformer, and frequency domain convolution to acquire short-term dynamic patterns, long-term trends, and potential abnormal patterns from sensor data such as vibration, pressure, temperature, and current, serving as temporal features. Cross-modal alignment networks can map modal features such as images, text, and time series to a unified high-dimensional semantic space through self-attention mechanisms, gated fusion units (GFU), and modal co-training loss, enabling different modalities to achieve comparability and composability around the same process semantics.

[0045] S140. Based on the trained student model, the samples in the multimodal sample dataset are automatically labeled to obtain a multimodal labeled dataset; the student model is a model that performs knowledge transfer based on the teacher model.

[0046] The student model can be a model that transfers knowledge from the teacher model. Specifically, a general large model can be used as the teacher model, and corpora from fields such as industrial equipment documents, process specifications, historical work orders, protocol descriptions, and alarm databases can be introduced to fine-tune the large model and supplement its knowledge to obtain the teacher model. At the same time, a lightweight student model is constructed, which learns the representation distribution and decision boundary of the teacher model in the industrial scenario through knowledge distillation, and then deploys the student model in edge or field environments to achieve low-latency inference for industrial scenarios.

[0047] In this embodiment, a trained student model can be used to automatically label samples in a multimodal sample dataset to obtain a multimodal labeled dataset. For example, a student model that has completed knowledge transfer can be used to generate candidate labels and label confidence scores for unlabeled or weakly labeled image, time series, and text data in the multimodal sample dataset. A closed-loop labeling mechanism is constructed by combining contrastive learning, confidence filtering, and manual review. At the same time, a quality assessment network is used to automatically detect and correct the labeling results, forming a high-quality multimodal labeled dataset.

[0048] In the knowledge transfer and distillation stage, after the teacher model is pre-trained on a large-scale general corpus, it can be further trained using an industry-specific corpus. This enables the teacher model to recognize industrial protocol terms, equipment models, and the meaning of process parameters. The student model can employ a smaller parameter-scale network structure, deployed at the edge, with a parameter scale of thousands to millions. Knowledge distillation is performed using the teacher model's output as soft labels. Methods such as Kullback-Leibler Divergence (KL) divergence loss and mutual information preservation loss are used to maximize the preservation of industrial semantic capabilities. On the same batch of industrial tasks or samples, the teacher model's output is used as soft labels. Knowledge distillation is completed by minimizing the differences in output distribution, thereby maintaining expressive power and inference performance close to that of the teacher model while significantly reducing computational overhead.

[0049] By systematically introducing large-scale model knowledge transfer into industrial data fusion and interoperability scenarios, a complete technical loop is formed, encompassing a general semantic space, industrial-specific semantics, and lightweight on-site reasoning capabilities. Through industrial-domain fine-tuning of the general large-scale model and knowledge distillation, semantic capabilities are pushed down to edge-side small models. This enables the system to simultaneously understand the business meaning of multimodal data such as images, time series, text, and protocol messages, supporting unified expression and deep fusion across modalities, devices, and systems. Compared to traditional solutions relying on single models or manual rules, this embodiment significantly improves the semantic understanding depth and fusion accuracy of industrial data, providing a high-quality data and model foundation for subsequent state recognition, fault diagnosis, and process optimization.

[0050] In one embodiment, the automatic labeling of samples in the multimodal sample dataset based on a trained student model to obtain a multimodal labeled dataset includes: determining, based on the trained student model, the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector for each sample in the multimodal sample dataset; and, based on the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector for each sample, selecting samples in the multimodal sample dataset that meet a first preset condition as data in the multimodal labeled dataset; wherein, the first preset condition is that the confidence level of a sample is in the first confidence interval and passes the cross-modal consistency check, or the confidence level of a sample is in the second confidence interval and passes manual review, wherein the confidence level in the first confidence interval is higher than the confidence level in the second confidence interval.

[0051] Here, the candidate label set refers to the set of candidate labels for the samples predicted by the student model, and the label confidence distribution refers to the distribution of confidence scores for samples belonging to different labels. The multimodal consistency score indicates whether the outputs of different modalities are consistent when processing the same task in a multimodal manner. The semantic similarity vector indicates the degree of similarity between a sample and other samples. The first and second confidence intervals can be set according to actual conditions; this embodiment does not impose any limitations on them.

[0052] In this embodiment, the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector of each sample in the multimodal sample dataset can be determined based on the trained student model. Based on the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector of each sample, the samples in the multimodal sample dataset that meet the first preset condition can be used as data in the multimodal annotation dataset.

[0053] In the intelligent annotation stage, this embodiment utilizes a lightweight student model that has completed knowledge transfer to automatically generate labels and assess the quality of a large number of unlabeled or weakly labeled multimodal samples (including images, time-series signals, alarm texts, log fragments, protocol fields, etc.) in industrial scenarios. Compared with traditional methods that rely on manual labeling or weakly supervised rules, the closed-loop annotation mechanism constructed in this embodiment, which combines "large model inference-driven + confidence-level hierarchical management + rapid expert review + incremental continuous learning," can significantly reduce annotation costs and improve annotation consistency.

[0054] For example, a multimodal sample dataset is input into a student model that has undergone knowledge transfer. By inheriting the industrial semantic capabilities of the teacher model, the student model can perform multi-task reasoning on image defect morphology, abnormal patterns of time-series signals, equipment status, meaning of alarm text, semantics of protocol fields, etc., and output: Candidate Label Set, Label Confidence Distribution, Cross-Modal Consistency Score, and Semantic Similarity Vector. These outputs can be used to drive subsequent hierarchical processing and annotation quality control.

[0055] To ensure labeling accuracy, the labels generated by the model can be divided into three levels based on confidence level, thereby improving the model's ability to identify samples under complex working conditions. High-Confidence Zone: The first confidence interval is greater than or equal to Th_H. The confidence threshold Th_H can be customized, usually 0.8. Simultaneously, the sample needs to pass a cross-modal consistency check (e.g., visual defects are consistent with corresponding temporal vibration patterns). If the sample passes this check, it can be directly added to the training set as a "pseudo-label" and written to the annotation database to form a traceable record. This mechanism can cover more than 80% of repetitive and easily identifiable samples, significantly reducing manual intervention.

[0056] Medium-Confidence Zone: Confidence levels fall between the first confidence interval (Th_L) and Th_H, where Th_L can be 0.2 and Th_H can be 0.8. Samples belonging to the first confidence interval have some interpretability but exhibit ambiguity. These samples can be pushed to the "Expert Quick Review Interface," where the following data will be presented through visualization tools: model prediction labels and confidence levels, original samples (images / signals / text), model attention weight heatmap, relevant process context information, and historical similar sample comparison results. Experts can perform "one-click confirmation," "quick correction," or "supplementary labels." The system will record these comments and use them for subsequent model optimization.

[0057] Low-Confidence Zone: These samples lack sufficient confidence and typically correspond to complex defects, rare operating conditions, or sample noise. These samples will automatically enter the Hard Case Pool, which can be used for incremental learning of the model, prompt engineering optimization, hard sample clustering analysis, data augmentation strategy design, few-shot transfer learning, etc.

[0058] In one embodiment, the cross-modal consistency check includes one or more of the following checks: whether image defects correspond to time-series vibration anomalies, whether alarm text matches events in the device log, whether the change trend of protocol fields is consistent with the process status, and whether the process stage label corresponds to the production batch progress.

[0059] In this embodiment, cross-modal label consistency check can improve the reliability of the label. The cross-modal consistency check includes whether the image defects correspond to the timing vibration anomalies, whether the alarm text matches the events in the equipment log, whether the change trend of the protocol field is consistent with the process status, and whether the process stage label corresponds to the production batch progress. If semantic conflicts or modal inconsistencies occur, the sample may be problematic. The sample is automatically marked as a sample to be checked and enters the expert approval or difficult sample pool.

[0060] After labeling is completed, label quality can be automatically improved through the following methods: logical rule constraint verification, such as setting constraints like "temperature must increase during the heating stage"; topological consistency verification, such as ensuring that defect types in upstream and downstream processes are not contradictory; temporal smoothing strategies to avoid incorrect labeling caused by single-point noise; cluster-driven label correction, ensuring that labels in the same cluster are consistent; and a multi-model voting mechanism, with teacher model vs. student model vs. incremental model voting verification. If label deviation is automatically detected, an error correction process can be triggered.

[0061] The following data can be continuously used for incremental training: pseudo-labeled samples, medium-confidence samples confirmed by experts, key samples from the hard sample pool, and new working condition samples exhibiting label drift. Label drift refers to labels whose label distribution has changed significantly.

[0062] This embodiment enables the student model to continuously adapt to new equipment conditions, process changes, and data distribution changes by periodically fine-tuning it, thus achieving sustainable evolution.

[0063] The above approach forms a closed-loop system of "model annotation - expert review - automatic verification - incremental learning - model improvement", which can improve annotation efficiency by 5-50 times (depending on the scenario), reduce the proportion of manual intervention by more than 80%, significantly enhance label consistency and the model's ability to identify rare samples and long-tail scenarios, thereby solving the long-standing problems of labeling difficulties, high costs, and inconsistent quality in industrial scenarios. This makes multimodal data available and sustainable, providing a highly reliable semantic foundation for subsequent fusion, inference, and interoperability.

[0064] S150. Construct an industrial association graph based on a graph neural network and the multimodal labeled dataset, and process the industrial association graph to obtain a fused unified state representation.

[0065] Graph neural networks can be a type of deep learning model specifically designed for processing graph-structured data. Industrial relationship graphs can be the core graph-structured data that carries the physical topology, business logic, and multimodal data relationships of industrial systems.

[0066] In this embodiment, a multimodal labeled dataset can be processed based on a graph neural network to construct an industrial association graph. This graph can then be further processed to obtain a fused unified state representation. For example, an industrial association graph of multi-source industrial data can be constructed based on a graph neural network. Equipment, sensors, process sections, variables, and events in the industrial field are considered nodes in the graph, while physical connections, process dependencies, temporal correlations, and semantic relationships are considered edges. A multi-layer graph attention network is used to aggregate and fuse cross-modal features, and a joint loss function is used to balance supervised and unsupervised structural signals to obtain the fused unified state representation.

[0067] In one embodiment, the step of constructing an industrial association graph based on a graph neural network and the multimodal labeled dataset, and processing the industrial association graph to obtain a fused unified state representation, includes: constructing an industrial association graph based on the multimodal labeled dataset, wherein the graph nodes of the industrial association graph are field equipment, sensor nodes, workstation processes, process parameters, quality inspection results, and log events, and the graph edges are physical connection relationships, control logic dependencies, temporal sequence relationships, semantic relevance, and dynamic relevance of working conditions; aggregating cross-modal features of the industrial association graph through a multi-layer graph attention network to obtain aggregated features; and optimizing the aggregated features through a joint loss function to obtain a fused unified state representation.

[0068] In this context, "field equipment" refers to equipment deployed in industrial production sites, while "sensor nodes" refer to the smallest data acquisition unit composed of sensors, data acquisition modules, and communication modules, deployed in field equipment or the production environment. "Control logic dependency" refers to the dependency relationship between instructions and states formed between equipment, processes, and control units in an industrial system based on preset control programs or business rules. "Cross-modal feature aggregation" refers to the process of mapping features from different modalities such as text, images, and time series to a unified semantic space for alignment, and then integrating complementary features from each modality through weighted combination, attention filtering, tensor operations, etc., to generate a comprehensive feature representation that incorporates multimodal semantic information.

[0069] In this embodiment, an industrial association graph can be constructed based on a multimodal labeled dataset. A multi-layer graph attention network is used to aggregate cross-modal features from the industrial association graph, resulting in aggregated features. A joint loss function is then used to optimize the aggregated features, yielding a unified fused state representation. For example, in the multi-source fusion stage, an initial industrial association graph can be constructed based on equipment topology, production process flow, operating condition evolution trajectory, and historical statistical relationships. Field equipment, sensor nodes, workstation processes, process parameters, quality inspection results, log events, and other unified regions are abstracted as graph nodes. Physical connections (such as pipelines and energy flow), control logic dependencies (such as PLC trigger chains), temporal sequences (such as process A → process B), semantic relevance (such as vibration characteristics corresponding to defect types), and dynamic operating condition relevance are mapped as graph edges. Weighted directed or undirected edge structures enable a precise representation of the complex industrial system architecture. During the training of the graph neural network, node vectors are initialized with cross-modal representations in a unified semantic space. Edge features are composed of process dependency weights, equipment physical distance decay functions, mutual information of variables, and historical statistical confidence. The model dynamically calculates the influence strength between different nodes through a multi-layer graph attention mechanism (GAT / GTransformer), enabling the network to automatically adapt to structural changes under different production stages, load conditions, and operating scenarios, ultimately forming a high-dimensional fusion representation of the global industrial operating status. This fusion representation not only captures local features but also reflects global dependencies at the equipment, process, and system levels, providing a unified semantic foundation for subsequent protocol interoperability, decision reasoning, and anomaly detection.

[0070] S160. Semantic parsing of industrial protocols is performed using a semantic coding model, and cross-protocol mapping is performed in conjunction with the industrial association graph to obtain interoperability results.

[0071] The semantic encoding model can be a dedicated semantic representation model obtained through secondary pre-training on industrial corpora. Cross-protocol mapping refers to the conversion of data between different industrial protocols. Through cross-protocol mapping, data fields, control commands, status identifiers, and other content in the source protocol can be converted into a form that the target protocol can recognize and process. Interoperability results can be executable technical carriers that enable data interoperability between industrial heterogeneous protocol devices / systems.

[0072] In this embodiment, semantic encoding models can be used to perform semantic parsing of industrial protocols, and cross-protocol mapping can be performed in conjunction with industrial association graphs to obtain interoperability results. For example, industrial protocols can be OPC UA, Modbus, etc. The semantic encoding model can be an enhanced BERT-BiLSTM model. For node names, variable descriptions, and data packets of industrial protocols such as OPC UA and Modbus, the enhanced BERT-BiLSTM model can be used for semantic parsing. Combined with an industrial knowledge graph built based on a large model, semantically similar or equivalent fields between different protocols are automatically identified, generating cross-protocol mapping relationships and obtaining interoperability results, thereby achieving automatic variable alignment and mutual access between protocols.

[0073] In one embodiment, the step of semantically parsing the industrial protocol using a semantic encoding model and performing cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results includes: semantically parsing the data model description of the industrial protocol using a semantic encoding model to generate a protocol semantic embedding representation; the protocol semantic embedding representation includes semantic information of new protocol variables; the data model description includes node trees, register definitions, service methods, namespace descriptions, variable annotations, and access attributes; combining the entity nodes and relationship structure in the industrial association graph, performing entity alignment and relationship linking on the new protocol variables to form a protocol view under a unified semantic space; calculating the embedding vector similarity and semantic distance of the protocol semantic embedding representation, as well as the process constraint matching degree and historical mapping credibility of the new protocol variables; generating mapping relationships between different protocol variables based on the embedding vector similarity, semantic distance, process constraint matching degree, and historical mapping credibility, obtaining an executable mapping table, a set of conversion rules, and an automatically generated conversion function; the mapping relationships include equivalent mapping, similar mapping, combined mapping, and rule mapping; and using the executable mapping table, the set of conversion rules, and the automatically generated conversion function as the interoperability result.

[0074] Among these, protocol semantic embedding representation can refer to the matrix representation formed by mapping industrial protocols to a unified high-dimensional numerical vector space. Entity alignment can refer to the process of identifying and associating different names, identifiers, or descriptions pointing to the same real industrial entity. Relationship linking can refer to the process of establishing semantic associations for aligned industrial entities and completing and integrating scattered entity relationships. New protocol variables can be variable objects defined for newly added industrial protocols. Process constraint matching degree can be a quantitative indicator that measures the degree to which new protocol variables fit with preset constraints. Historical mapping credibility can refer to the degree of fit between the historical mapping relationships of entities, variables, protocols, etc., in different data sources and the current data association requirements, as well as the reliability of the mapping relationships themselves. An executable mapping table can refer to a set of association mapping relationships of heterogeneous protocols / data entities that can be directly invoked and executed. Equivalent mapping can refer to a one-to-one mapping relationship where the entities, variables, or protocol elements of the source and target ends are completely consistent in semantics, business meaning, and data attributes. Similar mapping can refer to association mapping relationships where the entities, variables, or protocol elements of the source and target ends are similar but not completely equivalent in core business semantics. Combination mapping refers to the association mapping relationship established between multiple source elements and a single target element after integrating them according to preset logic. Rule mapping refers to the association mapping relationship established between source and target elements only if they meet preset business rules or transformation logic.

[0075] In this embodiment, the data model description of the industrial protocol can be semantically parsed using a semantic coding model to generate a protocol semantic embedding representation. The protocol semantic embedding representation includes the semantic information of new protocol variables. Combined with the entity nodes and relationship structure in the industrial association graph, the new protocol variables can be entity aligned and relationship linked to form a protocol view under a unified semantic space. Then, the embedding vector similarity and semantic distance of the protocol semantic embedding representation, as well as the process constraint matching degree and historical mapping credibility of the new protocol variables, are calculated. Thus, based on the embedding vector similarity, semantic distance, process constraint matching degree, and historical mapping credibility, mapping relationships between different protocol variables can be generated, resulting in an executable mapping table, a set of transformation rules, and an automatically generated transformation function. The executable mapping table, the set of transformation rules, and the automatically generated transformation function can serve as interoperability results.

[0076] This embodiment, through a unified semantic space and knowledge graph, tightly integrates the multimodal feature fusion results with cross-protocol variable mapping, enabling integrated collaborative processing from the data layer, semantic layer, and protocol layer. The system can automatically parse the field semantics of various industrial protocols such as OPC UA and Modbus, generate cross-protocol mapping relationships based on a large model, and dynamically select appropriate interoperability paths and mapping strategies by combining the fused global state representation, thereby significantly reducing the workload of manually writing and maintaining rules. Compared with existing solutions that rely on manual configuration and static mapping, this embodiment can significantly reduce the configuration cost and operational threshold of protocol interfacing and system integration, and has the ability to adaptively adjust to changes in operating conditions, enhancing the automation and robustness of cross-system interconnection.

[0077] For example, in the protocol semantic parsing and interoperability phase, data model descriptions (including node trees, register definitions, service methods, namespace descriptions, variable annotations, access attributes, etc.) from industrial protocols such as OPC UA, Modbus, EtherNet / IP, and Profinet can be uniformly parsed. Through protocol specification document extraction, textual structure transformation, and natural language normalization, their structural information is uniformly converted into processable semantic text fragments, which are then input into a semantic encoding model enhanced by industrial corpus to generate a protocol semantic embedding representation. Subsequently, combined with entity nodes and relationship structures in the industrial knowledge graph (such as "temperature variable - belongs to → heating process" and "register address - maps to → equipment signal channel"), new protocol variables are automatically aligned to entities and linked to relationships to form a protocol view under a unified semantic space. By calculating multi-dimensional indicators such as embedding vector similarity, semantic distance, process constraint matching degree, and historical mapping credibility, the system automatically generates mapping relationships between variables of different protocols, including equivalent mapping, similarity mapping, combination mapping (multiple source variables constitute the target variable), and rule mapping (variables that require the application of transformation logic). Finally, these relationships are provided to the interoperability engine in the form of executable mapping tables, sets of transformation rules, or automatically generated transformation functions. This enables cross-protocol, cross-device, and cross-system data access, transformation, and sharing, greatly reducing the workload of manual protocol adaptation and improving the automation and adaptability of interoperability.

[0078] Furthermore, Figure 2 This is a flowchart illustrating an industrial data fusion and interoperability processing method provided in an embodiment of the present invention, as shown below. Figure 2As shown, in this embodiment, during the collaborative optimization phase, the data fusion module and the protocol interoperability module can be abstracted into two reinforcement learning agents, enabling them to autonomously adjust their strategies based on the business environment and system feedback during long-term operation. The fusion agent uses device status recognition accuracy, anomaly detection recall rate, prediction task loss value, and graph structure stability as reward signals, dynamically adjusting feature selection methods, node update strategies, attention allocation mechanisms, and graph structure update frequencies to achieve optimal adaptation of the fused representation to changes in operating conditions. The interoperability agent uses protocol conversion power, conversion latency, data loss rate, resource consumption, and cache hit rate as core indicators, exploring the optimal combination of different protocol parsing frequencies, mapping update cycles, caching strategies, and path selection schemes through reinforcement learning. The system uses a joint reward function to fuse the results of the two agents, ensuring that both maintain optimal local performance while considering the overall system's collaborative performance, achieving a balance between fusion accuracy, interoperability efficiency, system latency, and computational resource consumption. During long-term operation, the two agents can continuously learn and adaptively update their strategies, enabling the system to have self-evolutionary capabilities to cope with device changes, protocol upgrades, operating condition fluctuations, and changes in business requirements.

[0079] The dual-agent collaborative optimization mechanism proposed in this embodiment integrates data fusion strategies and protocol interoperability strategies into a unified reinforcement learning framework. By comprehensively considering indicators such as fusion accuracy, protocol conversion power, end-to-end latency, and computational resource consumption, it continuously adjusts and optimizes the behavior of each sub-module. The fusion agent is responsible for dynamically selecting feature dimensions, graph structure update frequency, and inference accuracy mode under different operating conditions; the interoperability agent is responsible for adjusting the protocol mapping update cycle, caching strategy, and conversion path, thereby maximizing the overall system performance while ensuring business reliability. Compared with traditional static configuration or single-objective optimization methods, this embodiment enables the system to maintain stable operation under high load, complex operating conditions, and changing scenarios, achieving a dynamic balance between fusion quality, interoperability efficiency, and resource consumption. It provides industrial enterprises with an engineerable, portable, and sustainably evolving solution for building future-oriented intelligent data infrastructure.

[0080] S170. The unified state representation and the interoperability results after fusion are provided to the upper-layer application through a microservice architecture.

[0081] Microservice architecture refers to a software architecture pattern that breaks down system functions into independent, loosely coupled service units and achieves service collaboration through standardized interfaces.

[0082] In this embodiment, the unified state representation and interoperability results can be provided to upper-layer applications through a service architecture. For example, the unified semantic representation and interoperability results can be exposed to the outside world through a microservice architecture, providing data services to upper-layer applications in the form of application programming interfaces (APIs) or message buses, and supporting various industrial application scenarios such as equipment monitoring, fault prediction, quality inspection, and energy efficiency analysis.

[0083] This embodiment can generate structured, time-series, or event-driven unified data service content based on the fused global semantic representation and the variable values ​​after protocol interoperability. It also includes auxiliary information such as model inference confidence, multimodal consistency weights, and protocol mapping credibility, providing support to upper-layer business modules through a unified data service interface. This interface is based on a microservice architecture and supports multiple delivery methods, including Representational State Transfer (REST), gRPC Remote Procedure Call (gRPC), Open Platform Communications Unified Architecture Server (OPC UA Server), and Message Queuing Telemetry Transport Topic (MQTT Topic). It also provides a streaming subscription mode to meet the needs of real-time monitoring, anomaly warning, and closed-loop control scenarios. For tasks requiring historical analysis, model training, or batch computation, the interface also supports on-demand querying of high-quality, fused, and interoperable data snapshots, and can automatically switch data views according to application requirements, such as device-level views, process-level views, production batch views, or system-level global views. This enables a closed-loop system that integrates data access, intelligent understanding, deep fusion, and protocol interoperability, allowing the system to meet the low latency requirements of real-time industrial control scenarios while supporting multi-scenario and multi-level intelligent manufacturing business applications.

[0084] This invention provides an industrial data fusion and interoperability processing method, comprising: acquiring multi-source industrial data from an industrial environment; preprocessing the multi-source industrial data using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; extracting features from the standardized input dataset using a cross-modal coding framework, mapping features of different modalities to a unified semantic space to obtain a unified multimodal sample dataset; automatically labeling samples in the multimodal sample dataset based on a trained student model to obtain a multimodal labeled dataset; wherein the student model is a model that performs knowledge transfer based on a teacher model; constructing an industrial association graph based on a graph neural network and the multimodal labeled dataset, processing the industrial association graph to obtain a fused unified state representation; performing semantic parsing of industrial protocols using a semantic coding model, performing cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results; and providing the fused unified state representation and the interoperability results to upper-layer applications through a microservice architecture. This method, taking into account the characteristics of "multi-source, multi-modal, strongly correlated, and highly specialized" industrial sites, constructs a unified technical framework that can operate in a closed loop, integrating data access, semantic extraction, knowledge transfer, intelligent annotation, graph structure fusion, protocol semantic alignment, and reinforcement learning-driven collaborative optimization. This not only significantly improves the semantic understanding and fusion depth of industrial data, but also establishes dynamic interoperability capabilities across protocols and systems, solving the problems of low efficiency and long cycle time in processing industrial data in existing technologies.

[0085] This embodiment addresses common problems in existing industrial data processing, such as multi-source heterogeneity, semantic fragmentation, high annotation costs, limited fusion effects, and low protocol interoperability efficiency. Based on large-scale pre-trained models, it constructs a multi-layered knowledge transfer chain: "General Large Model → Industrial Domain Large Model → Lightweight Inference Model." This deeply integrates general semantic understanding capabilities with industrial expertise, forming a unified semantic representation space for industrial scenarios. The design encompasses key technology modules including multimodal intelligent annotation, data quality optimization, graph neural network fusion, protocol semantic parsing and cross-protocol mapping, and reinforcement learning-driven fusion-interoperability collaborative optimization. It builds an integrated technical system from data access, semantic modeling, fusion decision-making to protocol interoperability, enabling intelligent processing of industrial data across the entire "acquisition-understanding-fusion-interoperability" chain. This enables the introduction of knowledge transfer capabilities from large models into industrial scenarios, achieving unified representation and adaptive understanding of multi-source heterogeneous industrial data; reducing industrial data annotation costs, improving annotation quality and consistency, and providing high-quality training samples for subsequent intelligent analysis; enhancing the accuracy and robustness of industrial data fusion, enabling deep collaboration of multimodal features in a unified space; achieving automatic semantic mapping and efficient interoperability between multiple protocols and systems, reducing reliance on manual configuration; achieving dynamic balance and continuous optimization between fusion quality and interoperability efficiency through a dual-agent reinforcement learning mechanism; and forming an industrial data fusion and interoperability platform that can be engineered and migrated across industries, supporting the deep application of intelligent manufacturing and the industrial internet.

[0086] Example 2 Figure 3 This is a schematic diagram of an industrial data fusion and interoperability processing device provided in Embodiment 2 of the present invention. The device is applicable to the processing of industrial data obtained from an industrial field environment. The device can be implemented by software and / or hardware and is generally integrated on an electronic device.

[0087] like Figure 3 As shown, the device includes: Acquisition module 210 is used to acquire multi-source industrial data from the industrial field environment; Processing module 220 is used to preprocess the multi-source industrial data by combining a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; The extraction module 230 is used to extract features from the standardized input dataset through a cross-modal coding framework, mapping features of different modalities to a unified semantic space to obtain a multimodal sample dataset with a unified representation. The annotation module 240 is used to automatically annotate the samples in the multimodal sample dataset based on the trained student model to obtain a multimodal annotated dataset; the student model is a model that performs knowledge transfer based on the teacher model; The construction module 250 is used to construct an industrial association graph based on a graph neural network combined with the multimodal labeled dataset, and to process the industrial association graph to obtain a fused unified state representation; Parsing module 260 is used to perform semantic parsing of industrial protocols through a semantic coding model, and to perform cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results; Module 270 is provided to provide the integrated unified state representation and the interoperability results to upper-layer applications through a microservice architecture.

[0088] This embodiment provides an industrial data fusion and interoperability processing device, comprising: an acquisition module for acquiring multi-source industrial data from an industrial site environment; a processing module for preprocessing the multi-source industrial data using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; an extraction module for extracting features from the standardized input dataset using a cross-modal coding framework, mapping features of different modalities to a unified semantic space to obtain a unified representation of a multimodal sample dataset; an annotation module for automatically annotating samples in the multimodal sample dataset based on a trained student model to obtain a multimodal annotation dataset; wherein the student model is a model based on a teacher model for knowledge transfer; a construction module for constructing an industrial association graph based on a graph neural network and the multimodal annotation dataset, processing the industrial association graph to obtain a fused unified state representation; a parsing module for semantically parsing industrial protocols using a semantic coding model, performing cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results; and a providing module for providing the fused unified state representation and the interoperability results to upper-layer applications through a microservice architecture.

[0089] Furthermore, the multi-source industrial data includes at least high-frequency data, low-frequency data, event-type data, and image-type data. Correspondingly, the processing module 220 includes: The high-frequency data is subjected to window aggregation and low-pass filtering to obtain multi-scale time slices; The missing segments in the low-frequency data are recovered by interpolation, migration interpolation and morphological smoothing to obtain the recovered low-frequency data. The low-frequency data is time-aligned with the multi-scale time slice to obtain aligned low-frequency data. The event data is timestamped and bound to a context, and then inserted into the multi-scale time slice to obtain the embedded event data. Generate association indexes for image data with the working conditions and variable states at corresponding moments in multi-scale time slices, and obtain image data with association indexes; The multi-scale time slices, aligned low-frequency data, embedded event data, and image data with associated indexes are used as data in the standardized input dataset.

[0090] Furthermore, the cross-modal coding framework includes a text coding sub-model, an image coding sub-model, a temporal coding sub-model, and a cross-modal alignment network. The data types in the standardized input dataset include text, image, and temporal types. Correspondingly, the extraction module 230 includes: The text features are obtained by semantically extracting the text type data using the text encoding sub-model. Spatial features of the image type data are extracted using the image coding sub-model. The temporal characteristics of the time-series data are obtained through the temporal coding sub-model. The cross-modal alignment network maps the text features, spatial features, and temporal features to a unified semantic space, resulting in a unified representation of the multimodal sample dataset. The text encoding sub-model is a large language model that has been trained twice on industrial corpora. The image encoding sub-model adopts a fusion structure of deep convolutional network and visual Transformer. The temporal encoding sub-model adopts a combination structure of bidirectional long short-term memory network, temporal Transformer and frequency domain convolution.

[0091] Furthermore, the annotation module 240 includes: Based on the trained student model, determine the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector for each sample in the multimodal sample dataset; Based on the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector of each sample, the samples in the multimodal sample dataset that meet the first preset condition are taken as the data in the multimodal annotation dataset. The first preset condition is that the confidence level of the sample is in the first confidence level interval and passes the cross-modal consistency check, or the confidence level of the sample is in the second confidence level interval and passes the manual review, wherein the confidence level in the first confidence level interval is higher than the confidence level in the second confidence level interval.

[0092] Furthermore, the cross-modal consistency check includes one or more of the following checks: whether image defects correspond to time-series vibration anomalies, whether alarm text matches events in the equipment log, whether the change trend of protocol fields is consistent with the process status, and whether the process stage label corresponds to the production batch progress.

[0093] Furthermore, building module 250 includes: An industrial association graph is constructed based on the multimodal labeled dataset. The graph nodes of the industrial association graph are field equipment, sensor nodes, workstation processes, process parameters, quality inspection results and log events. The graph edges are physical connection relationships, control logic dependencies, temporal sequence relationships, semantic relevance and dynamic relevance of working conditions. The industrial association graph is subjected to cross-modal feature aggregation through a multi-layer graph attention network to obtain aggregated features; The aggregated features are optimized using a joint loss function to obtain a unified state representation after fusion.

[0094] Furthermore, the parsing module 260 includes: The data model description of the industrial protocol is semantically parsed using a semantic coding model to generate a protocol semantic embedding representation. The protocol semantic embedding representation includes the semantic information of new protocol variables. The data model description includes a node tree, register definitions, service methods, namespace descriptions, variable annotations, and access attributes. By combining the entity nodes and relationship structure in the industrial association diagram, the new protocol variables are aligned with entities and linked with relationships to form a protocol view under a unified semantic space. Calculate the embedding vector similarity and semantic distance of the semantic embedding representation of the protocol, as well as the process constraint matching degree and historical mapping credibility of the new protocol variable; Based on the embedding vector similarity, semantic distance, process constraint matching degree, and historical mapping credibility, a mapping relationship between different protocol variables is generated, resulting in an executable mapping table, a set of conversion rules, and an automatically generated conversion function; the mapping relationship includes equivalent mapping, similarity mapping, combination mapping, and rule mapping; The executable mapping table, the set of transformation rules, and the automatically generated transformation functions are used as the interoperability result.

[0095] The aforementioned industrial data fusion and interoperability processing device can execute the industrial data fusion and interoperability processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0096] Example 3 Figure 4A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0097] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0098] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0099] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as industrial data fusion and interoperability processing methods.

[0100] In some embodiments, the industrial data fusion and interoperability processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the industrial data fusion and interoperability processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the industrial data fusion and interoperability processing method by any other suitable means (e.g., by means of firmware).

[0101] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0103] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0105] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0106] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0107] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0108] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for industrial data fusion and interoperability processing, characterized in that, The method includes: Acquire multi-source industrial data from industrial field environments; The multi-source industrial data is preprocessed using a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; By extracting features from the standardized input dataset using a cross-modal coding framework, features from different modalities are mapped to a unified semantic space, resulting in a multimodal sample dataset with a unified representation. The samples in the multimodal sample dataset are automatically labeled based on the trained student model to obtain a multimodal labeled dataset; the student model is a model that performs knowledge transfer based on the teacher model; An industrial association graph is constructed based on a graph neural network combined with the multimodal labeled dataset. The industrial association graph is then processed to obtain a fused unified state representation. The semantic encoding model is used to perform semantic parsing of industrial protocols, and cross-protocol mapping is performed in combination with the industrial association graph to obtain interoperability results; The unified state representation and interoperability results are provided to upper-layer applications through a microservice architecture.

2. The method according to claim 1, characterized in that, The multi-source industrial data includes at least high-frequency data, low-frequency data, event-type data, and image data. Correspondingly, the multi-scale time-axis reconstruction mechanism is used to preprocess the multi-source industrial data to obtain a standardized input dataset, including: The high-frequency data is subjected to window aggregation and low-pass filtering to obtain multi-scale time slices; The missing segments in the low-frequency data are recovered by interpolation, migration interpolation and morphological smoothing to obtain the recovered low-frequency data. The low-frequency data is time-aligned with the multi-scale time slice to obtain aligned low-frequency data. The event data is timestamped and bound to a context, and then inserted into the multi-scale time slice to obtain the embedded event data. Generate association indexes for image data with the working conditions and variable states at corresponding moments in multi-scale time slices, and obtain image data with association indexes; The multi-scale time slices, aligned low-frequency data, embedded event data, and image data with associated indexes are used as data in the standardized input dataset.

3. The method according to claim 1, characterized in that, The cross-modal coding framework includes a text coding sub-model, an image coding sub-model, a temporal coding sub-model, and a cross-modal alignment network. The data types in the standardized input dataset include text, image, and temporal types. Correspondingly, the cross-modal coding framework is used to extract features from the standardized input dataset, mapping features from different modalities to a unified semantic space to obtain a unified representation of the multimodal sample dataset, including: The text features are obtained by semantically extracting the text type data using the text encoding sub-model. Spatial features of the image type data are extracted using the image coding sub-model. The temporal characteristics of the time-series data are obtained through the temporal coding sub-model. The cross-modal alignment network maps the text features, spatial features, and temporal features to a unified semantic space, resulting in a unified representation of the multimodal sample dataset. The text encoding sub-model is a large language model that has been trained twice on industrial corpora. The image encoding sub-model adopts a fusion structure of deep convolutional network and visual Transformer. The temporal encoding sub-model adopts a combination structure of bidirectional long short-term memory network, temporal Transformer and frequency domain convolution.

4. The method according to claim 1, characterized in that, The process of automatically labeling samples in the multimodal sample dataset based on the trained student model yields a multimodal labeled dataset, including: Based on the trained student model, determine the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector for each sample in the multimodal sample dataset; Based on the candidate label set, label confidence distribution, multimodal consistency score, and semantic similarity vector of each sample, the samples in the multimodal sample dataset that meet the first preset condition are taken as the data in the multimodal annotation dataset. The first preset condition is that the confidence level of the sample is in the first confidence level interval and passes the cross-modal consistency check, or the confidence level of the sample is in the second confidence level interval and passes the manual review, wherein the confidence level in the first confidence level interval is higher than the confidence level in the second confidence level interval.

5. The method according to claim 4, characterized in that, The cross-modal consistency check includes one or more of the following checks: whether image defects correspond to time-series vibration anomalies, whether alarm text matches events in the equipment log, whether the change trend of protocol fields is consistent with the process status, and whether the process stage label corresponds to the production batch progress.

6. The method according to claim 1, characterized in that, The process of constructing an industrial association graph based on a graph neural network and the multimodal labeled dataset, and processing the industrial association graph to obtain a fused unified state representation includes: An industrial association graph is constructed based on the multimodal labeled dataset. The graph nodes of the industrial association graph are field equipment, sensor nodes, workstation processes, process parameters, quality inspection results and log events. The graph edges are physical connection relationships, control logic dependencies, temporal sequence relationships, semantic relevance and dynamic relevance of working conditions. The industrial association graph is subjected to cross-modal feature aggregation through a multi-layer graph attention network to obtain aggregated features; The aggregated features are optimized using a joint loss function to obtain a unified state representation after fusion.

7. The method according to claim 1, characterized in that, The process of semantically parsing industrial protocols using a semantic encoding model and performing cross-protocol mapping based on the industrial association graph to obtain interoperability results includes: The data model description of the industrial protocol is semantically parsed using a semantic coding model to generate a protocol semantic embedding representation. The protocol semantic embedding representation includes the semantic information of new protocol variables. The data model description includes a node tree, register definitions, service methods, namespace descriptions, variable annotations, and access attributes. By combining the entity nodes and relationship structure in the industrial association diagram, the new protocol variables are aligned with entities and linked with relationships to form a protocol view under a unified semantic space. Calculate the embedding vector similarity and semantic distance of the semantic embedding representation of the protocol, as well as the process constraint matching degree and historical mapping credibility of the new protocol variable; Based on the embedding vector similarity, semantic distance, process constraint matching degree, and historical mapping credibility, a mapping relationship between different protocol variables is generated, resulting in an executable mapping table, a set of conversion rules, and an automatically generated conversion function; the mapping relationship includes equivalent mapping, similarity mapping, combination mapping, and rule mapping; The executable mapping table, the set of transformation rules, and the automatically generated transformation functions are used as the interoperability result.

8. An industrial data fusion and interoperability processing device, characterized in that, The device includes: The acquisition module is used to acquire multi-source industrial data from the industrial field environment; The processing module is used to preprocess the multi-source industrial data by combining a multi-scale time axis reconstruction mechanism to obtain a standardized input dataset; The extraction module is used to extract features from the standardized input dataset through a cross-modal coding framework, mapping features from different modalities to a unified semantic space to obtain a multimodal sample dataset with a unified representation. The annotation module is used to automatically annotate the samples in the multimodal sample dataset based on the trained student model, so as to obtain a multimodal annotated dataset; the student model is a model that performs knowledge transfer based on the teacher model; The construction module is used to construct an industrial association graph based on a graph neural network combined with the multimodal labeled dataset, and to process the industrial association graph to obtain a fused unified state representation; The parsing module is used to perform semantic parsing of industrial protocols through a semantic coding model, and to perform cross-protocol mapping in conjunction with the industrial association graph to obtain interoperability results. A module is provided to provide the integrated unified state representation and the interoperability results to upper-layer applications through a microservice architecture.

9. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the industrial data fusion and interoperability processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the industrial data fusion and interoperability processing method according to any one of claims 1-7.