Table data feature representation and migration method for industrial intelligent scene
By building a tabular pre-training framework and feature-level semantic alignment mechanism, the problem of difficult to migrate new features in industrial intelligence scenarios is solved, efficient model migration and deployment is achieved, and the adaptability and efficiency of industrial intelligence systems are improved.
Patent Information
- Application Number
- CN202510520762.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
In the industrial intelligent scenarios, new features are difficult to migrate, and the cost of model reuse is high, resulting in high training costs and low deployment efficiency, which has become a bottleneck restricting large-scale implementation.
By building a tabular pre-training framework that characterizes reusable and structural transferability, a feature-level semantic alignment and migration mechanism is adopted, and a class-center-guided comparison regularization constraint is used to achieve reuse of shared features and fast adaptation of new features, and fine-tune it in combination with the Transformer model.
It significantly improves the migration performance and deployment efficiency of the model under new tasks, realizes rapid adaptation and efficient model migration, and is suitable for dynamic modeling and model reuse in multi-source heterogeneous industrial data environment.
Smart Images

Figure CN120338043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for representing and migrating tabular data features for industrial intelligent scenarios, which is suitable for dynamic modeling and model reuse scenarios in multi-source heterogeneous industrial data environments, and is particularly suitable for new task modeling requirements that continuously introduce new features in enterprise-level intelligent systems, and belongs to the technical field of machine learning and tabular data modeling. Background Art
[0002] Tabular data is one of the most common forms of data in the real world and is widely used in many industries such as finance, healthcare, energy, and industrial manufacturing. Especially in the process of industrial intelligence, a large number of monitoring, control, and operational decision-making tasks rely on modeling and analysis of structured tabular data. In recent years, deep learning methods have made breakthrough progress in the fields of image and natural language processing. Inspired by this, academia and industry have also begun to apply deep models to tabular data tasks, aiming to mine the complex feature interaction structure in tabular data.
[0003] The current mainstream table learning methods are divided into two categories: one is the Boosting algorithm based on decision trees, which is robust in small sample and high noise environments, but lacks in high-order feature modeling capabilities; the other is a modeling method based on deep neural networks, which can automatically learn feature interactions and have certain migration capabilities. However, since table data often exhibits high heterogeneity in practical applications, that is, the inconsistency of feature names, types and semantics between different tasks, this poses a great challenge to model pre-training and migration.
[0004] Especially in industrial intelligence scenarios, new tasks are constantly generated, and often contain new features that are completely different from previous tasks. This inconsistency in feature space makes it difficult to directly migrate and reuse existing models. The traditional "pre-training-fine-tuning" paradigm is facing failure. The high training cost and low deployment efficiency have become bottlenecks restricting the large-scale implementation of industrial AI systems. Summary of the invention
[0005] Purpose of the invention: In view of the problems of "difficulty in migrating new features" and "high model reuse cost" faced by the prior art in processing multi-task table modeling, the present invention proposes a table data feature representation and migration method for industrial intelligent scenarios, and is committed to building a table pre-training framework with "reusable representation", "migratable structure" and "efficient fine-tuning" capabilities. This method achieves the reuse of shared feature representations and the rapid adaptation of new features through feature-level semantic alignment and migration mechanisms, thereby significantly improving the migration performance and engineering deployment efficiency of the model under the condition of continuously adding new features to new tasks.
[0006] Technical solution: A method for feature representation and migration of tabular data for industrial intelligent scenarios. The feature representation and migration method is applicable to scenarios where new features are continuously introduced in multi-source industrial tasks such as industrial equipment fault warning, production process optimization, energy consumption monitoring, and multi-site data collaborative analysis. The method includes a dataset construction and representation pre-training process, a new task and pre-trained model matching process, and a model migration and system deployment process.
[0007] In the dataset construction and representation pre-training process, heterogeneous tabular data from multiple domains is collected, and after standardization cleaning and quality screening, multiple pre-training datasets are constructed; In the new task and pre-trained model matching process, the feature representation module and the Transformer model are trained respectively, and a class center-guided contrast regularization constraint is introduced to enhance the semantic structure and migration ability of the representation; In the model migration and system deployment process, according to the feature names or semantic information in the new task, the representation and model structure matching the shared features of the new task are automatically retrieved from the pre-trained model library; after directly using the existing feature representations, only the representations of the newly added unseen features and the Transformer structure need to be fine-tuned to achieve rapid adaptation and migration deployment.
[0008] The dataset construction and representation pre-training process includes the following steps: Step 100, collect tabular data from multiple domains from public Internet data sources or the user's internal historical database; Step 101, perform semantic alignment and field normalization on the tabular data based on a unified feature naming and field standard, and complete data cleaning using methods such as missing value processing, outlier detection, and type verification; Step 102, perform sampling and quality screening strategies based on the cleaned data to construct a set of pre-training datasets with complete structures and unified semantics; Step 103, construct the feature representation for each sample in the pre-training dataset set. Categorical features are obtained through look-up tables, numerical features are mapped using linear projection, and a feature averaging strategy is used to generate sample representations of a unified dimension; Step 104, construct class centers based on sample labels or pseudo-labels, and introduce class center-guided contrast regularization constraints to make the representations of the same class samples gather; Step 105, through joint optimization of the main loss function of the new task and the class center contrast regularization constraint, train the feature representation module and the Transformer structure to obtain feature representations with consistent semantics and migration capabilities.
[0009] The new task and pre-trained model matching process includes the following steps: Step 200: Based on the pre-trained Transformer structure and feature representation module from the existing dataset, construct a pre-trained model library; Step 201: When a new task arrives, the system retrieves the corresponding representation of the shared features from the pre-trained model library according to the new task feature name or semantic information; Step 202: For the newly added unseen features, initialize them with the mean of the pre-trained representations and set the newly added feature representations as learnable parameters. The representations of the remaining shared features are kept frozen to reduce the computational overhead; Step 203: Initialize the Transformer structure parameters as the pre-trained model parameters, and only perform short-term training on the Transformer layer and the newly added representations to complete the transfer adaptation.
[0010] The model transfer and system deployment process includes the following steps: Step 300: Directly reuse the representations corresponding to the shared features in the pre-trained model to reduce redundant training overhead and utilize the knowledge in the existing tasks; Step 301: Perform fine-tuning operations on the target task data, and only update the representations of the unseen features and the adjustable parameters in the Transformer structure; Step 302: After fine-tuning, deploy the model to the industrial system. The system supports dynamically identifying new features through the feature name matching mechanism and automatically triggering the transfer process; Step 303: The system supports periodic updating of the pre-trained model library and model performance evaluation, and can continuously optimize the feature representations and model structure in combination with real-time feedback.
[0011] In the above, the contrast regularization constraint method adopts the strategy of minimizing the Euclidean distance based on the class center to enhance the quality and semantic discriminability of the feature representations and improve the cross-task transfer ability.
[0012] A table data feature representation and transfer device for industrial intelligent scenarios includes a dataset construction and representation pre-training module, a new task and pre-trained model matching module, and a model transfer and system deployment module; The dataset construction and representation pre-training module collects multi-domain heterogeneous table data, and constructs multiple pre-trained datasets through standardized cleaning and quality screening; The new task and pre-trained model matching module trains the feature representation module and the Transformer model respectively, and introduces class center-guided contrast regularization constraints to enhance the semantic structure and transfer ability of the representations; The model migration and system deployment module automatically retrieves the representations and model structures that match the shared features from the pre-trained model library according to the feature names or semantic information in the new task; after directly using the existing feature representations, only the representations of the newly added unseen features and the Transformer structure need to be fine-tuned to achieve rapid adaptation and migration deployment.
[0013] The implementation process and method of the device are the same and will not be elaborated here.
[0014] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for feature representation and migration of tabular data for industrial intelligent scenarios are implemented.
[0015] A computer-readable storage medium stores a computer program for performing the above-mentioned feature representation and migration of tabular data for industrial intelligent scenarios. Description of the Drawings
[0016] Figure 1 It is the flowchart of dataset construction and representation pre-training according to an embodiment of the present invention; Figure 2 It is the flowchart of matching between new tasks and pre-trained models according to an embodiment of the present invention; Figure 3 It is the flowchart of model migration and system deployment according to an embodiment of the present invention. Detailed Embodiments
[0017] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.
[0018] This embodiment is described with the industrial equipment fault prediction task as a specific application scenario. Industrial intelligent systems usually need to monitor the real-time operating status of multiple types of equipment and give early warnings of fault risks. The collected data presents a typical structured tabular form. Each row represents the equipment operating status record within a time period, and the columns are different monitoring indicators or meta-information. Common features include equipment number, operating current, voltage, temperature, vibration amplitude, equipment type, production line location, last maintenance time, etc. There may be differences in available features between different equipment types. For example, large cooling systems may have exclusive features such as "cooling water flow rate" and "refrigerant pressure", while small conveyor equipment may include unique indicators such as "belt tension" and "roller speed". Such feature differences between different equipment constitute significant heterogeneity in tabular data.
[0019] To establish a general-purpose equipment status prediction model, the system first collects equipment operation data from multiple industrial parks. Some of the data comes from the user's internal historical database, and some comes from open-source equipment monitoring platforms on the industrial Internet. There are problems such as inconsistent formats, chaotic field naming, and a large number of missing values in the original tabular data during the collection stage. The system unifies the field naming specifications through an automated cleaning process, implements anomaly detection and completion strategies, and performs field normalization using domain rules, finally constructing multiple high-quality tabular datasets for pre-training.
[0020] In the pre-training stage, the system trains the feature extraction module and the Transformer structure for different tasks respectively. The features of each sample are first generated into preliminary representations through categorical embedding or numerical mapping, and then a sequence-insensitive and dimension-aligned unified sample representation is obtained through the feature averaging strategy. The system further introduces a class center-guided contrast regularization strategy, dynamically constructs class centers in each training batch, and promotes the clustering of semantically similar samples in the feature space, thereby obtaining general representations with semantic structures and easy to transfer. At the same time, the system jointly trains the Transformer network based on high-quality feature representations to improve the model's representation ability under complex tasks.
[0021] In the deployment stage, assume that a new equipment type, "intelligent hydraulic pump", is introduced in a factory. This equipment adds multiple unseen features such as "hydraulic pressure", "return oil temperature", and "pump valve opening degree" in the tabular data. The system automatically recognizes that this task has the same general features such as "current", "voltage", "temperature", and "equipment type" as the existing pre-training tasks and can be matched. It immediately extracts the representations and Transformer parameters corresponding to the shared features between tasks from the model library, and initializes the new feature representations to the mean of the existing representations. Subsequently, only the representations and Transformer models of the new features need to be trained on the new task for a short time to complete transfer learning. This process effectively avoids the computing power waste and time delay caused by training the model from scratch, and realizes the online deployment of the new task in a short time.
[0022] Through the above implementation methods, the present invention can provide an efficient and scalable model transfer solution in an industrial environment with highly heterogeneous tabular features and dynamically changing task requirements, significantly improving the generalization ability, deployment efficiency, and practical application value of the tabular modeling system, and is applicable to multiple typical intelligent manufacturing scenarios such as large-scale industrial equipment management, condition diagnosis, and operation and maintenance decision-making.
[0023] Obviously, those skilled in the art should understand that each step of the method for table data feature characterization and migration for industrial intelligent scenarios in the above embodiments of the present invention, or each module of the table data feature characterization and migration for industrial intelligent scenarios, can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A method for feature representation and migration of tabular data for industrial intelligent scenarios, characterized in that It includes a dataset construction and representation pre-training process, a new task and pre-trained model matching process, and a model migration and system deployment process; In the dataset construction and representation pre-training process, multi-domain heterogeneous tabular data is collected, and after standardization cleaning and quality screening, multiple pre-training datasets are constructed; In the new task and pre-trained model matching process, a feature representation module and a Transformer model are trained respectively, and a class center-guided contrast regularization constraint is introduced to enhance the semantic structure and migration ability of the representation; In the model migration and system deployment process, according to the feature names or semantic information in the new task, the representation and model structure matching the shared features with the new task are automatically retrieved from the pre-trained model library; after directly using the existing feature representations, only the representations of the newly added unseen features and the Transformer structure need to be fine-tuned to achieve fast adaptation and migration deployment.
2. The method for characterizing and migrating tabular data features for industrial intelligent scenarios according to claim 1, wherein The feature representation and migration method is applicable to scenarios where new features are continuously introduced in new tasks of multi-source industrial tasks such as industrial equipment fault warning, production process optimization, energy consumption monitoring, and multi-site data collaborative analysis.
3. The method for characterizing and migrating tabular data features for industrial intelligent scenarios according to claim 1, wherein The dataset construction and representation pre-training process includes the following steps: Step 100, collect tabular data in multiple domains from public Internet data sources or the user's internal historical database; Step 101, perform semantic alignment and field normalization on the tabular data based on a unified feature naming and field standard, and complete data cleaning by using methods such as missing value processing, outlier detection, and type verification; Step 102, execute sampling and quality screening strategies based on the cleaned data to construct a set of pre-training datasets with complete structure and unified semantics; Step 103, construct the feature representation for each sample in the set of pre-training datasets. For categorical features, the embedding is obtained by using a lookup table method, for numerical features, they are mapped by using a linear projection method, and a feature averaging strategy is adopted to generate a sample representation with a unified dimension; Step 104, construct class centers based on sample labels or pseudo-labels, and introduce a class center-guided contrast regularization constraint to make the representations of the same-class samples gather; Step 105, through jointly optimizing the main loss function of the new task and the class center contrast regularization constraint, train the feature representation module and the Transformer structure to obtain feature representations with consistent semantics and migration ability.
4. The method for characterizing and migrating tabular data features for industrial intelligent scenarios according to claim 1, wherein The new task and pre-trained model matching process includes the following steps: Step 200, based on the pre-trained Transformer structure and feature representation module of the existing dataset, construct a pre-trained model library; Step 201, when a new task arrives, retrieve the corresponding representation of the shared features from the pre-trained model library according to the feature names or semantic information of the new task; Step 202, for the newly added unseen features, initialize them with the mean of the pre-trained representations, and set the representations of the newly added features as learnable parameters; the representations of the remaining shared features remain frozen; Step 203, initialize the Transformer structure parameters as the pre-trained model parameters, and only fine-tune the Transformer layer and the newly added representations, that is, train for a short time, so as to complete the migration adaptation.
5. The method for characterizing and migrating tabular data features for industrial intelligent scenarios according to claim 1, wherein The model migration and system deployment process includes the following steps: Step 300: directly reuse the representations corresponding to the features shared with the current task in the pre-trained model; Step 301: perform a fine-tuning operation on the target task data, and only update the representations of unseen features and the adjustable parameters in the Transformer structure; Step 302: after fine-tuning, deploy the model to the industrial system, and the system supports dynamically identifying new features through a feature name matching mechanism and automatically triggering the migration process; Step 303: the system supports periodic updates of the pre-trained model library and model performance evaluation, and can continuously optimize the feature representations and model structure in combination with real-time feedback.
6. The method for characterizing and migrating tabular data features for industrial intelligent scenarios according to claim 1, wherein The contrast regularization constraint method adopts a strategy of minimizing the Euclidean distance based on the class center.
7. A table data feature characterization and migration device for industrial intelligent scenarios, characterized in that It includes a dataset construction and representation pre-training module, a new task and pre-trained model matching module, and a model migration and system deployment module; The dataset construction and representation pre-training module collects multi-domain heterogeneous tabular data, and constructs multiple pre-trained datasets through standardized cleaning and quality screening; The new task and pre-trained model matching module trains the feature representation module and the Transformer model respectively, and introduces a contrast regularization constraint guided by the class center to enhance the semantic structure and migration ability of the representation; The model migration and system deployment module automatically retrieves the representations and model structures that match the shared features from the pre-trained model library according to the feature names or semantic information in the new task; after directly using the existing feature representations, only the representations of the newly added unseen features and the Transformer structure need to be fine-tuned to achieve rapid adaptation and migration deployment.
8. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for feature representation and migration of tabular data for industrial intelligent scenarios as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for performing the feature representation and migration of tabular data for industrial intelligent scenarios as described in any one of claims 1-7.