A semi-supervised table regression method and system based on training geometric portrait gating
Patent Information
- Application Number
- CN202611006676.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
但这类路由机制多面向文本请求、传感器场景或特定业务规则,尚未形成一种面向通用表格回归、且仅依赖训练侧信息的训练几何画像门控方式
[0029]与现有技术相比,本发明至少具有以下有益效果。
Smart Images

Figure CN122817871A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and structured data modeling technology, specifically relating to a method and system for model training, state selection, prediction output and diagnostic record generation for tabular regression tasks under the condition that the number of labeled samples is limited and unlabeled covariate samples can be obtained during the training phase. Background Technology
[0002] Tabular regression tasks are widely used in scenarios such as industrial process forecasting, equipment condition estimation, energy modeling, medical monitoring data analysis, real estate valuation, and performance index prediction. This type of data is typically stored in a row-column structure, with the target variable being a continuous numerical value. In practical applications, target labels often rely on detection, experiments, expert annotation, or long-term operational records, which are costly to obtain; in contrast, covariate samples containing only input features are usually easier to accumulate.
[0003] For the problem of tabular regression with few labels, existing solutions can be roughly divided into the following categories.
[0004] One approach employs a semi-supervised pseudo-labeling strategy: first, an initial model is trained using a small number of labeled samples; then, pseudo-labels are generated for the unlabeled samples, and some of these pseudo-labeled samples are selected for subsequent training. This approach is straightforward, but the quality of the pseudo-labels is significantly affected by the initial model. When the initial model exhibits systematic biases in local regions, subsequent training can easily amplify these biases; this problem is particularly pronounced on tabular data with non-smooth boundaries, heteroscedasticity, or obvious local cluster structures.
[0005] Another approach focuses on tabular neural networks or few-shot deep learning. Neural models excel at representing overall smooth trends and high-order feature interactions, but their training stability is often affected by data distribution when the number of labels is limited. Simply increasing the network capacity cannot determine whether the current task is better suited to neural models, tree models, or a certain intermediate prediction state after introducing unlabeled covariates.
[0006] Another category of solutions employs tree models, ensemble models, or automated machine learning. Gradient boosting trees, random forests, and extreme random trees are sensitive to local partitioning and axial boundaries, while automated machine learning systems often improve performance through model library search, validation set selection, and stacked learning. These methods are quite effective in engineering applications, but their model selection is often based on validation performance or meta-learner output, making it difficult to directly explain why the geometry of the training data leads to a particular predicted state.
[0007] In addition, some solutions employ model routing, confidence ensemble, or residual correction to switch between different models or compensate for prediction errors. However, these routing mechanisms are mostly geared towards text requests, sensor scenarios, or specific business rules, and have not yet formed a training geometric profile gating method that is geared towards general tabular regression and relies solely on training-side information.
[0008] Therefore, it is still necessary to provide a solution that avoids simply treating unlabeled samples as real labels while utilizing unlabeled covariates on the training side; selects the prediction state based on the morphology of the training data itself while preserving the complementary bias between the neural model and the tree model; and does not read test set information during the state selection process, so that the prediction results and their sources have a clearer basis for reproduction and review. Summary of the Invention
[0009] The technical problem to be solved by this invention is: in a few-label tabular regression task, how to improve the stability and interpretability of prediction results by utilizing unlabeled covariates on the training side and heterogeneous biases between neural models and tree models without using test labels or test covariates, and to reduce the uncertainty caused by direct pseudo-label expansion, unified bagging or black-box stacking.
[0010] To address the above problems, this invention provides a semi-supervised table regression method based on training geometric profile gating, comprising:
[0011] Obtain the labeled training set and the unlabeled covariate set from the training side;
[0012] The neural anchor model and tree anchor model are trained based on the labeled training set;
[0013] Anchor responses are generated based on the neural anchor model and the tree anchor model;
[0014] An unlabeled projection state is constructed based on the unlabeled covariates on the training side and the anchor response;
[0015] Computational training geometric profiling;
[0016] Select the target prediction state from the candidate prediction states based on the trained geometric profile;
[0017] Based on the target predicted state, output regression prediction values and diagnostic records.
[0018] This invention also provides a system for a semi-supervised table regression method based on training geometric profile gating, comprising:
[0019] The training data access module is used to receive labeled training sets and unlabeled covariate sets from the training side;
[0020] The training-side preprocessing module is used to preprocess the training-side data.
[0021] A neural anchor training module is used to train a neural anchor model based on the labeled training set.
[0022] The tree anchor training module is used to train the tree anchor model based on the labeled training set.
[0023] The unlabeled projection module is used to generate anchor responses based on the neural anchor model and the tree anchor model, and to construct unlabeled projection states based on the training-side unlabeled covariates and the anchor responses.
[0024] The training geometric profile module is used to calculate the training geometric profile;
[0025] The gated routing module is used to select the target prediction state from the candidate prediction states based on the trained geometric profile;
[0026] A candidate prediction state library is used to store candidate prediction states.
[0027] The residual correction module is used to perform residual correction after the target prediction state is determined;
[0028] The prediction and diagnosis output module is used to output regression prediction values and diagnostic records.
[0029] Compared with the prior art, the present invention has at least the following beneficial effects.
[0030] First, unlabeled covariates are not directly equivalent to labeled samples. Instead, they participate in the learning process through anchor conditional responses and training-side coverage information, thereby reducing the risk of repeated reinforcement of pseudo-label errors from the source.
[0031] Secondly, the model state selection is driven by the training geometric profile. Profile parameters such as smoothness, axial splitting evidence, anchor point risk, and coverage consistency jointly characterize the task morphology, giving the state selection a clear training-side basis.
[0032] Third, the complementary biases of the neural model and the tree model are incorporated into the same training process. The neural anchor points focus on the overall continuous representation, while the tree anchor points focus on local partitioning and boundary responses. Together, they provide the information required for gated routing and unlabeled projection.
[0033] Fourth, candidate prediction states are triggered based on the training profile. Different prediction states can be enabled for smooth tasks, tasks with strong local changes, tasks with relatively stable tree self-training, and tasks with relatively stable neural anchors; when the enhanced state does not meet the triggering conditions, it can revert to the basic state.
[0034] Fifth, the system synchronously generates diagnostic records. The diagnostic records can specify the selected prediction status, the number of triggered images, the usage of labelless projection, and the residual correction strength, which facilitates the reproduction of experiments, the investigation of anomalies, and the explanation of the specific implementation process. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the system composition of the present invention.
[0036] Figure 2 This is a schematic diagram of the method flow of the present invention.
[0037] Figure 3 This is a schematic diagram of the training geometric image gating of the present invention.
[0038] Figure 4 This is a schematic diagram illustrating the overall technical effect of one embodiment of the present invention.
[0039] Figure 5 This is a comparison chart of typical application data prediction effects in one embodiment of the present invention. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the technical concept and feasible path of the present invention, and do not constitute a limitation on the scope of protection.
[0041] The basic process of this invention is as follows: neural anchors and tree anchors are established on the labeled training set, and semi-supervised projection information is formed by using the responses of the two types of anchors on the unlabeled covariates on the training side; then, the training geometric profile is calculated, and the candidate prediction state is triggered by the profile; after the prediction state is determined, residual correction is performed as needed, and finally the prediction result and diagnosis record are output.
[0042] A semi-supervised table regression method based on training geometric profile gating may include the following steps:
[0043] 1. Obtain the labeled training set and the unlabeled covariate set for the training side.
[0044] 2. Preprocess the training data, where the standardization of the target variable is done only on the labeled training set.
[0045] 3. Train the neural anchor model and the tree anchor model on the labeled training set respectively.
[0046] 4. Anchor responses are generated from neural anchors and tree anchors, enabling unlabeled covariates to participate in learning through anchor conditional responses.
[0047] 5. Calculate a training geometric profile, wherein the training geometric profile includes two or more of the following: smoothness, axial splitting evidence, neural anchor risk, tree anchor risk, risk ratio, anchor consistency, or unlabeled coverage.
[0048] 6. Perform gated routing based on the training geometric profile and select the target predicted state from the candidate predicted state library.
[0049] 7. After the target prediction state is determined, perform optional residual correction.
[0050] 8. Output regression prediction values and generate diagnostic records that include selected status, number of profiles, unlabeled usage, and runtime.
[0051] In one implementation, the neural anchor prediction value is denoted as f. N (x), let f be the predicted value of the tree anchor point. T (x), the anchor point response can be written as equation (1):
[0052] (1)
[0053] Where α is the anchor point combination coefficient; q A (x) is not the true label of the unlabeled sample, but rather the conditional response given by the two anchors on the training side.
[0054] The training geometric profile can be represented by equation (2):
[0055] (2)
[0056] Where S represents the smoothness index, A represents evidence of axial splitting, and R... N Indicates the risk of neural anchors, R T Indicates tree anchor risk, ratio NT This represents the risk ratio between the two types of anchor points, where ρ represents anchor point consistency or unlabeled coverage.
[0057] Gated routing can be expressed as equation (3):
[0058] (3)
[0059] Where G is the training-side gating map, and state is the selected candidate prediction state. This gating map does not read test labels, test covariates, or test set evaluation results.
[0060] The final predicted value can be expressed as equation (4):
[0061] (4)
[0062] Among them, f state(x) represents the output of the selected prediction state, h R (x) represents the output of the residual correction model, λ R This indicates the strength of the residual correction. Residual correction is performed only after the state is determined and does not participate in the selection of candidate predicted states.
[0063] like Figure 1 As shown, the system may include a training data access module, a training-side preprocessing module, a neural anchor training module, a tree anchor training module, a labelless projection module, a training geometric profiling module, a gated routing module, a candidate prediction state library, a residual correction module, and a prediction and diagnosis output module.
[0064] The training data access module receives a labeled training set and an unlabeled covariate set from the training side; the labeled training set contains sample features and continuous target values, while the unlabeled covariate set only contains sample features.
[0065] The training-side preprocessing module performs missing value handling, class encoding, feature scaling, and target standardization. Feature preprocessing can fit both labeled training features and unlabeled covariates simultaneously, while target standardization only fits the labeled training target.
[0066] The neural anchor training module is used to train neural network-type regressors to provide overall smooth trends and high-order feature interaction responses; the tree anchor training module is used to train tree model-type regressors to provide local partitioning, axial splitting, and nonlinear boundary responses.
[0067] The unlabeled projection module generates anchor conditional responses on unlabeled covariates during training, enabling unlabeled samples to participate in function morphology learning. This module does not treat the response values of unlabeled samples as true labels.
[0068] The training geometric profiling module is used to calculate profiling parameters such as smoothness, axial split evidence, anchor risk, risk ratio, and coverage consistency; the gated routing module selects the target prediction state based on these profiling parameters.
[0069] The candidate prediction state library can store one or more of the following: basic neural tree states, unlabeled projection states, stable neural states, tree self-training bag states, high-capacity tree states, random forest-like states, or extremely random tree-like states. The residual correction module is used to correct the residual error after the target prediction state is fixed, and the prediction and diagnosis output module is used to provide the final prediction value and diagnostic record.
[0070] like Figure 2 As shown, one implementation process of the method is as follows.
[0071] S101, Obtain the labeled training set D L ={(x i ,yi And the training-side unlabeled covariate set D U ={x u}; where x is the table feature and y is the continuous target value.
[0072] S102, preprocess the training data. Feature transformation can be based on a combination of labeled features and unlabeled covariates from the training side; target standardization is based solely on a fitting of the labeled training target.
[0073] S103, Training the neural anchor model f N Tree anchor point model f T Neural anchors are used to capture overall trends, while tree anchors are used to capture localized fragmented structures.
[0074] S104, generate the anchor point response q according to equation (1). A (x). For unlabeled covariates, only the anchor conditional response is calculated, and this response is not used as the true observation label.
[0075] S105, train unlabeled projective states based on unlabeled covariates and anchor conditional responses. This step allows the distribution of training-side covariates to participate in determining the shape of the model function without expanding the true label set.
[0076] S106, Calculate the training geometric profile. The profile may be derived from cross-validation, out-of-bag estimation, anchor point differences, neighborhood consistency, or tree split statistics.
[0077] S107 selects the target prediction state through a gating mapping G (profile). The gating rule can be a rule table, a segmented threshold, a monotonic scoring function, or an equivalent mapping determined by the training data.
[0078] S108, Fit or call the selected target prediction state to obtain the base predictor f state .
[0079] S109, Fitting the residual correction model h after fixing the state. R Residual correction intensity λ R It can be determined by training-side cross-validation, leave-one-out estimation, or conservative contraction rules.
[0080] S110, output the final predicted value ŷ(x) according to equation (4), and output the diagnostic record simultaneously.
[0081] like Figure 3 As shown, the training geometric profile gating is used to map the training data format to candidate prediction states. This gating does not select the model after the fact based on the test set performance, but rather completes the state selection during the training phase based on the profile formed by the training set and the unlabeled covariates on the training side.
[0082] In one implementation, candidate prediction states include the following types:
[0083] Basic neural tree state: the basic output of the neural anchor and tree anchor;
[0084] Unlabeled projection state: a semi-supervised state obtained using unlabeled covariates from the training side and anchor conditional responses;
[0085] Neural stability: Enabled when low axis, overall smoothness, or more stable neural anchor points;
[0086] Tree self-training bagging state: Enabled when tree anchor points are stable and label coverage is relatively consistent;
[0087] High-capacity tree state: Enabled when local boundaries are strong and tree bias is effective;
[0088] Forest-type supplementary states: used for scenarios with high noise or more stable random partitioning.
[0089] In one specific embodiment, when the training profile shows that the objective function is relatively smooth and the tree self-training process is stable, the gating selects the tree self-training bag state; when the training profile shows weak evidence of axial splitting and low risk of neural anchors, the gating selects the neural stable state; when the training profile shows significant local changes and low risk of the tree model, the gating selects the high-capacity tree state; when the enhancement state does not meet the triggering conditions, the gating reverts to the basic neural tree state or the unlabeled projection state.
[0090] The above gating conditions are merely examples. In actual deployment, the threshold or scoring function can be replaced with equivalent ones based on the scale of the training data, feature types, computational budget, and target stability requirements.
[0091] In one embodiment, the method of the present invention is applied to 30 tabular regression tasks and compared with methods such as AutoGluon, LightGBM-SelfTrain, CatBoost-SelfTrain, XGBoost-SelfTrain, ExtraTrees, CatBoost, TabM, XGBoost, RandomForest, LightGBM, HistGradientBoosting, and Ridge under the same evaluation criteria. Evaluation metrics include average ranking, number of times a single best result is achieved, number of times a result is in the top three, average R², and average computation time. Table 1 lists the overall ranking in this embodiment.
[0092] Table 1 Overall Ranking
[0093] method Average ranking Single-item optimal number Top three times Average R2 Average time Method of the present invention 2.033333 21 25 0.785099 4.923 seconds AutoGluon 4.266667 1 15 0.765871 17.965 seconds LightGBM-SelfTrain 6.766667 0 4 0.740593 0.329 seconds CatBoost-SelfTrain 7.233333 2 6 0.745447 0.286 seconds XGBoost-SelfTrain 7.266667 1 2 0.746463 0.257 seconds ExtraTrees 7.600000 1 7 0.742504 0.062 seconds CatBoost 7.900000 0 3 0.743967 0.112 seconds TabM 8.433333 0 10 0.723026 2.060 seconds XGBoost 8.466667 0 2 0.742129 0.120 seconds RandomForest 9.166667 0 1 0.735655 0.104 seconds LightGBM 9.233333 0 1 0.721055 0.122 seconds HistGradientBoosting 9.633333 0 0 0.725489 0.497 seconds Ridge 12.066667 1 2 0.551124 0.001 seconds
[0094] From Table 1 and Figure 4 As can be seen, in this embodiment, the average ranking of the method of the present invention is 2.033333, the number of times a single item is optimal is 21, the number of times it ranks in the top three is 25, the average R² is 0.785099, and the average computation time is 4.923 seconds. Compared with AutoGluon, the method of the present invention achieves a higher average R² and a higher average ranking under this evaluation metric, while having a shorter average computation time; compared with the listed self-trained tree model and single model, the method of the present invention has a wider coverage in terms of the number of times a single item is optimal and the number of times it ranks in the top three.
[0095] The values in Table 1 are only used to illustrate the technical effects of a specific embodiment and do not limit the invention to achieving the same values under all data, device, or parameter conditions.
[0096] Table 2 lists the trigger states and comparison results for several typical datasets. For ease of explanation, the data names retain only the application identifier and do not include file extensions, source library names, or evaluation platform identifiers.
[0097] Table 2 Typical Cases
[0098] CPU small High smooth tree self-training bagization 0.972166 AutoGluon 0.968848 The training profile points to a relatively stable tree self-training path, and the bagged state achieves better results. Space GA Low-axis neural stability 0.619928 AutoGluon0.595059; TabM0.618368 Gated transitions to a neural steady state yields slightly better results than the standalone neural model. BikeSharingHour High-capacity tree state 0.925401 AutoGluon 0.923370 Local changes exist in the hourly sample region, thus triggering a high-capacity tree state. Elevators Single-neuron stable state 0.904623 AutoGluon0.882208; TabM0.897725 The gating did not employ a wide bagging path, thus preserving a more stable neural anchor response. Parkinsons Telemonitoring Neural tree complementarity 0.860695 AutoGluon 0.764239 In monitoring table data, the complementarity between neural anchors and tree anchors brings significant improvement. NavalPropulsion High-precision physical system boundary case 0.955215 AutoGluon 0.957131 The method of this invention is similar to, but does not exceed, the comparative method, and the diagnostic records can reflect this type of boundary situation.
[0099] Combined with Table 2 and Figure 5 For different types of data, neural anchor models can use multilayer perceptrons, TabM, Transformer-like table models or other table neural networks; tree anchor models can use gradient boosting trees, random forests, extreme random trees, CatBoost, LightGBM, XGBoost or other tree models.
[0100] Smoothness metrics in training geometric profiles can be represented by neighborhood consistency, kernel smoothing error, neural anchor cross-validation error, or other equivalent quantities; axial splitting evidence can be represented by tree splitting gain, feature splitting stability, or local residual distribution; coverage consistency can be represented by neighborhood distance between unlabeled and labeled samples, anchor response difference, or density coverage.
[0101] Gating rules can employ fixed thresholds, segmented rules, monotonic scoring functions, lightweight classifiers, or other equivalent mappings determined by training data. As long as it utilizes the training geometric profile to select the predicted state of the semi-supervised tabular regression without using test set information for this state selection, it can be considered an equivalent implementation of the present invention.
Claims
1. A semi-supervised table regression method based on training geometric profile gating, characterized in that, include: Obtain the labeled training set and the unlabeled covariate set from the training side; The neural anchor model and tree anchor model are trained based on the labeled training set; Anchor responses are generated based on the neural anchor model and the tree anchor model; An unlabeled projection state is constructed based on the unlabeled covariates on the training side and the anchor response; Computational training geometric profiling; Select the target prediction state from the candidate prediction states based on the trained geometric profile; Based on the target predicted state, output regression prediction values and diagnostic records.
2. The method according to claim 1, characterized in that, The training geometric profile includes at least two of the following: smoothness, axial splitting evidence, neural anchor risk, tree anchor risk, risk ratio, anchor consistency, and unlabeled coverage.
3. The method according to claim 1, characterized in that, The anchor point response is obtained by weighting the neural anchor point prediction value and the tree anchor point prediction value.
4. The method according to claim 1, characterized in that, The unlabeled projection state uses the anchor conditional response of the unlabeled covariates on the training side, and does not take the response value of the unlabeled sample as the true label.
5. The method according to claim 1, characterized in that, The candidate predicted states include at least two of the following: basic neural tree state, unlabeled projection state, neural stable state, tree self-training bag state, high-capacity tree state, random forest state, and extremely random tree state.
6. The method according to claim 1, characterized in that, The gated route that selects the target prediction state from the candidate prediction states based on the trained geometric profile does not use test labels, test covariates, or test set evaluation results.
7. The method according to claim 1, characterized in that, Also includes: After the target prediction state is determined, residual correction is performed, and the residual correction is not used for candidate prediction state selection.
8. The method according to claim 1, characterized in that, The diagnostic record includes at least one of the following: selected state, training geometric profile, unlabeled usage, residual correction strength, runtime, and results from relative comparison methods.
9. A system for implementing the semi-supervised table regression method based on training geometric profile gating as described in any one of claims 1-8, characterized in that, include: The training data access module is used to receive labeled training sets and unlabeled covariate sets from the training side; The training-side preprocessing module is used to preprocess the training-side data. A neural anchor training module is used to train a neural anchor model based on the labeled training set. The tree anchor training module is used to train the tree anchor model based on the labeled training set; The unlabeled projection module is used to generate anchor responses based on the neural anchor model and the tree anchor model, and to construct unlabeled projection states based on the training-side unlabeled covariates and the anchor responses. The training geometric profile module is used to calculate the training geometric profile; The gated routing module is used to select the target prediction state from the candidate prediction states based on the trained geometric profile; A candidate prediction state library is used to store candidate prediction states. The residual correction module is used to perform residual correction after the target prediction state is determined; The prediction and diagnosis output module is used to output regression prediction values and diagnostic records.