Structured data modeling and analysis method based on Transformer
By using the Transformer-based dual-stream feature encoder, hierarchical attention mechanism and feature affinity graph and other technical means in structured data modeling analysis, the problem of insufficient feature heterogeneity processing and interpretability in the existing technology is solved, and more efficient structured data modeling and model interpretability are achieved.
Patent Information
- Application Number
- CN202510131948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The prior art is difficult to effectively deal with the heterogeneity of numerical and class features in structured data modeling and analysis. The feature encoding method is too simple to fully utilize the correlation information between features and lacks interpretability support.
The structured data modeling and analysis method based on Transformer is adopted, and the numerical and class features are adaptively encoded through the dual-stream feature encoder, and multi-scale feature interaction modeling is achieved using hierarchical attention mechanism and feature affinity graph, and an information bottleneck feature filter and a dual-branch prediction network are constructed to provide an interpretable feature importance evaluation mechanism.
It improves the modeling effect of structured data, enhances the interpretability and credibility of the model, can effectively process heterogeneous features and make full use of the correlation information between features.
Smart Images

Figure CN119577402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a Transformer-based structured data modeling and analysis method. Background Art
[0002] With the advent of the big data era, the analysis and modeling of structured data has become an important research direction in the field of artificial intelligence. Traditional structured data analysis methods mainly rely on feature engineering and shallow machine learning models, such as decision trees and logistic regression. In recent years, the Transformer architecture has made breakthrough progress in the field of natural language processing, and its self-attention mechanism can effectively capture long-range dependencies in sequence data. However, when applying Transformer to structured data analysis, key issues such as heterogeneous processing of numerical and categorical features, feature interaction modeling, and model interpretability need to be addressed.
[0003] However, existing technologies still have problems in the field of structured data modeling and analysis. Traditional methods cannot effectively handle the heterogeneity of numerical features and categorical features. The feature encoding method is too simple and cannot fully utilize the correlation information between features. Existing feature extraction methods mainly focus on local feature interactions and lack the ability to model global feature dependencies, resulting in limited feature representation capabilities. The feature selection process often adopts a static scoring mechanism, ignoring the dynamic correlation between features and lacks interpretability support. The model prediction results lack a fine-grained explanation mechanism, making it difficult to quantify the contribution of each feature to the prediction results.
[0004] In summary, there is an urgent need for a Transformer-based structured data modeling and analysis method. By designing a dual-stream feature encoder, numerical features and categorical features are adaptively encoded respectively; hierarchical attention mechanism and feature affinity graph are used to realize multi-scale feature interaction modeling; information bottleneck feature filter and dual-branch prediction network are constructed to provide an interpretable feature importance evaluation mechanism while ensuring model performance; finally, feature contribution is quantitatively analyzed through the structured gradient method to provide fine-grained explanation support for model prediction results. The present invention not only improves the modeling effect of structured data, but also enhances the interpretability and credibility of the model, and can solve the problems in the prior art. Summary of the invention
[0005] The embodiment of the present invention provides a Transformer-based structured data modeling and analysis method, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention,
[0007] Provides Transformer-based structured data modeling and analysis methods, including:
[0008] Receive structured raw data, divide the structured raw data into numerical features and categorical features, perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; perform adaptive weight allocation through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; perform dynamic sampling based on the importance scores of the fused features to generate an enhanced training sample set;
[0009] The enhanced training sample set is input into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through an attention fusion mechanism across subgraphs to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation;
[0010] An information bottleneck feature filter is constructed to filter the deep feature representation by maximizing task-related information and minimizing redundant information to obtain a key feature subset; the key feature subset is input into a dual-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on an attention mechanism to obtain a feature weight vector, and generates a prediction result based on the feature weight vector, and the auxiliary reconstruction branch provides a regularization constraint by reconstructing the key feature subset; based on the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result to generate an analysis result.
[0011] In an optional embodiment,
[0012] Receiving structured raw data, dividing the structured raw data into numerical features and categorical features, performing piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, performing adaptive quantum coding on the categorical features to obtain continuous value mapping features; performing adaptive weight allocation through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; performing dynamic sampling based on the importance score of the fused features to generate an enhanced training sample set, including:
[0013] Receiving structured raw data, dividing the structured raw data into numerical features and categorical features; inputting the numerical features and categorical features into a dual-stream feature encoder for feature processing, wherein the dual-stream feature encoder includes a feature trunk stream and a feature mapping stream;
[0014] In the feature trunk stream, the data distribution of the numerical features is counted to obtain the data density, the position of the optimal segmentation point is determined based on the data density, the numerical features are divided into a plurality of segmentation intervals according to the optimal segmentation point, a linear transformation parameter matrix is constructed in each segmentation interval, the feature difference values between adjacent segmentation intervals are calculated, the optimized linear transformation parameter matrix of each segmentation interval is determined by minimizing the feature difference values, and the optimized linear transformation parameter matrix is used to perform linear transformation on the numerical features of each segmentation interval to obtain normalized trunk features;
[0015] In the feature mapping flow, each category value in the category feature is mapped to a corresponding quantum ground state, an initial amplitude parameter and a phase parameter are assigned to each quantum ground state, a quantum gate sequence including a rotation gate and a control gate is constructed, the quantum gate sequence is sequentially applied to the quantum ground state, a transformed quantum state is obtained through quantum state transformation, an entropy value of the transformed quantum state is calculated, and the transformed quantum state is projected to a continuous space based on the entropy value to obtain a continuous value mapping feature;
[0016] Respectively calculating the correlation coefficients between the normalized backbone features and the dimensions of the continuous value mapping features, combining the correlation coefficients to construct a feature correlation coefficient matrix, training a preset attention network based on the feature correlation coefficient matrix to obtain an importance vector of the feature dimension, and using the importance vector to perform a weighted combination of the normalized backbone features and the continuous value mapping features to obtain a fusion feature;
[0017] The contribution of each dimension of the fusion feature to the target task is calculated, and a feature importance score of each dimension is generated based on the contribution. A sampling probability distribution function is constructed according to the feature importance score, and the sampling probability distribution function is used to perform importance sampling on the training samples to generate an enhanced training sample set.
[0018] In an optional embodiment,
[0019] Counting the data distribution of the numerical feature to obtain data density, and determining the position of the optimal segmentation point based on the data density includes:
[0020] Based on the interquartile range and the number of samples of the numerical feature, a basic interval width is calculated, and based on the basic interval width, the value range of the numerical feature is divided into a plurality of basic statistical intervals, and the number of samples in each basic statistical interval is counted to obtain an interval frequency sequence;
[0021] Calculating the frequency ratios of adjacent basic statistical intervals in the interval frequency sequence to obtain interval similarity, merging adjacent basic statistical intervals whose interval similarity is greater than a preset merging threshold into new statistical intervals to obtain a merged frequency distribution sequence;
[0022] Calculating the cumulative frequency distribution according to the combined frequency distribution sequence, and generating a data density function based on the cumulative frequency distribution; selecting candidate segmentation positions in the combined frequency distribution sequence, calculating the number of samples and the numerical mean of the intervals on both sides of each candidate segmentation position, and calculating the interval variance value of the candidate segmentation position using the sample number and the numerical mean;
[0023] Selecting a position with a maximum interval variance value in the current interval to be segmented as a pending segmentation point, calculating the ratio of the interval variance value of the pending segmentation point to the overall variance of the numerical feature, and determining the variance contribution ratio of the pending segmentation point;
[0024] The local distribution skewness is calculated for the intervals on both sides of the pending segmentation point. When the difference of the local distribution skewness is greater than the preset skewness threshold and the variance contribution ratio is greater than the preset segmentation threshold, the pending segmentation point is determined as the optimal segmentation point. The process is repeated to determine the positions of all optimal segmentation points.
[0025] In an optional embodiment,
[0026] The enhanced training sample set is input into a feature extraction network including a hierarchical attention module and a self-calibration module. The hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through an attention fusion mechanism across subgraphs to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation, including:
[0027] Inputting the enhanced training sample set into the convolutional neural network to obtain an initial feature map, calculating the dot product of any two sample features in the initial feature map, and dividing the dot product by the module-length product of the corresponding feature vectors to obtain a feature similarity matrix; constructing a feature affinity graph based on the feature similarity matrix, wherein the vertices of the feature affinity graph correspond to the sample features, and the edge weights of the feature affinity graph correspond to the similarity values in the feature similarity matrix, and dividing the feature affinity graph into a plurality of feature subgraphs by a spectral clustering algorithm;
[0028] In each of the feature subgraphs, the feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction results of the query feature sequence and the key-value pair feature sequence are masked by using the temporal causal mask matrix to obtain an initial attention weight containing only historical feature information; a spatial distance matrix of sample features in each of the feature subgraphs is calculated, and the spatial distribution entropy of the feature subgraph is calculated based on the spatial distance matrix, and the initial attention weight is adaptively adjusted according to the spatial distribution entropy to obtain an adaptive attention weight; the key-value pair feature sequence is weighted based on the adaptive attention weight, and the position encoding information is superimposed to obtain a local feature representation of each of the feature subgraphs;
[0029] Based on a pre-built multi-layer perceptron network, a concatenated vector of local feature representations of adjacent feature subgraphs is input, and the concatenated vector is forward-propagated through the multi-layer perceptron network to obtain association weights between feature subgraphs, and the local feature representations are weighted summed based on the association weights to obtain a global feature representation;
[0030] The mean and standard deviation of the global feature representation are calculated in the feature dimension respectively; based on the pre-constructed dynamic normalization network, the global feature representation is input into the dynamic normalization network, and the feature transformation coefficient and feature offset are obtained through forward propagation; the global feature representation is subtracted from the mean and divided by the standard deviation to obtain a standardized feature, the standardized feature is multiplied by the feature transformation coefficient, and the feature offset is added to obtain a transformed feature, and the transformed feature is added to the global feature representation to obtain a self-calibration feature;
[0031] The self-calibration features are subjected to layer normalization processing and nonlinear transformation through a ReLU activation function to obtain a deep feature representation.
[0032] In an optional embodiment,
[0033] The feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction result of the query feature sequence and the key-value pair feature sequence is masked by using the temporal causal mask matrix to obtain the initial attention weight containing only the historical feature information, including:
[0034] Based on the similarity of the features of adjacent time steps in the feature subgraph, the feature sequence of the feature subgraph is grouped to obtain a query feature sequence and a key-value pair feature sequence, wherein the time step features whose similarity of the features of adjacent time steps is greater than a dynamic threshold determined based on the global statistical characteristics of the feature sequence are divided into the query feature sequence, and the remaining features are divided into the key-value pair feature sequence, to obtain the feature division results of the query feature sequence and the key-value pair feature sequence;
[0035] Based on the feature division result, the temporal correlation strength of the query feature sequence and the key-value pair feature sequence are respectively calculated by a recursive neural network, wherein the hidden state of the recursive neural network contains the long-term and short-term dependency information of the query feature sequence and the key-value pair feature sequence, and the temporal correlation strength of the query feature sequence and the temporal correlation strength of the key-value pair feature sequence are obtained;
[0036] Based on the temporal position relationship of the feature sequence, a first mask matrix is constructed as a temporal causal mask matrix, and the temporal association strength of the query feature sequence and the temporal association strength of the key-value pair feature sequence are converted into soft mask weights to obtain a second mask matrix;
[0037] Performing matrix multiplication on the query feature sequence and the key-value pair feature sequence to obtain an original attention score; performing addition operation on the original attention score and the first mask matrix to obtain a first mask result; performing element-wise multiplication operation on the first mask result and the second mask matrix to obtain a double-masked attention score; and normalizing the double-masked attention score to obtain an initial attention weight that only contains historical feature information.
[0038] In an optional embodiment,
[0039] An information bottleneck feature filter is constructed to filter the deep feature representation by maximizing task-related information and minimizing redundant information, and the key feature subsets obtained include:
[0040] Constructing an information bottleneck encoder, the information bottleneck encoder includes a three-layer perceptron structure, each layer of the perceptron structure is connected to a normalization layer, the deep feature representation is input into the information bottleneck encoder, and the output layer of the information bottleneck encoder generates probability distribution parameters of the compressed feature space; encoding and mapping the deep feature representation based on the probability distribution parameters to obtain a compressed feature representation;
[0041] Constructing a mutual information evaluation module, the mutual information evaluation module includes a three-layer perceptron structure, inputting a concatenated vector of the compressed feature representation and the label of the target task into the mutual information evaluation module, and outputting the mutual information amount between the compressed feature representation and the target task;
[0042] Calculating an autocorrelation matrix of the compressed feature representation, wherein each element in the autocorrelation matrix is the product of the corresponding feature vector minus the corresponding mean divided by the product of the corresponding standard deviation; calculating feature redundancy based on the difference between the autocorrelation matrix and the identity matrix;
[0043] Constructing a feature selection network, the feature selection network comprising a first fully connected layer and a second fully connected layer, inputting the compressed feature representation into the first fully connected layer, the output of the first fully connected layer being input into the second fully connected layer after passing through an activation function, the output of the second fully connected layer being passed through a gating function to obtain a gating value, and performing element-wise multiplication of the gating value and the compressed feature representation to obtain an initial screening feature;
[0044] Adjusting the output of the gating function based on a temperature parameter, wherein the temperature parameter decreases with the number of training rounds based on a preset decreasing amount, and performing element-by-element multiplication of the adjusted gating value and the compressed feature representation to obtain an adjusted screening feature;
[0045] Constructing an objective function by linearly combining the inverse of the mutual information, the feature redundancy, and the norm of the adjusted screening feature;
[0046] By iteratively optimizing the objective function, the optimization parameters of the information bottleneck encoder and the optimization parameters of the feature selection network are obtained, and an optimized information bottleneck encoder and an optimized feature selection network are obtained;
[0047] The deep feature representation is input into an optimized information bottleneck encoder to obtain an optimized compressed feature representation, and the optimized compressed feature representation is input into an optimized feature selection network to obtain a key feature subset.
[0048] In an optional embodiment,
[0049] According to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result, and the generated analysis results include:
[0050] Based on the feature weight vector and the prediction result, the correlation coefficient between each pair of features in the feature weight vector is calculated to generate a feature correlation matrix; based on the feature correlation matrix, a feature dependency graph is constructed, each feature in the key feature subset is used as a node in the feature dependency graph, and the correlation coefficient in the feature correlation matrix is used as the edge weight between corresponding nodes;
[0051] Constructing a graph Laplacian matrix according to the feature dependency graph, taking the sum of the edge weights of each node and the corresponding connected node as the diagonal element value of the graph Laplacian matrix, and taking the negative value of the edge weight between nodes as the non-diagonal element value of the graph Laplacian matrix;
[0052] Calculating the Jacobian matrix of the prediction result with respect to the key feature subset to obtain an initial gradient; using the graph Laplacian matrix as a structured regularization term, and performing a weighted combination of the structured regularization term and the initial gradient to obtain a structured gradient that takes feature dependencies into account;
[0053] Calculate the structured gradient magnitude of each feature in the key feature subset, the interaction gradient value of each feature with other features, and the centrality measure of each feature in the feature dependency graph;
[0054] The structured gradient amplitude, the interaction gradient value and the centrality measure are weightedly combined to obtain feature contribution; the feature contribution is normalized to obtain the standardized contribution of each feature in the key feature subset to the prediction result; the features in the key feature subset are ranked in importance based on the standardized contribution to generate an analysis result.
[0055] According to a second aspect of the embodiments of the present invention,
[0056] Provides a Transformer-based structured data modeling and analysis system, including:
[0057] The first unit is used to receive structured raw data, divide the structured raw data into numerical features and categorical features, perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, and perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; adaptively allocate weights through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; and dynamically sample based on the importance scores of the fused features to generate an enhanced training sample set;
[0058] The second unit is used to input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and a plurality of local feature representations are integrated through an attention fusion mechanism across subgraphs to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation;
[0059] The third unit is used to construct an information bottleneck feature filter, which filters the deep feature representation by maximizing task-related information and minimizing redundant information to obtain a key feature subset; the key feature subset is input into a two-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on the attention mechanism to obtain a feature weight vector, and generates a prediction result according to the feature weight vector, and the auxiliary reconstruction branch provides regularization constraints by reconstructing the key feature subset; according to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result to generate an analysis result.
[0060] According to a third aspect of the embodiments of the present invention,
[0061] An electronic device is provided, comprising:
[0062] processor;
[0063] a memory for storing processor-executable instructions;
[0064] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0065] According to a fourth aspect of the embodiments of the present invention,
[0066] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0067] In an embodiment of the present invention, numerical and categorical features are processed separately by a dual-stream feature encoder, and adaptive fusion is performed using a learnable gating mechanism, thereby effectively improving the processing capability of different types of data; dynamic sampling is performed based on the importance score of the fused features to generate an enhanced training sample set, which significantly improves the generalization ability and robustness of the model; a feature extraction network using a hierarchical attention module and a self-calibration module is used to capture the complex relationship between features through a feature affinity graph and a graph segmentation algorithm, and a global feature representation is extracted using a causal attention mechanism and cross-subgraph attention fusion; the self-calibration module further optimizes the feature representation through residual connections and dynamic normalization, thereby effectively improving the accuracy of feature extraction and the expressiveness of the model; an information bottleneck feature filter is introduced to minimize redundant information while retaining task-related information, thereby improving the efficiency of feature selection; the design of a dual-branch network not only ensures the accuracy of the prediction, but also provides regularization constraints through an auxiliary reconstruction branch, thereby enhancing the stability of the model; a structured gradient method is used to calculate the feature contribution, thereby providing explainability for model decisions, and helping users understand and trust the prediction results of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flow chart of a structured data modeling and analysis method based on Transformer according to an embodiment of the present invention;
[0069] Figure 2 It is a schematic diagram of the structure of a Transformer-based structured data modeling and analysis system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0071] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0072] Figure 1 FIG. 1 is a flow chart of a structured data modeling and analysis method based on Transformer according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0073] S101. Receive structured raw data, divide the structured raw data into numerical features and categorical features, perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, and perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; perform adaptive weight allocation through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; perform dynamic sampling based on the importance score of the fused features to generate an enhanced training sample set;
[0074] In this embodiment, numerical features and categorical features are processed separately through a dual-stream feature encoder, and adaptive encoding methods can be used for different feature types to improve the accuracy of feature expression; adaptive weights are assigned to different feature types through a learnable gating mechanism to achieve effective fusion of numerical features and categorical features, make full use of multi-source information, and improve the expressiveness of the model; dynamic sample optimization: dynamic sampling based on the importance score of the fused feature helps to balance the sample distribution and highlight key samples, thereby generating a more representative enhanced training sample set and improving the training effect and generalization ability of the model; the overall method is optimized in the data preprocessing and feature representation stages to provide more accurate and efficient feature input for downstream tasks, further improving the learning efficiency and performance of the model.
[0075] S102. Input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through a cross-subgraph attention fusion mechanism to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation;
[0076] In this embodiment, a feature affinity graph is constructed and divided into multiple feature subgraphs through a hierarchical attention module, which can capture the structured associations between features and improve the modeling ability of internal relationships of samples; the causal attention mechanism is used to extract local feature representations, and the local features are integrated through the attention fusion mechanism across subgraphs, which can effectively combine local details and global patterns and improve the comprehensiveness of feature expression; the calibration module realizes nonlinear transformation and adaptive adjustment of global feature representation through residual connection and dynamic normalization, which helps to strengthen key features and suppress redundant information, and improve the representation ability of deep features; the overall network architecture combined with hierarchical attention and self-calibration modules can more efficiently extract multi-dimensional features of complex samples and provide deep feature representations with high recognition for downstream tasks; the feature extraction network enhances the modeling ability of complex data relationships through graph segmentation, attention mechanism and dynamic calibration, and improves the adaptability and generalization ability of the model to diversified inputs.
[0077] S103. Construct an information bottleneck feature filter, filter the deep feature representation by maximizing task-related information and minimizing redundant information, and obtain a key feature subset; input the key feature subset into a two-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on the attention mechanism to obtain a feature weight vector, and generates a prediction result according to the feature weight vector, and the auxiliary reconstruction branch provides a regularization constraint by reconstructing the key feature subset; according to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result, and an analysis result is generated.
[0078] First, an information bottleneck feature filter is constructed to analyze and filter the deep feature representation. The mutual information value between each feature dimension and the prediction target is calculated. The higher the mutual information value, the stronger the relevance of the dimension feature to the prediction task. At the same time, the redundant information between feature dimensions is calculated to discover and remove redundant features with overlapping information. Through an iterative optimization process, a subset of key features is filtered out from the complete deep feature representation under the constraints of maximizing task-related information and minimizing feature redundant information.
[0079] The selected subset of key features is input into the dual-branch network for processing. In the main prediction branch, a multi-head attention mechanism is used to weight the features. The multi-head attention layer contains multiple parallel attention calculation units, and each attention head independently calculates the feature weight. In this way, the model can capture the correlation pattern between features from different angles. The outputs of all attention heads are merged to obtain a comprehensive feature weight vector. Based on this weight vector, the key features are weighted and combined, and the prediction results are generated through the fully connected layer.
[0080] In the auxiliary reconstruction branch, a symmetric encoder-decoder structure is constructed. The encoder maps key features to a low-dimensional latent space through a multi-layer fully connected network, and the decoder reconstructs the latent space representation back to the original feature space. This reconstruction task serves as an auxiliary supervisory signal to constrain the model's learning process by minimizing the reconstruction error. The reconstruction loss and the prediction loss together constitute the optimization target of the model, so that the learned feature representation is both beneficial to the prediction task and maintains the structural information of the original data.
[0081] Finally, the structured gradient method is used to interpret and analyze the prediction results. Combined with the feature weight vector obtained from the main branch, the contribution of each input feature to the prediction result is calculated. The feature contribution reflects the importance of different features in the decision-making process and can be used to understand the prediction basis of the model. Based on the contribution analysis, an interpretable result report is generated to help users understand the key influencing factors behind the prediction.
[0082] In this embodiment, the information bottleneck feature filter is used to maximize task-related information and minimize redundant information, extract key feature subsets from the deep feature representation, and significantly improve the efficiency and accuracy of feature selection; the main prediction branch assigns weights to the key feature subsets based on the attention mechanism, and generates more accurate prediction results by focusing on the features that have a greater impact on the prediction task; the auxiliary reconstruction branch introduces regularization constraints by reconstructing the key feature subsets to prevent model overfitting, while ensuring the integrity and expressiveness of the key feature subsets; the structured gradient method is used to quantify the contribution of each feature in the key feature subset to the prediction result, provide support for model interpretability and explainability, and help understand the decision-making basis of the model; through screening and optimization, features that are highly relevant to the task objectives are retained and strengthened, thereby improving the task adaptability and practical application value of the model.
[0083] In an optional implementation, receiving structured raw data, dividing the structured raw data into numerical features and categorical features, performing piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, and performing adaptive quantum coding on the categorical features to obtain continuous value mapping features; performing adaptive weight allocation through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; and performing dynamic sampling based on the importance score of the fused features to generate an enhanced training sample set includes:
[0084] Receiving structured raw data, dividing the structured raw data into numerical features and categorical features; inputting the numerical features and categorical features into a dual-stream feature encoder for feature processing, wherein the dual-stream feature encoder includes a feature trunk stream and a feature mapping stream;
[0085] In the feature trunk stream, the data distribution of the numerical features is counted to obtain the data density, the position of the optimal segmentation point is determined based on the data density, the numerical features are divided into a plurality of segmentation intervals according to the optimal segmentation point, a linear transformation parameter matrix is constructed in each segmentation interval, the feature difference values between adjacent segmentation intervals are calculated, the optimized linear transformation parameter matrix of each segmentation interval is determined by minimizing the feature difference values, and the optimized linear transformation parameter matrix is used to perform linear transformation on the numerical features of each segmentation interval to obtain normalized trunk features;
[0086] In the feature mapping flow, each category value in the category feature is mapped to a corresponding quantum ground state, an initial amplitude parameter and a phase parameter are assigned to each quantum ground state, a quantum gate sequence including a rotation gate and a control gate is constructed, the quantum gate sequence is sequentially applied to the quantum ground state, a transformed quantum state is obtained through quantum state transformation, an entropy value of the transformed quantum state is calculated, and the transformed quantum state is projected to a continuous space based on the entropy value to obtain a continuous value mapping feature;
[0087] Respectively calculating the correlation coefficients between the normalized backbone features and the dimensions of the continuous value mapping features, combining the correlation coefficients to construct a feature correlation coefficient matrix, training a preset attention network based on the feature correlation coefficient matrix to obtain an importance vector of the feature dimension, and using the importance vector to perform a weighted combination of the normalized backbone features and the continuous value mapping features to obtain a fusion feature;
[0088] The contribution of each dimension of the fusion feature to the target task is calculated, and a feature importance score of each dimension is generated based on the contribution. A sampling probability distribution function is constructed according to the feature importance score, and the sampling probability distribution function is used to perform importance sampling on the training samples to generate an enhanced training sample set.
[0089] In a specific implementation, firstly, structured raw data is received, and features are divided according to the data type. Taking the user purchase intention prediction task of an e-commerce platform as an example, the raw data includes user basic information, behavior data, and transaction records. Numerical features include: age (18-80), number of visits in the past 30 days (0-120), total consumption in the past 90 days (0-50000), number of days since the last purchase (0-365), number of favorite items (0-100), number of shopping cart items (0-50), etc.; category features include: gender (male / female), membership level (ordinary / silver / gold / diamond), commonly used device types (PC / Android / iOS), preferred shopping categories (clothing / digital / beauty / books, etc.), delivery area, etc.
[0090] In the feature main stream, the numerical features are first distributed and segmented: Taking the user age feature as an example, based on 500,000 user samples, the data distribution density of age is obtained. By setting the basic statistical interval width to 2 years, it is initially divided into 31 basic intervals (18-20, 20-22, ..., 78-80). The sample frequency ratio of adjacent intervals is calculated, and the intervals are merged when the ratio is greater than 0.85. For example, the frequencies of the three intervals 20-22, 22-24, and 24-26 are 15000, 14500, and 14800 respectively, and the two groups of adjacent ratios are 0.97 and 0.98 respectively, which meet the merging conditions and are merged into the 20-26 interval. After merging, the seven segmented intervals of 18-20, 20-26, 26-35, 35-45, 45-55, 55-65, and 65-80 are finally obtained.
[0091] Construct a linear transformation parameter matrix for each segmented interval: Take the 20-26 age interval as an example, construct a 2×2 linear transformation parameter matrix, and the initial value is the unit matrix. Calculate the feature difference value with the adjacent interval 26-35, and make the difference value lower than the threshold value of 0.01 through iterative optimization. The final transformation matrix is applied to the age value in this interval to achieve the normalized transformation of the feature. For example, the original age of 24 years old is transformed to a normalized value of 0.37. The same method is used for other numerical features to obtain a complete normalized backbone feature set.
[0092] In the feature mapping flow, quantum coding conversion is performed on the categorical features: Taking the membership level feature as an example, the four levels are first mapped to the corresponding quantum ground states: ordinary members are mapped to |00>, silver members are mapped to |01>, gold members are mapped to |10>, and diamond members are mapped to |11>. Assign initial parameters to each ground state: the amplitude is set to 0.5, and the phase angle is set to 0. Construct a sequence of 8 quantum gates, which are: X gate, CNOT gate, Z gate, X gate, CNOT gate, Z gate, X gate, CNOT gate. Apply this gate sequence to the ground state in turn to obtain the transformed quantum state. Calculate the entropy value of the quantum state. Ordinary members get an entropy value of 0.25, silver members get 0.45, gold members get 0.65, and diamond members get 0.85. Based on the entropy value, the quantum state is projected into a two-dimensional continuous space to obtain a continuous value representation of the membership level. For example, ordinary members are mapped to (0.2, 0.3), silver members are mapped to (0.4, 0.5), etc.
[0093] Perform feature correlation analysis and fusion: Calculate the correlation coefficients between all features. For example, the correlation coefficient between age and total consumption is 0.35, the correlation coefficient between membership level and number of visits is 0.62, etc. Organize these correlation coefficients into a correlation matrix with the matrix dimension of total number of features x total number of features. Based on this matrix, train an attention network with a three-layer structure, with 128, 64, and 32 hidden units, respectively, and use the ReLU activation function. The importance vector of the feature dimension is obtained through network training, for example, the importance of the age feature is 0.28, the importance of the total consumption is 0.35, the importance of the membership level is 0.31, etc. Use these importance values to perform a weighted combination of the normalized backbone features and the continuous value mapping features to obtain the final fused features.
[0094] Finally, the sample enhancement process is performed: the contribution of each dimension in the fusion feature to the purchase intention prediction is calculated. Based on 500,000 training samples, the contribution values of each feature dimension are calculated: age feature 0.25, total consumption 0.38, membership level 0.32, number of visits 0.28, etc. These contribution values are normalized as sampling probabilities, and the sampling probability distribution function is constructed. Importance sampling is performed on the original training sample set. Samples with a sampling probability exceeding 0.8 are replicated 3 times, samples with a probability between 0.5-0.8 are replicated 2 times, and samples with a probability between 0.3-0.5 are replicated 1 time. In this way, an enhanced training sample set is generated, and the final sample size reaches 1.8 times that of the original data set, about 900,000 sample records.
[0095] In this embodiment, the numerical features and categorical features are processed separately by a dual-stream feature encoder, thereby achieving efficient encoding of different types of features. The piecewise linear transformation maintains the continuity and local features of the numerical features, while the adaptive quantum coding makes full use of the advantages of quantum computing and improves the expressiveness of categorical features. The learnable gating mechanism is used for feature fusion to achieve dynamic weight allocation between features. The feature correlation is learned through the attention network, and the importance of different features is adaptively adjusted, thereby improving the adaptability of the model to different feature combinations. The dynamic sampling strategy based on feature importance effectively improves the quality of training samples. By calculating the contribution of features to the target task, a reasonable sampling probability distribution is constructed, and more representative training samples are generated, thereby improving the generalization ability and robustness of the model.
[0096] In an optional implementation, statistically analyzing the data distribution of the numerical feature to obtain data density, and determining the position of the optimal segmentation point based on the data density includes:
[0097] Based on the interquartile range and the number of samples of the numerical feature, a basic interval width is calculated, and based on the basic interval width, the value range of the numerical feature is divided into a plurality of basic statistical intervals, and the number of samples in each basic statistical interval is counted to obtain an interval frequency sequence;
[0098] Calculating the frequency ratios of adjacent basic statistical intervals in the interval frequency sequence to obtain interval similarity, merging adjacent basic statistical intervals whose interval similarity is greater than a preset merging threshold into new statistical intervals to obtain a merged frequency distribution sequence;
[0099] Calculating the cumulative frequency distribution according to the combined frequency distribution sequence, and generating a data density function based on the cumulative frequency distribution; selecting candidate segmentation positions in the combined frequency distribution sequence, calculating the number of samples and the numerical mean of the intervals on both sides of each candidate segmentation position, and calculating the interval variance value of the candidate segmentation position using the sample number and the numerical mean;
[0100] Selecting a position with a maximum interval variance value in the current interval to be segmented as a pending segmentation point, calculating the ratio of the interval variance value of the pending segmentation point to the overall variance of the numerical feature, and determining the variance contribution ratio of the pending segmentation point;
[0101] The local distribution skewness is calculated for the intervals on both sides of the pending segmentation point. When the difference of the local distribution skewness is greater than the preset skewness threshold and the variance contribution ratio is greater than the preset segmentation threshold, the pending segmentation point is determined as the optimal segmentation point. The process is repeated to determine the positions of all optimal segmentation points.
[0102] The local distribution skewness specifically refers to an indicator that measures whether the data distribution in a certain interval is symmetrical, and is used to describe the skew direction and degree of the data distribution in the interval. If the data distribution is symmetrical, the skewness is zero; if the data distribution is skewed to the right, it means that larger values appear less frequently but the tail is elongated, and the skewness is positive; if the data distribution is skewed to the left, it means that smaller values are rarer but form a longer tail, and the skewness is negative. The local distribution skewness focuses on the distribution characteristics within a certain interval, rather than the distribution of the overall data, so it can help to more carefully analyze the distribution characteristics of data in different intervals.
[0103] In a specific implementation, firstly, a statistical analysis is performed on the data distribution of the numerical feature. The overall value range is determined by calculating the maximum and minimum values of the numerical feature. The width of the basic statistical interval is set based on the interquartile range of the data and the total sample size. For example, for a data set with a sample size of 10,000, where the interquartile range of the age feature is 20 years old, the basic interval width can be set to five years old. According to this basic interval width, the value range is divided into several basic statistical intervals, and the number of samples in each interval is counted to obtain the initial interval frequency sequence.
[0104] Next, analyze the data distribution similarity of adjacent basic statistical intervals. Calculate the frequency ratio of adjacent intervals and set the merge threshold to 0.8. Taking age characteristics as an example, if the number of samples in the interval of 20 to 25 is 1,000, and the number of samples in the interval of 25 to 30 is 900, then the similarity of these two intervals is 0.9, which is greater than the merge threshold, and these two intervals can be merged into a new statistical interval. Process all adjacent intervals in turn to obtain the merged frequency distribution sequence.
[0105] Based on the combined frequency distribution sequence, the cumulative frequency of each interval is calculated and the data density function is constructed. The positions with obvious frequency changes in the frequency distribution sequence are selected as candidate segmentation positions. For each candidate position, the number of samples and the numerical mean of the intervals on the left and right sides are calculated respectively. The interval variance value of the position is calculated in this way to characterize the distinguishing effect of the position as a segmentation point.
[0106] In the interval that needs to be segmented, select the position with the maximum interval variance value as the pending segmentation point. Calculate the ratio of the interval variance value corresponding to the pending segmentation point to the variance of the feature population to obtain the variance contribution ratio. At the same time, calculate the data distribution skewness of the intervals on both sides of the pending segmentation point. Set the skewness threshold to 0.5 and the segmentation threshold to 0.3. When the skewness difference between the intervals on both sides is greater than the skewness threshold, and the variance contribution ratio is greater than the segmentation threshold, confirm that the pending segmentation point is the optimal segmentation point.
[0107] Repeat the above analysis process until the segmentation effect of all intervals does not meet the threshold requirement, and the positions of all optimal segmentation points can be obtained. Taking age as an example, possible optimal segmentation points include 18 years old, 35 years old, 50 years old, etc. These segmentation points divide the sample data into multiple intervals with significant statistical feature differences.
[0108] In this embodiment, intervals are merged by calculating the frequency ratio of adjacent intervals, which avoids the interference of data sparse intervals on the segmentation results and improves the stability and reliability of the segmentation results; interval variance value and variance contribution ratio are used as segmentation point evaluation indicators, combined with the difference constraint of interval distribution skewness, to ensure the statistical significance of the segmentation point position, so that the segmentation results have good discrimination; the optimal segmentation point is adaptively determined based on the data distribution characteristics, without the need to manually specify the number of segments, avoiding the influence of subjective factors and improving the applicability of the segmentation method in different scenarios.
[0109] In an optional implementation, the enhanced training sample set is input into a feature extraction network including a hierarchical attention module and a self-calibration module, the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through an attention fusion mechanism across subgraphs to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation, including:
[0110] Inputting the enhanced training sample set into the convolutional neural network to obtain an initial feature map, calculating the dot product of any two sample features in the initial feature map, and dividing the dot product by the module-length product of the corresponding feature vectors to obtain a feature similarity matrix; constructing a feature affinity graph based on the feature similarity matrix, wherein the vertices of the feature affinity graph correspond to the sample features, and the edge weights of the feature affinity graph correspond to the similarity values in the feature similarity matrix, and dividing the feature affinity graph into a plurality of feature subgraphs by a spectral clustering algorithm;
[0111] In each of the feature subgraphs, the feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction results of the query feature sequence and the key-value pair feature sequence are masked by using the temporal causal mask matrix to obtain an initial attention weight containing only historical feature information; a spatial distance matrix of sample features in each of the feature subgraphs is calculated, and the spatial distribution entropy of the feature subgraph is calculated based on the spatial distance matrix, and the initial attention weight is adaptively adjusted according to the spatial distribution entropy to obtain an adaptive attention weight; the key-value pair feature sequence is weighted based on the adaptive attention weight, and the position encoding information is superimposed to obtain a local feature representation of each of the feature subgraphs;
[0112] Based on a pre-built multi-layer perceptron network, a concatenated vector of local feature representations of adjacent feature subgraphs is input, and the concatenated vector is forward-propagated through the multi-layer perceptron network to obtain association weights between feature subgraphs, and the local feature representations are weighted summed based on the association weights to obtain a global feature representation;
[0113] The mean and standard deviation of the global feature representation are calculated in the feature dimension respectively; based on the pre-constructed dynamic normalization network, the global feature representation is input into the dynamic normalization network, and the feature transformation coefficient and feature offset are obtained through forward propagation; the global feature representation is subtracted from the mean and divided by the standard deviation to obtain a standardized feature, the standardized feature is multiplied by the feature transformation coefficient, and the feature offset is added to obtain a transformed feature, and the transformed feature is added to the global feature representation to obtain a self-calibration feature;
[0114] The self-calibration features are subjected to layer normalization processing and nonlinear transformation through a ReLU activation function to obtain a deep feature representation.
[0115] In a specific implementation, first, data enhancement is performed on the input training sample set. Data enhancement includes random cropping, horizontal flipping, brightness and contrast adjustment, etc., to generate an enhanced training sample set. The number of samples after enhancement is 4 times that of the original samples, and each original sample generates 3 new samples through different enhancement methods.
[0116] The enhanced training samples are input into the convolutional neural network for feature extraction. ResNet50 is used as the basic network to extract multi-scale features. The size of the feature map output in the last convolutional layer is 14x14, and the number of channels is 2048. For each sample, the feature map is flattened into a one-dimensional vector to obtain the initial feature map.
[0117] When calculating the feature similarity matrix, dot product operations are performed on any two sample feature vectors and normalized by dividing them by the product of their respective modulus lengths. For example, for two 2048-dimensional feature vectors, their cosine similarity is calculated, and the obtained similarity value ranges from -1 to 1. Based on the similarity matrix, a feature affinity graph is constructed, in which the vertices represent sample features and the edge weights are the corresponding similarity values.
[0118] The feature affinity graph is divided using the spectral clustering algorithm. First, the Laplacian matrix of the graph is calculated, and the eigenvectors corresponding to its smallest k eigenvalues are selected, where k is the preset number of subgraphs, usually set between 4 and 8. K-means clustering is performed on the eigenvectors to divide the samples into different feature subgraphs.
[0119] Implement the causal attention mechanism in each feature subgraph. Divide the feature sequence into a query sequence and a key-value pair sequence. Construct an upper triangular causal mask matrix to ensure that only historical information can be accessed at the current moment. Calculate the Euclidean distance between features to construct a spatial distance matrix, and calculate the normalized spatial distribution entropy. Adaptively adjust the attention weight according to the entropy value. The larger the entropy value, the more dispersed the feature distribution, and the smaller the corresponding attention weight.
[0120] For attention fusion between feature subgraphs, a three-layer multi-layer perceptron network is constructed. The hidden layer dimensions are 1024, 512, and 256 respectively. The local feature concatenation vectors of adjacent subgraphs are input and the association weights between them are output. All local features are weighted summed according to the association weights to obtain the global feature representation.
[0121] In the self-calibration module, the statistics of the global features in the feature dimension are first calculated. A two-layer dynamic normalization network is constructed with a hidden layer dimension of 1024. The network outputs feature transformation coefficients and offsets for adaptive adjustment of the standardized features. The adjusted features are added to the original features to achieve residual connection.
[0122] Finally, the self-calibration features are layer normalized and the final deep feature representation is obtained through the ReLU activation function. The feature dimension is kept at 2048 for easy use in subsequent tasks.
[0123] For example, taking the task of predicting the purchase intention of users on an e-commerce platform as an example, 900,000 enhanced training samples are first input into the convolutional neural network. The convolutional network contains 3 convolutional layers, with convolution kernel sizes of 3×3, 3×3, and 5×5, and output channels of 64, 128, and 256, respectively. After being processed by the convolutional network, an initial feature map with a dimension of 256 is obtained. The dot product similarity is calculated for any two sample features, and normalized to obtain a feature similarity matrix of 900,000×900,000.
[0124] Construct a feature affinity graph based on the feature similarity matrix. Each sample feature is used as a vertex of the graph, and the similarity value between samples is used as the weight of the edge. For example, if the similarity between the feature vector of user A and the feature vector of user B is 0.85, the edge weight between these two vertices in the affinity graph is 0.85. Use the spectral clustering algorithm to divide the feature affinity graph into 8 feature subgraphs, each of which contains a group of users with similar behavior patterns.
[0125] Feature sequence processing is performed inside the feature subgraph. Taking a subgraph containing 100,000 users as an example, the feature sequences are arranged in chronological order, and the sequences are divided into query feature sequences and key-value feature sequences based on the feature similarity threshold of 0.7. A temporal causal mask matrix of size 100,000 × 100,000 is constructed, in which only historical information before the current moment is retained and information at future moments is masked. The mask matrix is applied to the interaction results of the query sequence and the key-value sequence to obtain the initial attention weight.
[0126] Calculate the spatial distance matrix of sample features within the subgraph. For each pair of samples, calculate the Euclidean distance based on their feature vectors and construct a distance matrix of size 100,000 × 100,000. Calculate the spatial distribution entropy based on the distance matrix. The entropy value reflects the uniformity of feature distribution. For example, the spatial distribution entropy of a subgraph is 0.75, indicating that the feature distribution is relatively uniform; the entropy value of another subgraph is 0.45, indicating that the feature distribution is relatively concentrated. Adjust the initial attention weight according to the entropy value: for subgraphs with uniform distribution, keep the initial weight; for subgraphs with concentrated distribution, increase the attention weight of the local area. The obtained adaptive attention weight is used to weight the key-value pair feature sequence.
[0127] In order to preserve the position information of the feature, a position encoding vector is constructed. The dimension of the position encoding is the same as the feature dimension, which is 256. The position encoding is added to the weighted feature sequence to obtain the local feature representation of each subgraph. For example, the local feature representation of a subgraph captures the common features of the user group in recent shopping behavior.
[0128] A four-layer multilayer perceptron network is constructed, with the number of neurons in each layer being 512, 256, 128, and 64. The concatenated vector of the local feature representation of two adjacent feature subgraphs is input, and the association weight between the subgraphs is obtained through network forward propagation. For example, the association weight between the subgraph representing the young user group and the subgraph representing the middle-aged user group is 0.6, indicating that the shopping behaviors of these two groups are somewhat correlated. The local feature representations of all subgraphs are weighted summed based on the association weights to obtain a global feature representation with a dimension of 256.
[0129] The global feature representation is self-calibrated. First, the mean and standard deviation are calculated on the feature dimension. A dynamic normalization network is constructed, which contains two fully connected layers with 128 and 256 neurons respectively. The global feature representation is input into the network to obtain the feature transformation coefficient and feature offset. After the global feature representation is normalized, it is multiplied by the transformation coefficient and added with the offset to obtain the transformed feature. The transformed feature is added to the original global feature representation to obtain the self-calibrated feature.
[0130] Finally, the self-calibrated features are normalized and transformed nonlinearly through the ReLU activation function to obtain the final deep feature representation. The deep feature representation dimension is kept at 256, where each dimension contains the information of the original features after multi-level transformation and fusion. These features not only retain the behavioral characteristics of local user groups, but also contain the association information between different groups, providing high-quality feature representation for subsequent feature selection and prediction tasks.
[0131] In this embodiment, through the hierarchical attention mechanism and feature subgraph division, the local correlation between samples can be better captured and the discriminative ability of feature representation can be improved. The causal attention within the subgraph ensures the effective use of temporal information and avoids information leakage; the introduction of spatial distribution entropy to adaptively adjust the attention weight makes the attention allocation more reasonable and can highlight the contribution of important features. The correlation weight learning between feature subgraphs realizes the effective fusion of multi-scale features; the self-calibration module enhances the nonlinear expression ability of features through dynamic normalization and residual connection while maintaining the original information. This design improves the generalization performance of the model and makes the extracted features more robust.
[0132] In an optional implementation, the feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction result of the query feature sequence and the key-value pair feature sequence is masked by using the temporal causal mask matrix, so that the initial attention weight containing only historical feature information is obtained, including:
[0133] Based on the similarity of the features of adjacent time steps in the feature subgraph, the feature sequence of the feature subgraph is grouped to obtain a query feature sequence and a key-value pair feature sequence, wherein the time step features whose similarity of the features of adjacent time steps is greater than a dynamic threshold determined based on the global statistical characteristics of the feature sequence are divided into the query feature sequence, and the remaining features are divided into the key-value pair feature sequence, to obtain the feature division results of the query feature sequence and the key-value pair feature sequence;
[0134] Based on the feature division result, the temporal correlation strength of the query feature sequence and the key-value pair feature sequence are respectively calculated by a recursive neural network, wherein the hidden state of the recursive neural network contains the long-term and short-term dependency information of the query feature sequence and the key-value pair feature sequence, and the temporal correlation strength of the query feature sequence and the temporal correlation strength of the key-value pair feature sequence are obtained;
[0135] Based on the temporal position relationship of the feature sequence, a first mask matrix is constructed as a temporal causal mask matrix, and the temporal association strength of the query feature sequence and the temporal association strength of the key-value pair feature sequence are converted into soft mask weights to obtain a second mask matrix;
[0136] Performing matrix multiplication on the query feature sequence and the key-value pair feature sequence to obtain an original attention score; performing addition operation on the original attention score and the first mask matrix to obtain a first mask result; performing element-wise multiplication operation on the first mask result and the second mask matrix to obtain a double-masked attention score; and normalizing the double-masked attention score to obtain an initial attention weight that only contains historical feature information.
[0137] The soft mask weight specifically refers to a weight distribution method that is dynamically adjusted based on the importance of the feature, which is used to emphasize the contribution of different features to the overall calculation when processing feature sequences. It combines the temporal correlation strength of the feature and other contextual information to generate a weight between 0 and 1 to indicate the degree to which a feature should be paid attention to in the current calculation.
[0138] The soft mask weight is characterized by continuous adjustment. Compared with the hard mask (completely shielded or completely passed), it can partially reduce or enhance the influence of specific features in a flexible way. This dynamic weight can adaptively allocate attention resources during the model operation process according to the characteristics of the data, thereby more accurately capturing the temporal relationship and dependency between features.
[0139] In a specific implementation, the feature sequence in the feature subgraph is first dynamically divided. The similarity of the features is measured by calculating the cosine similarity between the features of adjacent time steps. Specifically, for any two adjacent time step features in the feature sequence, their cosine similarity values are calculated. At the same time, a dynamic threshold is set based on the statistical characteristics of the mean and standard deviation of the feature sequence. When the similarity of the features of adjacent time steps is greater than the dynamic threshold, these features are divided into the query feature sequence, and the remaining features are divided into the key-value pair feature sequence. For example, for a feature sequence with a length of 10, if the feature similarity of the 2nd-4th and 7th-8th time steps is higher than the threshold value 0.8, the features of these time steps are divided into the query feature sequence, and the features of the remaining time steps are divided into the key-value pair feature sequence.
[0140] Then, the long short-term memory network is used to process the query feature sequence and the key-value pair feature sequence respectively to extract their temporal association information. The long short-term memory network selectively retains and updates the feature information through a gating mechanism, and its hidden state contains the long-term dependency and short-term correlation of the feature sequence. For the query feature sequence, the hidden state output of the network represents its temporal association strength; for the key-value pair feature sequence, the hidden state output of the network represents its temporal association strength.
[0141] Then a double masking mechanism is constructed. The first masking constructs a causal mask matrix based on the temporal position relationship of the features to ensure that only historical information can be accessed at the current time step. In the specific implementation, the value of the upper triangular area is set to negative infinity, and the value of the lower triangular area is set to 0. The second masking converts the temporal correlation strength of the query sequence and the key-value pair sequence into a soft mask weight between 0 and 1 to form a soft mask matrix.
[0142] Finally, the attention weight is calculated. Perform matrix multiplication on the query feature sequence and the key-value pair feature sequence to obtain the original attention score, and add it to the causal mask matrix to obtain the first mask result. Multiply the first mask result by the soft mask matrix element by element to obtain the double-masked attention score. Softmax normalize the score to obtain the final attention weight. This attention weight only contains historical time series information and can be used for subsequent feature fusion.
[0143] For example, taking the analysis of e-commerce user behavior sequence as an example, the user behavior sequence in the feature subgraph is processed. First, the similarity of the features of adjacent time steps in the feature sequence is calculated. Assume that a user's behavior sequence is within 30 consecutive days, and the dimension of the daily behavior feature vector is 256. For the behavior features of any two consecutive days, their cosine similarity values are calculated.
[0144] Determine the dynamic threshold based on the global statistical characteristics of the feature sequence. Count the mean and standard deviation of the feature similarity of all adjacent time steps within 30 days, for example, the mean is 0.65 and the standard deviation is 0.15. The mean plus 0.5 times the standard deviation is used as the dynamic threshold. In this example, the threshold value is 0.725. Compare the similarity values of adjacent time steps with the threshold, and divide the time step features with similarity greater than the threshold into the query feature sequence, and the remaining features into the key-value pair feature sequence.
[0145] For example, the similarity of the behavior features on the 3rd and 4th days is 0.78, which is greater than the threshold of 0.725, so the features of the 4th day are classified into the query feature sequence; the similarity of the features on the 15th and 16th days is 0.68, which is less than the threshold, so the features of the 16th day are classified into the key-value feature sequence. After the division, a query feature sequence containing 12 time steps and a key-value feature sequence containing 18 time steps are obtained.
[0146] A bidirectional long short-term memory network (Bi-LSTM) is constructed as a recurrent neural network to calculate the temporal association strength. The network contains a two-layer Bi-LSTM structure with 128 hidden units. The query feature sequence is input into the network, and the hidden state of the network contains the long-term and short-term dependency information of the feature sequence. For the query feature sequence of 12 time steps, a temporal association strength matrix with a dimension of 12×128 is obtained. Similarly, the key-value pair feature sequence of 18 time steps is input into the network, and a temporal association strength matrix with a dimension of 18×128 is obtained.
[0147] The temporal causal mask matrix is constructed based on the temporal position relationship of the feature sequence. The matrix size is 30×30, corresponding to the complete 30-day sequence. In the matrix, only the information of the current time step and the previous time step is retained, and the position of the subsequent time step is marked as negative infinity as the first mask matrix. For example, for the prediction of the 10th day, only the feature information from the 1st to the 10th day is considered, and the information from the 11th to the 30th day is masked.
[0148] The temporal correlation strength between the query feature sequence and the key-value pair feature sequence is converted into a soft mask weight. The correlation strength is normalized by Softmax to obtain the second mask matrix. This matrix reflects the degree of correlation between features at different time steps. The time step with greater correlation strength has a greater mask weight.
[0149] The query feature sequence and the key-value pair feature sequence are matrix multiplied to obtain the original attention score matrix. The matrix dimension is the query sequence length multiplied by the key-value pair sequence length, that is, 12×18. This matrix is added to the first mask matrix to implement temporal causal masking, ensuring that only historical information is used when predicting. The masked result is then element-wise multiplied with the second mask matrix to obtain a dual mask attention score that considers both temporal causality and feature association strength.
[0150] Finally, the double-masked attention scores are normalized by Softmax to obtain the initial attention weights. These weights reflect the degree of correlation between each time step in the query sequence and the historical time step in the key-value sequence. For example, the attention weight of the query feature on the 10th day and the key-value feature on the 8th day is 0.15, indicating that the features of these two time steps have a strong correlation.
[0151] In this embodiment, a dynamic feature sequence division mechanism is used to adaptively divide the query sequence and the key-value pair sequence according to the feature similarity, thereby improving the discrimination of the feature representation and enhancing the model's perception of important features. A double mask mechanism is adopted to ensure strict temporal causality and introduce a flexible feature selection mechanism through soft mask weights, thereby effectively suppressing the interference of irrelevant features and improving the robustness of the model. The temporal correlation strength is extracted based on a recursive neural network, making full use of the long-term and short-term dependency information of the feature sequence, enhancing the model's ability to model temporal patterns, and improving the temporal consistency of the feature representation.
[0152] In an optional implementation, an information bottleneck feature filter is constructed to filter the deep feature representation by maximizing task-related information and minimizing redundant information, and the key feature subset obtained includes:
[0153] Constructing an information bottleneck encoder, the information bottleneck encoder includes a three-layer perceptron structure, each layer of the perceptron structure is connected to a normalization layer, the deep feature representation is input into the information bottleneck encoder, and the output layer of the information bottleneck encoder generates probability distribution parameters of the compressed feature space; encoding and mapping the deep feature representation based on the probability distribution parameters to obtain a compressed feature representation;
[0154] Constructing a mutual information evaluation module, the mutual information evaluation module includes a three-layer perceptron structure, inputting a concatenated vector of the compressed feature representation and the label of the target task into the mutual information evaluation module, and outputting the mutual information amount between the compressed feature representation and the target task;
[0155] Calculating an autocorrelation matrix of the compressed feature representation, wherein each element in the autocorrelation matrix is the product of the corresponding feature vector minus the corresponding mean divided by the product of the corresponding standard deviation; calculating feature redundancy based on the difference between the autocorrelation matrix and the identity matrix;
[0156] Constructing a feature selection network, the feature selection network comprising a first fully connected layer and a second fully connected layer, inputting the compressed feature representation into the first fully connected layer, the output of the first fully connected layer being input into the second fully connected layer after passing through an activation function, the output of the second fully connected layer being passed through a gating function to obtain a gating value, and performing element-wise multiplication of the gating value and the compressed feature representation to obtain an initial screening feature;
[0157] Adjusting the output of the gating function based on a temperature parameter, wherein the temperature parameter decreases with the number of training rounds based on a preset decreasing amount, and performing element-by-element multiplication of the adjusted gating value and the compressed feature representation to obtain an adjusted screening feature;
[0158] Constructing an objective function by linearly combining the inverse of the mutual information, the feature redundancy, and the norm of the adjusted screening feature;
[0159] By iteratively optimizing the objective function, the optimization parameters of the information bottleneck encoder and the optimization parameters of the feature selection network are obtained, and an optimized information bottleneck encoder and an optimized feature selection network are obtained;
[0160] The deep feature representation is input into an optimized information bottleneck encoder to obtain an optimized compressed feature representation, and the optimized compressed feature representation is input into an optimized feature selection network to obtain a key feature subset.
[0161] In a specific implementation, the construction process of the information bottleneck feature filter first requires the construction of an information bottleneck encoder. The encoder adopts a three-layer perceptron structure, and the number of neurons in each layer of perceptrons is 1024, 512 and 256 respectively, using ReLU as the activation function. Each layer of perceptrons is followed by a normalization layer to accelerate training convergence. The input deep feature representation dimension is 2048, and after being processed by the encoder, two sets of parameters are output: a mean vector and a variance vector, which are used to construct the probability distribution of the compressed feature space. Based on the reparameterization technology, a compressed feature representation with a dimension of 256 is sampled from the probability distribution.
[0162] The mutual information evaluation module also uses a three-layer perceptron structure with 512, 256, and 1 neurons. The compressed feature representation and the task label are concatenated as input, and the label is encoded in one-hot format. Through nonlinear transformation, a scalar value is finally output to represent the mutual information estimation between the compressed feature and the task.
[0163] The calculation of feature redundancy first requires the normalization of the compressed features. The mean and standard deviation of each feature dimension are calculated, and the feature is subtracted from the mean and divided by the standard deviation. Then the inner product between the standardized features is calculated to obtain the autocorrelation matrix. The sum of the squares of the difference between this matrix and the unit matrix is the feature redundancy measure.
[0164] The feature selection network contains two fully connected layers. The number of neurons in the first layer is 512, and the ReLU activation function is used. The number of neurons in the second layer is the same as the dimension of the compressed feature, that is, 256, and the Sigmoid function is used as the gating function. The network inputs the compressed feature representation, and the output gating value is between 0 and 1. The gating value is multiplied element by element with the compressed feature to achieve soft selection of the feature.
[0165] In order to make the feature selection more sparse, the temperature parameter is introduced to adjust the gating function. The initial value of the temperature parameter is set to 1.0, and it decreases by 0.1 every 100 rounds of training until it reaches the minimum value of 0.1. The smaller the temperature parameter, the closer the gating value is to 0 or 1, and the sparser the feature selection.
[0166] The objective function consists of three parts: mutual information loss, redundancy loss, and sparsity loss. Mutual information loss is the inverse of the mutual information estimate; redundancy loss is feature redundancy; sparsity loss is the L1 norm of the adjusted screening features. The weight ratio of the three losses is 1:0.5:0.1.
[0167] During the training process, the Adam optimizer is used to update the parameters, and the learning rate is set to 0.001. Each mini-batch contains 128 samples, and a total of 50 epochs are trained. After the training is completed, the optimized encoder and feature selection network are used to process the new deep features to obtain the final key feature subset.
[0168] In this embodiment, feature compression and information extraction are achieved through an information bottleneck encoder, which effectively reduces the feature dimension, improves the compactness and effectiveness of feature representation, and retains key task-related information; based on the dual constraints of mutual information evaluation and redundancy measurement, it is ensured that the screened features are highly relevant to the task and have low information redundancy, thereby improving the efficiency and discriminability of feature representation; a learnable feature selection network combined with a temperature regulation mechanism is used to achieve adaptive soft selection of features, which not only maintains the differentiability of the selection process, but also obtains sparse feature representation, thereby improving the interpretability and generalization ability of the model.
[0169] In an optional implementation, based on the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result, and the generated analysis result includes:
[0170] Based on the feature weight vector and the prediction result, the correlation coefficient between each pair of features in the feature weight vector is calculated to generate a feature correlation matrix; based on the feature correlation matrix, a feature dependency graph is constructed, each feature in the key feature subset is used as a node in the feature dependency graph, and the correlation coefficient in the feature correlation matrix is used as the edge weight between corresponding nodes;
[0171] Constructing a graph Laplacian matrix according to the feature dependency graph, taking the sum of the edge weights of each node and the corresponding connected node as the diagonal element value of the graph Laplacian matrix, and taking the negative value of the edge weight between nodes as the non-diagonal element value of the graph Laplacian matrix;
[0172] Calculating the Jacobian matrix of the prediction result with respect to the key feature subset to obtain an initial gradient; using the graph Laplacian matrix as a structured regularization term, and performing a weighted combination of the structured regularization term and the initial gradient to obtain a structured gradient that takes feature dependencies into account;
[0173] Calculate the structured gradient magnitude of each feature in the key feature subset, the interaction gradient value of each feature with other features, and the centrality measure of each feature in the feature dependency graph;
[0174] The structured gradient amplitude, the interaction gradient value and the centrality measure are weightedly combined to obtain feature contribution; the feature contribution is normalized to obtain the standardized contribution of each feature in the key feature subset to the prediction result; the features in the key feature subset are ranked in importance based on the standardized contribution to generate an analysis result.
[0175] In a specific implementation, firstly, a feature weight vector and prediction result data are obtained. For example, a data set containing ten key features includes environmental parameters such as temperature, humidity, and pressure. The prediction result is a quality indicator of a certain product.
[0176] Perform correlation analysis on the feature pairs in the feature weight vector. The feature correlation matrix is obtained by calculating the Pearson correlation coefficient. For example, the correlation coefficient between temperature and humidity is 0.7, indicating a strong positive correlation. After calculating the correlation coefficients of all feature pairs, the complete feature correlation matrix is obtained.
[0177] Construct a feature dependency graph based on the feature correlation matrix. Each feature is used as a node in the graph, and the correlation coefficient is used as the weight of the edge connecting the nodes. For example, the edge weight between the temperature node and the humidity node is 0.7. For feature pairs with correlation coefficients below the threshold, the edges can be left unconnected in the graph to highlight important feature dependencies.
[0178] When constructing the graph Laplacian matrix, the sum of the edge weights of each node and its connected nodes is calculated as the diagonal element. For example, if the sum of the edge weights between the temperature node and other nodes is 2.5, it is used as a diagonal element. The non-diagonal elements take the negative value of the edge weight between the corresponding nodes, and the corresponding positions of the nodes without edges are filled with zero.
[0179] Calculate the initial gradient of the prediction result for each feature. Use the finite difference method to make a small perturbation to each feature and observe the change in the prediction result. Perform a weighted combination of the graph Laplacian matrix and the initial gradient to obtain a structured gradient that takes into account feature dependencies. The weighting coefficient can be adjusted according to actual needs to balance the influence of feature independence and dependency.
[0180] Calculate the structured gradient amplitude of each feature to indicate the degree of direct influence of the feature on the prediction result. Calculate the interaction gradient between features to indicate the strength of feature synergy. Calculate the centrality of the feature in the dependency graph, including degree centrality, closeness centrality and other indicators, to indicate the importance of the feature in the overall feature network.
[0181] The above three indicators are weighted and combined to obtain the feature contribution. For example, the structured gradient amplitude weight can be set to 0.5, the interactive gradient weight to 0.3, and the centrality weight to 0.2. The feature contribution is normalized so that the sum of the contributions of all features is one. Finally, the features are sorted according to the standardized contribution and an analysis result report is generated.
[0182] In this embodiment, by constructing a feature dependency graph and a graph Laplacian matrix, the complex dependencies between features are effectively captured, the problem of traditional feature importance analysis methods ignoring feature interactions is avoided, and the accuracy of feature contribution assessment is improved; a structured gradient method is adopted to comprehensively consider the direct impact, interaction and network centrality of the features, and a multi-dimensional feature importance assessment is achieved, making the analysis results more comprehensive and reliable, and can better guide feature selection and optimization in practical applications; through standardized processing and sorting, an intuitive and clear feature importance sorting result is obtained, which is easy to understand and apply, while ensuring the comparability of analysis results in different scenarios, providing a reliable basis for subsequent model optimization and decision-making.
[0183] Figure 2 FIG. 1 is a schematic diagram of the structure of a structured data modeling and analysis system based on Transformer according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0184] The first unit is used to receive structured raw data, divide the structured raw data into numerical features and categorical features, perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, and perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; adaptively allocate weights through a learnable gating mechanism to fuse the normalized backbone features and the continuous value mapping features to obtain fused features; and dynamically sample based on the importance scores of the fused features to generate an enhanced training sample set;
[0185] The second unit is used to input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and a plurality of local feature representations are integrated through an attention fusion mechanism across subgraphs to obtain a global feature representation; the self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains a deep feature representation;
[0186] The third unit is used to construct an information bottleneck feature filter, which filters the deep feature representation by maximizing task-related information and minimizing redundant information to obtain a key feature subset; the key feature subset is input into a two-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on the attention mechanism to obtain a feature weight vector, and generates a prediction result according to the feature weight vector, and the auxiliary reconstruction branch provides regularization constraints by reconstructing the key feature subset; according to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result to generate an analysis result.
[0187] According to a third aspect of the embodiments of the present invention,
[0188] An electronic device is provided, comprising:
[0189] processor;
[0190] a memory for storing processor-executable instructions;
[0191] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0192] According to a fourth aspect of the embodiments of the present invention,
[0193] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0194] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. Transformer-based structured data modeling and analysis method, characterized by: include: Receive structured raw data, and divide the structured raw data into numerical features and categorical features, wherein the numerical features include user age, number of visits, total consumption, number of days since last purchase, number of favorite items, and number of items in shopping cart, and the categorical features include gender, membership level, device type, preferred shopping category, and delivery area; perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, and perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; perform adaptive weight allocation through a learnable gating mechanism, and fuse the normalized backbone features and the continuous value mapping features to obtain fused features; Perform dynamic sampling based on the importance score of the fusion feature to generate an enhanced training sample set; Input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through a cross-subgraph attention fusion mechanism to obtain a global feature representation; The self-calibration module receives the global feature representation, performs nonlinear transformation of the feature through residual connection and dynamic normalization, and obtains a deep feature representation; An information bottleneck feature filter is constructed to filter the deep feature representation by maximizing task-related information and minimizing redundant information to obtain a key feature subset; the key feature subset is input into a dual-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on an attention mechanism to obtain a feature weight vector, and generates a prediction result based on the feature weight vector, and the auxiliary reconstruction branch provides a regularization constraint by reconstructing the key feature subset; based on the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result, and an analysis result of the user purchase intention prediction is generated.
2. The method according to claim 1, characterized in that Receiving structured raw data, dividing the structured raw data into numerical features and categorical features, performing piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, performing adaptive quantum coding on the categorical features to obtain continuous value mapping features; performing adaptive weight allocation through a learnable gating mechanism, fusing the normalized backbone features and the continuous value mapping features to obtain fused features; Performing dynamic sampling based on the importance score of the fusion feature to generate an enhanced training sample set includes: Receiving structured raw data, dividing the structured raw data into numerical features and categorical features; inputting the numerical features and categorical features into a dual-stream feature encoder for feature processing, wherein the dual-stream feature encoder includes a feature trunk stream and a feature mapping stream; In the feature trunk stream, the data distribution of the numerical features is counted to obtain the data density, the position of the optimal segmentation point is determined based on the data density, the numerical features are divided into a plurality of segmentation intervals according to the optimal segmentation point, a linear transformation parameter matrix is constructed in each segmentation interval, the feature difference values between adjacent segmentation intervals are calculated, the optimized linear transformation parameter matrix of each segmentation interval is determined by minimizing the feature difference values, and the optimized linear transformation parameter matrix is used to perform linear transformation on the numerical features of each segmentation interval to obtain normalized trunk features; In the feature mapping flow, each category value in the category feature is mapped to a corresponding quantum ground state, an initial amplitude parameter and a phase parameter are assigned to each quantum ground state, a quantum gate sequence including a rotation gate and a control gate is constructed, the quantum gate sequence is sequentially applied to the quantum ground state, a transformed quantum state is obtained through quantum state transformation, an entropy value of the transformed quantum state is calculated, and the transformed quantum state is projected to a continuous space based on the entropy value to obtain a continuous value mapping feature; Respectively calculating the correlation coefficients between the normalized backbone features and the dimensions of the continuous value mapping features, combining the correlation coefficients to construct a feature correlation coefficient matrix, training a preset attention network based on the feature correlation coefficient matrix to obtain an importance vector of the feature dimension, and using the importance vector to perform a weighted combination of the normalized backbone features and the continuous value mapping features to obtain a fusion feature; The contribution of each dimension of the fusion feature to the target task is calculated, and a feature importance score of each dimension is generated based on the contribution. A sampling probability distribution function is constructed according to the feature importance score, and the sampling probability distribution function is used to perform importance sampling on the training samples to generate an enhanced training sample set.
3. The method according to claim 2, characterized in that Counting the data distribution of the numerical feature to obtain data density, and determining the position of the optimal segmentation point based on the data density includes: Based on the interquartile range and the number of samples of the numerical feature, a basic interval width is calculated, where the numerical feature includes the user's age; based on the basic interval width, the value range of the numerical feature is divided into a plurality of basic statistical intervals, and the number of samples in each basic statistical interval is counted to obtain an interval frequency sequence; Calculating the frequency ratios of adjacent basic statistical intervals in the interval frequency sequence to obtain interval similarity, merging adjacent basic statistical intervals whose interval similarity is greater than a preset merging threshold into new statistical intervals to obtain a merged frequency distribution sequence; Calculating the cumulative frequency distribution according to the combined frequency distribution sequence, and generating a data density function based on the cumulative frequency distribution; selecting candidate segmentation positions in the combined frequency distribution sequence, calculating the number of samples and the numerical mean of the intervals on both sides of each candidate segmentation position, and calculating the interval variance value of the candidate segmentation position using the sample number and the numerical mean; Selecting a position with a maximum interval variance value in the current interval to be segmented as a pending segmentation point, calculating the ratio of the interval variance value of the pending segmentation point to the overall variance of the numerical feature, and determining the variance contribution ratio of the pending segmentation point; The local distribution skewness is calculated for the intervals on both sides of the pending segmentation point. When the difference of the local distribution skewness is greater than the preset skewness threshold and the variance contribution ratio is greater than the preset segmentation threshold, the pending segmentation point is determined as the optimal segmentation point. The process is repeated to determine the positions of all optimal segmentation points.
4. The method according to claim 1, characterized in that Input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract local feature representations, and multiple local feature representations are integrated through a cross-subgraph attention fusion mechanism to obtain a global feature representation; The self-calibration module receives the global feature representation, performs feature nonlinear transformation through residual connection and dynamic normalization, and obtains the deep feature representation, including: Inputting the enhanced training sample set into the convolutional neural network to obtain an initial feature map, calculating the dot product of any two sample features in the initial feature map, and dividing the dot product by the module-length product of the corresponding feature vectors to obtain a feature similarity matrix; constructing a feature affinity graph based on the feature similarity matrix, wherein the vertices of the feature affinity graph correspond to the sample features, and the edge weights of the feature affinity graph correspond to the similarity values in the feature similarity matrix, and dividing the feature affinity graph into a plurality of feature subgraphs by a spectral clustering algorithm; In each of the feature subgraphs, the feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction results of the query feature sequence and the key-value pair feature sequence are masked by using the temporal causal mask matrix to obtain an initial attention weight containing only historical feature information; a spatial distance matrix of sample features in each of the feature subgraphs is calculated, and the spatial distribution entropy of the feature subgraph is calculated based on the spatial distance matrix, and the initial attention weight is adaptively adjusted according to the spatial distribution entropy to obtain an adaptive attention weight; the key-value pair feature sequence is weighted based on the adaptive attention weight, and the position encoding information is superimposed to obtain a local feature representation of each of the feature subgraphs; Based on a pre-built multi-layer perceptron network, a concatenated vector of local feature representations of adjacent feature subgraphs is input, and the concatenated vector is forward-propagated through the multi-layer perceptron network to obtain association weights between feature subgraphs, and the local feature representations are weighted summed based on the association weights to obtain a global feature representation; The mean and standard deviation of the global feature representation are calculated in the feature dimension respectively; based on the pre-constructed dynamic normalization network, the global feature representation is input into the dynamic normalization network, and the feature transformation coefficient and feature offset are obtained through forward propagation; the global feature representation is subtracted from the mean and divided by the standard deviation to obtain a standardized feature, the standardized feature is multiplied by the feature transformation coefficient, and the feature offset is added to obtain a transformed feature, and the transformed feature is added to the global feature representation to obtain a self-calibration feature; The self-calibration features are subjected to layer normalization processing and nonlinear transformation through a ReLU activation function to obtain a deep feature representation.
5. The method according to claim 4, characterized in that The feature sequence of the feature subgraph is divided into a query feature sequence and a key-value pair feature sequence, a temporal causal mask matrix is constructed, and the interaction result of the query feature sequence and the key-value pair feature sequence is masked by using the temporal causal mask matrix to obtain the initial attention weight containing only the historical feature information, including: The feature subgraph includes the e-commerce user behavior of the user for 30 consecutive days, which constitutes a feature sequence of the user behavior, calculates the similarity of the features of adjacent time steps in the feature sequence, groups the feature sequence, and obtains a query feature sequence and a key-value pair feature sequence, wherein the time step features whose similarity of the features of adjacent time steps is greater than a dynamic threshold determined based on the global statistical characteristics of the feature sequence are divided into the query feature sequence, and the remaining features are divided into the key-value pair feature sequence, and the feature division results of the query feature sequence and the key-value pair feature sequence are obtained; Based on the feature division result, the temporal correlation strength of the query feature sequence and the key-value pair feature sequence are respectively calculated by a recursive neural network, wherein the hidden state of the recursive neural network contains the long-term and short-term dependency information of the query feature sequence and the key-value pair feature sequence, and the temporal correlation strength of the query feature sequence and the temporal correlation strength of the key-value pair feature sequence are obtained; Based on the temporal position relationship of the feature sequence, a first mask matrix is constructed as a temporal causal mask matrix, and the temporal association strength of the query feature sequence and the temporal association strength of the key-value pair feature sequence are converted into soft mask weights to obtain a second mask matrix; Performing matrix multiplication on the query feature sequence and the key-value pair feature sequence to obtain an original attention score; performing addition operation on the original attention score and the first mask matrix to obtain a first mask result; performing element-wise multiplication operation on the first mask result and the second mask matrix to obtain a double-masked attention score; and normalizing the double-masked attention score to obtain an initial attention weight that only contains historical feature information.
6. The method according to claim 1, characterized in that An information bottleneck feature filter is constructed to filter the deep feature representation by maximizing task-related information and minimizing redundant information, and the key feature subsets obtained include: Constructing an information bottleneck encoder, the information bottleneck encoder includes a three-layer perceptron structure, each layer of the perceptron structure is connected to a normalization layer, the deep feature representation is input into the information bottleneck encoder, and the output layer of the information bottleneck encoder generates probability distribution parameters of the compressed feature space; encoding and mapping the deep feature representation based on the probability distribution parameters to obtain a compressed feature representation; Constructing a mutual information evaluation module, the mutual information evaluation module includes a three-layer perceptron structure, inputting a concatenated vector of the compressed feature representation and the label of the target task into the mutual information evaluation module, and outputting the mutual information amount between the compressed feature representation and the target task; Calculating an autocorrelation matrix of the compressed feature representation, wherein each element in the autocorrelation matrix is the product of the corresponding feature vector minus the corresponding mean divided by the product of the corresponding standard deviation; calculating feature redundancy based on the difference between the autocorrelation matrix and the identity matrix; Constructing a feature selection network, the feature selection network comprising a first fully connected layer and a second fully connected layer, inputting the compressed feature representation into the first fully connected layer, the output of the first fully connected layer being input into the second fully connected layer after passing through an activation function, the output of the second fully connected layer being passed through a gating function to obtain a gating value, and performing element-wise multiplication of the gating value and the compressed feature representation to obtain an initial screening feature; Adjusting the output of the gating function based on a temperature parameter, wherein the temperature parameter decreases with the number of training rounds based on a preset decreasing amount, and performing element-by-element multiplication of the adjusted gating value and the compressed feature representation to obtain an adjusted screening feature; Constructing an objective function by linearly combining the inverse of the mutual information, the feature redundancy, and the norm of the adjusted screening feature; By iteratively optimizing the objective function, the optimization parameters of the information bottleneck encoder and the optimization parameters of the feature selection network are obtained, and an optimized information bottleneck encoder and an optimized feature selection network are obtained; The deep feature representation is input into an optimized information bottleneck encoder to obtain an optimized compressed feature representation, and the optimized compressed feature representation is input into an optimized feature selection network to obtain a key feature subset.
7. The method according to claim 1, characterized in that According to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result, and the analysis result of the user purchase intention prediction is generated, including: Based on the feature weight vector and the prediction result, the correlation coefficient between each pair of features in the feature weight vector is calculated to generate a feature correlation matrix; based on the feature correlation matrix, a feature dependency graph is constructed, each feature in the key feature subset is used as a node in the feature dependency graph, and the correlation coefficient in the feature correlation matrix is used as the edge weight between corresponding nodes; Constructing a graph Laplacian matrix according to the feature dependency graph, taking the sum of the edge weights of each node and the corresponding connected node as the diagonal element value of the graph Laplacian matrix, and taking the negative value of the edge weight between nodes as the non-diagonal element value of the graph Laplacian matrix; Calculating the Jacobian matrix of the prediction result with respect to the key feature subset to obtain an initial gradient; using the graph Laplacian matrix as a structured regularization term, and performing a weighted combination of the structured regularization term and the initial gradient to obtain a structured gradient that takes feature dependencies into account; Calculate the structured gradient magnitude of each feature in the key feature subset, the interaction gradient value of each feature with other features, and the centrality measure of each feature in the feature dependency graph; The structured gradient amplitude, the interaction gradient value and the centrality measure are weightedly combined to obtain feature contribution; the feature contribution is normalized to obtain the standardized contribution of each feature in the key feature subset to the prediction result; the features in the key feature subset are ranked in importance based on the standardized contribution to generate an analysis result of user purchase intention prediction.
8. A structured data modeling and analysis system based on Transformer, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to receive structured raw data, divide the structured raw data into numerical features and categorical features, perform piecewise linear transformation on the numerical features through a dual-stream feature encoder to obtain normalized backbone features, perform adaptive quantum coding on the categorical features to obtain continuous value mapping features; perform adaptive weight allocation through a learnable gating mechanism, fuse the normalized backbone features and the continuous value mapping features, and obtain fused features; Perform dynamic sampling based on the importance score of the fusion feature to generate an enhanced training sample set; The second unit is used to input the enhanced training sample set into a feature extraction network including a hierarchical attention module and a self-calibration module, wherein the hierarchical attention module calculates a similarity matrix of sample features, constructs a feature affinity graph based on the similarity matrix, and divides the feature affinity graph into multiple feature subgraphs through a graph segmentation algorithm; in each of the feature subgraphs, a causal attention mechanism is used to extract a local feature representation, and a cross-subgraph attention fusion mechanism is used to integrate multiple local feature representations to obtain a global feature representation; The self-calibration module receives the global feature representation, performs nonlinear transformation of the feature through residual connection and dynamic normalization, and obtains a deep feature representation; The third unit is used to construct an information bottleneck feature filter, which filters the deep feature representation by maximizing task-related information and minimizing redundant information to obtain a key feature subset; the key feature subset is input into a two-branch network including a main prediction branch and an auxiliary reconstruction branch, the main prediction branch processes the key feature subset based on the attention mechanism to obtain a feature weight vector, and generates a prediction result according to the feature weight vector, and the auxiliary reconstruction branch provides regularization constraints by reconstructing the key feature subset; according to the feature weight vector and the prediction result, a structured gradient method is used to calculate the contribution of each feature in the key feature subset to the prediction result to generate an analysis result.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Remote sensing image saliency detection method based on double-branch architecture network
CN117333770A
Abnormality analysis method and device, storage medium and electronic equipment
CN119097322A