Prediction method based on human-computer interaction feature mining
By employing a feature mining method based on the large language model GPT-4, and utilizing a bilinear attention mechanism and a dual-branch lightweight convolutional network, the problem of low prediction accuracy for multimodal and multi-source heterogeneous data with small sample sizes is solved, achieving efficient feature association mining and accurate prediction.
Patent Information
- Application Number
- CN202511084068.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to effectively handle multimodal and heterogeneous data from multiple sources in small sample sizes, resulting in low prediction accuracy in human-computer interaction.
We employ a feature mining method based on the large language model GPT-4. This method extracts features from multimodal and multi-source heterogeneous data through a bilinear attention mechanism and a dual-branch lightweight convolutional network. We then combine this with a multilayer perceptron with residual connections to perform feature fusion, thereby enabling the correlation mining of cross-modal and multi-source heterogeneous features.
It significantly improved prediction accuracy in small samples and enhanced the clinical interpretability and market acceptance of the prediction results.
Smart Images

Figure CN120974231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of human-computer interaction, and in particular to a prediction method based on human-computer interaction feature mining. BACKGROUND
[0002] There are various traditional machine learning models, such as a support vector machine model (SVM), a random forest model, a decision tree model, a naive Bayes model, a baseline random model, and a polynomial logistic regression model. Applying these models to human-computer interaction can extract features to a certain extent, but cannot handle multi-modal and multi-source heterogeneous data.
[0003] TabPFN deep models, Wide&Deep models, and convolutional neural network models (CNN) belong to deep learning models, which can automatically mine valuable features from a large amount of data and perform well in handling complex nonlinear relationships, but cannot extract features from multi-source heterogeneous data.
[0004] GPT-4, as a pre-trained large model, is also a multi-modal model that can handle modalities such as pictures and text, but cannot complete training under small sample conditions and cannot be fine-tuned.
[0005] To overcome the deficiencies of the prior art, under small sample (a few hundred to a few thousand samples) conditions, a prediction method based on human-computer interaction feature mining is proposed, which can better handle multi-modal and multi-source heterogeneous data classification and fitting tasks, and improve prediction accuracy. SUMMARY
[0006] To overcome the deficiencies of the prior art, the purpose of the present application is to provide a prediction method based on human-computer interaction feature mining.
[0007] The purpose of the present application is achieved by the following technical solutions:
[0008] A prediction method based on human-computer interaction feature mining comprises the following steps:
[0009] S1, input the original data and prompt words into a large language model GPT-4, the original data including modal data and / or multi-source heterogeneous data; wherein, modal features are extracted for each prompt word corresponding modal data, and multi-source heterogeneous features are extracted for each prompt word corresponding multi-source heterogeneous data; the modal features are structured, and the structured modal features and / or multi-source heterogeneous features constitute original features;
[0010] S2, original feature ID is used to construct an input matrix X, the input matrix X is sent into an embedding layer, different original features are mapped into a unified semantic space and mapped into a dense representation E; a bilinear attention mechanism is used to calculate the mutual relationship between any two original features to obtain a related matrix A;
[0011] S3, applying a reversible permutation matrix M to the matrix A to obtain a matrix P;
[0012] S4, the matrix A and the matrix P are processed on a double-branch lightweight convolutional network in parallel, local and global feature interactions between original features are extracted respectively, and then the outputs of the two branches are flattened and spliced to form a splicing vector;
[0013] S5, the splicing vector in step S4 is supplemented with overall statistical information through a multi-layer perceptron with residual connection, thereby outputting a prediction result.
[0014] Preferably, the original data includes modal data and multi-source heterogeneous data; the modal data is a stroke functional prognosis medical history text and a CT image; the multi-source heterogeneous data is a laboratory biochemical index, and the prompt word is listed by a clinical expert according to medical knowledge.
[0015] Preferably, after obtaining the modal features in step S1, a missing rate threshold filtering, a medical reasonableness review, a multiple collinearity detection and an information gain evaluation method are used to remove high-missing, high-redundancy and weakly-related modal features.
[0016] Preferably, the modal features in step S1 are structured, the sub-type features of the modal features retain the original classification form, and the numerical features are discretized according to medical thresholds.
[0017] Preferably, the prompt words in step S1 are unique identifier, age, gender, diabetes, hypertension, slurred speech, difficulty swallowing, constipation, sleep, history of trauma surgery, smoking, alcohol consumption, left upper limb Brunnstrom stage, left lower limb Brunnstrom stage, right upper limb Brunnstrom stage, right lower limb Brunnstrom stage, left upper limb Ashworth muscle tension score, left lower limb Ashworth muscle tension score, right upper limb Ashworth muscle tension score, right lower limb Ashworth muscle tension score, coronary heart disease, hemiplegia (left / right), cerebral hemorrhage, pontine, frontal lobe, temporal lobe, occipital lobe, frontal-temporal-parietal lobe, frontal-temporal lobe, ventricle, basal ganglia, cerebellum, cerebellar cortex, brainstem, thalamus, cognitive impairment, cerebral infarction, periventricular region, cerebellar hemisphere, internal capsule, insular lobe, cerebral hemisphere, middle cerebral artery territory, middle cerebral artery territory, semiovale center, high blood sugar, erythrocyte sedimentation rate (ESR), direct bilirubin, indirect bilirubin, low albumin, low total cholesterol, triglycerides, uric acid, high C-reactive protein, lipoprotein A1 assay, D-dimer, neutrophil count, pain limitation, lesion type, Babinski sign, fecal incontinence, urinary catheter, time orientation, time orientation and person orientation + calculation ability.
[0018] Preferably, the spliced vector in step S5 is supplemented with overall statistical information by a residual connection multi-layer perceptron, and then mapped to a five-dimensional soft-max probability to output FIM 0-4 functional grading results.
[0019] Preferably, the original data is modal data, the modal data is a cover original picture, and the cover original picture is classified as popular or unpopular; the prompt words are listed by a cover design expert according to cover design knowledge; and the large language model GPT-4 extracts modal features and quantifies scores by performing on each cover original picture under each prompt word to obtain structured modal features.
[0020] Preferably, the prompt words are color appeal, layout hierarchy, cover composition, portrait expressiveness, title text strategy, information density and reading efficiency, emotion communication, realism, atmosphere creation, mimicry, content challenge, recognition, altruism, contrast, emotional tension, psychological compensation, misleading, suggestive, cognitive dissonance, attention hook design, discussion, voyeurism, identity counterpoint, instant reward, platform aesthetic adaptation, thumbnail adaptability, and multi-picture collaborative closed loop.
[0021] Preferably, the original data is modal data, the modal data is a text and creative product picture, and the text and creative product picture is classified as high sales and low sales; the prompt words are listed by an expert according to text and creative industry knowledge; and the large language model GPT-4 extracts modal features and quantifies scores by performing on each text and creative product picture under each prompt word to obtain structured modal features.
[0022] Preferably, the prompt words are cultural value, story, artistic quality, aesthetic delicacy, uniqueness, practicality, interest / interaction, technology, emotional resonance and memory points, topicality, personalization, series correlation, channel attribute, environmental friendliness, price and platform.
[0023] Compared with the prior art, the present application has the beneficial effects that:
[0024] Through the input of prompt words and original data, after extracting original features by GPT-4 and constructing an input matrix X, it is mapped to an embedding layer with unified semantics, and then all the correlation patterns between the original features are fully exposed by using a double linear attention mechanism. The local and global original features are efficiently sampled by using a permutation matrix and a double-branch lightweight convolutional network, realizing the correlation mining of cross-modal and multi-source heterogeneous features of the original features. Finally, a multi-layer perceptron with residual connection is used for feature fusion and feature extraction, which can better handle the classification and fitting tasks of multi-modal and multi-source heterogeneous data, and improve the prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 The original feature correlation graph of the matrix A obtained by the double linear attention mechanism of Embodiment 1.
[0026] Figure 2 The schematic diagram of the double-branch lightweight convolutional network of Embodiment 1 for extracting interaction features. DETAILED DESCRIPTION
[0027] In order to make the technical problems solved by the present application, the technical solutions adopted and the technical effects achieved more clear, the technical solutions of the embodiments of the present application will be further described in detail below with reference to the drawings.
[0028] Embodiment 1
[0029] The present embodiment provides a prediction method based on human-computer interaction feature mining, comprising the following steps,
[0030] S1, input the medical history texts (modal data), CT images (modal data), laboratory biochemical indicators (multi-source heterogeneous data), and prompt words defined by clinical experts into the large language model GPT-4. With the advantages of the large language model GPT-4 in text understanding, feature recognition, and relationship extraction, the features of the medical history texts, CT images, and laboratory biochemical indicators corresponding to each prompt word are extracted to obtain corresponding modal features and multi-source heterogeneous features. The missing rate threshold filtering (such as removing if the missing rate is greater than 40%), medical reasonableness review, multiple collinearity detection, and information gain evaluation are used to remove high-missing, high-redundancy, or weakly related modal features and multi-source heterogeneous features. The remaining modal features are kept in their original classification form for discrete types, and numerical features are discretized according to medical thresholds for structured processing to balance granularity and discriminability. The structured modal features and multi-source heterogeneous features form the original features;
[0031] S2, the original features are mapped to integer ID form to construct an input matrix X with a shape of 751x64; the input matrix X is sent to the embedding layer to map different original features to a unified semantic space and represent them as dense representations E;
[0032] The embedding layer converts discrete integer ID original feature representations into continuous dense vector representation space, effectively preserving special semantics and spatial relationships, laying the foundation for subsequent deep interaction learning. Compared with independent hot code or simple numerical encoding, dense representation effectively alleviates the dimensionality explosion problem, reduces model complexity, and automatically learns the implicit relationship information between original features, thereby improving the generalization ability and prediction accuracy of the model.
[0033] The bilinear attention mechanism is used to calculate the mutual relationship between any two original features to obtain the related matrix A, A = Tanh(EE T ), Tanh(EE T ) is the hyperbolic tangent function, E is the coordinate table embedding each original feature into the same semantic space, E T is the transpose of E, and the inner product similarity of all original feature pairs is obtained by multiplying E and E T , and then mapped to a bounded correlation matrix A of (-1, 1) through the hyperbolic tangent function, as shown in Figure 1 The matrix A = Tanh(EE T ) calculated by the bilinear attention mechanism explicitly exposes all potential relationships between original feature pairs (gray blocks in Figure 1 );
[0034] In contrast to the traditional self-attention or dot-product attention mechanism which only captures linear or simple interactions, the bilinear attention mechanism models the bidirectional nonlinear interaction between any two original features, effectively capturing complex, high-order, and nonlinear cross-modal feature combination relationships. The precise capture of such interactions not only significantly improves prediction accuracy, but also enhances the clinical interpretability of the prediction results. Through Grad-CAM heat maps, the feature interaction patterns are clearly displayed, and clinical experts can clearly understand the key clinical factors that affect the prediction.
[0035] S3, apply the reversible permutation matrix M to the transpose of matrix A, thereby obtaining matrix P, P = MAM T ;
[0036] S4, parallel processing matrix A and matrix P on the double-branch lightweight convolutional network, respectively extracting local and global interaction features, then flattening and concatenating the outputs of the two branches to form a concatenated vector;
[0037] In practice, the spatial distance of original features (such as the adjacent relationship in the table) may not necessarily reflect the true correlation between them. Traditional methods generally fix the order of original features and lack flexibility in adjusting or reconstructing distance relationships, such as Figure 2 As shown, apply a double-branch lightweight convolutional network to matrix A and matrix P respectively, efficiently sample local blocks and cross-domain combinations under these two arrangements, extract strongly interacting original feature sets, capture local and non-local interaction relationships between original features, and obtain more comprehensive cross-modal information.
[0038] S5, the concatenated vector in step S4 is supplemented with overall statistical information by a residual connected multilayer perceptron, mapped to a five-dimensional soft-max probability, and outputs FIM 0-4 level functional grading results, with specific grade division as shown in Table 1, and the output prediction result.
[0039] Table 1: Functionality Independence Rating (FIM) score and functional dependency grade division standard
[0040] Dependence level Degree of dependence FIM score breakdown 0 Independent 90~125 1 Mildly dependent 72~89 2 Moderately dependent 54~71 3 Severely dependent 36~53 4 Totally dependent 18~35
[0041] The residual connected multilayer perceptron significantly enhances the deep learning capabilities of the model, making the network more effective in capturing deep feature interaction information. The residual connected multilayer perceptron also effectively alleviates the gradient vanishing problem, making the model converge more quickly and perform more stably.
[0042] The 64 dimensions of the retained prompt words 64 in step S1 are respectively unique identifier, age, gender, diabetes, hypertension, slurred speech, difficulty swallowing, constipation, sleep, history of trauma and surgery, smoking, alcohol consumption, left upper limb Brunnstrom stage, left lower limb Brunnstrom stage, right upper limb Brunnstrom stage, right lower limb Brunnstrom stage, left upper limb Ashworth muscle tension score, left lower limb Ashworth muscle tension score, right upper limb Ashworth muscle tension score, right lower limb Ashworth muscle tension score, coronary heart disease, hemiplegia (left / right), cerebral hemorrhage, pontine, frontal lobe, temporal lobe, occipital lobe, frontal-temporal-parietal lobe, frontal-temporal lobe, ventricle, basal ganglia, cerebellum, cerebellar cortex, brain stem, thalamus, cognitive impairment, cerebral infarction, periventricular region, cerebellar hemisphere, internal capsule, insular lobe, cerebral hemisphere, cerebral artery supply area, cerebral artery supply area, semiovale center, high blood sugar, erythrocyte sedimentation rate (ESR), direct bilirubin, indirect bilirubin, low albumin, low total cholesterol, triglyceride, uric acid, high C-reactive protein, lipoprotein A1 determination, D-dimer, neutrophil count, pain limitation, lesion type, Babinski sign, fecal incontinence, urinary catheter, time orientation, time orientation and person orientation + calculation ability. The original features corresponding to the 64 dimensions of the prompt words are divided into text features, image features and biochemical features according to the source, wherein the text features have 30 dimensions, the image features have 22 dimensions, and the biochemical features have 12 dimensions.
[0043] The stroke case text, CT image and laboratory biological index data set has 751 cases. The 751 stroke data sets come from the rehabilitation center of the affiliated hospital of Dali University.
[0044] Model training and prediction result evaluation
[0045] The five-fold cross-validation method is used to train the model: the entire original data is equally divided into 5 subsets, 4 of which are taken as the training set each time, and the remaining 1 is taken as the validation set, and the cycle is repeated 5 times until each subset is used as the validation.
[0046] The cross-entropy loss is defined for the output prediction result:
[0047]
[0048] Where, y k is the one-hot encoding of the true label, is the prediction probability output by the model.
[0049] AdamW (learning rate 3x10-4) is used in the training process, and weight decay λ=0.01 is applied to all learnable parameters.
[0050] Example 2
[0051] The embodiment provides a prediction method based on human-computer interaction feature mining, including the following steps,
[0052] S1, input 414 cover original drawings (modal data, including two modalities of images and texts) and prompt words defined by cover design experts into a large language model GPT-4, the 414 cover original drawings are divided into popular and unpopular two categories according to click rate or browsing volume and the like; the prompt words are 27 dimensions, respectively color appeal, layout level, cover composition, portrait expressiveness, title text strategy, information density and reading efficiency, emotion communication, realism, atmosphere creation, imitativeness, content challenge, recognition, altruism, contrast, emotional tension, psychological compensation, misleading, suggestive, cognitive dissonance, attention hook design, discussion, voyeurism, identity matching, instant reward, platform aesthetic adaptation, thumbnail adaptation and multi-picture collaborative closed loop; the performance of the 414 cover original drawings on the 27-dimensional prompt words is quantified by GPT-4 to obtain a structured modal feature, and an original feature composed of the modal feature is obtained,
[0053] S2, the modal feature is mapped into an integer ID form, and an input matrix X with a shape of 414x27 is constructed; the input matrix X is sent into an embedding layer, and different original features are mapped into a unified semantic space to obtain a dense representation E; a bilinear attention mechanism is used to calculate the mutual relationship between any two original features to obtain a related matrix A, A=Tanh(EE T ), Tanh(EE T ) is a hyperbolic tangent function, E is a coordinate table of embedding each original feature into the same semantic space, E T is the transpose of E, and the inner product similarity of all original feature pairs is obtained by multiplying E and E T , and then mapped into a bounded correlation matrix A of (-1, 1) through the hyperbolic tangent function; S3, applying a reversible permutation matrix M to the matrix A to obtain a matrix P, P=MAM T ;
[0054] S4, the matrix A and the matrix P are processed in parallel on a double-branch lightweight convolutional network, and local and global feature interactions between the original features are extracted respectively, and then the outputs of the two branches are flattened and spliced to form a splicing vector;
[0055] S5, the splicing vector in step S4 is supplemented with overall statistical information through a multilayer perceptron with residual connection, and a prediction result is output.
[0056] The cover original drawings come from a short video platform, and the cover original drawings include image parts and text parts.
[0057] Model training and prediction result evaluation
[0058] The model is trained using five-fold cross-validation: the entire original data is divided into 5 subsets, and each time 4 of them are taken as the training set and the remaining 1 is taken as the validation set, and the cycle is repeated 5 times until each subset is used as validation.
[0059] The cross-entropy loss is defined for the output prediction result:
[0060]
[0061] where y k is the one-hot encoding of the true label, is the prediction probability output by the model.
[0062] AdamW (learning rate 3x10-4) is used during training, and weight decay λ=0.01 is applied to all learnable parameters.
[0063] The 27-dimensional quantization of a single cover original image can be completed within 3 minutes, the prediction delay of a single cover original image is less than 40ms, it can be deployed on an 8GB single card or a mobile terminal, and the attention heat map intuitively displays the high-impact original feature combination; the platform course real-time interaction data dynamically refreshes the scoring threshold and fine-tunes the prediction model, realizing human-machine collaborative online evolution.
[0064] Embodiment 3
[0065] The embodiment provides a prediction method based on human-computer interaction feature mining, comprising the following steps,
[0066] S1, input 730 kinds of cultural and creative product pictures (including the graphics of the pictures and the words on the pictures) and prompt words defined by cultural and creative experts into a large language model GPT-4, and divide the 730 kinds of cultural and creative product pictures into high sales (471) and low sales (259) according to the sales; the prompt word is 16-dimensional, and respectively is cultural value, story, artistic quality, aesthetic fineness, uniqueness, practicality, interest / interaction, technological sense, emotional resonance and memory point, topic, individualization, series correlation, channel attribute, environmental friendliness, price and platform; the performance of the 730 cultural and creative product pictures on the 16-dimensional prompt word is quantized by GPT-4 to 1-10, and the original features are obtained;
[0067] S2, the modal feature is mapped into an integer ID form, and an input matrix X with a shape of 730x16 is constructed; the input matrix X is sent into an embedding layer, and different original features are mapped into a unified semantic space to obtain a dense representation E; a bilinear attention mechanism is used to calculate the mutual relationship between any two original features, and a related matrix A is obtained, A=Tanh(EE T ), Tanh(EE T) is a hyperbolic tangent function, E is a coordinate table embedding each original feature into the same semantic space, E T is the transpose of E, and the inner product similarity of all original feature pairs is obtained by multiplying E and E T , and then mapped to a bounded correlation matrix A in (-1, 1) by a hyperbolic tangent function; S3, applying a reversible permutation matrix M to the transpose of matrix A to obtain matrix P, P = MAM T ;
[0068] S4, processing matrix A and matrix P on the double-branch lightweight convolutional network in parallel, respectively extracting local and global feature interactions between original features, and then flattening and splicing the outputs of the two branches to form a splicing vector;
[0069] S5, the splicing vector in step S4 is supplemented with overall statistical information by a multi-layer perceptron with residual connection, and the output prediction result is output.
[0070] The picture of the cultural and creative product includes a text part and a picture part. The pictures of 730 cultural and creative products are all collected from e-commerce and social media platforms to ensure that the information comes from real products for sale.
[0071] Model training and prediction result evaluation
[0072] The five-fold cross-validation method is used to train the model: the total original data is divided into 5 subsets, 4 of which are taken as the training set each time, and the remaining 1 is taken as the validation set, and the cycle is repeated for 5 rounds until each subset is used as the validation set.
[0073] The cross-entropy loss is defined for the output prediction result:
[0074]
[0075] Where y k is the one-hot encoding of the true label, and y is the prediction probability output by the model.
[0076] AdamW (learning rate 3x10-4) is used during training, and weight decay λ = 0.01 is applied to all learnable parameters.
[0077] A prediction method for multi-modal feature mining is constructed to predict the sales of cultural and creative products, effectively reducing the risk of blind design and production of cultural and creative enterprises, and the feature heat map intuitively reflects the key feature combination of market-accepted products.
[0078] Comparative Example 1
[0079] This comparative example is basically the same as Example 1, except that the multi-layer perceptron with residual connection is not used in this comparative example.
[0080] Comparative Example 2
[0081] This comparative example is substantially the same as Example 1, except that this comparative example does not have the step of applying the invertible permutation matrix M to transpose the matrix A to obtain the matrix P.
[0082] Comparative Example 3
[0083] This comparative example is substantially the same as Example 1, except that this comparative example does not have the step of applying the invertible permutation matrix M to transpose the matrix A to obtain the matrix P, nor does it have the residual connected multi-layer perceptron.
[0084] Table 2. Prediction accuracy of Example 1 and Comparative Examples 1-3
[0085]
[0086] As can be seen from Table 2, based on Comparative Example 3, the residual connected multi-layer perceptron processing (i.e. Comparative Example 2) improves the accuracy by about 4%; the application of the invertible permutation matrix M to transpose the matrix A to obtain the matrix P (Comparative Example 1) improves the accuracy by less than 3%, but the scheme as described in Example 1, based on Comparative Example 3, the application of the invertible permutation matrix M to transpose the matrix A to obtain the matrix P, followed by the residual connected multi-layer perceptron processing, improves the accuracy by more than 13%, indicating that the combination of the two steps in the scheme of the present application improves the prediction accuracy better than the effect of the two steps applied individually.
[0087] Comparative Example 4
[0088] This comparative example generates the original features, and the subsequent machine learning model is an existing support vector machine model.
[0089] 1) Input features: generate 64-dimensional original features as in step S1 of Example 1;
[0090] 2) Feature preprocessing: input after standardization (StandardScaler);
[0091] 3) Kernel function: radial basis kernel function (RBF);
[0092] 4) Hyperparameter setting: penalty coefficient C = 1.0, kernel parameter γ = 0.01;
[0093] 5) Model training: five-fold cross-validation;
[0094] 6) Output prediction: use the decision function threshold 0.0 to classify.
[0095] Comparative Example 5
[0096] After generating the original features, the subsequent machine learning model is an existing random forest model. 1) Input features: 64-dimensional original features generated in step S1 of Example 1 are inputted;
[0097] 2) Feature preprocessing: the input is standardized (StandardScaler), and the category features are directly inputted;
[0098] 3) Model structure: composed of multiple decision trees;
[0099] 4) Hyperparameter setting: the number of decision trees n_estimators = 50, and the maximum depth of the tree max_depth = 15;
[0100] 5) Model training: five-fold cross-validation;
[0101] 6) Output prediction: all decision trees vote for the majority class.
[0102] Comparative Example 6
[0103] After generating the original features, the subsequent machine learning model is an existing decision tree model.
[0104] 1) Input features: 64-dimensional original features generated in step S1 of Example 1 are inputted;
[0105] 2) Feature preprocessing: the input is standardized (StandardScaler), and the category features are directly inputted;
[0106] 3) Model structure: composed of multiple decision trees;
[0107] 4) Hyperparameter setting: the number of decision trees n_estimators = 50, and the maximum depth of the tree max_depth = 15;
[0108] 5) Model training: five-fold cross-validation;
[0109] 6) Output prediction: all decision trees vote for the majority class.
[0110] Comparative Example 7
[0111] After generating the original features, the subsequent machine learning model is an existing decision tree model.
[0112] 1) Input features: 64-dimensional original features generated in step S1 of Example 1 are inputted;
[0113] 2) Feature preprocessing: the input is standardized (StandardScaler), and the category features are directly inputted;
[0114] 3) Model structure: composed of multiple decision trees;
[0115] 4) Hyperparameter setting: the number of decision trees n_estimators = 50, and the maximum depth of the tree max_depth = 15;
[0116] 5) Model training: Compute the conditional probability of each class given the features;
[0117] 6) Output prediction: Predict the class with the maximum posterior probability.
[0118] Comparative Example 8
[0119] After generating the original features, the subsequent machine learning model is the existing polynomial logistic regression model.
[0120] 1) Input features: Same as Step S1 of Example 1, generate 64-dimensional original features;
[0121] 2) Feature preprocessing: Normalize the numerical features using standardization, and then use second-order or higher-order polynomial feature cross-explicit construction to build combined features;
[0122] 3) Model structure: Logistic regression model (multi-class, using Multinomial Logistic Regression);
[0123] 4) Hyperparameter setting: Polynomial cross-feature order (degree = 4), and the optimization solver uses the LBFGS algorithm;
[0124] 5) Model training: Use five-fold cross-validation;
[0125] 6) Output prediction: Select the class corresponding to the maximum probability value.
[0126] Comparative Example 9
[0127] After generating the original features, the subsequent machine learning model is the existing Wide&Deep model.
[0128] 1) Input features: Same as Step S1 of Example 1, generate 64-dimensional original features;
[0129] 2) Feature preprocessing: Standardize all original features uniformly;
[0130] 3) Model structure: Wide part (explicit interaction): Use a linear model to capture low-order explicit combined interaction relationships; Deep part (implicit interaction): Use a multi-layer feedforward neural network (MLP) to capture high-order nonlinear interaction relationships;
[0131] 4) Hyperparameter setting: The Deep part uses a 2-layer MLP (128 and 64 hidden units), with ReLU as the activation function and a Dropout ratio of 0.2; 2. The model optimizer uses the Adam optimizer with a learning rate set to 0.001;
[0132] 5) Model training: five-fold cross-validation was used, and the maximum training rounds were 50. The early stopping strategy was to stop when the validation accuracy did not improve for 5 rounds.
[0133] 6) Output prediction: the final classification was output by a softmax function to predict the category.
[0134] The accuracy of Example 1 and Comparative Examples 4-9 was tested, and the results are shown in Table 3.
[0135] Table 3: Prediction accuracy of Example 1 and Comparative Examples 4-9
[0136]
[0137] Comparative Example 10
[0138] After generating the original features in this comparative example, the subsequent machine learning model is an existing support vector machine model.
[0139] 1) Input features: 27-dimensional original features were generated according to step S1 of Example 2;
[0140] 2) Feature preprocessing: input after standardization (StandardScaler);
[0141] 3) Kernel function: radial basis kernel function (RBF);
[0142] 4) Hyperparameter setting: penalty coefficient C = 1.0, kernel parameter γ = 0.01;
[0143] 5) Model training: five-fold cross-validation;
[0144] 6) Output prediction: use the decision function threshold 0.0 to classify categories.
[0145] Comparative Example 11
[0146] After generating the original features in this comparative example, the subsequent machine learning model is an existing random forest model.
[0147] 1) Input features: 27-dimensional original features were generated according to step S1 of Example 2;
[0148] 2) Feature preprocessing: input after standardization (StandardScaler), and category features were directly input;
[0149] 3) Model structure: composed of multiple decision trees;
[0150] 4) Hyperparameter setting: number of decision trees n_estimators = 50, maximum tree depth max_depth = 15;
[0151] 5) Model training: five-fold cross-validation;
[0152] 6) Output prediction: the majority class voted by all decision trees.
[0153] Comparative Example 12
[0154] After generating the original features, the subsequent machine learning model is an existing decision tree model.
[0155] 1) Input features: same as step S1 of Example 2, 27-dimensional original features are generated;
[0156] 2) Feature preprocessing: both categorical features and numerical features are directly inputted;
[0157] 3) Model structure: single classification decision tree;
[0158] 4) Hyperparameter setting: Gini coefficient is used as the feature splitting standard, and the maximum tree depth is 10;
[0159] 5) Model training: five-fold cross-validation;
[0160] 6) Output prediction: classified according to tree nodes.
[0161] Comparative Example 13
[0162] After generating the original features, the subsequent machine learning model is an existing Naive Bayes model.
[0163] 1) Input features: same as step S1 of Example 2, 27-dimensional original features are generated;
[0164] 2) Feature preprocessing: standardize continuous features, and directly use categorical features;
[0165] 3) Model assumption: independent distribution between features;
[0166] 4) Model type: Gaussian Naive Bayes;
[0167] 5) Model training: calculate the conditional probability of each category feature;
[0168] 6) Output prediction: predict the maximum category according to the posterior probability.
[0169] Comparative Example 14
[0170] After generating the original features, the subsequent machine learning model is an existing multinomial logistic regression model.
[0171] 1) Input features: same as step S1 of Example 2, 27-dimensional original features are generated;
[0172] 2) Feature preprocessing: standardize the numerical features by standardization, and then use second-order or higher-order multinomial feature cross-explicit construction to build combined features;
[0173] 3) Model structure: Logistic Regression model (multi-class, using Multinomial Logistic Regression);
[0174] 4) Hyperparameter setting: polynomial cross-feature degree = 4, using LBFGS algorithm for optimization solver;
[0175] 5) Model training: using five-fold cross-validation;
[0176] 6) Output prediction: selecting the class corresponding to the maximum probability.
[0177] Comparative Example 15
[0178] After generating the original features, the subsequent machine learning model is the existing Wide&Deep model.
[0179] 1) Input features: generate 27-dimensional original features as in step S1 of Example 2;
[0180] 2) Feature preprocessing: standardize all original features uniformly;
[0181] 3) Model structure: Wide part (explicit interaction): use a linear model to capture low-order explicit combination interaction; Deep part (implicit interaction): use a multi-layer feedforward neural network (MLP) to capture high-order nonlinear interaction;
[0182] 4) Hyperparameter setting: 2-layer MLP (128 and 64 hidden units) for the Deep part, with ReLU activation function and Dropout ratio of 0.2; 2. Model optimizer uses Adam optimizer with learning rate set to 0.001;
[0183] 5) Model training: using five-fold cross-validation, training a maximum of 50 rounds per fold, and using early stopping strategy with 5 rounds of no improvement in validation accuracy to stop;
[0184] 6) Output prediction: the final classification outputs the predicted class through the softmax function.
[0185] Comparative Example 16
[0186] This comparative example uses the GPT-4 model
[0187] 1) Input features: directly input the 414 cover original images of Example 2;
[0188] 2) Feature preprocessing: the cover original image is converted into visual embedding by image encoder (Vision Encoder, such as CLIP-ViT), and the text on the cover original image is processed into semantic embedding by the built-in word segmentation and context coding of GPT-4; no additional manual feature construction is required;
[0189] 3) Model structure: a multi-modal GPT-4 model (with visual-linguistic joint representation learning capability) is used, and cross-modal interaction is realized through multi-layer self-attention mechanism inside; a two-class output layer (softmax) is connected at the top;
[0190] 4) Hyperparameter setting: the maximum context length is 512, and the visual input resolution is 224x224; the AdamW optimizer is used during training, the learning rate is 1e-5, and the batch size is 16;
[0191] 5) Model training: no cross-validation is used, and one-time fine-tuning is directly performed on the entire training set;
[0192] 6) Output prediction: the model outputs the probability distribution of two classes (0 / 1), and directly selects the class corresponding to the maximum probability as the final binary classification prediction result.
[0193] The accuracy of Example 2 and Comparative Examples 10-16 was tested, and the results are shown in Table 4.
[0194] Table 4: Prediction accuracy of Example 2 and Comparative Examples 10-16
[0195]
[0196] Comparative Example 17
[0197] After generating the original features in this comparative example, the subsequent machine learning model is an existing support vector machine model.
[0198] 1) Input features: generate 16-dimensional original features as in step S1 of Example 3;
[0199] 2) Feature preprocessing: input after standardization (StandardScaler);
[0200] 3) Kernel function: radial basis kernel function (RBF);
[0201] 4) Hyperparameter setting: penalty coefficient C = 1.0, kernel parameter γ = 0.01;
[0202] 5) Model training: five-fold cross-validation;
[0203] 6) Output prediction: use the decision function threshold 0.0 to classify.
[0204] Comparative Example 18
[0205] The subsequent machine learning model of this comparative example is an existing random forest model after generating the original features.
[0206] 1) Input features: 16-dimensional original features generated in step S1 of Example 3;
[0207] 2) Feature preprocessing: input after standardization (StandardScaler), and directly input category features;
[0208] 3) Model structure: composed of multiple decision trees;
[0209] 4) Hyperparameter setting: number of decision trees n_estimators = 50, and maximum tree depth max_depth = 15;
[0210] 5) Model training: five-fold cross-validation;
[0211] 6) Output prediction: majority category voting of all decision trees.
[0212] Comparative Example 19
[0213] The subsequent machine learning model of this comparative example is an existing decision tree model after generating the original features.
[0214] 1) Input features: 16-dimensional original features generated in step S1 of Example 3;
[0215] 2) Feature preprocessing: directly input category features and numerical features;
[0216] 3) Model structure: single classification decision tree;
[0217] 4) Hyperparameter setting: Gini coefficient (Gini) is used as the feature splitting standard, and the maximum tree depth is 10;
[0218] 5) Model training: five-fold cross-validation;
[0219] 6) Output prediction: classification according to tree nodes.
[0220] Comparative Example 20
[0221] The subsequent machine learning model of this comparative example is an existing Naive Bayes model after generating the original features.
[0222] 1) Input features: 16-dimensional original features generated in step S1 of Example 3;
[0223] 2) Feature preprocessing: standardize continuous features, and directly use category features;
[0224] 3) Model assumption: independent distribution between features;
[0225] 4) Model type: Gaussian Naive Bayes;
[0226] 5) Model training: Compute the conditional probability of each feature for each class;
[0227] 6) Output prediction: Predict the class with the maximum posterior probability.
[0228] Comparative Example 21
[0229] After generating the original features, the subsequent machine learning model is the existing polynomial logistic regression model.
[0230] 1) Input features: Same as Step S1 of Example 3, generate 16-dimensional original features;
[0231] 2) Feature preprocessing: Standardize the numerical features using standardization, and then use second-order or higher-order polynomial feature cross-explicit construction to build combined features;
[0232] 3) Model structure: Logistic regression model (multi-class, using Multinomial Logistic Regression);
[0233] 4) Hyperparameter setting: Polynomial cross-feature order (degree = 4), and the optimization solver uses the LBFGS algorithm;
[0234] 5) Model training: Use five-fold cross-validation;
[0235] 6) Output prediction: Select the class corresponding to the maximum probability value.
[0236] Comparative Example 22
[0237] After generating the original features, the subsequent machine learning model is the existing Wide&Deep model.
[0238] 1) Input features: Same as Step S1 of Example 3, generate 16-dimensional original features;
[0239] 2) Feature preprocessing: Standardize all original features uniformly;
[0240] 3) Model structure: Wide part (explicit interaction): Use a linear model to capture low-order explicit combined interaction relationships; Deep part (implicit interaction): Use a multi-layer feedforward neural network (MLP) to capture high-order nonlinear interaction relationships;
[0241] 4) Hyperparameters: Deep part adopts 2-layer MLP (128 and 64 hidden units) with ReLU activation function and 0.2 Dropout ratio; 2. Model optimizer adopts Adam optimizer with learning rate set to 0.001;
[0242] 5) Model training: Five-fold cross-validation is adopted, with a maximum of 50 rounds of training per fold, and an early stopping strategy of stopping when the validation accuracy does not improve for 5 rounds;
[0243] 6) Output prediction: The final classification is output by the softmax function to predict the category.
[0244] Example 23
[0245] The comparative example adopts a GPT-4 model.
[0246] 1) Input features: directly input the 414 cover original images of Example 2;
[0247] 2) Feature preprocessing: the cover original images are converted into visual embeddings by the image encoder (Vision Encoder, such as CLIP-ViT), and the text on the cover original images is processed into semantic embeddings by the built-in word segmentation and context encoding of GPT-4; no additional manual feature construction is required;
[0248] 3) Model structure: a multi-modal GPT-4 model (with visual-linguistic joint representation learning capability) is adopted, which realizes cross-modal interaction through multi-layer self-attention mechanism; a binary classification output layer (softmax) is connected at the top layer;
[0249] 4) Hyperparameters: the maximum context length is 512, and the visual input resolution is 224x224; during training, the AdamW optimizer is adopted with a learning rate of 1e-5 and a batch size of 16;
[0250] 5) Model training: no cross-validation is adopted, and one-time fine-tuning is directly performed on the entire training set;
[0251] 6) Output prediction: the model outputs the probability distribution of two categories (0 / 1), and directly selects the category corresponding to the maximum probability as the final binary classification prediction result.
[0252] The accuracy of Examples 3 and Comparative Examples 17-23 is tested respectively, and the results are shown in Table 4.
[0253] Table 4 Prediction accuracy of Examples 3 and Comparative Examples 17-23
[0254]
[0255] Comparative Example 24
[0256] This comparative example is substantially the same as Example 2, except that this comparative example does not go through the residual connected multi-layer perceptron.
[0257] Comparative Example 25
[0258] This comparative example is substantially the same as Example 2, except that this comparative example does not go through the step of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P.
[0259] Comparative Example 26
[0260] This comparative example is substantially the same as Example 2, except that this comparative example does not go through the step of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P, nor does it go through the residual connected multi-layer perceptron.
[0261] Table 5. Prediction accuracy of Example 2 and Comparative Examples 24-26
[0262]
[0263]
[0264] As can be seen from Table 5, on the basis of Comparative Example 26, the accuracy rate is improved by about 2.6% through the residual connected multi-layer perceptron processing (i.e. Comparative Example 25); the accuracy rate is improved by about 3% through the step of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P (Comparative Example 24), but as described in Example 2, on the basis of Comparative Example 26, the accuracy rate is improved by more than 8% through the steps of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P and then going through the residual connected multi-layer perceptron processing, which shows that the combination of the two steps in the scheme of the present application is better than the effect of the two steps applied separately in improving the prediction accuracy.
[0265] Comparative Example 27
[0266] This comparative example is substantially the same as Example 3, except that this comparative example does not go through the residual connected multi-layer perceptron.
[0267] Comparative Example 28
[0268] This comparative example is substantially the same as Example 3, except that this comparative example does not go through the step of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P.
[0269] Comparative Example 29
[0270] This comparative example is substantially the same as Example 3, except that this comparative example does not go through the step of applying the invertible permutation matrix M to the transpose of matrix A to obtain matrix P, nor does it go through the residual connected multi-layer perceptron.
[0271] Table 6 prediction accuracy of example 3 and comparative examples 27-29
[0272]
[0273] From table 6, it can be seen that, based on comparative example 29, the accuracy rate is improved by about 4.3% by residual connection multilayer perceptron processing (i.e. comparative example 28); the accuracy rate is improved by about 2.7% by applying the reversible permutation matrix M to transpose the matrix A to obtain the matrix P (comparative example 27), but the scheme as described in example 3, based on comparative example 29, the accuracy rate is improved by more than 8% by applying the reversible permutation matrix M to transpose the matrix A to obtain the matrix P and then performing residual connection multilayer perceptron processing, which shows that the combination of the two steps in the scheme of the present application is better than the effect of the two steps applied alone in improving the prediction accuracy.
[0274] The above embodiments are only some preferred embodiments of the present application, and cannot be used to limit the scope of protection of the present application. Any non-essential changes and substitutions made by those skilled in the art on the basis of the present application shall fall within the scope of protection of the present application.
Claims
1. A prediction method based on human-computer interaction feature mining, characterized in that, Includes the following steps: S1. Input the raw data and prompt words into the large language model GPT-4. The raw data includes modal data and / or multi-source heterogeneous data. Specifically, extract modal features from the modal data corresponding to each prompt word, and extract multi-source heterogeneous features from the multi-source heterogeneous data corresponding to each prompt word. Structure the modal features, and the structured modal features and / or the multi-source heterogeneous features constitute the original features. S2. The original features are ID-ized to construct an input matrix X. The input matrix X is fed into the embedding layer to map different original features to a unified semantic space, which is mapped to a dense representation E. A bilinear attention mechanism is used to calculate the relationship between any two original features to obtain the correlation matrix A. S3. Transpose the matrix A by applying the invertible permutation matrix M to obtain the matrix P; S4. Process the matrix A and matrix P in parallel on a dual-branch lightweight convolutional network, extract the local and global feature interactions between the original features respectively, and then flatten and concatenate the outputs of the two branches to form a concatenated vector. S5. The spliced vector in step S4 is supplemented with overall statistical information by a multilayer perceptron with residual connections, thereby outputting the prediction result.
2. The prediction method based on human-computer interaction feature mining according to claim 1, characterized in that, The raw data includes the modal data and the multi-source heterogeneous data; the modal data consists of stroke functional prognosis medical history texts and CT images; the multi-source heterogeneous data consists of laboratory biochemical indicators; and the prompts are listed by clinical experts based on medical knowledge.
3. The prediction method based on human-computer interaction feature mining according to claim 2, characterized in that, After obtaining the modal features in step S1, the modal features with high missing values, high redundancy, and weak correlation are removed by using methods such as missing rate threshold filtering, medical rationality verification, multicollinearity detection, and information gain evaluation.
4. The prediction method based on human-computer interaction feature mining according to claim 3, characterized in that, The modal features described in step S1 are structured, the categorical features of the modal features retain their original classification form, and the numerical features are discretized based on medical thresholds.
5. The prediction method based on human-computer interaction feature mining according to claim 4, characterized in that, The prompt words mentioned in step S1 are: unique identifier, age, gender, diabetes, hypertension, slurred speech, dysphagia, constipation, sleep, history of trauma or surgery, smoking, alcohol consumption, Brunnstrom staging of the left upper limb, Brunnstrom staging of the left lower limb, Brunnstrom staging of the right upper limb, Brunnstrom staging of the right lower limb, Ashworth muscle tone score of the left upper limb, Ashworth muscle tone score of the left lower limb, Ashworth muscle tone score of the right upper limb, Ashworth muscle tone score of the right lower limb, coronary artery disease, hemiplegia (left / right), cerebral hemorrhage, pons, frontal lobe, temporal lobe. Occipital lobe, frontotemporal-parietal lobe, frontotemporal lobe, ventricles, basal ganglia, cerebellum, cerebellar cortex, brainstem, thalamus, cognitive impairment, cerebral infarction, periventricular region, cerebellar hemisphere, internal capsule, insula, cerebral hemisphere, middle cerebral artery supply area, centrum semiovale, elevated blood glucose, erythrocyte sedimentation rate (ESR), direct bilirubin, indirect bilirubin, low albumin, low total cholesterol, triglycerides, uric acid, elevated C-reactive protein, lipoprotein A1 measurement, D-dimer, neutrophil count, pain limitation, lesion type, Babinski sign, fecal incontinence, urinary catheter, time orientation, time orientation and person orientation + calculation ability.
6. The prediction method based on human-computer interaction feature mining according to claim 5, characterized in that, In step S5, the spliced vector is supplemented with overall statistical information by a multilayer perceptron with residual connections, and then mapped to a five-dimensional soft-max probability to output the functional classification results of FIM levels 0-4.
7. The prediction method based on human-computer interaction feature mining according to claim 1, characterized in that, The original data is modal data, which is the original cover image, and the original cover image is classified as popular or unpopular; the prompt words are listed by cover design experts based on cover design knowledge; the large language model GPT-4 extracts the modal features by evaluating the performance of each original cover image on each prompt word, and performs scoring and quantification to obtain structured modal features.
8. The prediction method based on human-computer interaction feature mining according to claim 7, characterized in that, The prompts include: color appeal, layout hierarchy, cover composition, portrait expressiveness, title text strategy, information density and reading efficiency, emotional communication, realism, atmosphere creation, imitativeness, content challenge, recognizability, altruism, contrast, emotional tension, psychological compensation, misleading, suggestive, cognitive dissonance, attention hook design, discussion potential, voyeuristic desire, sense of identity alignment, sense of instant reward, platform aesthetic compatibility, thumbnail adaptability, and multi-image collaborative closed loop.
9. The prediction method based on human-computer interaction feature mining according to claim 1, characterized in that... The original data is the modal data, which consists of images of cultural and creative products, categorized into high-selling and low-selling products. The prompts are listed by experts based on their knowledge of the cultural and creative industry. The large language model GPT-4 extracts the modal features by evaluating the performance of each prompt on each image of a cultural and creative product, and then scores and quantifies these features to obtain structured modal features.
10. The prediction method based on human-computer interaction feature mining according to claim 9, characterized in that... The prompts include cultural value, storytelling, artistic quality, aesthetic refinement, uniqueness, practicality, fun / interactivity, technological feel, emotional resonance and memorability, topicality, personalization, series relevance, channel attributes, environmental friendliness, price, and platform.