Questionnaire index system construction method and device, equipment, medium and product

By combining the improved multiple dichotomy method and random forest algorithm with R-type clustering, the subjectivity and redundancy problems in the construction of questionnaire indicator systems in existing technologies are solved, realizing the construction of efficient and scientific indicator systems and improving the accuracy and interpretability of questionnaire data processing.

CN121745248APending Publication Date: 2026-03-27SHENYANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing methods for constructing questionnaire indicator systems suffer from strong subjectivity, insufficient capture of complex interactive relationships in questionnaire data, lack of efficient feature selection mechanisms, poor interpretability of the constructed indicator systems, and a tendency to overfit or lose information when dealing with high-dimensional questionnaire data.

Method used

An improved multivariate bisection method and random forest algorithm were used to process the questionnaire data. The data was constructed through numerical coding and sparse matrix, followed by R-type clustering to form a clear indicator system. The Cronbach's α coefficient and KMO-Bartlett test were used to ensure data quality. The random forest algorithm was used to select key questions, and R-type clustering was used to reduce redundancy among indicators.

Benefits of technology

It enables objective assessment of the importance of questionnaire variables, captures complex interactions between variables, reduces data dimensionality, avoids overfitting and information loss, and improves the scientificity and interpretability of the indicator system, making it suitable for processing high-dimensional questionnaire data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745248A_ABST
    Figure CN121745248A_ABST
Patent Text Reader

Abstract

The invention discloses a questionnaire index system construction method, a questionnaire index system construction device, equipment, a medium and a product, and relates to the field of questionnaire indexes. According to the questionnaire result, processing each option in each multi-choice question by adopting an improved multiple dichotomy to obtain a sparse matrix corresponding to each multi-choice question; performing data inspection on the codes corresponding to the options in the single-choice questions and the sparse matrixes corresponding to the multiple-choice questions, and correcting or eliminating the questions according to an inspection result; inputting the questions into a random forest model to obtain an initial index system; and R-type clustering is performed on the initial index system, and the index system of the target questionnaire is constructed according to the clustering result, so that human prejudice interference can be avoided, complex interaction among questionnaire data can be captured, the efficiency can be improved, the interpretability can be enhanced, and the situation of over-fitting or information loss can be avoided in the face of high-dimensional questionnaire data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of questionnaire indicators, and in particular to a method, apparatus, equipment, medium and product for constructing a questionnaire indicator system. Background Technology

[0002] Questionnaires are an important tool for researching and collecting feedback from specific groups. An effective questionnaire needs to be scientific, logical, and practical. Typically, questionnaires need to be quantitatively evaluated based on an indicator system to ensure their scientific validity. However, existing methods for constructing questionnaire indicator systems mainly include traditional approaches such as expert experience and statistical analysis. Expert experience relies on researchers' subjective judgment to select indicators and determine the structure, which is easily influenced by personal biases and lacks objective data support. Traditional statistical analysis methods, such as factor analysis and principal component analysis, while data-driven, have limited ability to handle nonlinear relationships between variables and struggle to effectively address multicollinearity. Furthermore, existing methods often handle indicator selection and system construction separately, lacking systematic integration, resulting in a final indicator system that is either highly redundant or insufficiently representative. These methods share the following drawbacks: 1) They are highly subjective, and the results are limited by expert knowledge; 2) They fail to capture the complex interactions in questionnaire data; 3) They lack efficient feature selection mechanisms; 4) The constructed indicator systems have poor interpretability; 5) When faced with high-dimensional questionnaire data, they are prone to overfitting or information loss, making it difficult to accurately identify truly important indicators or reasonably reflect the hierarchical relationships between indicators. These shortcomings prevent accurate evaluation of the questionnaire. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, equipment, medium and product for constructing a questionnaire indicator system, which can avoid interference from human bias, capture complex interactions between questionnaire data, improve efficiency, enhance interpretability, and prevent overfitting or information loss when dealing with high-dimensional questionnaire data.

[0004] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for constructing a questionnaire indicator system, including: obtaining multiple questionnaire results corresponding to the target questionnaire.

[0005] Each option in each single-choice question in the target survey questionnaire is numerically coded to obtain the code corresponding to each option in each single-choice question.

[0006] For any questionnaire result, based on the questionnaire result, an improved multiple dichotomy method is used to process each option in each multiple-choice question to obtain the sparse matrix corresponding to each multiple-choice question.

[0007] Data validation is performed on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

[0008] The retained questions from the target survey questionnaire corresponding to the results of each questionnaire are input into the random forest model to obtain the initial indicator system.

[0009] The initial indicator system is subjected to R-type clustering to obtain clustering results, and the indicator system of the target questionnaire is constructed based on the clustering results.

[0010] Secondly, this application provides a device for constructing a questionnaire indicator system, including: an acquisition module for acquiring multiple questionnaire results corresponding to a target questionnaire.

[0011] The single-choice question processing module is used to numerically encode each option in each single-choice question in the target survey questionnaire, thereby obtaining the code corresponding to each option in each single-choice question.

[0012] The multiple-choice processing module is used to process each option in each multiple-choice question using an improved multiple binary search method to obtain the sparse matrix corresponding to each multiple-choice question, based on the questionnaire results.

[0013] The data validation module is used to validate the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

[0014] The initial indicator system determination module is used to input the retained items of the target survey questionnaire corresponding to the results of each questionnaire into the random forest model to obtain the initial indicator system.

[0015] The indicator system construction module is used to perform R-type clustering on the initial indicator system to obtain clustering results, and to construct the indicator system of the target questionnaire based on the clustering results.

[0016] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for constructing a questionnaire indicator system.

[0017] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for constructing a questionnaire indicator system.

[0018] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for constructing a questionnaire indicator system.

[0019] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method, apparatus, device, medium, and product for constructing a questionnaire indicator system. The ensemble learning characteristics of the random forest algorithm can objectively assess the importance of each variable. By constructing a large number of decision trees for cross-validation, it avoids the interference of human bias. At the same time, its excellent nonlinear processing capability can capture the complex interactions between variables. R-type clustering scientifically groups the screened questions, effectively solving the problem of indicator redundancy and achieving efficient feature selection. Simultaneously, this method performs R-type clustering on the initial indicators, and hierarchical clustering based on the correlation between indicators, grouping strongly correlated variables into the same category, forming a clear hierarchical structure of indicators, and intuitively improving interpretability. The indicator system obtained after R-type clustering, based on the clustering results, serves as the representative indicator, significantly reducing data dimensionality, reducing redundant noise, and avoiding overfitting and information loss. The correlation coefficient matrix objectively determines the degree of correlation between indicators. The organic combination of the two methods forms a closed-loop system: random forest provides the first round of objective screening to ensure the inclusion of key questions; clustering analysis establishes the structural relationship between indicators based on dimensionality reduction, and the final indicator system is both comprehensive and concise. This application specifically addresses the dual dilemmas of subjective screening being prone to bias and simple statistics being inaccurate in traditional techniques. Through a two-stage processing approach, it ensures the scientific nature of the indicator system construction while enhancing the stability and interpretability of the results, making it particularly suitable for processing modern questionnaire data containing a large number of potentially related variables. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a method for constructing a questionnaire indicator system according to an embodiment of this application.

[0022] Figure 2 A schematic diagram illustrating the principle of a questionnaire index system construction method provided in an embodiment of this application.

[0023] Figure 3 This is a schematic diagram of the functional modules of a questionnaire index system construction device provided in an embodiment of this application.

[0024] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] This application provides a method for constructing a questionnaire indicator system for questionnaires containing only multiple-choice questions. In an exemplary embodiment, such as... Figure 1 As shown, the method for constructing the questionnaire indicator system includes: Step 201: Obtaining multiple questionnaire results corresponding to the target questionnaire.

[0028] Step 202: Number each option in each single-choice question of the target survey questionnaire to obtain the code corresponding to each option in each single-choice question.

[0029] Step 203: For any questionnaire result, based on the questionnaire result, use the improved multiple dichotomy method to process each option in each multiple-choice question to obtain the sparse matrix corresponding to each multiple-choice question.

[0030] Step 204: Perform data verification on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the verification results, modify or remove the questions in the target questionnaire to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

[0031] Step 205: Input the retained items of the target survey questionnaire corresponding to each questionnaire result into the random forest model to obtain the initial indicator system.

[0032] Step 206: Perform R-type clustering on the initial indicator system to obtain clustering results, and construct the indicator system for the target questionnaire based on the clustering results. The proposed indicator system can quantitatively evaluate the questionnaire results answered by respondents.

[0033] In practical applications, each option in the single-choice questions of the target survey questionnaire is numerically coded to obtain the code corresponding to each option in each single-choice question. Specifically, questionnaires containing only multiple-choice questions generally have two types of questions: single-choice questions and multiple-choice questions. Among them, most single-choice questions have hierarchical relationships between the options, such as education level, income range, etc. For this, the direct coding method is used for data preprocessing. In this process, the logical characteristics of the classification structure and the data analysis needs must be taken into account. First, the coding method is determined according to the variable type: nominal variable types (such as occupation type) can be one-hot coded, while ordinal variables (such as satisfaction rating) are mapped to continuous integer values ​​according to the hierarchical order (the inherent logical order or intensity gradient between the classification options within the ordinal variable) to retain potential trend correlation. The hierarchical order reflects the progressive relationship or intensity difference between the options. For example: satisfaction rating: very dissatisfied (1) → dissatisfied (2) → neutral (3) → satisfied (4) → very satisfied (5), then the five options are quantified into 1-5, 5 numbers according to the hierarchical order. The purpose of this quantification is to: preserve the order information of the options, ensuring that the numerical values ​​reflect the hierarchical relationship between categories (3>2>1); and to be compatible with statistical modeling, allowing algorithms such as regression or clustering to identify potential trends (e.g., higher satisfaction leads to lower turnover). However, it is important to note that gaps cannot be quantified. For options with implicit numerical meanings (e.g., age groups divided by numerical range), this application can directly use the median or lower limit of the interval for categorical quantification. However, it is crucial to avoid oversimplification during the processing (e.g., forcibly converting "once a week" and "once a month" to 1 and 0.25 times), to prevent distorting the data distribution. Furthermore, if the hierarchical design is unreasonable (e.g., income groupings have excessively large or overlapping spans), the coded data may still amplify classification bias. This can be addressed by combining chi-square tests or analysis of variance to verify the effectiveness of the grouping.

[0034] In another exemplary embodiment of this application, based on the questionnaire results, an improved multiple dichotomy method is used to process each option in each multiple-choice question to obtain a sparse matrix corresponding to each multiple-choice question. Specifically, this includes: based on the questionnaire results, an improved multiple dichotomy method is used to process each option in each multiple-choice question to obtain a binary variable corresponding to each option in each multiple-choice question.

[0035] Construct a sparse matrix for each multiple-choice question based on the binary variables corresponding to each option.

[0036] In practical applications, data processing for multiple-choice questions often lacks explicit hierarchical relationships or mutually exclusive frameworks between options. For example, in interest-related questions, each option may be independent and parallel. This application employs an improved multiple binary classification method to construct a sparse matrix for quantification. Specifically, each multiple-choice question is decomposed into several dummy variables, each independently represented by a "0-1" binary identifier: 0 for unselected and 1 for selected. Then, the binary variables corresponding to all options are combined into a sparse matrix. This method, by deconstructing multidimensional composite selection into a univariate binary classification problem, avoids the dimensional redundancy and matrix sparsity dilemma caused by the explosion of option permutations and combinations in traditional multiple classification coding, while also being compatible with the standardization requirements of machine learning models for feature inputs.

[0037] In another exemplary embodiment of this application, the options in the single-choice questions include options of numerical variable type, options of nominal variable type, and options of ordinal variable type. Before performing data verification on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question, and correcting or eliminating questions in the target questionnaire based on the verification results to obtain the retained questions of the target questionnaire corresponding to the questionnaire results, the process further includes: data cleaning based on the IQR algorithm on the codes corresponding to the numerical variable type options and ordinal variable type options in each single-choice question to obtain the cleaned codes corresponding to each option in each single-choice question. The IQR algorithm is used to detect outliers of numerical variables. During the cleaning process, if the questionnaire option is a numerical variable (such as age, income, rating, etc.), the original value is directly IQR cleaned (without considering the coding). If the questionnaire option is an ordinal variable type, it needs to be coded into a numerical value first, and then the reasonableness of the coded value is determined.

[0038] In practical applications, the IQR algorithm is used to clean the data for the codes corresponding to numerical and ordinal variables in each multiple-choice question, resulting in the cleaned codes for each option. Specifically, the interquartile range (IQR) algorithm is a robust outlier detection method based on the non-parametric characteristics of data distribution. Its core logic lies in dividing the data into four segments using quartiles. The difference between the third quartile (Q3, 75th percentile) and the first quartile (Q1, 25th percentile) is the IQR (Q3-Q1). This method defines the non-outlier interval as [Q1-1.5IQR, Q3+1.5IQR], and considers extreme values ​​exceeding this interval as potential outliers. Compared to mean-based detection methods such as Z-score, IQR is not limited by the data distribution pattern and suppresses the interference of a few extreme values ​​on the cleaning results by discarding data far from the median.

[0039] When using this method for data cleaning, outliers are initially identified and missing values ​​are removed or filled to ensure overall data integrity. Then, Q1 (25th quantile, the threshold below 25% of the data) and Q3 (75th quantile, the threshold above 75% of the data) are calculated using quantile functions, and the IQR is calculated. The specific expression is as follows: .

[0040] Then, outlier boundaries are defined, with LB and UB representing the lower and upper limits of the normal range, respectively, and their expressions are as follows: .

[0041] The next step is to identify and handle outliers. There are many ways to handle data that exceeds the range [LB, UB], but three main types are commonly used.

[0042] (1) Deletion: Directly remove abnormal data. This method requires evaluation of sample size loss.

[0043] (2) Boundary: Replace values ​​that exceed the upper or lower limit with the corresponding boundary values.

[0044] (3) Manual verification: For extreme values ​​that are abnormal but have physical significance, such as age, manual secondary verification is required in conjunction with logical rules.

[0045] In another exemplary embodiment of this application, after data cleaning is performed on the codes corresponding to the numerical variable type options and the ordinal variable type options in each multiple-choice question based on the IQR algorithm to obtain the cleaned codes corresponding to each option in each multiple-choice question, the method further includes: verifying and adjusting the cleaned codes corresponding to each option in each multiple-choice question.

[0046] In practical applications, the cleaning codes corresponding to each option in each single-choice question are verified and adjusted, specifically: (1) Distribution reanalysis: draw a box plot of the cleaned data to confirm whether outliers have been reasonably eliminated.

[0047] (2) Dynamic parameter tuning: For fields with strong skewed distribution (data distribution deviates significantly from the normal distribution, resulting in data concentrated at one extreme while the other extreme has a long tail phenomenon), such as income range, the threshold coefficient can be adjusted in combination with the characteristics of the variable (e.g., change to 2×IQR to reduce false deletions). The adjusted threshold coefficient can also be used for outlier detection and processing.

[0048] In practical applications, it is important to avoid mechanically applying IQR-based algorithms for data cleaning. Instead, a comprehensive judgment should be made in conjunction with questionnaire design logic and domain knowledge. For outlier data, multivariate correlation checks and cross-validation should be performed to prevent accidental deletion of data. Manual review should be conducted to verify any potential questionnaire data entry errors. Furthermore, a detailed record of the processing should be kept, and a list of outlier samples should be maintained for future traceability and analysis.

[0049] In another exemplary embodiment of this application, data verification is performed on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the verification results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results. Specifically, data verification is performed on the cleaned codes corresponding to each option in each single-choice question, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question. Based on the verification results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

[0050] In another exemplary embodiment of this application, data verification is performed on the cleaned codes corresponding to each option in each single-choice question, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question. Based on the verification results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results. Specifically, this includes: calculating the Cronbach's α coefficient of the target questionnaire corresponding to the questionnaire results based on the cleaned codes corresponding to each option in each single-choice question, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question; correcting or removing the questions in the target questionnaire to obtain the retained target questionnaire corresponding to the questionnaire results.

[0051] Based on the cleaned codes corresponding to each option in each single-choice question of the retention target questionnaire corresponding to the questionnaire results, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question, calculate the KMO value and Bartlett P value of each item in the retention target questionnaire corresponding to the questionnaire results.

[0052] Based on the KMO value and Bartlett P value of each item in the target questionnaire corresponding to the questionnaire results, the items in the target questionnaire corresponding to the questionnaire results are modified or removed to obtain the retained items of the target questionnaire corresponding to the questionnaire results.

[0053] In practical applications, Cronbach's α coefficient is the core indicator for measuring the internal consistency of a scale or questionnaire. It is used to assess the reliability of a questionnaire by calculating the correlation between items. Usually, a Cronbach's α coefficient ≥ 0.7 is required. If the Cronbach's α coefficient ≥ 0.7 (0.6 is the minimum tolerance threshold), it indicates that the items have high homogeneity and the data reliability meets the analysis requirements. If the Cronbach's α coefficient < 0.7, it indicates poor consistency between items, which may be measuring different constructs or the item descriptions may be problematic. The items that affect consistency the most need to be corrected or removed. Specific process: (1) Removal: For each item, delete the item and recalculate the Cronbach's α coefficient. Observe whether the Cronbach's α coefficient of the questionnaire increases. If the Cronbach's α coefficient of the questionnaire increases significantly after deleting an item, the item should be removed first. (2) Correction: If some items are correlated with the total score of other items (Corrected Item-Total Correlation, CITC) < 0.3, it indicates that the item measurement is inconsistent and should be restated or deleted. This operation can be performed using SPSS by running "Reliability Analysis". Check the "Cronbach's α if Item Deleted" column and remove the items that have the greatest impact on the Cronbach's α coefficient.

[0054] However, the Cronbach's α coefficient method has limitations, including its sensitivity to the number of questions; a higher number of questions may result in an artificially inflated Cronbach's α coefficient. Furthermore, this method cannot reflect complexity beyond a single dimension. Therefore, the KMO test and Bartlett's test are combined to address this issue.

[0055] The KMO test and Bartlett's test of sphericity are used to verify whether data is suitable for dimensionality reduction methods such as factor analysis. The KMO value (0-1) measures partial correlation between variables, typically requiring a KMO ≥ 0.6 (ideally above 0.8). The Bartlett's test of sphericity requires a significance level of Bartlett p < 0.05, indicating a significant correlation between variables, thus demonstrating data validity and allowing for the extraction of common factors. Note that a low KMO may indicate insufficient sample size or irrelevance between items; in such cases, the scale structure needs to be adjusted based on domain knowledge. Using both tests together allows for a systematic assessment of data reliability and analytical suitability. When KMO ≥ 0.6 and Bartlettp < 0.05, the data is suitable for factor analysis, and dimensionality reduction can be directly performed on the variables (items). If KMO < 0.6 but Bartlettp < 0.05, low-correlation variables (MSA < 0.5) need to be deleted or high-correlation variables need to be merged and retested. If Bartlettp is not significant (Bartlett ≥ 0.05), it indicates that there is insufficient correlation between variables, and factor analysis should be abandoned, replaced by dimensional analysis or adjustment of the questionnaire design. If the sample size is insufficient (N < number of variables × 5), the number of variables can be reduced or PLS regression can be used instead. Finally, the rationality of the factor structure needs to be verified in conjunction with business logic. Specifically...

[0056] (1) KMO standard (Kaiser-Meyer-Olkin Measure).

[0057] ①KMO ≥ 0.8: Very suitable for factor analysis.

[0058] ② 0.7 ≤ KMO < 0.8: acceptable, but some problems may need optimization.

[0059] ③ KMO < 0.7: Insufficient information sharing among items; consider revising or removing items with low KMO. If some items have extremely low relevance, they may be measuring different dimensions and should be deleted. Revision mainly considers reclassifying items; for example, if an item covers multiple constructs (such as simultaneously measuring "satisfaction" and "loyalty"), the questionnaire needs to be split.

[0060] (2) Bartlett's Test of Sphericity.

[0061] Hypothesis test: H0 (null hypothesis): The items are independent of each other (not suitable for factor analysis).

[0062] H1 (Alternative Hypothesis): There is a correlation between the items (suitable for factor analysis).

[0063] Judgment criterion: p<0.05 → reject H0, the question can proceed to factor analysis.

[0064] p≥0.05 → This indicates that there is no significant correlation between the questions, and some questions should be redesigned or removed.

[0065] Specifically, based on the KMO value and Bartlett's P value of each item, items in the target survey questionnaire are modified or eliminated. The specific adjustment method is as follows: If KMO < 0.7 and p ≥ 0.05, it indicates a poor overall item structure, and the most irrelevant items (e.g., correlation with other items < 0.2) should be eliminated. The item wording should be adjusted (e.g., to increase clarity and avoid ambiguity). Data should be collected again (if the number of items is insufficient, it may be necessary to expand the collection).

[0066] In practical applications, the Random Forest (RF) algorithm is a statistical learning theory that uses bootstrap resampling to extract multiple samples from the original data (each sample corresponds to a single questionnaire result for the target survey). A decision tree model is built for each bootstrap sample, and stable statistics such as the mean or Gini index of the corresponding attributes are estimated. The predictions from multiple decision trees are combined, and a final result is obtained through voting. Index selection based on the Random Forest algorithm can optimize questionnaire design, improve data quality, and enhance analytical efficiency. In questionnaire data, Random Forest can calculate the feature importance of each item, filtering out the key issues that have the greatest impact on the research objective, thereby simplifying redundant questions, shortening questionnaire length, and improving the data quality of core questions. Furthermore, Random Forest can detect abnormal responses (such as randomly filled options), assisting in data cleaning and ensuring the credibility of the analysis results. Ultimately, questionnaires optimized based on this method can reduce the burden on respondents and enhance the accuracy of analytical models (such as prediction and classification), making it suitable for scenarios such as market research and psychological assessment.

[0067] Random forest models, with their ensemble learning advantages and feature processing mechanisms, have become an ideal tool for dealing with high-dimensional data problems. This model significantly reduces the risk of a single decision tree overfitting high-dimensional sparse data by constructing multiple decision trees and aggregating the results. Simultaneously, it overcomes the curse of dimensionality by using bootstrap resampling and random subspace methods (randomly selecting some features to generate tree nodes). Its inherent feature importance assessment can automatically filter key variables, quickly identifying strongly correlated features in thousands of dimensions of data, improving model interpretability. Furthermore, random forests are highly robust to complex interactions, noise interference, and missing values ​​common in high-dimensional data, and support parallel computation, making them particularly suitable for research fields such as bioinformatics and text mining that require processing tens of thousands of features. By balancing efficiency and accuracy, this model demonstrates superior performance compared to traditional linear models and non-ensemble tree models in high-dimensional scenarios. The specific steps are as follows: First, combine the codes corresponding to each option of the single-choice questions and the sparse matrices corresponding to the multiple-choice questions to construct the feature matrix X. Each row of this matrix represents a questionnaire result (i.e., a sample), and each column represents a question option (i.e., a feature variable). For example, a retention questionnaire containing 5 single-choice questions and 1 multiple-choice question (with 4 options) will generate a feature matrix with 5 + 4 = 9 columns. Assume a series of decision trees... The resulting forest is denoted as {h(x)}.

[0068] Definition 1: Edge function: .

[0069] in, Let I() represent the average, I() represent the indicator function, and max represent the maximum value; Y represents the correct classification vector; and j represents the incorrect classification vector. The marginal function represents the degree to which the number of votes for the correct classification exceeds the maximum number of votes for the incorrect classification; it can be concluded that the larger the marginal function, the higher the confidence of the classification.

[0070] Definition 2: Generalization error .

[0071] in, It is a mathematical symbol for probability, and the subscripts X and Y represent the definition space of probability.

[0072] Definition 3: The edge function of a random forest is .

[0073] in, To determine the probability of the correct classification; The maximum probability of misclassification for other categories, i.e., the target variable has two or more categories, where c represents the number of misclassifications.

[0074] For each decision tree, during the construction of the random forest, there is an initial dataset and an unsampled dataset. Let the unsampled dataset be denoted as... . For an input random vector x, in The categories of voting in China are: The proportion is then: .

[0075] In the above formula, the numerator is the sum of the number of correct classifications for each decision tree and its corresponding unsampled dataset; the denominator is the sum of the number of samples in all unsampled datasets.

[0076] Definition 4: The strength of a random forest {h(x)} is the expectation of the marginal function of the random forest, that is: .

[0077] Where n represents the number of categories, This represents an estimate of the probability of a random forest correctly classifying a data type. This represents the maximum probability of a random forest correctly classifying a data type.

[0078] Definition 5: The average correlation between trees in a random forest is defined as the variance of the marginal function. Divide by the square of the forest's standard deviation ,Right now: .

[0079] Where k represents the total number of decision trees, As OOB estimation; yes The OOB estimation. Its specific expression is: and .

[0080] In the formula, x i Let i represent the i-th sample. Let y represent the sample probability and y represent the class. In order to make in the training set The class with the largest estimated value that is different from the class of y. Its expression is: .

[0081] In practical applications, particularly in data-driven research, constructing evaluation indicator systems often faces the challenge of excessive complexity. Overly granular indicator classifications can lead to semantic overlap and redundant information, increasing data storage costs and significantly raising computational and analytical loads due to multicollinearity affecting model accuracy. For scenarios like questionnaire surveys with high data dimensionality and large sample sizes, R-type clustering can achieve efficient dimensionality reduction by mining the correlation structure between indicators. This method uses all evaluation indicators in the initial indicator system as the analysis object. First, it calculates the correlation matrix between each pair of variables (evaluation indicators) using correlation coefficients (calculated based on the codes corresponding to the options in the questions or a sparse matrix). In the initial stage of clustering, each indicator is considered an independent subclass; subsequently, it iterative "aggregation-reorganization" operations are performed: in each round, the two groups of indicators with the largest absolute values ​​of their current correlation coefficients are selected for priority merging, gradually forming a hierarchical nested classification tree structure, ultimately converging until all indicators are classified into one large category. In this process, the optimal number of clusters can be determined using criteria such as silhouette coefficient and intragroup correlation coefficient, combined with the elbow algorithm of dendritic diagrams or statistical tests. Classifications are then truncated at a pre-set confidence level. Dendritic diagram visualization can help identify core indicator groups, alleviating the problem of important indicators being overwhelmed in high-dimensional data, and providing a reasonable structured foundation for subsequent analysis. The specific steps are as follows.

[0082] Assuming the original index set There are a total of n indicators. The sample set is... There are m samples in total, each sample is n-dimensional. Let the i-th sample and the j-th sample be represented as follows: .

[0083] Where, x in This represents the encoding, sparse matrix, or original value corresponding to the options in the n questions of the i-th sample.

[0084] Step (1): Initialize each metric as an independent subclass, denoted as The correlation between clusters is quantitatively described by the average correlation coefficient of the clusters, i.e., for two clusters... , Class average correlation coefficient for: .

[0085] in The Pearson correlation coefficient is expressed as follows: .

[0086] N represents the total number of samples. This indicates that the r-th sample is in index s i The value on, This indicates that the r-th sample is in index s j The value on.

[0087] The class mean correlation coefficient matrix is ​​calculated based on the class mean correlation coefficient between each pair of classes. Initialize to: i≠j and .

[0088] Step (2): Find Let the largest term in the middle be denoted as . .Will , Merge them into one class, and update the class average correlation coefficient matrix according to the calculation formula in step (1). .

[0089] Step (3): Repeat step (2) until all indicators are clustered into a single cluster. Single clustering means that all indicators to be analyzed are grouped into the same clustering algorithm at once, without pre-stratification or segmentation.

[0090] Step (4): Select an appropriate confidence level and output the clustering results. .

[0091] in, Indicates clustering results The l-th category. R-type clustering analysis can objectively reflect the inherent relationships between indicators. After R-type clustering of questionnaire data, the structure of the indicator system becomes more intuitive, reducing the difficulty of data processing.

[0092] In practice, questionnaires typically include four types of questions: objective multiple-choice questions to quickly capture common viewpoints; scale-based questions to deeply quantify respondents' attitudes; open-ended questions to explore the potential personalized needs of the research group; and mixed question types to balance efficiency and information richness. Among these, questionnaires containing only multiple-choice questions have significant efficiency advantages: on the one hand, respondents do not need to spend time organizing their answers, resulting in high efficiency, and the structured data can be directly encoded and statistically analyzed, avoiding the time cost of analyzing open-ended text; on the other hand, they reduce bias by pre-setting standardized answer ranges for options, avoiding ambiguity in subjective expressions or errors in interpreting open-ended questions, and improving data consistency; in addition, this type of questionnaire is suitable for large-scale sampling and mobile questionnaires, and can still cover diverse needs through multiple-option design, meeting the reliability and validity requirements of quantitative research, making it particularly suitable for commercial or policy research scenarios that require rapid decision-making.

[0093] It is important to note that while questionnaires consisting solely of objective multiple-choice questions have many advantages, they still face significant challenges in data analysis. The most significant challenge lies in the limitations of the data dimensions; while seemingly highly structured options facilitate statistical analysis, they may lead to conclusions that deviate from reality. Furthermore, without pre-optimized logical validation rules, the data cleaning stage requires manual removal of contradictory options, resulting in significant cost and inefficiency. Implicit biases may infiltrate option combinations, necessitating calibration using advanced statistical models or external data, further increasing analytical complexity. Improper data processing can lead to inaccurate results, consequently impacting decision-making and research conclusions. Therefore, such questionnaires often require the use of mixed research methods or tools such as Bayesian networks to compensate for the interpretability limitations of closed-ended data and achieve scientifically sound data processing.

[0094] To address this, this application first performs structured numerical transformations for single-choice and multiple-choice questions based on their characteristics: single-choice questions utilize direct coding to achieve targeted quantification based on the hierarchical relationship between options; multiple-choice questions employ an improved multiple dichotomy method to construct a binary matrix representing the multiple-choice features. Subsequently, a questionnaire data processing method based on the IQR algorithm is used for data cleaning, and the reliability is verified using Cronbach's α coefficient test and the validity is validated using the KMO-Bartlett test. Based on this, the random forest method is used to screen core indicators. After calculation, the indicators are ranked from most important to least important, and a subset is selected as evaluation indicators as needed. R-type clustering is then used to hierarchically divide these indicators, establishing a multi-level comprehensive evaluation indicator system. For example... Figure 2 As shown, the evaluation system is formed through data preprocessing (numericalization), data cleaning, data verification, and then the selection and clustering of evaluation indicators. These steps effectively improve the accuracy, efficiency, and adaptability of questionnaire data processing, providing a scientific and reasonable solution for data analysis.

[0095] This application proposes a method for numerical processing of multimodal questionnaire data. It uses ordered coding and an improved multiple dichotomy method to convert the option data of single-choice and multiple-choice questions in the questionnaire into numerical values, ensuring the uniformity of data types for different types of questions and the feasibility of statistical analysis.

[0096] This application proposes a questionnaire evaluation index selection method based on the random forest algorithm. By combining the prediction results of multiple decision trees, it effectively reduces the overfitting risk of single-tree models and enhances the stability of the evaluation system. Furthermore, the random forest algorithm is highly adaptable to high-dimensional nonlinear data and can efficiently handle multiple types of data and complex interaction relationships in questionnaires, improving the objectivity and scientific rigor of the index system and ultimately making the evaluation results more interpretable and valuable for application.

[0097] This application proposes an evaluation index classification method based on R-type clustering. By analyzing the correlation coefficients between indicators, highly correlated indicators are aggregated into the same category, thereby eliminating redundant information and avoiding the problem of repeatedly measuring the same potential dimension. Furthermore, cluster analysis can intuitively identify the inherent relationship structure between indicators, selecting the most representative core indicators from each category. This significantly reduces the number of indicators and computational complexity while ensuring that key information is not lost, thus significantly improving the scientific rigor and simplicity of the questionnaire evaluation index system.

[0098] Based on the same inventive concept, this application also provides a questionnaire indicator system construction device for implementing the questionnaire indicator system construction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the questionnaire indicator system construction device provided below can be found in the limitations of the questionnaire indicator system construction method described above, and will not be repeated here.

[0099] In one exemplary embodiment, such as Figure 3 As shown, a device for constructing a questionnaire indicator system is provided, including: an acquisition module for acquiring multiple questionnaire results corresponding to a target questionnaire.

[0100] The single-choice question processing module is used to numerically encode each option in each single-choice question in the target survey questionnaire, thereby obtaining the code corresponding to each option in each single-choice question.

[0101] The multiple-choice processing module is used to process each option in each multiple-choice question using an improved multiple binary search method to obtain the sparse matrix corresponding to each multiple-choice question, based on the questionnaire results.

[0102] The data validation module is used to validate the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

[0103] The initial indicator system determination module is used to input the retained items of the target survey questionnaire corresponding to the results of each questionnaire into the random forest model to obtain the initial indicator system.

[0104] The indicator system construction module is used to perform R-type clustering on the initial indicator system to obtain clustering results, and to construct the indicator system of the target questionnaire based on the clustering results.

[0105] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores data for constructing a questionnaire indicator system. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements a method for constructing a questionnaire indicator system.

[0106] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0107] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method embodiments.

[0108] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method embodiments.

[0109] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method embodiments.

[0110] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0111] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0112] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.

[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for constructing a questionnaire indicator system, characterized in that, The method for constructing the questionnaire indicator system includes: Obtain the results of multiple questionnaires corresponding to the target survey questionnaire; Each option in each single-choice question in the target survey questionnaire is numerically coded to obtain the code corresponding to each option in each single-choice question; For any questionnaire result, based on the questionnaire result, the improved multiple dichotomy method is used to process each option in each multiple-choice question to obtain the sparse matrix corresponding to each multiple-choice question. Data validation is performed on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results. Input the retained questions from the target survey questionnaire corresponding to the results of each questionnaire into the random forest model to obtain the initial indicator system; The initial indicator system is subjected to R-type clustering to obtain clustering results, and the indicator system of the target questionnaire is constructed based on the clustering results.

2. The method for constructing a questionnaire indicator system according to claim 1, characterized in that, The options in the single-choice questions include options of numerical variable type, nominal variable type, and ordinal variable type. Data validation is performed on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are revised or eliminated to obtain the retained questions of the target questionnaire corresponding to the questionnaire results. This also includes: The IQR algorithm is used to clean the codes corresponding to the numerical and ordinal variable types of options in each multiple-choice question, resulting in the cleaned codes for each option in each multiple-choice question.

3. The method for constructing a questionnaire indicator system according to claim 2, characterized in that, Data validation is performed on the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results, specifically: Data validation is performed on the cleaned codes corresponding to each option in each single-choice question, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results.

4. The method for constructing a questionnaire indicator system according to claim 1, characterized in that, Based on the questionnaire results, an improved multiple binary search method was used to process each option in each multiple-choice question to obtain the sparse matrix corresponding to each multiple-choice question, specifically including: Based on the questionnaire results, an improved multiple dichotomy method was used to process each option in each multiple-choice question to obtain the binary variables corresponding to each option in each multiple-choice question. Construct a sparse matrix for each multiple-choice question based on the binary variables corresponding to each option.

5. The method for constructing a questionnaire indicator system according to claim 3, characterized in that, Data validation is performed on the cleaned codes corresponding to each option in each single-choice question, the codes corresponding to each option of nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results, specifically including: Based on the cleaning code corresponding to each option in each single-choice question, the code corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question, calculate the Cronbach's α coefficient of the target questionnaire corresponding to the questionnaire results, modify or remove the items in the target questionnaire, and obtain the retained target questionnaire corresponding to the questionnaire results. Based on the cleaned codes corresponding to each option in each single-choice question of the retention target questionnaire corresponding to the questionnaire results, the codes corresponding to each option of the nominal variable type in each single-choice question, and the sparse matrix corresponding to each multiple-choice question, calculate the KMO value and Bartlett P value of each item in the retention target questionnaire corresponding to the questionnaire results. Based on the KMO value and Bartlett P value of each item in the target questionnaire corresponding to the questionnaire results, the items in the target questionnaire corresponding to the questionnaire results are modified or removed to obtain the retained items of the target questionnaire corresponding to the questionnaire results.

6. The method for constructing a questionnaire indicator system according to claim 2, characterized in that, After data cleaning based on the IQR algorithm to obtain the cleaned codes for each option in each multiple-choice question, the following steps are also included: The cleaning codes corresponding to each option in each multiple-choice question were verified and adjusted.

7. A device for constructing a questionnaire indicator system, characterized in that, The questionnaire indicator system construction device includes: The acquisition module is used to acquire the results of multiple questionnaires corresponding to the target questionnaire; The single-choice question processing module is used to numerically encode each option in each single-choice question in the target survey questionnaire, so as to obtain the code corresponding to each option in each single-choice question; The multiple-choice processing module is used to process each option in each multiple-choice question using an improved multiple binary search method to obtain the sparse matrix corresponding to each multiple-choice question for any questionnaire result. The data validation module is used to validate the codes corresponding to each option in each single-choice question and the sparse matrix corresponding to each multiple-choice question. Based on the validation results, the questions in the target questionnaire are corrected or removed to obtain the retained questions of the target questionnaire corresponding to the questionnaire results. The initial indicator system determination module is used to input the retained items of the target survey questionnaire corresponding to the results of each questionnaire into the random forest model to obtain the initial indicator system; The indicator system construction module is used to perform R-type clustering on the initial indicator system to obtain clustering results, and to construct the indicator system of the target questionnaire based on the clustering results.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the questionnaire indicator system construction method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the questionnaire indicator system construction method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the questionnaire indicator system construction method as described in any one of claims 1-6.