Questionnaire analysis method and system based on graph mining, electronic device and medium
By constructing and analyzing a core attribute network based on graph mining, the problem of mining deep-level core needs attributes in questionnaire data in existing technologies is solved, and questionnaire analysis with high credibility and interpretability is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2022-10-14
- Publication Date
- 2026-04-17
AI Technical Summary
Existing questionnaire analysis methods lack highly reliable in-depth core demand attribute mining analysis, principal component analysis has poor effectiveness and interpretability, and cannot effectively reduce the impact of extreme data.
We employ a graph mining-based approach to analyze the influence of different core attributes by establishing a questionnaire matrix, cleaning data, calculating attribute similarity, constructing a core attribute network, and performing graph mining.
It improves the interpretability of survey data, can intuitively show the relationships between various dimensions of data, reduces the impact of extreme data, and has high credibility and applicability.
Smart Images

Figure CN115544316B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to a graph mining-based questionnaire analysis method, system, electronic device, and medium. Background Technology
[0002] Currently, the mainstream analytical methods for questionnaire surveys include descriptive statistical analysis, reliability analysis, and validity analysis. These methods are mainly used for basic frequency analysis, data reliability analysis, and data validity analysis of questionnaire data. However, there is a lack of highly reliable analytical methods for mining and analyzing the deep-seated core needs attributes of questionnaire surveys.
[0003] The development of artificial intelligence technologies, represented by machine learning and data mining, has brought new tools to the analysis of questionnaires. The latest trend in mining and analyzing the deep-seated core needs attributes of questionnaires is to apply principal component analysis (PCA) to the analysis, reducing the dimensionality of the questionnaire data to identify the factors with the greatest influence. However, the effectiveness and interpretability of PCA are often unsatisfactory, and it cannot effectively preserve data information during transformation. Furthermore, PCA often exhibits ambiguity in interpreting the meaning of latent variables, resulting in poor interpretability. Summary of the Invention
[0004] The purpose of this invention is to provide a graph mining-based questionnaire analysis method, system, electronic device, and medium that can reduce the impact of extreme data in questionnaires and improve the interpretability of questionnaire data.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A graph mining-based questionnaire analysis method, the method comprising:
[0007] Obtain the survey questionnaire data to be analyzed;
[0008] Based on the survey data to be analyzed, a survey matrix is established; the row vectors of the survey matrix represent the ratings of the same user for different attributes; the column vectors of the survey matrix represent the ratings of different users for the same attribute.
[0009] The questionnaire matrix was cleaned to obtain the scoring matrix;
[0010] Based on the rating matrix, the similarity between different attribute ratings is calculated, and a core attribute adjacency matrix is established.
[0011] A core attribute network is established using the attributes in the core attribute adjacency matrix as nodes in the network.
[0012] Graph mining is performed on the core attribute network to obtain the minimum core attribute network;
[0013] Based on the aforementioned minimal core attribute network, the influence of different core attributes is analyzed to obtain the survey questionnaire analysis results.
[0014] Optionally, establishing a questionnaire matrix based on the questionnaire data to be analyzed specifically includes:
[0015] When the attribute ratings in the questionnaire to be analyzed correspond to multiple-choice questions, each option in the multiple-choice questions is converted into a Boolean single-choice question;
[0016] Based on the Boolean value of the single-choice question, determine the attribute score corresponding to the multiple-choice question;
[0017] When the item corresponding to the attribute score in the questionnaire to be analyzed is a multiple-choice question, the attribute score corresponding to the multiple-choice question is determined based on the attribute value corresponding to the multiple-choice question.
[0018] A questionnaire matrix is established based on the attribute scores corresponding to the multiple-choice questions and the attribute scores corresponding to the single-choice questions.
[0019] Optionally, the step of cleaning the questionnaire matrix to obtain the scoring matrix specifically includes:
[0020] The range of values for the attributes in the questionnaire matrix is normalized by applying min-max normalization to obtain the normalized questionnaire matrix.
[0021] Based on the normalized questionnaire matrix, different attributes are coded to obtain attribute codes;
[0022] The missing values in the attribute code are filled in to obtain the filled attribute code;
[0023] The outliers and noise data in the imputed attribute codes are removed to obtain the scoring matrix.
[0024] Optionally, the step of calculating the similarity between different attribute scores based on the rating matrix and establishing a core attribute adjacency matrix specifically includes:
[0025] Based on the rating matrix, the similarity between different attribute ratings is calculated to obtain the correlation coefficient matrix of each attribute.
[0026] From the correlation coefficient matrix of each attribute, attributes that meet the set threshold are selected to obtain the core user requirement attributes;
[0027] Based on the similarity between the user's core requirement attributes, a core attribute adjacency matrix is established.
[0028] Optionally, Pearson similarity can be used to calculate the similarity between different attribute scores.
[0029] Optionally, the step of performing graph mining on the core attribute network to obtain the minimum core attribute network specifically includes:
[0030] Recursively delete the node with the minimum degree in the core attribute network to obtain the current core attribute network;
[0031] Based on the current core attribute network, update the core attribute adjacency matrix to obtain the updated core attribute adjacency matrix;
[0032] Based on the updated core attribute adjacency matrix, establish the updated core attribute network;
[0033] When the node with the minimum degree in the updated core attribute network is deleted, the updated core attribute network becomes an empty network, and the current core attribute network becomes the minimum core attribute network.
[0034] A graph mining-based questionnaire analysis system, applied to the aforementioned graph mining-based questionnaire analysis method, the system comprising:
[0035] The acquisition module is used to acquire the survey questionnaire data to be analyzed.
[0036] The questionnaire matrix building module is used to build a questionnaire matrix based on the questionnaire data to be analyzed; the row vector of the questionnaire matrix represents the ratings of the same user on different attributes; the column vector of the questionnaire matrix represents the ratings of different users on the same attribute.
[0037] The data cleaning module is used to clean the questionnaire matrix to obtain a scoring matrix.
[0038] The core attribute adjacency matrix building module is used to calculate the similarity between different attribute scores based on the scoring matrix and build the core attribute adjacency matrix.
[0039] The core attribute network establishment module is used to establish a core attribute network by using the attributes in the core attribute adjacency matrix as nodes in the network.
[0040] The graph mining module is used to perform graph mining on the core attribute network to obtain the minimum core attribute network.
[0041] The analysis module is used to analyze the influence of different core attributes based on the minimum core attribute network and obtain the survey questionnaire analysis results.
[0042] Optionally, the core attribute adjacency matrix establishment module includes:
[0043] The similarity calculation submodule is used to calculate the similarity between different attribute scores based on the rating matrix, and obtain the correlation coefficient matrix of each attribute.
[0044] The filtering submodule is used to filter out attributes that meet the set threshold from the attribute correlation coefficient matrix to obtain the core user requirement attributes.
[0045] A submodule is established to create a core attribute adjacency matrix based on the similarity between the user's core requirement attributes.
[0046] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor running the computer program to enable the electronic device to perform the graph mining-based questionnaire analysis method described above.
[0047] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described graph mining-based questionnaire analysis method.
[0048] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0049] This invention provides a graph mining-based questionnaire analysis method, comprising: acquiring questionnaire data to be analyzed; establishing a questionnaire matrix based on the questionnaire data; the row vectors of the questionnaire matrix representing the ratings of the same user on different attributes; the column vectors of the questionnaire matrix representing the ratings of different users on the same attribute; cleaning the questionnaire matrix to obtain a rating matrix; calculating the similarity between ratings of different attributes based on the rating matrix to establish a core attribute adjacency matrix; using the attributes in the core attribute adjacency matrix as nodes in the network to establish a core attribute network; performing graph mining on the core attribute network to obtain a minimal core attribute network; and analyzing the influence of different core attributes based on the minimal core attribute network to obtain the questionnaire analysis results. This invention, through graph mining-based analysis, can intuitively show the relationships between various dimensions of the data. The graph mining algorithm continuously narrows down the core requirement attribute graph, ultimately obtaining the minimal core requirement attribute graph. Furthermore, this method is highly intuitive, has good interpretability, is less affected by extreme data in the questionnaire, and has strong applicability. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 A flowchart of the graph mining-based questionnaire analysis method provided by this invention;
[0052] Figure 2 An analysis flowchart of a specific embodiment of the present invention;
[0053] Figure 3 This is a schematic diagram of the user core requirement attribute network provided by the present invention;
[0054] Figure 4 This is a module diagram of the graph mining-based questionnaire analysis system provided by the present invention.
[0055] Symbol explanation:
[0056] 1-Acquisition Module, 2-Questionnaire Matrix Building Module, 3-Cleaning Module, 4-Core Attribute Adjacency Matrix Building Module, 5-Core Attribute Network Building Module, 6-Graph Mining Module, 7-Analysis Module. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] The purpose of this invention is to provide a graph mining-based questionnaire analysis method, system, electronic device, and medium that can reduce the impact of extreme data in questionnaires and improve the interpretability of questionnaire data.
[0059] Extensive research has been conducted on existing questionnaire analysis methods, including descriptive statistical analysis and reliability coefficient analysis. However, a highly reliable analytical method for uncovering the deep-seated core needs of questionnaire users remains lacking. This invention utilizes graph mining algorithms from data mining to mine and analyze the deep-seated core needs of questionnaires. To deploy this graph mining algorithm, a specific questionnaire is used as the source dataset. First, the correlation coefficients between items in each module of the source dataset are calculated. Then, based on graph mining analysis methods, the core needs attributes of the target dataset are mined and analyzed. Through graph mining analysis, researchers can effectively uncover the deep-seated core needs attributes in questionnaires, demonstrating high reliability, interpretability, and wide applicability.
[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] Example 1
[0062] like Figure 1 and Figure 2 As shown, this invention provides a graph mining-based questionnaire analysis method, the method comprising:
[0063] Step S1: Obtain the survey questionnaire data to be analyzed.
[0064] Step S2: Based on the questionnaire data to be analyzed, establish a questionnaire matrix; the row vector of the questionnaire matrix represents the ratings of the same user for different attributes; the column vector of the questionnaire matrix represents the ratings of different users for the same attribute.
[0065] S2 specifically includes:
[0066] Step S21: When the item corresponding to the attribute score in the questionnaire to be analyzed is a multiple-choice question, convert each option in the multiple-choice question into a Boolean single-choice question.
[0067] Step S22: Determine the attribute score corresponding to the multiple-choice question based on the Boolean value of the single-choice question.
[0068] Step S23: When the item corresponding to the attribute score in the questionnaire to be analyzed is a single-choice question, determine the attribute score corresponding to the single-choice question based on the attribute value corresponding to the single-choice question.
[0069] Step S24: Establish a questionnaire matrix based on the attribute scores corresponding to the multiple-choice questions and the attribute scores corresponding to the single-choice questions.
[0070] In practical applications, the obtained survey data to be analyzed is used as the source dataset, and a survey matrix is built based on the source dataset. Each row of the matrix represents a user's rating vector, and each value in the matrix represents a user's rating of a certain attribute. The range of attribute values in the matrix is not entirely the same; for example, some single-choice questions correspond to attribute values ranging from natural numbers between 1 and 6, while some single-choice questions correspond to attribute values only 1 and 2. Multiple-choice questions in the survey are converted into several single-choice questions, and the attribute value corresponding to each single-choice question is of Boolean type, that is, a value of 0 or 1. For example, if the user's options for "What sports do you usually like?" are "Playing ball" and "Swimming", this question is converted into two single-choice questions: "Do you like playing ball?" and "Do you like swimming?" The option value corresponding to each single-choice question is of Boolean type, that is, 0 or 1.
[0071] Step S3: Clean the questionnaire matrix to obtain the scoring matrix.
[0072] S3 specifically includes:
[0073] Step S31: Apply min-max standardization to normalize the value range of the attributes in the questionnaire matrix to obtain the normalized questionnaire matrix.
[0074] Step S32: Based on the normalized questionnaire matrix, encode the different attributes to obtain attribute codes.
[0075] Step S33: Fill in the missing values in the attribute code to obtain the filled attribute code.
[0076] Step S34: Remove outliers and noise data from the imputed attribute encoding to obtain the scoring matrix.
[0077] In practical applications, the source dataset is encoded according to the questionnaire matrix in step S2. Different encoding methods are used for different item attributes, such as discrete encoding for ordinal attributes. Then, missing values are filled in, continuous attributes are filled in with mean, and ordinal attributes are filled in with mode. Finally, noise analysis is performed on the data to remove outliers and noisy data.
[0078] Because different attributes have different value ranges, min-max standardization is used to normalize the value ranges of all selected attributes. Therefore, step S34 includes applying a transformation function to normalize the matrix after removing outliers and noise data from the imputed attribute encoding. The transformation function is:
[0079]
[0080] Where, x i It is the original value of the i-th attribute, x max and x min These represent the maximum and minimum values of the attribute in the dataset, respectively. norm This represents the normalized value of the attribute. After normalization using a transformation function, the resulting matrix is the rating matrix.
[0081] Step S4: Based on the rating matrix, calculate the similarity between different attribute ratings and establish a core attribute adjacency matrix. Pearson similarity is used to calculate the similarity between different attribute ratings.
[0082] S4 specifically includes:
[0083] Step S41: Based on the rating matrix, calculate the similarity between different attribute ratings to obtain the correlation coefficient matrix of each attribute.
[0084] In practical applications, the similarity between user ratings of different attributes is calculated based on the rating matrix obtained in step S3. Among all attributes in the questionnaire, some represent the core potential needs of all users. Different users show high similarity in their ratings of these core needs attributes. However, for non-core attributes, different users give significantly different ratings due to their individual characteristics, resulting in lower similarity. Calculating the similarity of user ratings for different attributes can uncover users' core needs attributes. Since each user has a different rating benchmark, Pearson similarity is used to calculate user attribute similarity. Pearson similarity reflects the degree of linear correlation between two variables, and its value ranges from -1 to 1. When the linear relationship between two variables strengthens, the correlation coefficient tends to 1 or -1; when one variable increases and the other also increases, it indicates a positive correlation, and the correlation coefficient is greater than 0; if one variable increases but the other decreases, it indicates a negative correlation, and the correlation coefficient is less than 0; if the correlation coefficient is equal to 0, it indicates that there is no linear correlation between them.
[0085] Let X and Y represent the rating vectors obtained from two attributes in the user rating matrix, respectively. The Pearson similarity between these two attributes can be calculated as follows:
[0086]
[0087] Where cov(X,Y) is the covariance of X and Y, and σ X and σ Y Let X and Y represent the standard deviations, respectively, and E(.) represent the expected value.
[0088] Step S42: Select attributes that meet the set threshold from the attribute correlation coefficient matrix to obtain the core user requirement attributes.
[0089] In practical applications, based on the correlation coefficient matrix of each attribute constructed in step S41, since the Pearson similarity value ranges from -1 to 1, a larger value indicates that different users have more consistent opinions on the two attributes. A similarity threshold is set to filter out the attributes with the highest similarity values. Although individual users differ significantly, the ratings given by all users for these filtered attributes with high similarity are extremely similar; these attributes represent the users' core needs. The threshold should generally be greater than 0.8; the larger the threshold, the higher the similarity between the filtered attributes. Therefore, values with a threshold greater than 0.8 are selected.
[0090] Step S43: Establish a core attribute adjacency matrix based on the similarity between the user's core requirement attributes.
[0091] In practical applications, a core attribute adjacency matrix is established based on the similarity between the user's core requirement attributes selected in step S42. The row vectors in the matrix represent the similarity vectors between a given core attribute and all other core attributes. The elements in the adjacency matrix represent the Pearson similarity between two core attributes, and the similarity between a core attribute and itself is set to 1.
[0092] Step S5: Use the attributes in the core attribute adjacency matrix as nodes in the network to establish a core attribute network.
[0093] In practical applications, a core attribute network is established based on the core attribute adjacency matrix created in step S4. The construction of the core attribute network helps to better analyze the relationships between different attributes, avoiding the information barriers caused by analyzing attributes in isolation. Each core attribute serves as a node in the network. In the core attribute adjacency matrix, if two core attributes have a similarity, an edge is established between these two core attributes in the attribute network (the node does not connect to itself).
[0094] Step S6: Perform graph mining on the core attribute network to obtain the minimum core attribute network.
[0095] S6 specifically includes:
[0096] Step S61: Recursively delete the node with the minimum degree in the core attribute network to obtain the current core attribute network.
[0097] Step S62: Update the core attribute adjacency matrix according to the current core attribute network to obtain the updated core attribute adjacency matrix.
[0098] Step S63: Establish the updated core attribute network based on the updated core attribute adjacency matrix.
[0099] Step S64: After deleting the node with the minimum degree in the updated core attribute network, the updated core attribute network is an empty network, and the current core attribute network is the minimum core attribute network.
[0100] In practical applications, graph mining is performed on the core attribute network constructed in step S5, and the following steps are executed: 1. Recursively delete the node with the minimum degree; 2. Update the node table; 3. Update the adjacency matrix; 4. Repeat the above operations. If the core attribute network becomes empty after deleting a node with the minimum degree, then the core attribute network before deleting that node is the minimum core requirement attribute graph.
[0101] Step S7: Based on the minimum core attribute network, analyze the influence of different core attributes to obtain the survey questionnaire analysis results.
[0102] After constructing the minimum core attribute network, the influence of each attribute is determined by its degree; the higher the degree, the greater the influence. Taking a human-computer trust survey questionnaire as an example, which has 26 questions and was completed by 41 users, the questionnaire data is first cleaned and preprocessed. The similarity between different attributes is calculated, and values with a threshold greater than 0.8 are selected. An adjacency matrix is then constructed for the selected attributes, similar attributes are merged, and graph mining analysis is used to obtain the final user core need attribute network, as shown below. Figure 3 As shown. Figure 3 The nodes in the table represent different attributes in the questionnaire, and the corresponding node table is shown in Table 1.
[0103] Table 1 Node Table
[0104]
[0105]
[0106] Based on the user's core needs attribute diagram, we can analyze the influence of different attributes. From this diagram, we can draw the following conclusions:
[0107] 1. This user group generally has a relatively positive attitude towards autonomous driving technology.
[0108] 2. This user group's attitude towards human-machine trust is more easily influenced by the degree of acceptance of autonomous driving technology by those around them, followed by the influence of media reports.
[0109] 3. The key factors affecting human-machine trust among this user group are whether autonomous driving technology can make up for the shortcomings of traditional technology and whether it can improve the utilization rate of time in the car.
[0110] Therefore, the graph mining-based questionnaire analysis method provided by this invention enables data researchers to intuitively see the relationships between various dimensions of the data. According to the algorithm steps of graph mining, the core demand attribute graph can be continuously reduced to obtain the smallest core demand attribute graph. Then, the core demand attributes of the user group can be quickly and accurately analyzed. In addition, the method is highly intuitive, less affected by extreme data in the questionnaire, and has high accuracy, applicability and interpretability.
[0111] Example 2
[0112] To implement the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, a graph mining-based questionnaire analysis system is provided below, such as... Figure 4 As shown, the system includes:
[0113] Module 1 is used to acquire the survey questionnaire data to be analyzed.
[0114] The questionnaire matrix building module 2 is used to build a questionnaire matrix based on the questionnaire data to be analyzed. The row vectors of the questionnaire matrix represent the ratings of the same user for different attributes. The column vectors of the questionnaire matrix represent the ratings of different users for the same attribute.
[0115] The cleaning module 3 is used to clean the questionnaire matrix to obtain a scoring matrix.
[0116] The core attribute adjacency matrix building module 4 is used to calculate the similarity between different attribute scores based on the scoring matrix and build the core attribute adjacency matrix.
[0117] The core attribute network establishment module 5 is used to establish a core attribute network by using the attributes in the core attribute adjacency matrix as nodes in the network.
[0118] Graph mining module 6 is used to perform graph mining on the core attribute network to obtain the minimum core attribute network.
[0119] Analysis module 7 is used to analyze the influence of different core attributes based on the minimum core attribute network and obtain the questionnaire analysis results.
[0120] The core attribute adjacency matrix establishment module 4 includes:
[0121] The similarity calculation submodule is used to calculate the similarity between different attribute scores based on the rating matrix, and obtain the correlation coefficient matrix of each attribute.
[0122] The filtering submodule is used to filter out attributes that meet the set threshold from the attribute correlation coefficient matrix to obtain the core user requirement attributes.
[0123] A submodule is established to create a core attribute adjacency matrix based on the similarity between the user's core requirement attributes.
[0124] Example 3
[0125] This invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor runs the computer program to enable the electronic device to perform the graph mining-based questionnaire analysis method of Embodiment 1.
[0126] Alternatively, the aforementioned electronic device may be a server.
[0127] In addition, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the graph mining-based questionnaire analysis method of Embodiment 1.
[0128] In the graph mining-based questionnaire analysis method, system, electronic device, and medium provided by this invention, the graph mining-based questionnaire analysis method can effectively obtain users' core needs attributes, thereby enabling targeted strategies for different user groups. First, a user rating matrix is constructed based on the questionnaire. To deploy the graph mining algorithm, the data needs to be cleaned (specifically, the data needs to be encoded according to different attribute types, followed by missing value imputation and noise removal, and finally, normalization of each attribute). Then, the pairwise similarity between different attributes is calculated, and values with a threshold greater than 0.8 are selected to obtain the attributes with the most similar user ratings. Next, a core attribute adjacency matrix is established for these attributes, and a core attribute network is constructed based on the adjacency matrix. Finally, the graph mining method is used to analyze the network structure and the influence of different core needs.
[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0130] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A questionnaire analysis method based on graph mining, characterized in that, The method includes: Obtain the survey questionnaire data to be analyzed; Based on the survey data to be analyzed, a survey matrix is established; the row vectors of the survey matrix represent the ratings of the same user for different attributes; the column vectors of the survey matrix represent the ratings of different users for the same attribute; the attribute is a description of the questions in the survey. The questionnaire matrix is cleaned to obtain a scoring matrix, which specifically includes: The range of values for the attributes in the questionnaire matrix is normalized by applying min-max normalization to obtain the normalized questionnaire matrix. Based on the normalized questionnaire matrix, different attributes are coded to obtain attribute codes; The missing values in the attribute code are filled in to obtain the filled attribute code; Remove outliers and noise data from the imputed attribute codes to obtain the scoring matrix; Based on the rating matrix, the similarity between different attribute ratings is calculated, and a core attribute adjacency matrix is established, specifically including: Based on the rating matrix, the similarity between different attribute ratings is calculated to obtain the correlation coefficient matrix of each attribute. From the correlation coefficient matrix of each attribute, attributes that meet the set threshold are selected to obtain the core user requirement attributes; Based on the similarity between the user's core needs attributes, establish a core attribute adjacency matrix; A core attribute network is established using the attributes in the core attribute adjacency matrix as nodes in the network. Graph mining is performed on the core attribute network to obtain the minimum core attribute network, specifically including: Recursively delete the node with the minimum degree in the core attribute network to obtain the current core attribute network; Based on the current core attribute network, update the core attribute adjacency matrix to obtain the updated core attribute adjacency matrix; Based on the updated core attribute adjacency matrix, establish the updated core attribute network; When the node with the minimum degree in the updated core attribute network is deleted, the updated core attribute network becomes an empty network, and the current core attribute network becomes the minimum core attribute network. Based on the aforementioned minimal core attribute network, the influence of different core attributes is analyzed to obtain the survey questionnaire analysis results.
2. The questionnaire analysis method based on graph mining according to claim 1, characterized in that, The step of establishing a questionnaire matrix based on the questionnaire data to be analyzed specifically includes: When the attribute ratings in the questionnaire to be analyzed correspond to multiple-choice questions, each option in the multiple-choice questions is converted into a Boolean single-choice question; Based on the Boolean value of the single-choice question, determine the attribute score corresponding to the multiple-choice question; When the item corresponding to the attribute score in the questionnaire to be analyzed is a multiple-choice question, the attribute score corresponding to the multiple-choice question is determined based on the attribute value corresponding to the multiple-choice question. A questionnaire matrix is established based on the attribute scores corresponding to the multiple-choice questions and the attribute scores corresponding to the single-choice questions.
3. The questionnaire analysis method based on graph mining according to claim 1, characterized in that, Pearson similarity was used to calculate the similarity between different attribute scores.
4. A questionnaire analysis system based on graph mining, characterized in that, The system includes: The acquisition module is used to acquire the survey questionnaire data to be analyzed. The questionnaire matrix building module is used to build a questionnaire matrix based on the questionnaire data to be analyzed; the row vectors of the questionnaire matrix represent the ratings of the same user on different attributes; the column vectors of the questionnaire matrix represent the ratings of different users on the same attribute; the attribute is a description of the questions in the questionnaire. The data cleaning module is used to clean the questionnaire matrix to obtain a scoring matrix, specifically including: The range of values for the attributes in the questionnaire matrix is normalized by applying min-max normalization to obtain the normalized questionnaire matrix. Based on the normalized questionnaire matrix, different attributes are coded to obtain attribute codes; The missing values in the attribute code are filled in to obtain the filled attribute code; Remove outliers and noise data from the imputed attribute codes to obtain the scoring matrix; A core attribute adjacency matrix building module is used to calculate the similarity between different attribute scores based on the scoring matrix and build a core attribute adjacency matrix; the core attribute adjacency matrix building module includes: The similarity calculation submodule is used to calculate the similarity between different attribute scores based on the rating matrix, and obtain the correlation coefficient matrix of each attribute. The filtering submodule is used to filter out attributes that meet the set threshold from the attribute correlation coefficient matrix to obtain the core user requirement attributes. A submodule is established to create a core attribute adjacency matrix based on the similarity between the user's core requirement attributes. The core attribute network establishment module is used to establish a core attribute network by using the attributes in the core attribute adjacency matrix as nodes in the network. The graph mining module is used to perform graph mining on the core attribute network to obtain the minimum core attribute network, specifically including: Recursively delete the node with the minimum degree in the core attribute network to obtain the current core attribute network; Based on the current core attribute network, update the core attribute adjacency matrix to obtain the updated core attribute adjacency matrix; Based on the updated core attribute adjacency matrix, establish the updated core attribute network; When the node with the minimum degree in the updated core attribute network is deleted, the updated core attribute network becomes an empty network, and the current core attribute network becomes the minimum core attribute network. The analysis module is used to analyze the influence of different core attributes based on the minimum core attribute network and obtain the survey questionnaire analysis results.
5. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to enable the electronic device to perform the graph mining-based questionnaire analysis method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the graph mining-based questionnaire analysis method as described in any one of claims 1 to 3.