Psychological questionnaire data analysis method and system based on hidden variable model
By constructing an undirected graph and using the transformation of independent noise conditions to separate single-factor sets, the problem of accurate identification of latent causal relationships in psychological questionnaire data was solved, and effective support for psychological assessment decision-making was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANTOU UNIV MEDICAL COLLEGE
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods cannot effectively address the potential interactions between observed indicators and the problem of insufficient measurement indicators when identifying latent causal relationships in psychological questionnaire data. This results in insufficient accuracy in identifying causal structures and an inability to provide reliable decision support.
By acquiring and standardizing psychological questionnaire data, an undirected graph was constructed and single-factor sets were separated. The latent variables and their causal directions were determined by transforming the independent noise conditions, thus forming a causal structure graph.
It improves the accuracy of identifying latent variable causal structures, provides effective technical support for psychological assessment decision-making, and reduces human analysis bias.
Smart Images

Figure CN121964019A_ABST
Abstract
Description
A Method and System for Psychological Questionnaire Data Analysis Based on Latent Variable Model Technical Field
[0001] This application relates to the field of data analysis technology, and in particular to a method and system for analyzing psychological questionnaire data based on latent variable models. Background Technology
[0002] In research and applications in psychology, healthcare, and other fields, the discovery of causal relationships between latent variables is currently primarily based on linear latent variable models. These models typically rely on the pure measurement hypothesis, using pure measurement variables as proxy variables for corresponding latent variables to identify the causal direction between them. While existing methods show promising theoretical promise, they require specific constraints, such as each latent variable needing a sufficient number of measurement variables and no direct causal interference between observed indicators, to ensure the accuracy of causal structure identification. However, in practical applications, such as market survey questionnaire data analysis and clinical case data analysis, it is often impossible to ensure in advance that there are no potential interactions between observed indicators, nor is it possible to guarantee that each latent variable has a sufficient number of pure measurement indicators. This makes it difficult for existing methods to learn enough accurate causal structure information, thus failing to provide reliable support for decision-making. Summary of the Invention
[0003] The main purpose of this application is to propose a method and system for analyzing psychological questionnaire data based on a latent variable model, aiming to improve the accuracy of identifying the causal structure between latent variables in psychology.
[0004] To achieve the above objectives, one aspect of this application proposes a method for analyzing psychological questionnaire data based on a latent variable model. The method includes: obtaining a first questionnaire dataset containing multiple valid completion results for a pre-defined psychological questionnaire, the pre-defined psychological questionnaire containing multiple psychological assessment items and multiple grade score options corresponding to each psychological assessment item; standardizing the first questionnaire dataset to obtain a second questionnaire dataset; constructing an undirected graph between the multiple psychological assessment items as multiple observed variables; separating the undirected graph according to the second questionnaire dataset and a first pre-defined transformation independent noise condition to form all single-factor sets, and determining the latent variables corresponding to each single-factor set; and determining the causal direction between all latent variables corresponding to all single-factor sets according to the second questionnaire dataset and the second pre-defined transformation independent noise condition to form a causal structure graph.
[0005] Further, the first questionnaire dataset is obtained by: acquiring an original questionnaire dataset, which contains several original completion results for the preset psychology questionnaire, each original completion result containing a grade score selection result corresponding to each psychology assessment item, and each original completion result carrying the response time; removing all original completion results that meet the invalid response conditions from the several original completion results to form the first questionnaire dataset; wherein, the invalid response conditions include at least one of the following: the response time carried is lower than the average response time of a preset proportion, at least one grade score selection result corresponding to at least one of the psychology assessment item is empty, and at least one of the following: a preset number of consecutive selected grade score options corresponding to the psychology assessment items are in the same position, and the average response time is the average of the several response times carried by the several original completion results.
[0006] Further, the first preset transformation independent noise condition includes a first sub-condition, a second sub-condition, a third sub-condition, and a fourth sub-condition; the step of separating the undirected graph to form all single-factor sets based on the second questionnaire dataset and the first preset transformation independent noise condition includes: according to the second questionnaire dataset, deleting the connection edges between every two observed variables in the undirected graph that have a connection relationship and satisfy the first sub-condition to form several initial sub-clusters; according to the second questionnaire dataset, merging every two initial sub-clusters that have an intersection relationship and satisfy the second sub-condition to form multiple first sub-clusters; according to the second questionnaire dataset, deleting each of the multiple first sub-clusters that satisfies the third sub-condition to obtain all valid sub-clusters; and according to the second questionnaire dataset and the fourth sub-condition, extracting relevant psychological assessment items from each valid sub-cluster to form a corresponding single-factor set.
[0007] Further, the step of extracting relevant psychological assessment items for each effective subgroup to form a corresponding single-factor set based on the second questionnaire dataset and the fourth sub-condition includes: pairwise combining all observed variables contained in each effective subgroup to obtain multiple observed variable combinations; for each observed variable combination: based on the second questionnaire dataset, counting the number of all observed variables contained in other effective subgroups that make the observed variable combination satisfy the fourth sub-condition; selecting the maximum number from the multiple numbers corresponding to the multiple observed variable combinations, and then merging and deduplicating all observed variable combinations whose corresponding number is equal to the maximum number to obtain the corresponding single-factor set.
[0008] Furthermore, the first sub-condition includes: In the formula, Refers to the transformation of independent noise functions. and Let the undirected graph contain two observed variables that have a connection relationship. and The other two observed variables contained in the undirected graph; the second sub-condition includes: In the formula, The observed variables contained in the intersection of two initial subclusters that have an intersection relationship. and These are the two observed variables contained in the two initial sub-clusters that do not fall within the intersection; the third sub-condition includes: In the formula, and For the two observed variables contained in the first sub-cluster, The fourth sub-condition includes the observed variables contained in the other first sub-clusters; In the formula, For the combination of observed variables determined based on effective subclusters, The observed variables included in other effective subgroups.
[0009] Further, the second preset transformation independent noise condition includes a fifth sub-condition and a sixth sub-condition; determining the causal direction between all latent variables corresponding to all single factor sets based on the second questionnaire dataset and the second preset transformation independent noise condition includes: selecting observed variables that satisfy the fifth sub-condition from each single factor set based on the second questionnaire dataset and using them as root observed variables; and determining the causal direction between the two latent variables corresponding to each of the two single factor sets when the relationship between the two root observed variables and the two other observed variables contained in each of the two single factor sets satisfies the sixth sub-condition based on the second questionnaire dataset.
[0010] Furthermore, the fifth sub-condition includes: In the formula, For the root observation variables contained in the single-factor set, and The sixth sub-condition includes: two other observed variables contained in the single-factor set; In the formula, and These are the two root observation variables contained in the two single-factor sets. and These are two other observed variables contained in the two single-factor sets, respectively.
[0011] To achieve the above objectives, another aspect of this application proposes a psychological questionnaire data analysis system based on a latent variable model. The system includes: a first module for acquiring a first questionnaire dataset containing multiple valid completion results for a preset psychological questionnaire, the preset psychological questionnaire containing multiple psychological assessment items and multiple grade score options corresponding to each psychological assessment item; a second module for standardizing the first questionnaire dataset to obtain a second questionnaire dataset; a third module for constructing an undirected graph between the multiple psychological assessment items as multiple observed variables, and then separating the undirected graph according to the second questionnaire dataset and a first preset transformation independent noise condition to form all single-factor sets, determining the latent variables corresponding to each single-factor set; and a fourth module for determining the causal direction between all latent variables corresponding to all single-factor sets according to the second questionnaire dataset and the second preset transformation independent noise condition, to form a causal structure graph.
[0012] To achieve the above objectives, another aspect of this application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for analyzing psychological questionnaire data based on a latent variable model.
[0013] To achieve the above objectives, another aspect of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for analyzing psychological questionnaire data based on a latent variable model.
[0014] This application includes at least the following beneficial effects: First, a first questionnaire dataset is standardized to obtain a second questionnaire dataset. The first questionnaire dataset contains multiple valid completion results for a pre-defined psychological questionnaire, which includes multiple psychological assessment items and multiple grade score options corresponding to each item. Then, multiple psychological assessment items are treated as multiple observed variables, and an undirected graph is constructed between them. Subsequently, based on the second questionnaire dataset and the first pre-defined transformed independent noise condition, the undirected graph is separated to form all single-factor sets, and the latent variables corresponding to each single-factor set are determined. Finally, based on the second questionnaire dataset and the second pre-defined transformed independent noise condition, the causal direction between all latent variables corresponding to all single-factor sets is determined to form a causal structure graph. This can provide effective technical support in the field of psychological assessment decision-making. Furthermore, by introducing relevant transformed independent noise conditions for data-assisted analysis, the accuracy of identifying the causal structure between psychological latent variables can be improved. Attached Figure Description
[0015] Figure 1 is a flowchart illustrating a method for analyzing psychological questionnaire data based on a latent variable model according to an embodiment of this application; Figure 2 is a schematic diagram of a causal structure diagram according to an embodiment of this application; Figure 3 is a schematic diagram of the composition of a psychological questionnaire data analysis system based on a latent variable model according to an embodiment of this application; Figure 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0017] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0018] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] In research and applications in psychology, healthcare, and other fields, core research subjects such as personality traits, psychological states, and disease risks are often latent variables that cannot be directly observed. These latent variables need to be indirectly represented through several observable indicators. For example, in the "Big Five" personality model in psychology, "conscientiousness" needs to be measured through questionnaire items such as "task completion rigor" and "plan execution ability." In the healthcare field, "cardiovascular disease risk" needs to be indirectly reflected through data such as "blood pressure monitoring values," "blood lipid indicators," and "self-reported exercise frequency." Uncovering the causal relationships between latent variables is the core of this type of research. Relying solely on correlation cannot support scientific decision-making. For instance, a significant correlation may be observed between "high scores in low mood" and "poor sleep quality," but blindly improving sleep through "sleep aid interventions" may be ineffective because it fails to address the core latent variable of "depressive tendency." Therefore, techniques for discovering causal relationships between latent variables based on observational data are crucial.
[0021] Current mainstream methods for discovering causal relationships in latent variables all rely on the "pure measurement hypothesis" as their core premise. This hypothesis requires that for each latent variable, there exists a "pure measurement subset" of observed indicators that are causally independent of each other. However, in real-world scenarios, direct causal influences between observed indicators are common, leading to frequent violations of the "pure subset hypothesis," which has become a key bottleneck restricting the practicality of these methods. For example, in psychological research, the observed indicator of "depressive tendency," "low mood," directly reduces "sleep onset efficiency," thereby affecting another observed indicator, "self-rated sleep quality." Similarly, the observed indicator of "extroversion," "social activity," increases "frequency of proactive communication," causing the "pure subset hypothesis" to fail.
[0022] Currently, the discovery of causal relationships between latent variables is mainly based on linear latent variable models, which typically rely on the pure measurement assumption. Pure measurement variables are used as proxy variables for corresponding latent variables to identify the causal direction between them. For example, Cai et al. proposed the Triad constraint method, which, by constructing pseudo-regression residuals between measurement variables and testing their independence, first clearly demonstrated the identifiability of causal directions between latent variables. Chen et al. proposed the tensor rank conditional method, which, by exploring the correlation between the tensor rank of the contingency table of observed variable sets and the d-separation set support, first effectively identified the causal structure of discrete latent variables, solving the problem of causal discovery under nonlinear relationships and complex latent structures in discrete data. Although these existing methods have good theoretical prospects, they all require specific constraints, such as each latent variable needing a sufficient number of measurement variables and no direct causal interference between observed indicators, to ensure the accuracy of causal structure identification. However, in practical applications, such as market survey questionnaire data analysis and clinical case data analysis, it is often impossible to ensure in advance that there are no potential interactions between observed indicators, nor is it possible to guarantee that each latent variable has a sufficient number of pure measurement indicators. This makes it difficult for existing methods to learn enough correct causal structure information and provide reliable support for decision-making.
[0023] In view of this, this application provides a method and system for analyzing psychological questionnaire data based on a latent variable model. The scheme proposes to first standardize a first questionnaire dataset to obtain a second questionnaire dataset. The first questionnaire dataset contains multiple valid completion results for a pre-defined psychological questionnaire, which includes multiple psychological assessment items and multiple grade score options corresponding to each item. These multiple psychological assessment items are then treated as multiple observed variables, and an undirected graph is constructed between them. Subsequently, based on the second questionnaire dataset and a first pre-defined transformed independent noise condition, the undirected graph is separated to form all single-factor sets, and the latent variables corresponding to each single-factor set are determined. Finally, based on the second questionnaire dataset and the second pre-defined transformed independent noise condition, the causal direction between all latent variables corresponding to all single-factor sets is determined to form a causal structure graph. This can provide effective technical support in the field of psychological assessment decision-making, and by introducing relevant transformed independent noise conditions for data-assisted analysis, the accuracy of identifying the causal structure between psychological latent variables can be improved. It is understood that this scheme is used to analyze the causal relationships of latent variables in a pre-defined psychological questionnaire, reducing human analysis bias.
[0024] The psychological questionnaire data analysis method based on latent variable models provided in this application relates to the field of data analysis technology. It can be applied to terminals, servers, or software running on terminals or servers. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application that implements the above methods, but is not limited to the above forms.
[0025] Please refer to Figure 1. Figure 1 is an optional flowchart of a psychological questionnaire data analysis method based on a latent variable model provided in an embodiment of this application. The method may include, but is not limited to, the following steps S101 to S104: Step S101: Obtain a first questionnaire dataset, which contains multiple valid completion results for a preset psychological questionnaire. The preset psychological questionnaire contains multiple psychological assessment items and multiple grade score options corresponding to each psychological assessment item; Step S102: Standardize the first questionnaire dataset to obtain a second questionnaire dataset; Step S103: Use the multiple psychological assessment items as multiple observation variables to construct an undirected graph between the multiple observation variables. Then, based on the second questionnaire dataset and the first preset transformation independent noise condition, separate the undirected graph to form all single-factor sets and determine the latent variables corresponding to each single-factor set; Step S104: Based on the second questionnaire dataset and the second preset transformation independent noise condition, determine the causal direction between all latent variables corresponding to all single-factor sets to form a causal structure graph.
[0026] Steps S101 to S104 as shown in the embodiments of this application, by introducing relevant transformation independent noise conditions to analyze psychological questionnaire data, can improve the accuracy of identifying the causal structure between psychological latent variables, thereby providing effective data support in the field of psychological assessment and decision-making.
[0027] In step S101 of some embodiments, the preset psychological questionnaire may include, but is not limited to, items indicating "prone to anxiety" (using...). (Indicates) Items indicating "significant emotional fluctuations" (using) (Indicates) Items indicating "sensitive to negative evaluations" (using) (Indicates) Items indicating "willingness to listen to others' needs" (using...) (Indicates) Items indicating "being considerate of others' situations" (using) (Indicates) Items indicating "inclination towards compromise and concession" (using) (Indicates) Items indicating "actively participating in social activities" (using) (Indicates) Items indicating "enjoyment of making new friends" (using) (Indicates) Items indicating "enjoying group interaction" (using) (Indicates) Items indicating "quickly integrating into new social circles" (using) (Indicates) Items indicating "properly handling interpersonal conflicts" (using) (indicated by) and items indicating "feeling comfortable in social situations" (using) (This is indicated by the text). The multiple level score options for each psychological assessment item are preferably developed using a 5-point Likert scoring method. Each level score option corresponds to a different scoring standard, which may include option A (1 point), option B (2 points), option C (3 points), option D (4 points), and option E (5 points). Higher scores indicate a higher degree of agreement.
[0028] In step S101 of some embodiments, the first questionnaire dataset can be obtained in the following way: First, an original questionnaire dataset is obtained, which contains several original completion results for the preset psychological questionnaire. Each original completion result contains the grade score selection result corresponding to each psychological assessment item. The grade score selection result is 1 point, 2 points, 3 points, 4 points, 5 points, or a blank value, and each original completion result carries the corresponding answer time. It is understood that the several original completion results are obtained by distributing standardized questionnaires to target groups such as ordinary adults and people with specific psychological characteristics for online completion and statistical processing. Then, all original completion results that meet the invalid answer conditions are removed from the several original completion results to form the first questionnaire data. The invalid answer conditions include at least one of the following: the answer time is too short, the answer follows a habitual pattern, and the key item is not answered. The answer time is too short, which means that the answer time of the original answer is lower than the average answer time of a preset proportion. The average answer time is the average of the answer times of several original answers. The preset proportion is preferably set to 1 / 3. The answer follows a habitual pattern, which means that the selected level score options for a consecutive preset number of psychological assessment items are in the same position, that is, the same option (such as option A) is selected for multiple consecutive psychological assessment items. The preset number is preferably set to 10. The key item is not answered, which means that the level score selection result for at least one psychological assessment item is empty, that is, no level score option is selected for the psychological assessment item.
[0029] In this application, by removing invalid data from the obtained original questionnaire dataset according to preset criteria to generate the first questionnaire dataset, the reliability of subsequent analysis of the relevant questionnaire data can be improved.
[0030] In step S102 of some embodiments, regarding the standardization of the first questionnaire dataset to obtain the second questionnaire dataset, the corresponding implementation method includes: for each psychological assessment item included in the preset psychological questionnaire, extracting multiple non-empty level score selection results corresponding to the psychological assessment item from multiple valid completion results, and then performing Z-score standardization on it to obtain multiple first level score selection results corresponding to the psychological assessment item. Subsequently, the second questionnaire dataset is formed by summarizing to eliminate the impact of dimensional differences on subsequent questionnaire data analysis.
[0031] In step S103 of some embodiments, the construction of an undirected graph among multiple observed variables can be implemented, but is not limited to, using the classic FindPattern algorithm. This involves analyzing the relationships between multiple observed variables using second-order covariance information and Tetrad constraints (CS1-CS3) to form the undirected graph, laying the foundation for subsequent separation of single-factor sets. It should be noted that each observed variable can be understood as a node in the undirected graph. If two observed variables are correlated, they are represented in the undirected graph by connecting the corresponding two nodes with edges. Optionally, each observed variable can be connected to each of the other observed variables to form the undirected graph.
[0032] In step S103 of some embodiments, the first preset transformation independent noise condition includes a first sub-condition, a second sub-condition, a third sub-condition, and a fourth sub-condition; regarding the separation of the undirected graph to form all single-factor sets based on the second questionnaire dataset and the first preset transformation independent noise condition, the corresponding implementation may include, but is not limited to, the following steps S201 to S204.
[0033] Step S201: Based on the second questionnaire dataset, delete the edges between every two observed variables in the undirected graph that have a connection relationship and satisfy the first sub-condition, to form several independent initial sub-clusters; wherein, the first sub-condition includes: In the formula, Refers to the transformation of independent noise functions. and Let the two observed variables in this undirected graph be connected. and Let these be the other two observed variables contained in the undirected graph; it is understandable that for and If it exists The first sub-condition is satisfied. If the set of all observed variables contained in the undirected graph is considered, then the determination is made. and Belonging to different single-factor sets, i.e. and Nodes that do not share the same hidden parent set should be deleted from the undirected graph. and The connecting edges formed between them.
[0034] Step S202: Based on the second questionnaire dataset, merge every two initial sub-clusters that have an intersection relationship and satisfy the second sub-condition to form multiple first sub-clusters; wherein, the second sub-condition includes: In the formula, The observed variables contained in the intersection of two initial subclusters that have an intersection relationship. and These are two observed variables contained in two initial subclusters that have an intersection relationship but do not fall within their intersection; it can be understood that for two initial subclusters that have an intersection relationship... and Determine the intersection between them. If it exists , and If the second sub-condition is satisfied, then the intersection is determined. For a single-factor set, two initial sub-clusters need to be... and The observed variables are merged into a composite group and used as the first subgroup to aggregate the information of the same latent variable. In addition, initial subgroups that do not intersect with other initial subgroups can be directly used as the first subgroup, and two initial subgroups that intersect but do not meet the second subcondition can each be directly used as the first subgroup.
[0035] Step S203: Based on the second questionnaire dataset, delete each of the multiple first subgroups that satisfy the third sub-condition, to obtain all valid subgroups; wherein, the third sub-condition includes: In the formula, and For the two observed variables contained in the first sub-cluster, For the observed variables contained in other first sub-clusters; it is understandable that if for all ,exist And for any All satisfy the third sub-condition. Refers to a set formed by multiple first sub-clusters, and and For set If it contains two first sub-clusters, then determine Multi-factor clusters are identified and eliminated. In this step, it can be considered that each first sub-cluster that does not satisfy the third sub-condition is directly regarded as a valid sub-cluster.
[0036] Step S204: Based on the second questionnaire dataset and the fourth sub-condition, extract relevant psychological assessment items for each valid subgroup to form a corresponding single-factor set. Specifically, taking one valid subgroup as an example, the implementation method for determining the corresponding single-factor set based on the valid subgroup is explained as follows: First, combine all observed variables contained in the valid subgroup in pairs to obtain multiple combinations of observed variables; then, for each combination of observed variables: based on the second questionnaire dataset, count the number of all observed variables contained in other valid subgroups that make the combination of observed variables satisfy the fourth sub-condition; wherein, the fourth sub-condition includes: In the formula, For a combination of observed variables determined based on this effective subcluster, The observed variables are contained in other effective subgroups. For example, suppose there are three effective subgroups, denoted as the first effective subgroup, the second effective subgroup, and the third effective subgroup. For a target observed variable combination determined based on the first effective subgroup, if it is determined that the second effective subgroup contains two observed variables that make the target observed variable combination satisfy the fourth subcondition, and if it is determined that the third effective subgroup contains only one observed variable that makes the target observed variable combination satisfy the fourth subcondition, then the number corresponding to the target observed variable combination is statistically determined to be 3. Finally, the maximum number is selected from the multiple numbers corresponding to multiple observed variable combinations, and all observed variable combinations with a corresponding number equal to the maximum number are merged and deduplicated to obtain the corresponding single factor set.
[0037] In this application, by adopting the relevant Transformed Independent Noise Condition (TIN), the single-factor set and the multi-factor set can be effectively separated, and the purity of the separated single-factor set can be ensured. This adapts to the complex correlation of real evaluation data, thereby improving the accuracy of causal structure recognition in impure scenarios and avoiding distortion in causal structure learning. Compared with existing methods, this can break through the limitation of the "pure subsupposition".
[0038] In step S103 of some embodiments, the corresponding implementation methods for determining the latent variables corresponding to each single factor set may include, but are not limited to: calling a pre-determined latent variable database, which records several latent variables related to the field of psychology and the specific definition of each latent variable; for each single factor set, performing a matching query in the latent variable database based on the semantic features of each observed variable contained in the single factor set to obtain the latent variables corresponding to the single factor set.
[0039] For example, the latent variable database records at least one latent variable representing "neuroticism" (using... (represented by), representing the latent variable of "acceptability" (using) (represented by) latent variables representing "extroversion" (using) (represented by) and latent variables representing "social adaptability" (using) If, through steps S201 to S204 above, it is determined that there are a total of four single-factor sets, denoted as the first single-factor set, the second single-factor set, the third single-factor set, and the fourth single-factor set, respectively, the first single-factor set includes items indicating "easily feeling anxious," items indicating "significant mood swings," and items indicating "sensitivity to negative evaluations." The second single-factor set includes items indicating "willingness to listen to others' demands," items indicating "good at understanding others' situations," and items indicating "tendency to compromise and yield." The third single-factor set includes items indicating "actively participating in social activities," items indicating "enjoying making new friends," and items indicating "eagerness to make new friends." The fourth single-factor set includes items representing "enjoying group interaction," items representing "quickly integrating into new social circles," items representing "properly handling interpersonal conflicts," and items representing "feeling comfortable in social situations." By performing semantic feature matching queries on the above single-factor sets in the latent variable database, it was determined that the latent variable corresponding to the first single-factor set is a latent variable representing "neuroticism," the latent variable corresponding to the second single-factor set is a latent variable representing "agreeableness," the latent variable corresponding to the third single-factor set is a latent variable representing "extraversion," and the latent variable corresponding to the fourth single-factor set is a latent variable representing "social adaptability."
[0040] In step S104 of some embodiments, the second preset transformation independent noise condition includes a fifth sub-condition and a sixth sub-condition; regarding the determination of the causal direction between all latent variables corresponding to all single factor sets based on the second questionnaire dataset and the second preset transformation independent noise condition, the corresponding implementation may include, but is not limited to, the following steps S301 to S302.
[0041] Step S301: Based on the second questionnaire dataset, select the observed variables that satisfy the fifth sub-condition from each single-factor set and use them as root observed variables. These root observed variables can be defined as effective proxies for the latent variables corresponding to the single-factor set. The fifth sub-condition includes: In the formula, For the root observation variables contained in the single-factor set, and For the two other observed variables contained in this single-factor set; it is understandable that, for If all exist The fifth sub-condition is satisfied. If it is a single-factor set, then determine The ancestor set contains only itself and its corresponding latent variable, without interference from other observed variables. It can be used as the single factor set Effective proxy for the corresponding hidden variables.
[0042] Step S302: Based on the second questionnaire dataset, when determining that the relationship between the two root observed variables and the two other observed variables contained in each pair of single-factor sets satisfies the sixth sub-condition, determine the causal direction between the two latent variables corresponding to each pair of single-factor sets; wherein, the sixth sub-condition includes: In the formula, and These are the two root observation variables contained in the two single-factor sets. and These are two other observed variables contained in these two single-factor sets; it can be understood that for single-factor sets... Corresponding hidden variables and single factor set Corresponding hidden variables If hidden variables effective agent Latent variables effective agent Single factor set Other observed variables included and single factor set Other observed variables included If the sixth subcondition is satisfied, then the latent variables are determined according to the TIN orientation rule. and latent variables The causal direction between them is This can indicate latent variables. Negative influence on latent variables .
[0043] Considering that existing methods typically use randomly selected items as proxies for latent variables, which is susceptible to interference from direct associations between items, this application introduces relevant transformational independent noise conditions to screen root observed variables. This improves the effectiveness of proxies, effectively reduces interference from other observed variables, significantly enhances the accuracy of latent variable representation, and reduces the bias of random selection. Furthermore, considering that existing methods often rely on low-order statistics, which are susceptible to Markov equivalence class limitations, this application strengthens the ability to identify causal flows between latent variables by using non-Gaussian properties and TIN orientation rules. This fully captures deep dependencies between variables, avoids misjudging indirect effects as direct effects, provides more reliable causal evidence for psychological mechanism research (such as how neuroticism affects social adaptability), and offers effective technical support in the field of psychological assessment.
[0044] In some embodiments, the transform-independent noise function appears in the different sub-conditions described above. The following explanation is provided: Let... and Given two sets of observed variables, each containing at least one observed variable, and assuming the observed variables follow a LiNGAM (i.e., a linear non-Gaussian acyclic model), then define a transformation-independent noise function. The definition of is: In the formula, for The dimension of the set of observed variables. The number of all observed variables included. for a subspace and such that , It refers to and independent, For the weight vector, Generally refers to the transpose symbol. For subspace The dimension is specifically based on these two sets of observed variables. and The results of analyzing and determining the first-level scores corresponding to each observed variable are as follows: This is understandable, as it involves transforming independent noise functions. The input is these two sets of observed variables. and The return value is located in Integers within the range.
[0045] Among them, subspace dimension It can be obtained, but is not limited to, through the following methods: First, based on the Linear Non-Gaussian Acyclic Model (LiNGAM) and the Transform Independent Noise (TIN) condition, the weight vector is defined. The generation requires revolving around independent linear transformation subspaces. Expanding, this subspace is equivalent to a mixture matrix. zero space ,in, For generating weight vectors The set of observed variables, The set of observed variables used for independence testing, For the set of observed variables The set of ancestor nodes (and their corresponding set of observed variables) (the source of noise); in practice, if the mixing matrix is known... The weight vector can be obtained by calculating its null space using the SVD algorithm. The basis vectors are then linearly combined to generate several different weight vectors. Secondly, based on the set of observed variables... For each observed variable included, multiple first-level score selection results are used to calculate several weight vectors. The corresponding set of several variables Finally, each set of variables is verified one by one according to the HSIC independence test principle. With the set of observed variables Whether they are independent, and then filter them to find those that are related to the set of observed variables. Independent set of all variables and the set of all selected variables All corresponding weight vectors After summarizing and forming a set, matrix rank analysis is performed to obtain the subspace. dimension .
[0046] In step S104 of some embodiments, after clarifying all latent variables corresponding to all single-factor sets and the causal directions between all latent variables, it is known that each single-factor set contains multiple observed variables. By merging the mapping relationship between "observed variables and latent variables" and the causal structure relationship between "latent variables," a causal structure diagram of the latent variables of psychological assessment can be obtained, as shown in Figure 2. Furthermore, after generating this causal structure diagram, it can be visualized, allowing researchers to intuitively understand the complete causal structure system of the latent variables of psychological assessment.
[0047] Understandably, this causal structure graph is constructed based on the assumptions of a linear non-Gaussian acyclic latent variable model, and typically needs to satisfy the following three core assumptions: First, the causal structure graph is a directed acyclic graph and obeys the causal Markov assumption and the loyalty assumption; Second, each latent variable in the causal structure graph has at least one single-factor set, which consists of multiple observed variables and they share a unique latent parent node, and is independent of other single-factor sets after the corresponding latent variable is given; Third, the causal relationship between variables is linear, and the noise terms follow a non-Gaussian distribution and are independent of each other.
[0048] Please refer to Figure 3. Figure 3 is a schematic diagram of an optional structural composition of a psychological questionnaire data analysis system based on a latent variable model provided in an embodiment of this application. The system can implement the above-mentioned psychological questionnaire data analysis method based on a latent variable model. The system may include, but is not limited to, the following: a first module 401, used to obtain a first questionnaire dataset, which contains multiple valid completion results for a preset psychological questionnaire. The preset psychological questionnaire contains multiple psychological assessment items and multiple grade score options corresponding to each psychological assessment item; a second module 402, used to standardize the first questionnaire dataset to obtain a second questionnaire dataset; a third module 403, used to construct an undirected graph between multiple psychological assessment items as multiple observed variables, and then separate the undirected graph according to the second questionnaire dataset and the first preset transformation independent noise condition to form all single-factor sets, and determine the latent variables corresponding to each single-factor set; a fourth module 404, used to determine the causal direction between all latent variables corresponding to all single-factor sets according to the second questionnaire dataset and the second preset transformation independent noise condition, so as to form a causal structure graph.
[0049] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those implemented in the above method embodiments, and the beneficial effects achieved by this system embodiment are also the same as those achieved by the above method embodiments.
[0050] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for analyzing psychological questionnaire data based on a latent variable model. This electronic device can include any smart terminal such as a tablet computer or desktop computer.
[0051] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those implemented by the above method embodiments, and the beneficial effects achieved by the present device embodiments are also the same as those achieved by the above method embodiments.
[0052] Please refer to Figure 4, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes: a processor 501, which can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, for executing related programs to implement the technical solutions provided in the embodiments of this application; and a memory 502, which can be implemented using a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM), etc. The memory 502 can store the operating system and other applications. When the technical solutions provided in the embodiments of this application are implemented through software or firmware, the relevant program code is stored in the memory 502 and is called and executed by the processor 501. The input / output interface 503 is used to realize information input and output. The communication interface 504 is used to realize communication interaction between this device and other devices. Communication can be realized through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.). The bus 505 transmits information between various components of the device (such as processor 501, memory 502, input / output interface 503 and communication interface 504). The processor 501, memory 502, input / output interface 503 and communication interface 504 realize communication connection between each other within the device through the bus 505.
[0053] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for analyzing psychological questionnaire data based on a latent variable model.
[0054] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented by this storage medium embodiment are the same as those implemented by the above method embodiments, and the beneficial effects achieved by this storage medium embodiment are also the same as those achieved by the above method embodiments.
[0055] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0056] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0057] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0058] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0059] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0060] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0061] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0062] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0063] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0064] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0065] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0066] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for analyzing psychological questionnaire data based on a latent variable model, characterized in that, The method includes: acquiring a first questionnaire dataset, which contains multiple valid completion results for a preset psychology questionnaire, the preset psychology questionnaire containing multiple psychology assessment items and multiple grade score options corresponding to each psychology assessment item; standardizing the first questionnaire dataset to obtain a second questionnaire dataset; constructing an undirected graph between the multiple psychology assessment items as multiple observed variables, and then separating the undirected graph according to the second questionnaire dataset and a first preset transformation independent noise condition to form all single-factor sets, and determining the latent variables corresponding to each single-factor set; and determining the causal direction between all latent variables corresponding to all single-factor sets according to the second questionnaire dataset and the second preset transformation independent noise condition to form a causal structure graph.
2. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 1, characterized in that, The first questionnaire dataset is obtained through the following method: obtaining an original questionnaire dataset, which contains several original completion results for the preset psychology questionnaire, each original completion result containing the grade score selection result corresponding to each psychology assessment item, and each original completion result carrying the response time; removing all original completion results that meet the invalid response conditions from the several original completion results to form the first questionnaire dataset; wherein, the invalid response conditions include at least one of the following: the response time carried is lower than the average response time of a preset proportion, at least one grade score selection result corresponding to at least one of the psychology assessment item is empty, and at least one of the following: a preset number of consecutive selected grade score options corresponding to the psychology assessment items are in the same position, and the average response time is the average of the several response times carried by the several original completion results.
3. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 1, characterized in that, The first preset transformation independent noise condition includes a first sub-condition, a second sub-condition, a third sub-condition, and a fourth sub-condition. The step of separating the undirected graph to form all single-factor sets based on the second questionnaire dataset and the first preset transformation independent noise condition includes: deleting the connection edges between every two observed variables in the undirected graph that have a connection relationship and satisfy the first sub-condition, based on the second questionnaire dataset, to form several initial sub-clusters; merging every two initial sub-clusters that have an intersection relationship and satisfy the second sub-condition, based on the second questionnaire dataset, to form multiple first sub-clusters; deleting each first sub-cluster in the multiple first sub-clusters that satisfies the third sub-condition, based on the second questionnaire dataset, to obtain all valid sub-clusters; and extracting relevant psychological assessment items from each valid sub-cluster based on the second questionnaire dataset and the fourth sub-condition to form a corresponding single-factor set.
4. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 3, characterized in that, The step of extracting relevant psychological assessment items for each effective subgroup to form a corresponding single-factor set based on the second questionnaire dataset and the fourth sub-condition includes: pairwise combining all observed variables contained in each effective subgroup to obtain multiple observed variable combinations; for each observed variable combination: based on the second questionnaire dataset, counting the number of all observed variables contained in other effective subgroups that make the observed variable combination satisfy the fourth sub-condition; selecting the maximum number from the multiple numbers corresponding to the multiple observed variable combinations, and then merging and deduplicating all observed variable combinations whose corresponding number is equal to the maximum number to obtain the corresponding single-factor set.
5. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 4, characterized in that, The first sub-condition includes: In the formula, Refers to the transformation of independent noise functions. and Let the undirected graph contain two observed variables that have a connection relationship. and The other two observed variables contained in the undirected graph; the second sub-condition includes: In the formula, The observed variables contained in the intersection of two initial subclusters that have an intersection relationship. and These are the two observed variables contained in the two initial sub-clusters that do not fall within the intersection; the third sub-condition includes: In the formula, and For the two observed variables contained in the first sub-cluster, The fourth sub-condition includes the observed variables contained in the other first sub-clusters; In the formula, For the combination of observed variables determined based on effective subclusters, The observed variables included in other effective subgroups.
6. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 1, characterized in that, The second preset transformation independent noise condition includes a fifth sub-condition and a sixth sub-condition; determining the causal direction between all latent variables corresponding to all single factor sets based on the second questionnaire dataset and the second preset transformation independent noise condition includes: selecting observed variables that satisfy the fifth sub-condition from each single factor set and using them as root observed variables based on the second questionnaire dataset; and determining the causal direction between the two latent variables corresponding to each of the two single factor sets when the relationship between the two root observed variables and the two other observed variables contained in each of the two single factor sets satisfies the sixth sub-condition based on the second questionnaire dataset.
7. The method for analyzing psychological questionnaire data based on a latent variable model according to claim 6, characterized in that, The fifth sub-condition includes: In the formula, For the root observation variables contained in the single-factor set, and The sixth sub-condition includes: two other observed variables contained in the single-factor set; In the formula, and These are the two root observation variables contained in the two single-factor sets. and These are two other observed variables contained in the two single-factor sets, respectively.
8. A psychological questionnaire data analysis system based on a latent variable model, characterized in that, The system includes: a first module for acquiring a first questionnaire dataset, which contains multiple valid completion results for a preset psychological questionnaire, the preset psychological questionnaire containing multiple psychological assessment items and multiple grade score options corresponding to each psychological assessment item; a second module for standardizing the first questionnaire dataset to obtain a second questionnaire dataset; a third module for constructing an undirected graph between the multiple psychological assessment items as multiple observed variables, and then separating the undirected graph according to the second questionnaire dataset and a first preset transformed independent noise condition to form all single-factor sets, and determining the latent variables corresponding to each single-factor set; and a fourth module for determining the causal direction between all latent variables corresponding to all single-factor sets according to the second questionnaire dataset and the second preset transformed independent noise condition, to form a causal structure graph.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the psychological questionnaire data analysis method based on the latent variable model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the psychological questionnaire data analysis method based on the latent variable model as described in any one of claims 1 to 7.