Intelligent questionnaire splitting and combining method
By splitting the questionnaire according to the principles of modularity, orthogonality, and connectability, and combining user information tagging and algorithmic inference, the problems of repetitive questioning and logical discontinuity in existing questionnaire splitting systems are solved, achieving lightweight questionnaires and efficient data analysis.
Patent Information
- Application Number
- CN202511111920.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-21
AI Technical Summary
Existing questionnaire splitting systems cannot dynamically adjust based on question semantics or user behavior, leading to duplicate questions or logical gaps. They also lack automated analysis of the semantic relevance of questions, resulting in low user completion efficiency.
The questionnaire is split into two parts using an intelligent questionnaire splitting method. The questionnaire is split into parts based on the principles of modularity, orthogonality and connectability. The user information tag filtering and relevance matching algorithms are combined to push personalized questionnaire components. The Bayesian algorithm, monotonic interpolation method and Monte Carlo algorithm are used for data inference and completion.
It achieves lightweight questionnaires, improves user completion efficiency, expands the scope of analysis, reduces redundant investment, and fully leverages the value of data insights.
Smart Images

Figure CN120994809A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method for intelligent questionnaire splitting and combining. Background Technology
[0002] Survey questionnaires are a crucial bridge for transforming research objectives into quantifiable and analyzable data. By systematically asking pre-set questions, they provide researchers with objective evidence to describe phenomena, explain correlations, predict trends, support decision-making, and evaluate effects. Current technologies largely rely on manually preset static rules (such as categorization by question type), which cannot be dynamically adjusted based on question semantics or user behavior. Existing questionnaire splitting systems only support fixed template splitting, leading to repeated questioning or logical gaps. Mainstream tools (such as the XIAOJUSURVEY system) lack automated analysis of the semantic relevance of questions, resulting in redundancy or logical conflicts in the split sub-questionnaires, leading to decreased user completion efficiency. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides an intelligent questionnaire splitting and combining method, which solves the problems.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a questionnaire intelligent splitting and combining method, comprising the following steps: Step 1: Users enter their personal information on the platform according to their needs; Step 2: Divide the questionnaire items into several variations according to the principles of modularity, orthogonality, and connectability, and combine them into a questionnaire set of appropriate length; Step 3: Tag users based on their information, identify different audiences based on the tags, and use a relevance matching algorithm to filter users by tags in the audience database; Step 4: Extract different component questionnaires from the questionnaire collection and push them to users with different tags; Step 5: Use the different component questionnaires answered by users as fragment data and use algorithms to infer and complete them.
[0005] Preferably, the user's personal information in step one includes explicit information and implicit information. The explicit information includes demographics, industry, and job title, while the implicit information includes past questionnaire answering records, website / APP behavior, and activity participation records.
[0006] Preferably, in step two, the modular principle specifically means that the unit after the questionnaire is broken down is a module, and related questions are treated as a whole in the breakdown. A module consists of multiple related questions, or it can consist of a single question that is independent of other questions.
[0007] Preferably, the orthogonality principle in step two specifically means that each pair of modules appearing together once in a questionnaire is considered as a test point. The orthogonal design ensures that the combination of each module conforms to the following principles: each module appears with an even frequency throughout the entire experiment; each test point appears with an even frequency throughout the entire experiment; and each test point is independent of each other and unrelated throughout the entire experiment.
[0008] Preferably, the connectability principle in step two specifically means that, due to the orthogonality principle, any two modules in a questionnaire can be connected in a balanced way by the respondents, that is, each two modules will be answered simultaneously by a group of respondents of equal size.
[0009] Preferably, the relevance matching algorithm includes the following steps: Feature vectorization: Let the user feature vector be... The feature vector of the problem / module is ( (For problem / module index) Relevance matching algorithm: Basic method: Cosine similarity Variant: Dot product similarity Probabilistic approach: Softmax weighted ,in For the problem / module The probability of being selected. For temperature parameters, The total number of candidate questions; Decision rule: Threshold filtering , For the preset threshold; next question Dynamic sorting; module-level matching: ,in For module The set of questions included.
[0010] Preferably, the combination algorithm in step two is as follows: Problem modeling: Let the module set be... ; Problem Set Each problem has a predefined subset of modules. ; Connect variables Representation module and Whether directly connected; position variables Representation module Center coordinates; Dimensional parameters Representation module Width and height; Objective function: Minimize the weighted total length: , in For connection weights, ; Constraints: Orthogonality constraints Through large Linearization: ,in Indicate whether Axis alignment; overall coherence of the related questions , This is achieved by constraining all related modules to form a fully connected subgraph; anti-overlap constraints. .
[0011] Preferably, the algorithm in step five includes the Bayesian algorithm, the monotonic interpolation method, and the Monte Carlo algorithm. The formula for the Bayesian algorithm is as follows; ; in, For posterior probability (event) In the given (probability after occurrence) For the likelihood function (event) exist (probability of occurrence) For prior probability ( (initial probability) Marginal probability ( The total probability is usually calculated using the law of total probability: ).
[0012] Preferably, the formula for the monotonic interpolation method is as follows: Let the slope of adjacent points be... interpolation point derivative at point Must meet: like and Same number: (Weighted harmonic average) weight , ; like and Different signs or zero: (Preserve local extrema); interpolation function (in) (interval) The system consists of endpoint values , and derivative , The only certainty.
[0013] Preferably, the formula for the Monte Carlo algorithm is as follows: Estimated expected value: , , From the distribution Independently distributed samples drawn from the data; Estimate the definite integral: , ; High-dimensional integrals: , ,in It is the integration region The volume.
[0014] This invention provides an intelligent questionnaire splitting and combining method. Compared with the prior art, it has the following advantages: (1) The intelligent questionnaire splitting method breaks down the lengthy questionnaire into several variations according to the three principles of orthogonality, connectivity and modularity. After the data collection is completed, Bayes' theorem is used to analyze the questionnaire as a whole, so as to achieve the purpose of questionnaire lightweighting without compromising the reliability and validity of the analysis.
[0015] (2) The intelligent questionnaire splitting method uses experimental design thinking to infer the whole picture by statistical methods based on the incomplete information of each person. It is an application of big data thinking in the field of questionnaire survey.
[0016] (3) The intelligent questionnaire splitting method expands the scope of analysis by having different users participate in the same data collection process, enabling different users to share public data, share public data costs, reduce redundant investment, and give full play to the value of insights. Attached Figure Description
[0017] Figure 1 This is a schematic diagram illustrating the working principle of the present invention; Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0018] The technical solutions in the embodiments of the present invention have been clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Please see Figures 1 to 2 This invention provides a technical solution: a questionnaire intelligent splitting and combining method, comprising the following steps: Step 1: Users enter their personal information on the platform according to their own needs. The personal information of users in Step 1 includes explicit information and implicit information. Explicit information includes demographics, industry, and job title, while implicit information includes past questionnaire answering records, website / APP behavior, and activity participation records. Step Two: The questionnaire items are broken down into several variations according to the principles of modularity, orthogonality, and connectability. These variations are then combined into a questionnaire set of appropriate length. The modularity principle in Step Two specifically means that the units after questionnaire breakdown are modules, and the questions within a module are divided into mandatory and optional questions. Only optional questions participate in the breakdown. Related questions (i.e., prerequisite and follow-up questions) are treated as a single unit in the breakdown. A module consists of multiple related questions, or it can consist of a single question independent of other questions. The orthogonality principle in Step Two specifically means that any two modules appearing together once in a questionnaire are considered as one experimental point. Orthogonal design ensures that the combination of modules conforms to the following principles: the frequency of each module appearing throughout the experiment is balanced; the frequency of each experimental point appearing throughout the experiment is balanced; and the experimental points are independent and unrelated throughout the experiment. The connectability principle in Step Two specifically means that due to the orthogonality principle, any two modules in a questionnaire can be balancedly connected through respondents, meaning that every two modules will be answered simultaneously by a group of respondents of equal size. The combination algorithm in Step Two is as follows: Problem modeling: Let the module set be... ; Problem Set Each problem has a predefined subset of modules. ; Connect variables Representation module and Whether directly connected; position variables Representation module Center coordinates; Dimensional parameters Representation module Width and height; Objective function: Minimize the weighted total length: , in For connection weights, ; Constraints: Orthogonality constraints Through large Linearization: ,in Indicate whether Axis alignment; overall coherence of the related questions , This is achieved by constraining all related modules to form a fully connected subgraph; anti-overlap constraints. ; Step 3: Tag users based on their information, identify different audiences based on the tags, and use a relevance matching algorithm in the audience database to filter users by tag. The relevance matching algorithm includes the following steps: Feature vectorization: Let the user feature vector be... The feature vector of the problem / module is ( (For problem / module index) Relevance matching algorithm: Basic method: Cosine similarity Variant: Dot product similarity Probabilistic approach: Softmax weighted ,in For the problem / module The probability of being selected. For temperature parameters, The total number of candidate questions; Decision rule: Threshold filtering , For the preset threshold; next question Dynamic sorting; module-level matching: ,in For module The set of questions included; Step 4: Extract different component questionnaires from the questionnaire collection and push them to users with different tags; Step 5: The different components of the user-answered questionnaire are treated as fragment data. Questions from different categories appear in a single questionnaire through orthogonal combinations. This splitting provides a pathway to connect different categories, making it suitable for inference using techniques such as Bayesian methods. Algorithms are then used for inference completion. The algorithms in Step 5 include the Bayesian algorithm, monotonic interpolation, and Monte Carlo algorithm. The formula for the Bayesian algorithm is as follows: ; in, For posterior probability (event) In the given (probability after occurrence) For the likelihood function (event) exist (probability of occurrence) For prior probability ( (initial probability) Marginal probability ( The total probability is usually calculated using the law of total probability: ).
[0020] The formula for monotonic interpolation is as follows: Let the slope of adjacent points be... interpolation point derivative at point Must meet: like and Same number: (Weighted harmonic average) weight , ; like and Different signs or zero: (Preserve local extrema); interpolation function (in) (interval) The system consists of endpoint values , and derivative , The only certainty.
[0021] The formula for the Monte Carlo algorithm is as follows: Estimated expected value: , , From the distribution Independently distributed samples drawn from the data; Estimate the definite integral: , ; High-dimensional integrals: , ,in It is the integration region The volume.
[0022] By using Bayesian algorithm, monotonic interpolation and Monte Carlo algorithm together, the Monte Carlo algorithm solves the bottleneck of Bayesian high-dimensional integral calculation, the monotonic interpolation provides a surrogate model for Monte Carlo algorithm to be evaluated quickly, and the Bayesian algorithm is the "normalizer" of monotonic interpolation (injecting domain knowledge through a probabilistic framework). The three constitute a complete link of probabilistic modeling → function approximation → stochastic calculation.
[0023] In practice, users create projects on the platform according to their needs, determine the respondents' criteria, enter personal information, and write and publish questionnaires. The system breaks down the questionnaire into different modules, ensuring orthogonal design and connectivity between modules, and pushes appropriately sized questionnaire modules to different users for questioning and answering. After data collection, Bayes' theorem is used to analyze the questionnaire as a whole, achieving the goal of lightweight questionnaires without compromising the reliability and validity of the analysis. Using experimental design thinking, the system uses incomplete information from each individual to infer the whole picture through statistical methods, which is an application of big data thinking in the field of questionnaire surveys. The participation of different users in the same data collection process expands the scope of analysis, enabling different users to share public data, share public data costs, reduce redundant investment, and fully realize the value of insights.
[0024] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0025] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0026] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A questionnaire intelligent splitting and combining method, characterized in that: Includes the following steps: Step 1: Users enter their personal information on the platform according to their needs; Step 2: Divide the questionnaire items into several variations according to the principles of modularity, orthogonality, and connectability, and combine them into a questionnaire set of appropriate length; Step 3: Tag users based on their information, identify different audiences based on the tags, and use a relevance matching algorithm to filter users by tags in the audience database; Step 4: Extract different component questionnaires from the questionnaire collection and push them to users with different tags; Step 5: Use the different component questionnaires answered by users as fragment data and use algorithms to infer and complete them.
2. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: The user's personal information in step one includes explicit information and implicit information. The explicit information includes demographics, industry, and job title, while the implicit information includes past questionnaire answers, website / APP behavior, and activity participation records.
3. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: In step two, the modular principle specifically means that the unit after the questionnaire is broken down is a module, and related questions are treated as a whole in the breakdown. A module consists of multiple related questions, and can also consist of a single question that is independent of the other questions.
4. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: In step two, the orthogonality principle specifically means that each pair of modules appearing together once in a questionnaire is considered as one experimental point. The orthogonal design ensures that the combination of modules conforms to the following principles: the number of times each module appears in the whole experiment is balanced; the number of times each experimental point appears in the whole experiment is balanced; and the experimental points are independent of each other and have no correlation in the whole experiment.
5. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: The connectability principle in step two specifically means that, due to the orthogonality principle, any two modules in a questionnaire can be connected in a balanced way by the respondents, that is, each pair of modules will be answered simultaneously by a group of respondents of equal size.
6. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: The correlation matching algorithm includes the following steps: Feature vectorization: Let the user feature vector be... The feature vector of the problem / module is ( (For problem / module index) Relevance matching algorithms: Basic method: Cosine similarity ; Variant: Dot product similarity Probabilistic approach: Softmax weighted ,in For the problem / module The probability of being selected. For temperature parameters, The total number of candidate questions; Decision rule: Threshold filtering , For the preset threshold; next question Dynamic sorting; module-level matching: ,in For module The set of questions included.
7. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: The combination algorithm in step two is as follows: Problem modeling: Let the module set be... ; Problem Set Each problem has a predefined subset of modules. ; Connect variables Representation module and Whether directly connected; position variables Representation module Center coordinates; Dimensional parameters Representation module Width and height; Objective function: Minimize the weighted total length: , in For connection weights, ; Constraints: Orthogonality constraints Through large Linearization: ,in Indicate whether Axis alignment; overall coherence of the related questions , This is achieved by constraining all related modules to form a fully connected subgraph; anti-overlap constraints. .
8. The intelligent questionnaire splitting and combining method according to claim 1, characterized in that: The algorithms in step five include Bayesian algorithm, monotonic interpolation method and Monte Carlo algorithm. The formula of Bayesian algorithm is as follows; ; in, For posterior probability (event) In the given (probability after occurrence) For the likelihood function (event) exist (probability of occurrence) For prior probability ( (initial probability) Marginal probability ( The total probability is usually calculated using the law of total probability: ).
9. The intelligent questionnaire splitting and combining method according to claim 8, characterized in that: The formula for the monotonic interpolation method is as follows: Let the slope of adjacent points be... interpolation point derivative at point Must meet: like and Same number: (Weighted harmonic average) weight , ; like and Different signs or zero: (Preserve local extrema); interpolation function (in) (interval) The system consists of endpoint values , and derivative , The only certainty.
10. The intelligent questionnaire splitting and combining method according to claim 8, characterized in that: The formula for the Monte Carlo algorithm is as follows: Estimated expected value: , , From the distribution Independently distributed samples drawn from the data; Estimate the definite integral: , ; High-dimensional integrals: , ,in It is the integration region The volume.