Deep learning-based method and apparatus for generating user survey questionnaire
Through deep learning technology and pre-trained models, combined with cross-verification and hyperparameter tuning, the automated design of user survey questionnaires is realized, solving the problems of low degree of automation and low efficiency of questionnaires design in the existing technology, and improving the quality and design efficiency of questionnaires.
Patent Information
- Application Number
- PCT/CN2024/125158
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-19
AI Technical Summary
The existing user survey questionnaire design is low, the target positioning and labeling are inaccurate, the work efficiency is low, the technical threshold is high, and the operation is cumbersome.
Using a deep learning-based method, the LLaMA pre-trained model is used to learn language statistical rules and semantic representations, combined with MSE and K-fold cross-validation to evaluate the model performance, and hyperparameter tuning is used to automatically generate high-quality questionnaires that meet user needs.
It improves the quality and rationality of questionnaire design, reduces the difficulty of design, realizes the automation and efficiency of questionnaire generation, and is suitable for research needs of different industries and products.
Smart Images

Figure CN2024125158_19062025_PF_FP_ABST
Abstract
Description
A method and device for generating user survey questionnaires based on deep learning Technical Field
[0001] The present invention belongs to the technical field of questionnaire generation, and in particular relates to a method and device for generating user survey questionnaires based on deep learning. Background Art
[0002] Currently, user research is a crucial tool for evaluating products and services. For enterprise products, user research can save valuable time, development costs, and resources, leading to better and more successful products. For users, user research helps products better align with their real needs. By understanding our users, we can design features that are useful, easy to use, and powerful enough to solve real problems. Compared to user research methods like interviews, scenario experiments, and focus groups, questionnaires offer high efficiency, low cost, and objective, unified content. Furthermore, questionnaires are not limited to a specific number of participants, allowing for research targeting a wider range of users.
[0003] Currently, there are many difficulties in conducting user research using questionnaires:
[0004] The success of a questionnaire survey depends largely on whether the questionnaire design is reasonable. This kind of work often requires personnel to have professional knowledge such as statistics or sociology.
[0005] During the questionnaire design process, questionnaire designers need to ensure objectivity as much as possible.
[0006] Depending on the industry and product of the user survey, the questionnaire model and design method are different, and an appropriate model needs to be adopted for design statistics.
[0007] In this case, the questionnaire design for user research is still a task with high technical barriers for most teams or enterprises. Therefore, a method and device for generating user research questionnaires based on deep learning are proposed.
[0008] Summary of the Invention
[0009] In light of the above shortcomings of the existing technology, the present invention aims to provide a method and apparatus for generating user survey questionnaires based on deep learning. By leveraging a large language model for deep learning, this method automatically generates survey questionnaires based on simple user input of questionnaire design requirements. Based on survey questionnaire data collected for different industries, products, and survey objectives, a machine learning algorithm model learns question details and options to generate high-quality survey questionnaires that meet user needs. Algorithms are then used to optimize the performance of specific models to improve model efficiency. Finally, GPT is used to optimize the language of the questionnaire, resulting in a questionnaire that can be directly used for user surveys.
[0010] In a first aspect of the present invention, a method for generating a user survey questionnaire based on deep learning is proposed, comprising:
[0011] S1. A questionnaire survey model based on LLaMA pre-training for learning language statistical laws and semantic representations;
[0012] S2. Use MSE to perform K-fold cross validation on the trained model: Set the model performance evaluation threshold Y, use MSE to perform K-fold cross validation on the trained model, and perform the following steps based on the validation results:
[0013] The MES value of the model training result is less than the performance evaluation threshold Y, indicating that the model effect meets expectations;
[0014] The MES value of the model training result is not less than the performance evaluation threshold Y, and the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model;
[0015] S3. Use the DE algorithm to perform hyperparameter tuning on the questionnaire survey model, and re-compare the MES value of the model training result with the performance evaluation threshold Y after hyperparameter tuning until the model effect meets expectations;
[0016] S4. Input user needs into the questionnaire survey model to generate a user survey questionnaire.
[0017] Furthermore, the questionnaire survey model based on LLaMA pre-training for learning statistical laws and semantic representation of language includes:
[0018] Input the prepared training set data and test set data set for model training into the LLaMA model, learn the statistical laws and semantic representation of language through the training process, and form a questionnaire survey model;
[0019] Furthermore, before the questionnaire survey model based on LLaMA pre-training for learning the statistical laws and semantic representation of language is used, the following steps are also included:
[0020] Obtain sample data from existing research reports and perform data cleaning on the acquired data;
[0021] Data labeling is performed on the sample data of the survey report after data cleaning, including data labeling of the questions and options in the collected questionnaires;
[0022] Randomly select 70-80% of the total data as the data used for model training, and the remaining 20-30% of the total data as the test set;
[0023] The test set data samples are evenly divided into k data subsets for evaluation of subsequent training results.
[0024] Furthermore, when performing K-fold cross validation on the trained model using MSE, the following steps are included:
[0025] Determine the number of k in the k data subsets;
[0026] Randomly select one of the k subsets as the validation set, and the remaining k-1 subsets as the training set. Repeat k times, using different training and validation sets each time, until all the subsets in the k subsets have been used as training and validation sets.
[0027] Calculate the average of the MSE values of the k evaluation results as the performance evaluation threshold Y for the final model performance evaluation.
[0028] Furthermore, when calculating the average of the MSE values of the k evaluation results, the following formula is used: MSE = (1 / n) * ∑ (y_pred - y_true) 2 ;
[0029] Where y_pred is the predicted value, y_true is the true value, n is the number of samples, and ∑ represents the sum.
[0030] Furthermore, the data annotation of the questions and option contents in the collected questionnaire includes keyword information annotation, and the annotated data is used as a data sample.
[0031] Furthermore, when the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model, the following steps are included:
[0032] Initialize the population and DE parameters;
[0033] Calculate the MSE value in each population;
[0034] Set the maximum number of iterations T for algorithm tuning. When the algorithm cycles T times, the algorithm tuning is completed.
[0035] The following formula is used for mutation operation: V i (g+1)=X R1 (g)+F[X R2 (g)-X R3 (g)];
[0036] Where: R1, R2 and R3 are three random numbers in the interval [1, NP], F is the scaling factor, and g represents the gth generation;
[0037] The crossover operation is performed using the following formula:
[0038] Perform selection operation: By comparing the fitness values of the new individual and the target individual, the individual with high fitness is selected to remain in the population as the target individual of the next generation.
[0039] Furthermore, when initializing the population, the population initialization operation is performed by randomly generating question titles, options, and sequence parameters.
[0040] In a second aspect of the present invention, a device for implementing a method for generating a user survey questionnaire based on deep learning is proposed, which is used to implement the method for generating a user survey questionnaire based on deep learning as described above, and the device includes:
[0041] Acquisition module, used to obtain sample data of existing research reports;
[0042] The processing module is used to clean the sample data of the obtained survey report;
[0043] The annotation module is used to annotate the sample data of the survey report after data cleaning;
[0044] Pre-training module, which trains questionnaire survey models for learning language statistical laws and semantic representation;
[0045] Validation module, used to perform K-fold cross validation on the trained model using MSE;
[0046] A tuning module, used to perform hyperparameter tuning on the questionnaire survey model using a DE algorithm;
[0047] The generation module is used to generate user survey questionnaires based on user needs.
[0048] According to a third aspect of the present invention, an electronic device is provided, comprising:
[0049] at least one processor; and,
[0050] a memory communicatively connected to the at least one processor; wherein,
[0051] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a user survey questionnaire based on deep learning as described above.
[0052] In a fourth aspect of the present invention, a computer-readable storage medium is proposed, which stores computer instructions, and the computer instructions are used to enable a computer to execute the method of generating a user survey questionnaire based on deep learning as described above.
[0053] The beneficial effects of the present invention are as follows:
[0054] 1. The method and apparatus described in this invention determine the appropriate research questions and options based on the user's industry, product, and research intent, and automatically generate a questionnaire. This improves the quality and rationality of questionnaire design and reduces its difficulty.
[0055] 2. The present invention only requires users to input descriptive requirements, such as "conduct a user preference survey on newly produced dairy snacks in the food factory". The model used in the present invention will recognize the text entered by the user and generate a model that meets the user's needs through the trained questionnaire generation model.
[0056] 3. This paper uses the LLaMA model. Through deep learning, the model can generate a questionnaire that meets user needs based on the linguistic meaning of the industry, product, purpose, etc. The advantage of using this pre-trained model is that it achieves excellent performance with a small parameter scale.
[0057] 4. The present invention introduces the MSE and k-fold cross-validation methods to evaluate the training effect of the model, so that the questionnaire generation model can stably output accurate content.
[0058] 5. The DE algorithm is used in the present invention to perform model tuning on the model. The iteration of the algorithm continuously optimizes the hyperparameters in the model. In the present invention, the main optimized parameters of the model include learning rate, number of iterations, population size and other parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings described below are only some of the embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings.
[0060] FIG1 is a flow chart of a method for generating a user survey questionnaire based on deep learning according to an embodiment of the present invention.
[0061] FIG2 is a schematic diagram of k-fold cross-validation of a method for generating a user survey questionnaire based on deep learning according to an embodiment of the present invention.
[0062] FIG3 is a flowchart of an algorithm tuning method for generating a user survey questionnaire based on deep learning according to an embodiment of the present invention.
[0063] FIG4 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present invention.
[0065] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0066] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0067] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0068] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0069] The present invention proposes a method and device for generating user survey questionnaires based on deep learning, which solves the problems of low automation, inaccurate target positioning and labeling, low work efficiency, high technical requirements for testers, cumbersome operations and long time consumption in existing methods.
[0070] Before describing the embodiments of the present invention, the present invention provides an explanation of professional terms:
[0071] LLaMA (Large Language Model Meta AI): LLaMA is an open-source pre-trained language model released by MetaAI's Facebook Artificial Intelligence Lab (FAIR). The model was trained between November 2022 and February 2023.
[0072] MSE (Mean Square Error): Mean square error, which is a commonly used indicator to measure the prediction accuracy of data prediction models or regression models.
[0073] K-fold cross-validation: K-fold cross-validation is a commonly used model evaluation method used to evaluate and select the performance and generalization ability of machine learning models.
[0074] Differential Evolution (DE): A global optimization algorithm based on the principles of biological evolution, DE searches for optimal solutions by simulating mutation and inheritance. It boasts excellent global search and convergence performance, making it suitable for solving high-dimensional, nonlinear, and multimodal optimization problems and widely used in various fields.
[0075] Method Example
[0076] 1 , which shows the basic process steps of the method of the present invention, a method for generating a user survey questionnaire based on deep learning provided by the present invention includes the following steps:
[0077] S1. A questionnaire survey model based on LLaMA pre-training for learning language statistical laws and semantic representations;
[0078] In this step, a large amount of data is collected before pre-training to provide data support, including:
[0079] S1.1. Data collection: Collect a large number of research report samples, questionnaire samples, relevant product information and other data.
[0080] When collecting data, we collected various types of questionnaire questions and their corresponding options from the Internet, professional consulting companies or professional think tanks in various industries, such as:
[0081] Would you like to try a healthy dairy snack made from natural ingredients and without additives?
[0082] ·yes
[0083] ·no
[0084] To ensure the effectiveness of the model, the data collected is no less than 30,000.
[0085] S1.2 Data Cleaning: Clean the collected data, including removing noise data, processing missing values, correcting erroneous information, etc. This step is used to ensure the accuracy of subsequent model training and task generation.
[0086] Specifically, in actual operation, duplicate data in the collected data are deleted and missing values are processed: check whether there are missing values in the collected questionnaire questions and option data. If the missing values can be filled, fill the missing values; if there are too many missing values, delete the data with a large number of missing values.
[0087] S1.3 Data labeling: Label the questions and options in the collected questionnaires, including keyword information labeling: industry, manufacturer, product, category, etc.; question sentiment tendency: positive, negative, and use the labeled data as data samples. The number of data samples must not be less than 20,000.
[0088] In actual operation, the cleaned data is labeled (using BIESO), and the labeled data forms a sample set. The sample set data is stored in the database for subsequent model training and model effect evaluation;
[0089] The specific marking process is as follows:
[0090] B (Beginning): Indicates the starting character of a named entity.
[0091] I (Inside): Indicates the internal character of a named entity.
[0092] E (End): Indicates the end character of the named entity.
[0093] S (Single): represents a single named entity character.
[0094] O (Outside): indicates a non-named entity character.
[0095] For example, the following questionnaire questions and option data are marked:
[0096] Are you (S-person) willing to try (O) a (O) healthy (B) milk (B-product) snack (E-category) made from (O) natural (B-) ingredients (E-) and without (B-) additives (E-)?
[0097] Yes (S-option)
[0098] ·No(S-option)
[0099] The data is labeled with emotional types, and the emotional type atmosphere is positive (positive), neutral (neutral), and negative (negative).
[0100] Are you (S-person) willing to try (O) a (O) healthy (B-positive) (E-product) snack (B-category) made (O) from (O) natural (B-positive) raw materials (I-positive) and without (B-positive) additives (I-positive)?
[0101] Yes (S-option)
[0102] ·No (S-option).
[0103] S1.4. Divide the data set: It is used to divide the data in the database, distinguishing the training set and test set used for model training, as well as the training set and validation set used to verify the model effect.
[0104] S1.41. Randomly select a portion of the data as the data used for model training, of which 70% to 80% of the data is used as the training set, and the remaining 20% to 30% of the data is used as the test set.
[0105] S1.42. Divide the remaining processed data samples into k data subsets for subsequent evaluation of training results. The number of subsets K can be, for example, 10 / 15 / 20. Generally, it is 10 for the first operation. In the following content, k is taken as 10.
[0106] S1.5. Model Pre-training: Use the LLaMA pre-trained model for language modeling tasks. Input the prepared training and test datasets into the LLaMA model. Through the training process, it learns the statistical laws and semantic representations of language, forming a questionnaire generation model.
[0107] S2. As shown in Figure 2, it is a schematic diagram of k-fold cross validation. The trained model is subjected to k-fold cross validation using MSE: the model performance evaluation threshold Y is set, and the trained model is subjected to k-fold cross validation using MSE. The following steps are performed based on the validation results:
[0108] The MES value of the model training result is less than the performance evaluation threshold Y, indicating that the model effect meets expectations;
[0109] The MES value of the model training result is not less than the performance evaluation threshold Y, and the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model.
[0110] Step S2 is mainly to judge the model effect, which specifically includes the following steps:
[0111] S2.1. Determine the number of k in the k subsets (K is 5 to 10 according to the actual situation of the model).
[0112] S2.2. Randomly select one of the k subsets as the validation set, and the remaining k-1 subsets as the training set.
[0113] S2.3. Repeat this step k times, using different training and validation sets each time.
[0114] S2.4. The process ends when all the k subsets have been used as training sets and validation sets.
[0115] S2.5. Calculate the average MSE value of the k evaluation results and use it as the threshold for the final model performance evaluation. If the MSE value of the model training result is less than this threshold, it proves that the model effect meets expectations.
[0116] In actual operation, MSE is used to perform K-fold cross validation on the trained model. According to the actual situation of the model, the K-fold cross training is divided into 10 subsets to determine whether the model effect meets expectations. The detailed steps are as follows:
[0117] Randomly select one of the 10 subsets as the validation set, and the remaining 9 subsets as the training set;
[0118] This step was repeated 10 times, using different training and validation sets each time.
[0119] The process ends when all the 10 subsets have been used as training sets and validation sets.
[0120] Calculate the MSE value of 10 cross-validations and take the average of the MSE values of the 10 validations as the model effect threshold, which is used to judge the effect of model tuning; MSE = (1 / n) * ∑ (y_pred - y_true) 2 ;
[0121] Where y_pred is the predicted value, y_true is the true value, n is the number of samples, and ∑ represents the sum.
[0122] S3, referring to FIG3 , is the algorithm tuning process of the present invention, wherein the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model, and after the hyperparameter tuning, the MES value of the model training result is re-compared with the performance evaluation threshold Y until the model effect meets expectations;
[0123] The tuning process is to make the model generated by the questionnaire survey model more consistent with user psychological expectations. It specifically includes the following steps:
[0124] S3.1. Initialize the population and DE parameters: Initialize the population and DE parameters. In this invention, we initialize the population by randomly generating parameters such as question titles, options, and order.
[0125] S3.2. Calculate the MSE value in each population;
[0126] S3.3. Set the maximum number of iterations T for algorithm tuning. When the algorithm cycles T times, the algorithm tuning is complete.
[0127] In actual operation, the value of T can be 50 / 100 / 150 or other values. Here, T is set to 100. Performing one mutation operation, crossover operation, and selection operation is considered as one algorithm iteration. After this process is repeated 100 times, the algorithm tuning is terminated to verify the algorithm tuning effect. Through algorithm tuning, the optimal parameters of learning rate, number of iterations, and population size are found.
[0128] S3.4. Use the following formula to perform mutation operation: V i (g+1)=X R1 (g)+F[X R2 (g)-X R3 (g)];
[0129] Where: R1, R2 and R3 are three random numbers in the interval [1, NP], F is the scaling factor, and g represents the gth generation;
[0130] S3.5. Use the following formula to perform the crossover operation:
[0131] Among them, CR is the crossover probability, and new individuals are randomly generated by probability;
[0132] V i,j (g+1): This is a mutation operation of the differential evolution algorithm.
[0133] g+1 represents the next generation;
[0134] i and j represent the jth dimension of the i-th individual in the population. The algorithm generates new solution candidates by randomly perturbing the individuals, improving the diversity of the population, and constructs the next generation of individuals by randomly extracting three different individuals and performing linear operations on them.
[0135] x i,j (g): This is the value of the jth dimension of the i-th individual in the population at the g (current) generation.
[0136] U i,j (g+1): This represents the new solution obtained after the crossover operation. In the crossover operation, the mutation vector Vi and the target vector xi are recombined to generate a new solution Ui.
[0137] if rand(0,1): rand(0,1) here generates a random number between 0 and 1. The if function determines whether to use the mutation vector or the target vector based on the value of the random number generated by rand(0,1).
[0138] S3.6. Perform selection operation: By comparing the fitness values of the new individual and the target individual, select the individual with higher fitness to keep in the population as the target individual of the next generation.
[0139] After the above steps S1-S3 are completed, user needs can be input into the questionnaire survey model to generate a questionnaire, which includes the following steps:
[0140] S4. Inputting user needs into the questionnaire survey model to generate a user survey questionnaire, specifically including:
[0141] S4.1. Input requirement description: The user inputs a specific requirement description, which is passed as input to the tuned LLaMA model.
[0142] S4.2. Questionnaire Generation: The optimized questionnaire generation model generates a preliminary questionnaire based on the user's needs description. This questionnaire will be generated based on various factors in the user's needs, including relevant questions and options.
[0143] S4.3. Questionnaire Modification: Use the GPT model to further modify the language of the generated questionnaire. The GPT model has the ability to generate language, which can enhance the expressiveness and fluency of the questionnaire, ensuring that the questionnaire content is more consistent with the user's needs.
[0144] S4.4. Questionnaire output: The modified questionnaire is output to the user, which can be an editable questionnaire file or an online survey link, so that the user can further edit, share or use it.
[0145] In the present invention, users only need to input descriptive requirements, such as "conduct a user preference survey on newly produced dairy snacks in the food factory". The model used in the present invention will recognize the text input by the user and generate a model that meets the user's needs through the trained questionnaire generation model.
[0146] This paper uses the LLaMA model. Through deep learning, the model can generate a questionnaire that meets user needs based on the language meaning of the industry, product, purpose, etc. The advantage of using this pre-trained model is that it achieves excellent performance based on a smaller parameter scale.
[0147] The present invention introduces the MSE and k-fold cross-validation methods to evaluate the training effect of the model, so that the questionnaire generation model can stably output accurate content.
[0148] The DE algorithm is used in the present invention to perform model tuning on the model. The iteration of the algorithm continuously optimizes the hyperparameters in the model. In the present invention, the main optimized parameters of the model include learning rate, number of iterations, population size and other parameters.
[0149] Device embodiment
[0150] Another specific embodiment of the present invention discloses a device for generating user survey questionnaires based on deep learning, including:
[0151] Acquisition module, used to obtain sample data of existing research reports;
[0152] The processing module is used to clean the sample data of the obtained survey report;
[0153] The annotation module is used to annotate the sample data of the survey report after data cleaning;
[0154] Pre-training module, which trains questionnaire survey models for learning language statistical laws and semantic representation;
[0155] Validation module, used to perform K-fold cross validation on the trained model using MSE;
[0156] A tuning module, used to perform hyperparameter tuning on the questionnaire survey model using a DE algorithm;
[0157] The generation module is used to generate user survey questionnaires based on user needs.
[0158] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0159] On the other hand, the present invention provides an electronic device. FIG4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. It shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device shown in FIG4 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0160] As shown in Figure 4, electronic equipment may include a processor (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In RAM, various programs and data required for the operation of the electronic equipment are also stored. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0161] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as magnetic tapes and hard disks; and communication devices. The communication device can allow the electronic device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Although FIG4 shows an electronic device with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or provided instead.
[0162] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by the processor, all or part of the steps of a method for generating a user survey questionnaire based on deep learning in an embodiment of the present disclosure are executed.
[0163] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0164] According to an embodiment of the present disclosure, a computer-readable storage medium stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the method for generating a user survey questionnaire based on deep learning described in each embodiment of the present disclosure are performed.
[0165] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0166] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0167] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0168] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0169] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0170] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0171] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.
[0172] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein. The above description has been provided for the purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for generating a user survey questionnaire based on deep learning, characterized in that: include: S1. Questionnaire survey model based on LLaMA pre-training for learning statistical laws and semantic representation of language; S2. Use MSE to perform K-fold cross validation on the trained model: Set the model performance evaluation threshold Y, use MSE to perform K-fold cross validation on the trained model, and perform the following steps based on the validation results: The MES value of the model training result is less than the performance evaluation threshold Y, and it is determined that the model effect meets expectations; The MES value of the model training result is not less than the performance evaluation threshold Y, and the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model; S3. Use the DE algorithm to perform hyperparameter tuning on the questionnaire survey model, and re-compare the MES value of the model training result with the performance evaluation threshold Y after the hyperparameter tuning until the model effect meets expectations; S4. Input user needs into the questionnaire survey model to generate a user survey questionnaire.
2. The method for generating a user survey questionnaire based on deep learning according to claim 1, characterized in that: The questionnaire survey model based on LLaMA pre-training for learning the statistical laws and semantic representation of language includes: The prepared training set data and test set data sets for model training are input into the LLaMA model. The statistical laws and semantic representations of the language are learned through the training process to form a questionnaire survey model.
3. The method for generating a user survey questionnaire based on deep learning according to claim 2, characterized in that: Before the questionnaire survey model based on LLaMA pre-training for learning the statistical laws and semantic representation of language, the following steps are also included: Obtain sample data from existing research reports and clean the acquired data; Data labeling is performed on the sample data of the survey report after data cleaning, including data labeling of the questions and options in the questionnaire collection; Randomly select 70-80% of the total data as the data used for model training, and the remaining 20-30% of the total data as the test set; The test set data samples are evenly divided into k data subsets for evaluation of subsequent training results.
4. The method for generating a user survey questionnaire based on deep learning according to claim 3, characterized in that: When the MSE is used to perform K-fold cross validation on the trained model, the following steps are included: Determine the number of k in the k data subsets; Randomly select one of the k subsets as the validation set and the remaining k-1 subsets as the training set. Repeat this process k times, using different training and validation sets each time, until all the k subsets have been trained. End after training set and validation set; Calculate the average MSE value of k evaluation results as the performance evaluation threshold Y for the final model performance evaluation; When calculating the average value of the MSE value of the k evaluation results, the following formula is used: MSE=(1 / n)*∑(y_pred-y_true) 2 ; Among them, y_pred is the predicted value, y_true is the true value, n is the number of samples, and ∑ represents the sum.
5. The method for generating a user survey questionnaire based on deep learning according to claim 2, characterized in that: The data annotation of the questions and option contents in the collected questionnaire includes keyword information annotation, and the annotated data is used as a data sample.
6. The method for generating a user survey questionnaire based on deep learning according to claim 5, characterized in that: When the DE algorithm is used to perform hyperparameter tuning on the questionnaire survey model, the following steps are included: Initialize the population and DE parameters; Calculate the MSE value in each population; Set the maximum number of iterations of algorithm tuning to T. When the algorithm cycles to T times, the algorithm tuning is completed. The following formula is used for mutation operation: V i (g+1)=X R1 (g)+F[X R2 (g)-X R3 (g)]; Where: R1, R2 and R3 are three random numbers in the interval [1, NP], F is the scaling factor, and g represents the gth generation; The crossover operation is performed using the following formula: Perform selection operation: By comparing the fitness values of the new individual and the target individual, select the individual with high fitness to keep in the population as the target individual of the next generation.
7. The method for generating a user survey questionnaire based on deep learning according to claim 6, characterized in that: When the population is initialized, the population is initialized by randomly generating question titles, options, and sequence parameters.
8. A device for generating user survey questionnaires based on deep learning, characterized in that: For implementing the method for generating a user survey questionnaire based on deep learning as described in any one of claims 1 to 7, the device comprises: Acquisition module, used to obtain sample data of existing survey reports; A processing module is used to clean the sample data of the obtained survey report; The labeling module is used to label the sample data of the survey report after data cleaning; Pre-training module, which trains questionnaire survey models for learning statistical laws and semantic representation of languages; The validation module is used to perform K-fold cross validation on the trained model using MSE; A tuning module, used for performing hyperparameter tuning on the questionnaire survey model using a DE algorithm; The generation module is used to generate user survey questionnaires based on user needs.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a user survey questionnaire based on deep learning as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which can be loaded and executed by a processor as described in any one of claims 1-7 as a method for generating a user survey questionnaire based on deep learning.
Citation Information
Patent Citations
Questionnaire generation method, server and computer readable storage medium
CN108428152A
Model training method and device suitable for large language model, equipment and medium
CN116976424A
Method and device for generating user survey questionnaire based on deep learning
CN117670420A
System and method for generating, transmitting and using customized survey questionnaires
US20140180766A1
Cited By
Research questionnaire matching method based on AI behavior recognition
CN121981771A