Questionnaire generation method and system based on large language model and scene understanding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN HENGDAI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-07
AI Technical Summary
这些方法难以深入解析非结构化业务需求中的隐含逻辑,无法根据用户的时空环境、认知负荷和心理状态动态调整问卷策略,结果往往导致生成的问卷内容僵化、冗余,且与特定投放场景脱节
首先,在内容构建与语义对齐层面,本申请利用多源文本关键词的频次统计与关联强度分析,将非结构化的客户资料转化为量化的“文本表现程度”,并以此动态调整一级目标的权重。这种机制有效解决了传统LLM生成问卷时容易出现的“幻觉”或“重点偏差”问题,确保模型生成的提示词能够严格忠实于客户的核心诉求。通过将关键词映射表与权重计算相结合,为大语言模型提供了清晰的逻辑约束,使其生成的题目分布不仅符合统计学规律,更在语义层面上实现了与客户真实痛点的深度对齐,大幅提升了问卷的内容效度。
Smart Images

Figure CN122528853A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text generation technology, specifically to a method and system for generating questionnaires based on large language models and scene understanding. Background Technology
[0002] The importance of large language models in questionnaire generation technology lies in their superior semantic understanding and dynamic content generation capabilities, enabling them to transcend the limitations of traditional template-based generation. By transforming unstructured, multi-source business data into logically rigorous, context-appropriate structured questions, large models can effectively improve questionnaire quality. Furthermore, by incorporating context-aware weights, large language models can not only achieve personalized question customization to better match users' immediate psychological state, but also significantly reduce manual design costs while ensuring professional survey logic, thereby establishing an intelligent survey closed loop that achieves both high response rates and ensures data validity.
[0003] Current questionnaire generation technologies primarily rely on static template matching or traditional natural language processing methods, such as keyword extraction and fixed weight allocation. While these technologies can achieve basic question combinations, they have significant shortcomings in semantic understanding and context awareness. These methods struggle to deeply analyze the implicit logic within unstructured business needs and cannot dynamically adjust questionnaire strategies based on users' spatiotemporal environment, cognitive load, and psychological state. The result is often rigid, redundant, and disconnected questionnaire content from specific delivery scenarios. These issues reduce users' willingness to respond, distort feedback data, and fail to meet the growing demand for refined and personalized research and decision-making. Summary of the Invention
[0004] In view of the above, it is necessary to provide a questionnaire generation method and system based on large language models and scene understanding to solve the above problems.
[0005] The first aspect of this application provides a method for generating questionnaires based on large language models and scene understanding, the method comprising: Extract keywords from multi-source texts and categorize them into primary goals and secondary sub-goals according to the target mapping table preset in the questionnaire to be generated; Count the frequency of keywords for each primary objective in the text materials to obtain the key text for each primary objective; based on the number of keywords for each primary objective in all text materials, and combined with the relevance of the key text to the customer, determine the degree of textual performance of each primary objective in the questionnaire to be generated. Based on the difference in the textual representation of each primary objective and the other primary objectives in the questionnaire to be generated, and combined with the proportion of questions for the corresponding primary objectives in the reference questionnaire, the weights of each primary objective in the questionnaire to be generated are adjusted. Analyze the response rate, completion rate, and answer selection distribution of each secondary sub-target in the deployment test of each preset scenario, determine the reasonable performance of user feedback for each secondary sub-target in each deployment scenario, compare the reasonable performance of user feedback for each secondary sub-target with other secondary sub-targets in different deployment scenarios, and determine the weight of each secondary target in the questionnaire to be generated. Based on the weights of the primary and secondary objectives of the questionnaire to be generated, the final number of questions corresponding to each secondary sub-objective is obtained, and structured prompt words are used to call a large language model to generate the questionnaire.
[0006] Preferably, the key text for each primary objective is the text data with the highest number of keywords extracted from each primary objective.
[0007] Preferably, determining the degree of textual representation of each primary objective in the questionnaire to be generated specifically involves: The ratio of the number of keywords extracted from the text data for each primary objective to the total number of keywords extracted from the text data for all primary objectives is used as the keyword frequency for each primary objective. The ratio of the number of corresponding keywords in the key text of each primary objective to the association level of the key text is recorded as the association strength value. Based on the keyword frequency and the association strength value, the text performance level of each primary objective in the questionnaire to be generated is obtained; wherein, the text performance level is positively correlated with both the keyword frequency and the association strength value.
[0008] Preferably, the adjustment of the weights of each primary objective in the questionnaire to be generated is specifically as follows: Calculate the cumulative value of the differences in text performance between each primary objective and all other primary objectives to obtain the text performance differences of each primary objective; The difference between the textual representation level of each primary objective and the proportion of questions for the corresponding primary objective in the reference questionnaire is calculated and recorded as the questionnaire deviation degree of each primary objective. Based on the differences in textual performance of each primary objective and the deviation of the questionnaire, the degree of change in the requirement emphasis of each primary objective is determined and mapped to a preset range. Combined with the preset adjustment coefficient, the requirement emphasis of each primary objective is obtained. The actual weight of each primary objective is determined by the ratio of the weight of each primary objective to the sum of the weights of all primary objectives. This weight is then used as the weight of each primary objective in the questionnaire.
[0009] Preferably, the requirements for obtaining each primary objective are emphasized, specifically: For each primary objective, calculate the product of the preset adjustment coefficient and the mapping value of the required emphasis change, and then add it to the proportion of each primary objective in the reference questionnaire. The maximum value of the sum and the preset value is taken as the required emphasis of each primary objective.
[0010] Preferably, determining the reasonable performance of each secondary sub-target in user feedback across various deployment scenarios specifically involves: The dispersion of the number of each answer appearing in the question corresponding to each secondary sub-target in each delivery scenario is statistically analyzed, which is used as the option distribution of each secondary sub-target in each delivery scenario; the maximum value of the distribution of all options for each secondary sub-target in all delivery scenarios is obtained; Based on the response rate and completion rate of each secondary sub-target in each preset scenario, and combined with the negative correlation mapping result of the maximum value of all option distributions obtained for each secondary sub-target and the difference of option distributions in each delivery scenario, the reasonable performance of user feedback for each secondary sub-target in each delivery scenario is obtained.
[0011] Preferably, determining the weights of each secondary objective in the questionnaire to be generated specifically involves: By comparing the reasonable performance of user feedback for each secondary sub-target in different advertising scenarios and the reasonable performance of user feedback for different secondary sub-targets in the same advertising scenario, we can obtain the degree of emphasis of each secondary sub-target in each advertising scenario. The ratio of the degree of emphasis of each secondary sub-objective in each deployment scenario to the sum of the degree of emphasis of all secondary sub-objectives belonging to the same primary objective in each deployment scenario is used as the actual weight ratio of each secondary sub-objective, thus obtaining the weight of each secondary objective in the questionnaire.
[0012] Preferably, the specific process for obtaining the degree of emphasis of each secondary sub-target in each deployment scenario is as follows: Calculate the sum of the differences between the reasonable performance of user feedback for each secondary sub-objective in each delivery scenario and the reasonable performance of user feedback in all other delivery scenarios; Compare the difference between the reasonable performance of user feedback for each secondary sub-target under the same deployment scenario and the maximum reasonable performance of user feedback for all secondary sub-targets under the primary target. The negative correlation mapping result of the gap is positively fused with the sum of the differences to obtain the degree of emphasis of each secondary sub-target in each deployment scenario.
[0013] Preferably, the process of obtaining the final number of questions corresponding to each secondary sub-objective based on the primary objective weight and secondary objective weight of the questionnaire to be generated is specifically as follows: The integer value of the product of the total number of questions in the questionnaire to be generated, the weight of the primary objective, and the weight of each secondary sub-objective under the primary objective is taken as the final number of questions corresponding to each secondary sub-objective under the primary objective.
[0014] Secondly, embodiments of this application also provide a questionnaire generation system based on a large language model and scene understanding, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0015] This application has at least the following beneficial effects: Firstly, at the content construction and semantic alignment level, this application utilizes frequency statistics and association strength analysis of keywords from multi-source texts to transform unstructured customer data into quantifiable "textual representation levels," and dynamically adjusts the weights of primary objectives accordingly. This mechanism effectively solves the "illusion" or "focus bias" problems that easily occur when generating questionnaires using traditional LLM models, ensuring that the prompts generated by the model are strictly faithful to the core needs of customers. By combining the keyword mapping table with weight calculation, clear logical constraints are provided for the large language model, enabling the distribution of its generated questions to not only conform to statistical laws but also achieve deep semantic alignment with the real pain points of customers, significantly improving the content validity of the questionnaire.
[0016] Secondly, at the level of scenario understanding and interaction optimization, this application introduces historical delivery data (response rate, completion rate, distribution characteristics) to calibrate the weights of secondary sub-objectives, achieving a leap from "static generation" to "dynamic adaptation." By analyzing the rationality of user feedback in different scenarios, it is possible to identify which objectives are more easily accepted or understood by users in specific contexts, thereby guiding the large language model to adjust the priority and expression of questions. This feedback loop based on empirical data makes the generated questionnaire no longer a general template, but a customized tool with "scenario awareness," which can effectively reduce users' cognitive load, improve willingness to answer and data quality, and achieve deep integration of large model technology with actual business scenarios. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of a questionnaire generation method based on a large language model and scene understanding, provided as an embodiment of this application. Detailed Implementation
[0018] In the description of the embodiments in this application, the words "exemplary," "or," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary," "or," and "for example" is intended to present the relevant concepts in a specific manner.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application's specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0020] It should also be noted that the terms "first" and "second" in this application and its accompanying drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. The methods disclosed in the embodiments of this application or the methods shown in the flowcharts include one or more steps for implementing the method. Without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged, and some steps can also be deleted.
[0021] The following, in conjunction with the accompanying drawings, details the specific scheme of the questionnaire generation method and system based on large language models and scene understanding provided in this application.
[0022] Please see Figure 1 The diagram illustrates a flowchart of a questionnaire generation method based on a large language model and scene understanding, according to an embodiment of this application. The method includes the following steps: The first step is to extract keywords from multi-source texts and categorize them into primary goals and secondary sub-goals based on the target mapping table preset in the questionnaire to be generated.
[0023] We gather raw text data from business requirements, documents, meeting minutes, user feedback (customer service tickets, application reviews), historical questionnaires, and competitor reports. We then use automated scripts to integrate and clean the text—removing HTML tags and special symbols, correcting typos, and standardizing terminology—to lay the foundation for subsequent analysis.
[0024] Keyword extraction includes preprocessing and algorithm selection. The preprocessing stage involves word segmentation (using jieba in this example), removal of stop words, and optional part-of-speech filtering to retain content words. Subsequently, a large language model combined with prompt words is used to extract core concepts and implicit information from the text.
[0025] The extracted keywords are compared with a pre-defined classification sequence (including primary and secondary target mapping tables). If a keyword falls within the mapping range of a primary target, it is determined that the keyword belongs to the primary target. After manual verification, synonym merging, and weight assignment, the keyword set corresponding to each primary target and its distribution frequency are output. For example, if keywords (such as "payment failure" or "cashier") fall into the primary target... If the keyword falls within the mapping range of "payment experience" (e.g., "payment experience"), then it is determined that the keyword belongs to the primary target. It should be understood that the primary objectives, secondary sub-objectives, and their corresponding keyword mapping relationships of the questionnaire were pre-constructed before the text analysis was performed.
[0026] The second step is to count the frequency of keywords for each primary objective in the text materials to obtain the key text for each primary objective; based on the number of keywords for each primary objective in all text materials, and combined with the relevance of the key text to the customer, determine the degree of textual performance of each primary objective in the questionnaire to be generated.
[0027] Questionnaire design typically needs to simultaneously address multiple primary objectives within the current timeframe, and these primary objectives may have different emphases. To design a questionnaire that better reflects actual needs, this application accurately identifies the emphasis of different primary objectives based on relevant textual materials such as business requirements documents, meeting minutes, and user feedback.
[0028] To obtain the number of keywords in the text document for primary objective a. The number of keywords extracted from the text data for all primary objectives The ratio of these values is used as the keyword frequency for primary objective a, denoted as [missing value]. .
[0029] The relationships between text and customers are manually categorized to construct a multi-level classification system. For example, Level 1 strong association includes: customer service chat logs, audio / transcript of in-depth user interviews, and user-submitted work orders / complaint emails; Level 2 medium association includes: aggregated analysis reports of customer service work orders (such as summaries of frequently asked questions), and user interview minutes (names omitted, only opinions recorded); Level 3 weak association includes: descriptions of user reviews in competitor analysis reports, and text descriptions generated from user behavior logs (such as "User A frequently exits the page"). Each file is labeled with a corresponding association level to indicate the association level in subsequent formulas. or related weight Input.
[0030] Obtain the primary objective Based on the keyword distribution, identify the text materials containing the most of the target keyword, and use them as primary targets. The key text is denoted as ; Identify the key text The default association level to which it belongs serves as its association level. The higher the association level, the corresponding The smaller the value, the better.
[0031] Establish a weight allocation mechanism based on keyword density: Set a primary objective The number of relevant keywords in the key text is Analysis shows that, The proportion of keywords in the document is positively correlated, and directly maps the relationship between this primary objective and customer needs. The bigger The higher the proportion of demand Association level The stronger.
[0032] Therefore, the textual representation level of primary objective a can be obtained. The specific formula is as follows: ;in, Indicates the frequency of keywords for primary objective 'a'; Indicates the first-level objective The number of keywords in the corresponding key text; This indicates the association level of the key text corresponding to the primary objective 'a'; The correlation strength value is denoted as that of primary target a; This represents the normalization function; in this embodiment, the maximum-minimum normalization function is used. Furthermore, the textual representation levels of different primary objectives required by the questionnaire are obtained as a reference for questionnaire design, making the resulting questionnaire more reasonable.
[0033] It should be noted that all parameters involved in the formula calculation of this application are preprocessed using the maximum-minimum normalization algorithm before being substituted into the formula, and uniformly mapped to the dimensionless characteristic interval of [0,1] to eliminate the logical conflict of direct calculation of data with different dimensions.
[0034] The third step: Based on the difference in the textual representation of each primary objective in the questionnaire to be generated and the other primary objectives, and in conjunction with the proportion of questions for the corresponding primary objectives in the reference questionnaire, adjust the weight of each primary objective in the questionnaire to be generated.
[0035] To ensure a reasonable emphasis on different primary objectives in the questionnaire, an existing reference questionnaire was used as a basis, and each primary objective was compared against the others in terms of the textual representation of each objective. This process allowed for appropriate adjustments to the emphasis on each primary objective, resulting in a final questionnaire design that better meets practical needs.
[0036] Specifically, the text performance difference of a single primary objective 'a' is calculated by summing the differences in text performance between it and all other primary objectives, thus obtaining the text performance difference of primary objective 'a'. This difference is then expressed using the formula... Perform the calculation, where, This indicates the total number of primary objectives in the questionnaire; , These represent the textual representation levels of primary objective a and primary objective k, respectively. In this embodiment, the differences between variables are calculated using interpolation.
[0037] The difference between the textual representation level of primary objective a and the proportion of questions related to primary objective a in the reference questionnaire is calculated and denoted as the questionnaire deviation degree of primary objective a. This deviation is expressed using the formula... To obtain, in the formula, This indicates the degree of textual representation of primary objective a; This indicates the percentage of questions related to Level 1 Objective a in the reference questionnaire. It should be understood that if... , indicating a first-level objective The percentage of questions on the user side ( ) exceeded the pre-set questionnaire weight ( This indicates that the current questionnaire does not pay enough attention to this module and more questions need to be added; if , indicating a first-level objective The percentage of questions on the user side ( The weight of the questionnaire was lower than the pre-set weight. The fact that the current questionnaire focuses too much on this module suggests that the number of questions could be reduced.
[0038] When a single primary objective The greater the difference in textual expression, i.e., the higher the level of the primary objective... Compared to other objectives, it receives significantly more user attention. The greater the deviation of the questionnaire, the worse the current questionnaire design (...). This severely underestimates the actual needs of users. ).
[0039] Therefore, the degree of change in the requirements of primary objective 'a' in the current questionnaire design can be obtained, and the specific expression is as follows: ;in, This represents a preset adjustment coefficient used to limit the range of weight adjustment. In this embodiment, the value is 0.3, but the implementer can adjust it according to the actual situation. This represents the normalization function, specifically the max-min normalization method.
[0040] Processing with the max-min normalization method Obtain the intermediate value ,pass Mapping , so that its range is [-1, 1]. When When the value is larger and positive, the degree to which it emphasizes primary objective a in the reference questionnaire should increase; when The smaller the value, the greater the reduction in the degree of emphasis on primary objective a in the reference questionnaire should be.
[0041] Therefore, the requirements of primary objective a in questionnaire design are emphasized. The specific formula is as follows: In the formula, This represents the mapping value indicating the degree of change in the requirements of primary objective 'a'. Represents the maximum value function. This represents a preset minimum value constraint greater than zero, which is set to 0.01 in this embodiment.
[0042] The above method was used to calculate the required weight of different primary objectives in this questionnaire design. The actual required weight of primary objective a was then obtained by dividing the required weight of a single primary objective by the sum of the required weights of all types of primary objectives. , which serves as the primary objective weight of the questionnaire.
[0043] The fourth step is to analyze the response rate, completion rate, and answer selection distribution of each secondary sub-target in the deployment test of each preset scenario, determine the weight of each secondary sub-target in each deployment scenario, compare the reasonable performance of each secondary sub-target with the other secondary sub-targets in user feedback in different deployment scenarios, and determine the weight of each secondary target in the questionnaire to be generated.
[0044] Users' psychological state and memory clarity vary significantly depending on the time and context in which they receive survey requests, directly leading to biased feedback. To obtain truly effective data, the survey objectives must be precisely aligned with the user's current context. For example, to assess "payment process smoothness" (belonging to "product experience"), the best strategy is to trigger the questionnaire the instant the user enters the "payment completion page." Leveraging the high accessibility of user action memory minimizes cognitive load and guides users to provide accurate evaluations.
[0045] The questions were designed using LLM with a single secondary sub-objective, and the questions were deployed in multiple reasonable scenarios and at different times as experiments to obtain user feedback.
[0046] Obtain the response rate of the questions set for the secondary sub-target b in a single delivery scenario m. (Percentage of questionnaires that answered this question) Completion rate (The percentage of questions with complete answers) and calculate the distribution of answer choices. (Variance of the number of different answer choices). Response rate Completion rate The higher the value, the greater the user's attention to and willingness to participate in the issue; option distribution The larger the value, the wider the distribution of user feedback (not a single concentrated value), providing richer dimensions and sample information. Therefore, we can calculate the reasonable performance of user feedback for secondary sub-objective b in the deployment scenario m. The specific formula is as follows: In the formula, Secondary sub-objective The maximum value of the option distribution in different scenarios; Denote the preset minimum value, which is 1 in this embodiment. It represents the negative correlation mapping result of the difference between the maximum value of all option distributions obtained by the secondary sub-goal b and the option distribution in the delivery scenario m. It should be noted that if the secondary sub-goal lacks historical delivery data, its response rate, completion rate, and option distribution are given the mean value of the remaining known secondary sub-goals in the current delivery scenario.
[0047] Furthermore, calculate the secondary sub-goal in the delivery scenario of the reasonable performance of user feedback and its difference cumulative sum with the reasonable performance of user feedback in other delivery scenarios . The specific formula is: ; in the formula, represents the total number of delivery scenarios. The larger , it indicates that this delivery scenario is the "optimal delivery point" for the secondary sub-goal .
[0048] Compare the reasonable performance of user feedback of the secondary sub-goal in the same delivery scenario with the maximum reasonable performance of user feedback of all secondary sub-goals in this delivery scenario to obtain the gap : ; the smaller , it indicates that the competitiveness of the secondary sub-goal in this delivery scenario is stronger, and it may even be the optimal sub-goal.
[0049] Substitute the above two indicators into the weight allocation formula to obtain the set emphasis degree of the secondary sub-goal b of the primary goal a when the questionnaire delivery scenario is m: ; norm() represents the normalization function, and the maximum-minimum normalization method is adopted. Sub-goals with large advantages and small gaps will obtain significantly amplified weights and will preferentially occupy the core position and question quantity of the questionnaire.
[0050] For a single primary goal and all its subordinate secondary sub-goals , its set emphasis degree only represents the relative strength between sub-goals and is not the final weight. Through normalization processing (dividing by the sum of of all sub-goals under the same primary goal), is transformed into the secondary sub-goal of this primary goal a Actual weighting percentage (That is, the percentage of the number of questions in the primary objective) is used as the weight of the secondary objective in the questionnaire.
[0051] The fifth step is to obtain the final number of questions corresponding to each secondary sub-objective based on the weights of the primary and secondary objectives of the questionnaire to be generated, and to construct structured prompt words to call the large language model to generate the questionnaire.
[0052] Given the weights of the primary objectives and the secondary sub-objectives, designing a questionnaire using LLM first requires input processing and weight mapping. The inputs include the weights of each level, the preset total number of questions, and the expected proportion of question types for each sub-objective.
[0053] Specifically, the actual requirements for Level 1 Target 'a' are weighty and small. Determine the primary objective Percentage of total questions in the entire questionnaire; Secondary sub-objectives The actual weighting percentage determines the secondary sub-objectives Under the primary objective The specific proportion of questions allocated to each primary objective is determined. This leads to the final number of questions for each secondary sub-objective within the questionnaire for each primary objective: ;in, This indicates the total number of questions in the questionnaire to be generated, and round() indicates rounding to the nearest integer. It also specifies the question type requirements for each secondary sub-objective (in this example, single-choice and multiple-choice questions are included). It should be noted that if the calculated final number of questions is equal to 0, then the final number of questions for the corresponding secondary sub-objective will be forced to be 1.
[0054] The survey background description field, the target scenario feature field, the names of each primary / secondary objective, and the final number of questions for each sub-objective calculated are included. The questionnaire is integrated with the question type restrictions into a Prompt, which is then input into LLM to generate a draft questionnaire (JSON format). The generated JSON questionnaire is parsed, and the question distribution is checked to ensure that it conforms to the weight settings. The questionnaire is also manually reviewed for logical consistency (such as deduplication and exclusion). Based on feedback from small sample test data (response rate, completion rate), the question descriptions or local weights are fine-tuned to generate the final deployable questionnaire template, ensuring that the quantified weights are accurately converted into the actual questionnaire structure.
[0055] Based on the same inventive concept as the above methods, this application also provides a questionnaire generation system based on large language models and scene understanding, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described questionnaire generation methods based on large language models and scene understanding.
[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0057] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from its essential characteristics. Therefore, the embodiments described above should be considered exemplary and non-limiting in all respects; modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions of some technical features, without causing the essence of the corresponding technical solutions to deviate from the scope of the technical solutions in the embodiments of this application, should all be included within the protection scope of this application.
Claims
1. A questionnaire generation method based on large language models and scene understanding, characterized in that, The method includes the following steps: Extract keywords from multi-source texts and categorize them into primary goals and secondary sub-goals according to the target mapping table preset in the questionnaire to be generated; Count the frequency of keywords for each primary objective in the text materials to obtain the key text for each primary objective; based on the number of keywords for each primary objective in all text materials, and combined with the relevance of the key text to the customer, determine the degree of textual performance of each primary objective in the questionnaire to be generated. Based on the difference in the textual representation of each primary objective and the other primary objectives in the questionnaire to be generated, and combined with the proportion of questions for the corresponding primary objectives in the reference questionnaire, the weights of each primary objective in the questionnaire to be generated are adjusted. Analyze the response rate, completion rate, and answer selection distribution of each secondary sub-target in the deployment test of each preset scenario, determine the reasonable performance of user feedback for each secondary sub-target in each deployment scenario, compare the reasonable performance of user feedback for each secondary sub-target with other secondary sub-targets in different deployment scenarios, and determine the weight of each secondary target in the questionnaire to be generated. Based on the weights of the primary and secondary objectives of the questionnaire to be generated, the final number of questions corresponding to each secondary sub-objective is obtained, and structured prompt words are used to call a large language model to generate the questionnaire.
2. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The key text for each primary objective is specifically the text data with the highest number of keywords extracted based on each primary objective.
3. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The determination of the textual representation of each primary objective in the questionnaire to be generated specifically involves: The ratio of the number of keywords extracted from the text data for each primary objective to the total number of keywords extracted from the text data for all primary objectives is used as the keyword frequency for each primary objective. The ratio of the number of corresponding keywords in the key text of each primary objective to the association level of the key text is recorded as the association strength value. Based on the keyword frequency and the association strength value, the text performance level of each primary objective in the questionnaire to be generated is obtained; wherein, the text performance level is positively correlated with both the keyword frequency and the association strength value.
4. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The adjustment of the weights of each primary objective in the questionnaire to be generated is as follows: Calculate the cumulative value of the differences in text performance between each primary objective and all other primary objectives to obtain the text performance differences of each primary objective; The difference between the textual representation level of each primary objective and the proportion of questions for the corresponding primary objective in the reference questionnaire is calculated and recorded as the questionnaire deviation degree of each primary objective. Based on the differences in textual performance of each primary objective and the deviation of the questionnaire, the degree of change in the requirement emphasis of each primary objective is determined and mapped to a preset range. Combined with the preset adjustment coefficient, the requirement emphasis of each primary objective is obtained. The actual weight of each primary objective is determined by the ratio of the weight of each primary objective to the sum of the weights of all primary objectives. This weight is then used as the weight of each primary objective in the questionnaire.
5. The questionnaire generation method based on large language models and scene understanding as described in claim 4, characterized in that, The requirements for obtaining each primary objective are emphasized, specifically: For each primary objective, calculate the product of the preset adjustment coefficient and the mapping value of the required emphasis change, and then add it to the proportion of each primary objective in the reference questionnaire. The maximum value of the sum and the preset value is taken as the required emphasis of each primary objective.
6. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The determination of the reasonable performance of each secondary sub-target in user feedback across various deployment scenarios specifically includes: The dispersion of the number of each answer appearing in the question corresponding to each secondary sub-target in each delivery scenario is statistically analyzed, which is used as the option distribution of each secondary sub-target in each delivery scenario; the maximum value of the distribution of all options for each secondary sub-target in all delivery scenarios is obtained; Based on the response rate and completion rate of each secondary sub-target in each preset scenario, and combined with the negative correlation mapping result of the maximum value of all option distributions obtained for each secondary sub-target and the difference of option distributions in each delivery scenario, the reasonable performance of user feedback for each secondary sub-target in each delivery scenario is obtained.
7. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The determination of the weights of each secondary objective in the questionnaire to be generated is specifically as follows: By comparing the reasonable performance of user feedback for each secondary sub-target in different advertising scenarios and the reasonable performance of user feedback for different secondary sub-targets in the same advertising scenario, we can obtain the degree of emphasis of each secondary sub-target in each advertising scenario. The ratio of the degree of emphasis of each secondary sub-objective in each deployment scenario to the sum of the degree of emphasis of all secondary sub-objectives belonging to the same primary objective in each deployment scenario is used as the actual weight ratio of each secondary sub-objective, thus obtaining the weight of each secondary objective in the questionnaire.
8. The questionnaire generation method based on large language models and scene understanding as described in claim 7, characterized in that, The specific process for obtaining the degree of emphasis of each secondary sub-target in each deployment scenario is as follows: Calculate the sum of the differences between the reasonable performance of user feedback for each secondary sub-objective in each delivery scenario and the reasonable performance of user feedback in all other delivery scenarios; Under the same delivery scenario, obtain the difference between the reasonable performance of user feedback for each secondary sub-target and the maximum reasonable performance of user feedback for all secondary sub-targets under the primary target; The negative correlation mapping result of the gap is positively fused with the sum of the differences to obtain the degree of emphasis of each secondary sub-target in each deployment scenario.
9. The questionnaire generation method based on large language models and scene understanding as described in claim 1, characterized in that, The process of obtaining the final number of questions corresponding to each secondary sub-objective based on the weights of the primary and secondary objectives of the questionnaire to be generated is as follows: The integer value of the product of the total number of questions in the questionnaire to be generated, the weight of the primary objective, and the weight of each secondary sub-objective under the primary objective is taken as the final number of questions corresponding to each secondary sub-objective under the primary objective.
10. A questionnaire generation system based on large language models and scene understanding, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-9.