Questionnaire information structured data processing method, system and device
By classifying the questionnaire contents by word segmentation and text, filtering out the keyword combinations and analyzing their emotional labels, the problem of insufficiently detailed existing questionnaire analysis methods is solved, and more accurate user experience analysis and targeted service provision are achieved.
Patent Information
- Application Number
- CN202111492590.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing questionnaire content analysis methods are not meticulous enough to obtain the user's true feelings and cannot provide targeted services and guidance to enterprises.
By obtaining the answers to the collected questionnaire, performing word segmentation processing, dividing the long text into semantic vocabulary combinations, counting the number and frequency of each vocabulary combination, filtering out the target vocabulary combination, and inputting it into the pre-trained text classification algorithm model to obtain topic and emotional labels.
It realizes a more detailed analysis of the questionnaire content, obtains topics and emotional labels that users are concerned about, improves the accuracy of questionnaire information analysis, provides targeted services and guidance for enterprises, reduces repeated questionnaire analysis, and saves service resources.
Smart Images

Figure CN114282524B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology. Specifically, it relates to a method, system, electronic device, and storage medium for processing structured data of questionnaire information. Background Art
[0002] With the vigorous development of the Internet, conducting online questionnaires and service evaluations has become an important means for people to safeguard their rights and improve service quality. After the questionnaires are recovered, it is necessary to analyze the content of the users' questionnaire answers.
[0003] The existing analysis of questionnaire content is often not detailed enough, unable to obtain further information, unable to know the true feelings of users, and unable to provide targeted services and guidance for enterprises. Summary of the Invention
[0004] The first objective of the embodiments of this application is to provide a method for processing structured data of questionnaire information, aiming to solve at least one of the problems existing in the above-mentioned prior art.
[0005] The embodiments of this application are implemented as follows. A method for processing structured data of questionnaire information includes:
[0006] Obtain the answer content of the recovered questionnaire questions, perform word segmentation on the answer content, and split the long text of the answer content into several semantic-compliant vocabulary combinations;
[0007] Count the occurrence times of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and filter out the target vocabulary combinations according to the frequency of occurrence of each vocabulary combination;
[0008] Input the target vocabulary combinations into a pre-trained text classification algorithm model to obtain the topics corresponding to the target vocabulary combinations and the sentiment labels attached to the topics.
[0009] In one embodiment, the performing word segmentation on the answer content includes: converting the answer content into string-type data, and performing word segmentation on the string corresponding to the answer content using the NLP natural language processing algorithm.
[0010] In one embodiment, after splitting the long text of the answer content into several semantic-compliant vocabulary combinations and before counting the occurrence times of each vocabulary combination, it further includes: receiving new vocabulary added by the user and / or receiving a deletion request from the user for vocabulary combinations without business meaning among several vocabulary combinations, adding the new vocabulary to the vocabulary combinations and / or deleting the vocabulary combinations requested by the user to be deleted from the several vocabulary combinations.
[0011] In one embodiment, the training process of the text classification algorithm model includes: presetting multiple topics, hierarchically splitting each topic to obtain fine-grained topics under each topic, establishing sentiment tags for each topic and each fine-grained topic, marking the mapping relationship between the vocabulary combination and the topic, fine-grained topic and sentiment tag, and using the mapping relationship between the vocabulary combination and the topic, fine-grained topic and sentiment tag as the input and output of the text classification model to train and obtain the text classification algorithm model.
[0012] In one embodiment, the emotion tag includes positive, neutral and negative, and a mapping relationship is established between the emotion tag and the vocabulary combination. When the user clicks on the emotion tag, the user jumps to the answer content corresponding to the vocabulary combination mapped by the emotion tag.
[0013] In one embodiment, filtering out target vocabulary combinations according to the frequency of occurrence of each vocabulary combination includes: generating a word cloud visualization chart according to the frequency of occurrence of each vocabulary combination, and filtering out target vocabulary combinations based on preset rules according to the word cloud visualization chart, wherein the preset rule is that the vocabulary combinations are ranked in the top N in terms of text size in the word cloud visualization chart, and N is a positive integer.
[0014] Another object of the embodiment of the present application is to provide a questionnaire information structured data processing system, comprising:
[0015] A word segmentation module is used to obtain the answer content of the collected questionnaire questions, perform word segmentation processing on the answer content, and divide the long text of the answer content into a number of semantically consistent word combinations;
[0016] The statistical module is used to count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and screen out the target vocabulary combination according to the frequency of occurrence of each vocabulary combination;
[0017] The analysis and processing module is used to input the target vocabulary combination into a pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic.
[0018] In one embodiment, the word segmentation processing of the answer content includes: converting the answer content into string type data, and using an NLP natural language processing algorithm to perform word segmentation on the string corresponding to the answer content.
[0019] Another object of an embodiment of the present application is to provide an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the questionnaire information structured data processing method.
[0020] Another object of an embodiment of the present application is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the steps of the questionnaire information structured data processing method.
[0021] The embodiment of the present application provides a method, system, electronic device and storage medium for processing structured data of questionnaire information. The method for processing structured data of questionnaire information obtains the answer content of the recovered questionnaire questions, performs word segmentation processing on the answer content, and divides the long text of the answer content into several semantically consistent vocabulary combinations; counts the number of occurrences of each vocabulary combination, calculates the frequency of occurrence of each vocabulary combination, and selects the target vocabulary combination according to the frequency of occurrence of each vocabulary combination; inputs the target vocabulary combination into a pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. In this way, the recovered questionnaire content can be analyzed more carefully, the topics that users are concerned about and the emotional labels corresponding to the topics can be obtained, the accuracy of the questionnaire information analysis can be improved, and the effective content of the user's answers can be obtained, so as to provide targeted services and guidance for enterprises. Further, the repeated analysis of the recovered questionnaires can be reduced to reduce the waste of service resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The implementation process of the questionnaire information structured data processing method provided in one embodiment of the present application;
[0023] Figure 2 A schematic diagram of the main modules of a questionnaire information structured data processing system provided by an embodiment of the present application;
[0024] Figure 3 An exemplary system architecture diagram that can be applied to the embodiments of the present application;
[0025] Figure 4 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0028] It should be understood that although the terms first, second, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.
[0029] It should be pointed out that the embodiments in this application and the features in the embodiments may be combined with each other if there is no conflict.
[0030] In order to further explain the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific implementation method, structure, characteristics and effects of the present application are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0031] Figure 1 The implementation process of a questionnaire information structured data processing method provided by an embodiment of the present application is shown. For the convenience of description, only the part related to the embodiment of the present application is shown, which is described in detail as follows:
[0032] A method for processing structured data of questionnaire information comprises the following steps:
[0033] S101: Obtaining the answers to the collected questionnaire questions, performing word segmentation processing on the answers, and dividing the long text of the answers into a plurality of semantically consistent word combinations;
[0034] S102: Count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and select the target vocabulary combination according to the frequency of occurrence of each vocabulary combination;
[0035] S103: Input the target vocabulary combination into a pre-trained text classification algorithm model to obtain a topic corresponding to the target vocabulary combination and a sentiment tag attached to the topic.
[0036] In step S101: the answers to the collected questionnaire questions are obtained, the answers are segmented, and the long text of the answers is divided into a number of semantically consistent word combinations.
[0037] During a questionnaire survey, after the questionnaire is sent out and the user completes the questionnaire and collects it, the answers filled in by the user need to be analyzed, especially for the content of open-ended questionnaire questions, where the answers filled in by the user are long text content, so the questionnaire content needs to be analyzed. Questionnaire content analysis can parse the content of the questionnaire from thousands of questionnaires through language models and other methods, obtain the word frequency in the questionnaire content, and analyze the user's concerns based on the word frequency content. For example, taking the catering industry as an example, through word frequency, it can be understood that users may be more concerned about decoration, taste, dining environment, etc. However, such analysis cannot obtain further information, and it is impossible to provide targeted guidance to the company. In other words, it can only understand that the user is concerned about decoration, taste, etc., but it cannot further understand which part of the decoration the user is dissatisfied with, the taste of a specific day, and what the point of dissatisfaction is.
[0038] Here, the answers to the collected questionnaire questions can be obtained. For example, the answers to the open questions by the user are obtained to obtain the long text content of the questionnaire content, and the long text content is segmented to divide the long text of the answer content into several semantically consistent word combinations. This can facilitate subsequent analysis and processing. It should be noted that the semantically consistent word combinations can be segmented in a preset manner or using the word segmentation technology in the natural language processing algorithm to obtain semantically consistent word combinations, such as by training a word segmentation algorithm model.
[0039] For example, when the answer content is "The decoration style of the restaurant is not bad, the braised eggplant is delicious, but the waiter's attitude is not good", word segmentation can be performed, and the vocabulary combinations after word segmentation may include: restaurant's, decoration style, not bad, braised eggplant, delicious.
[0040] In one embodiment, the word segmentation processing of the answer content includes: converting the answer content into string type data, and using the NLP natural language processing algorithm to segment the string corresponding to the answer content. After obtaining the long text content of the answer content, the long text content of the answer content can be converted into string type data to facilitate word segmentation processing as an input of the natural language processing algorithm. Here, the NLP natural language processing algorithm can be used to perform word segmentation processing on the string corresponding to the answer content to obtain a number of semantically consistent vocabulary combinations. For example, by pre-training or selecting an existing NLP natural language processing algorithm model, the string corresponding to the long text content of the answer content can be input into the NLP algorithm model to obtain a semantically consistent vocabulary combination.
[0041] In one embodiment, after the long text of the answer content is divided into several semantically consistent word combinations, before counting the number of occurrences of each word combination, it also includes: receiving new words added by the user and / or receiving a user's request to delete a word combination that does not have business meaning in the several word combinations, adding the new words to the word combinations and / or deleting the word combinations requested to be deleted by the user from the several word combinations. Therefore, after obtaining semantically consistent word combinations through NLP algorithm model processing, users can also add new words that are not segmented by the word segmentation algorithm or delete the word combinations that are segmented by the word segmentation algorithm and do not conform to the business scenario and have no business meaning. Only words that conform to semantics and business scenarios are retained. , so as to improve the accuracy of subsequent statistical analysis, further improve the meticulousness of questionnaire content analysis, and obtain more information.
[0042] In step S102: count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and select the target vocabulary combination according to the frequency of occurrence of each vocabulary combination. After the long text content of the answer content is segmented, multiple vocabulary combinations are obtained. Among the multiple vocabulary combinations, the number of occurrences of each vocabulary combination may be one or more times. Therefore, after obtaining multiple vocabulary combinations, the number of occurrences of each vocabulary combination is counted, thereby calculating the word frequency of each vocabulary combination, and then the target vocabulary combination is selected according to the word frequency of each vocabulary combination as the dimension of subsequent opinion analysis. Here, the selection of the target vocabulary combination can be implemented based on preset rules or preset conditions. For example, when the frequency of occurrence of each vocabulary combination is greater than a certain threshold, the vocabulary combination is determined as the target vocabulary combination.
[0043] In one embodiment, the target vocabulary combination is screened out according to the frequency of occurrence of each vocabulary combination, including: generating a word cloud visualization chart according to the frequency of occurrence of each vocabulary combination, and screening out the target vocabulary combination based on the preset rules according to the word cloud visualization chart, wherein the preset rules are the vocabulary combinations whose text size in the word cloud visualization chart ranks as the top N, and N is a positive integer. Here, a word cloud visualization chart is generated according to the frequency of occurrence of each vocabulary combination, so that the vocabulary combination with a higher frequency can be clearly displayed according to the font size of each vocabulary combination in the word cloud visualization chart, so as to facilitate the determination of the target vocabulary combination. In the word cloud visualization chart, the larger the frequency of the vocabulary, the larger the font size, that is, the higher the user's attention, the larger the font size. For example, some points that customers pay attention to, such as I want to know whether he is talking about my decoration, my product, my service, or my transportation, by counting the frequency of occurrence of words such as decoration, products, and services, the font size in the word cloud chart reflects the user's attention.
[0044] It should be noted that users can also flexibly set the visualization effect of the word cloud, such as the compactness of word arrangement, the horizontal and vertical direction of text display, the word cloud shading style, the background color of the chart, and other personalized configurations. By setting different display forms of the word cloud, you can also adjust the preset rules for filtering the target word combination accordingly.
[0045] Exemplarily, preset rules can be set to filter target vocabulary combinations. For example, the top five vocabulary combinations in terms of font size in the word cloud visualization chart can be set as target vocabulary combinations. It should be noted that the word cloud visualization chart can also be pushed to the administrator client for display. The administrator selects the vocabulary combination he wants to set through the word cloud visualization chart, and then determines the vocabulary combination selected by the administrator as the target vocabulary combination after receiving it. This can improve the operability of questionnaire content information analysis.
[0046] In step S103: the target vocabulary combination is input into the pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. The target vocabulary combination is the evaluation dimension that the user is most concerned about, which is the experience that the user cares about most. The target vocabulary combination is input into the pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. The fine-grained evaluation information of the user's most concerned point in the answer content can be analyzed, and the emotional label of the user's most concerned topic can be obtained.
[0047] In one embodiment, the training process of the text classification algorithm model includes: presetting multiple topics, hierarchically splitting each topic to obtain fine-grained topics under each topic, establishing each topic and the emotion tag of each fine-grained topic, marking the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the emotion tag, and using the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the emotion tag as the input and output of the text classification model to train the text classification algorithm model. Thus, an accurate text classification algorithm model can be obtained, so as to process and output the fine-grained topic and its corresponding emotion tag for the target vocabulary combination. Then, the target vocabulary combination data is input, and the mapped topic and the emotion tag attached to the topic are output through the trained model, so as to learn the user's point of view and facilitate targeted guidance and service to the enterprise.
[0048] In one embodiment, the emotion tag includes positive, neutral and negative, and a mapping relationship is established between the emotion tag and the vocabulary combination. When the user clicks on the emotion tag, it jumps to the answer content corresponding to the vocabulary combination mapped by the emotion tag. In this way, the topic opinion statistical data can be used to identify what content users pay most attention to, and whether the text evaluation associated with a certain opinion is mostly positive or negative, so as to conduct targeted user public opinion research and improve business. At the same time, a public opinion analysis system from abstract to concrete can be realized. For example: the business party finds that the negative mentions of the topic of dish taste are high, and further can view all negative comment texts of the dish taste, and gain a more accurate insight into the user experience, thereby optimizing the business.
[0049] For example, the target vocabulary combination is high-frequency vocabulary such as quality, service, and price. After obtaining the target vocabulary combination, the topic is split into fine-grained hierarchical levels, such as price is split into price level, cost-effectiveness, and discount level. Based on the pre-trained NLP text classification algorithm model, the text of user feedback in the answer content is mapped to specific fine-grained topics, and sentiment analysis of positive, neutral, and negative labels is performed on fine-grained topics to achieve a system in which topics are attached with sentiment labels to form opinions. For example, the model can count the number of comments related to the topic of "price level" as 12, and further count the number of positive, negative, and neutral expressions about "price level" as 10, 5, and 4, respectively.
[0050] Therefore, based on this analysis, a data analysis page can be generated. The data analysis page includes the analysis result data of each topic, the fine-grained topics under each topic, and the emotional tags attached to the topic. Administrators can use the opinion statistics to identify what content users pay most attention to, and whether the text evaluations associated with a certain opinion are mostly positive or negative, and conduct targeted user public opinion research and judgment to improve business. For example, through opinion sentiment analysis, it was found that the most mentioned topic this month was "service attitude". Further analysis found that 80% of the evaluation texts were negative feedback on service attitude. Therefore, the business side can accurately perceive that the business problem is the poor attitude of service personnel, which leads to a decline in user experience, and can carry out targeted tasks such as service personnel training.
[0051] Therefore, the questionnaire information structured data processing method provided by the embodiment of the present application obtains the answer content of the recovered questionnaire questions, performs word segmentation processing on the answer content, and divides the long text of the answer content into several semantically consistent vocabulary combinations; counts the number of occurrences of each vocabulary combination, calculates the frequency of occurrence of each vocabulary combination, and screens out the target vocabulary combination according to the frequency of occurrence of each vocabulary combination; inputs the target vocabulary combination into a pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. In this way, the recovered questionnaire content can be analyzed more carefully, the topics that users are concerned about and the emotional labels corresponding to the topics can be obtained, the accuracy of the questionnaire information analysis can be improved, and the effective content of the user's answers can be obtained, so as to provide targeted services and guidance for enterprises. Further, the repeated analysis of the recovered questionnaires can be reduced to reduce the waste of service resources.
[0052] Figure 2 The main module schematic diagram of the questionnaire information structured data processing system provided by the embodiment of the present application is shown. For the convenience of explanation, only the part related to the embodiment of the present application is shown, which is described in detail as follows:
[0053] A questionnaire information structured data processing system 200, comprising:
[0054] The word segmentation module 201 is used to obtain the answer content of the collected questionnaire questions, perform word segmentation processing on the answer content, and divide the long text of the answer content into a plurality of semantically consistent word combinations;
[0055] A statistical module 202 is used to count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and screen out target vocabulary combinations according to the frequency of occurrence of each vocabulary combination;
[0056] The analysis and processing module 203 is used to input the target vocabulary combination into a pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the sentiment tag attached to the topic.
[0057] The word segmentation module 201 is used to obtain the answer content of the collected questionnaire questions, perform word segmentation processing on the answer content, and divide the long text of the answer content into a plurality of semantically consistent word combinations.
[0058] During a questionnaire survey, after the questionnaire is sent out and the user completes the questionnaire and collects it, the answers filled in by the user need to be analyzed, especially for the content of open-ended questionnaire questions, where the answers filled in by the user are long text content, so the questionnaire content needs to be analyzed. Questionnaire content analysis can parse the content of the questionnaire from thousands of questionnaires through language models and other methods, obtain the word frequency in the questionnaire content, and analyze the user's concerns based on the word frequency content. For example, taking the catering industry as an example, through word frequency, it can be understood that users may be more concerned about decoration, taste, dining environment, etc. However, such analysis cannot obtain further information, and it is impossible to provide targeted guidance to the company. In other words, it can only understand that the user is concerned about decoration, taste, etc., but it cannot further understand which part of the decoration the user is dissatisfied with, the taste of a specific day, and what the point of dissatisfaction is.
[0059] Here, the answers to the collected questionnaire questions can be obtained. For example, the answers to the open questions by the user are obtained to obtain the long text content of the questionnaire content, and the long text content is segmented to divide the long text of the answer content into several semantically consistent word combinations. This can facilitate subsequent analysis and processing. It should be noted that the semantically consistent word combinations can be segmented in a preset manner or using the word segmentation technology in the natural language processing algorithm to obtain semantically consistent word combinations, such as by training a word segmentation algorithm model.
[0060] For example, when the answer content is "The decoration style of the restaurant is not bad, the braised eggplant is delicious, but the waiter's attitude is not good", word segmentation can be performed, and the vocabulary combinations after word segmentation may include: restaurant's, decoration style, not bad, braised eggplant, delicious.
[0061] In one embodiment, the word segmentation processing of the answer content includes: converting the answer content into string type data, and using the NLP natural language processing algorithm to segment the string corresponding to the answer content. After obtaining the long text content of the answer content, the long text content of the answer content can be converted into string type data to facilitate word segmentation processing as an input of the natural language processing algorithm. Here, the NLP natural language processing algorithm can be used to perform word segmentation processing on the string corresponding to the answer content to obtain a number of semantically consistent vocabulary combinations. For example, by pre-training or selecting an existing NLP natural language processing algorithm model, the string corresponding to the long text content of the answer content can be input into the NLP algorithm model to obtain a semantically consistent vocabulary combination.
[0062] In one embodiment, after the long text of the answer content is divided into several semantically consistent word combinations, before counting the number of occurrences of each word combination, it also includes: receiving new words added by the user and / or receiving a user's request to delete a word combination that does not have business meaning in the several word combinations, adding the new words to the word combinations and / or deleting the word combinations requested to be deleted by the user from the several word combinations. Therefore, after obtaining semantically consistent word combinations through NLP algorithm model processing, users can also add new words that are not segmented by the word segmentation algorithm or delete the word combinations that are segmented by the word segmentation algorithm and do not conform to the business scenario and have no business meaning. Only words that conform to semantics and business scenarios are retained. , so as to improve the accuracy of subsequent statistical analysis, further improve the meticulousness of questionnaire content analysis, and obtain more information.
[0063] For the statistical module 202: it is used to count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and select the target vocabulary combination according to the frequency of occurrence of each vocabulary combination. After the long text content of the answer content is segmented, multiple vocabulary combinations are obtained. Among the multiple vocabulary combinations, the number of occurrences of each vocabulary combination may be one or more times. Therefore, after obtaining multiple vocabulary combinations, the number of occurrences of each vocabulary combination is counted, thereby calculating the word frequency of each vocabulary combination, and then the target vocabulary combination is selected according to the word frequency of each vocabulary combination as the dimension of subsequent opinion analysis. Here, the screening of the target vocabulary combination can be implemented based on preset rules or preset conditions. For example, when the frequency of occurrence of each vocabulary combination is greater than a certain threshold, the vocabulary combination is determined as the target vocabulary combination.
[0064] In one embodiment, the target vocabulary combination is screened out according to the frequency of occurrence of each vocabulary combination, including: generating a word cloud visualization chart according to the frequency of occurrence of each vocabulary combination, and screening out the target vocabulary combination based on the preset rules according to the word cloud visualization chart, wherein the preset rules are the vocabulary combinations whose text size in the word cloud visualization chart ranks as the top N, and N is a positive integer. Here, a word cloud visualization chart is generated according to the frequency of occurrence of each vocabulary combination, so that the vocabulary combination with a higher frequency can be clearly displayed according to the font size of each vocabulary combination in the word cloud visualization chart, so as to facilitate the determination of the target vocabulary combination. In the word cloud visualization chart, the larger the frequency of the vocabulary, the larger the font size, that is, the higher the user's attention, the larger the font size. For example, some points that customers pay attention to, such as I want to know whether he is talking about my decoration, my product, my service, or my transportation, by counting the frequency of occurrence of words such as decoration, products, and services, the font size in the word cloud chart reflects the user's attention.
[0065] It should be noted that users can also flexibly set the visualization effect of the word cloud, such as the compactness of word arrangement, the horizontal and vertical direction of text display, the word cloud shading style, the background color of the chart, and other personalized configurations. By setting different display forms of the word cloud, you can also adjust the preset rules for filtering the target word combination accordingly.
[0066] Exemplarily, preset rules can be set to filter target vocabulary combinations. For example, the top five vocabulary combinations in terms of font size in the word cloud visualization chart can be set as target vocabulary combinations. It should be noted that the word cloud visualization chart can also be pushed to the administrator client for display. The administrator selects the vocabulary combination he wants to set through the word cloud visualization chart, and then determines the vocabulary combination selected by the administrator as the target vocabulary combination after receiving it. This can improve the operability of questionnaire content information analysis.
[0067] For the analysis and processing module 203: it is used to input the target vocabulary combination into the pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. The target vocabulary combination is the evaluation dimension that the user is most concerned about, which is the experience that the user cares about most. The target vocabulary combination is input into the pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. The fine-grained evaluation information of the user's most concerned point in the answer content can be analyzed, and the emotional label of the user's most concerned topic can be obtained.
[0068] In one embodiment, the training process of the text classification algorithm model includes: presetting multiple topics, hierarchically splitting each topic to obtain fine-grained topics under each topic, establishing each topic and the emotion tag of each fine-grained topic, marking the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the emotion tag, and using the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the emotion tag as the input and output of the text classification model to train the text classification algorithm model. Thus, an accurate text classification algorithm model can be obtained, so as to process and output the fine-grained topic and its corresponding emotion tag for the target vocabulary combination. Then, the target vocabulary combination data is input, and the mapped topic and the emotion tag attached to the topic are output through the trained model, so as to learn the user's point of view and facilitate targeted guidance and service to the enterprise.
[0069] In one embodiment, the emotion tag includes positive, neutral and negative, and a mapping relationship is established between the emotion tag and the vocabulary combination. When the user clicks on the emotion tag, it jumps to the answer content corresponding to the vocabulary combination mapped by the emotion tag. In this way, the topic opinion statistical data can be used to identify what content users pay most attention to, and whether the text evaluation associated with a certain opinion is mostly positive or negative, so as to conduct targeted user public opinion research and improve business. At the same time, a public opinion analysis system from abstract to concrete can be realized. For example: the business party finds that the negative mentions of the topic of dish taste are high, and further can view all negative comment texts of the dish taste, and gain a more accurate insight into the user experience, thereby optimizing the business.
[0070] For example, the target vocabulary combination is high-frequency vocabulary such as quality, service, and price. After obtaining the target vocabulary combination, the topic is split into fine-grained hierarchical levels, such as price is split into price level, cost-effectiveness, and discount level. Based on the pre-trained NLP text classification algorithm model, the text of user feedback in the answer content is mapped to specific fine-grained topics, and sentiment analysis of positive, neutral, and negative labels is performed on fine-grained topics to achieve a system in which topics are attached with sentiment labels to form opinions. For example, the model can count the number of comments related to the topic of "price level" as 12, and further count the number of positive, negative, and neutral expressions about "price level" as 10, 5, and 4, respectively.
[0071] Therefore, based on this analysis, a data analysis page can be generated. The data analysis page includes the analysis result data of each topic, the fine-grained topics under each topic, and the emotional tags attached to the topic. Administrators can use the opinion statistics to identify what content users pay most attention to, and whether the text evaluations associated with a certain opinion are mostly positive or negative, and conduct targeted user public opinion research and judgment to improve business. For example, through opinion sentiment analysis, it was found that the most mentioned topic this month was "service attitude". Further analysis found that 80% of the evaluation texts were negative feedback on service attitude. Therefore, the business side can accurately perceive that the business problem is the poor attitude of service personnel, which leads to a decline in user experience, and can carry out targeted tasks such as service personnel training.
[0072] Therefore, the questionnaire information structured data processing system provided by the embodiment of the present application obtains the answer content of the recovered questionnaire questions, performs word segmentation processing on the answer content, and divides the long text of the answer content into several semantically consistent vocabulary combinations; counts the number of occurrences of each vocabulary combination, calculates the frequency of occurrence of each vocabulary combination, and screens out the target vocabulary combination according to the frequency of occurrence of each vocabulary combination; inputs the target vocabulary combination into a pre-trained text classification algorithm model to obtain the topic corresponding to the target vocabulary combination and the emotional label attached to the topic. In this way, the recovered questionnaire content can be analyzed more carefully, the topics that users are concerned about and the emotional labels corresponding to the topics can be obtained, the accuracy of the questionnaire information analysis can be improved, and the effective content of the user's answers can be obtained, so as to provide targeted services and guidance for enterprises. Further, the repeated analysis of the recovered questionnaires can be reduced to reduce the waste of service resources.
[0073] An embodiment of the present application also provides an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by one or more processors, the one or more processors implement the questionnaire information structured data processing method of the embodiment of the present application.
[0074] The embodiment of the present application also provides a computer-readable medium on which a computer program is stored. When the program is executed by a processor, the method for processing structured data of questionnaire information of the embodiment of the present application is implemented.
[0075] Figure 3 An exemplary system architecture 300 is shown to which the questionnaire information structured data processing method or system according to the embodiment of the present application can be applied.
[0076] like Figure 3 As shown, the system architecture 300 may include terminal devices 301, 302, 303, a network 304 and a server 305. The network 304 is used to provide a medium for communication links between the terminal devices 301, 302, 303 and the server 305. The network 304 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0077] Users can use terminal devices 301, 302, 303 to interact with server 305 through network 304 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 301, 302, 303, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0078] The terminal devices 301 , 302 , and 303 may be various electronic devices having a display screen and supporting web browsing, including but not limited to a vehicle-mounted smart screen, a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.
[0079] The server 305 may be a server providing various services, such as a background management server providing support for the messages sent by the users using the terminal devices 301, 302, 303. The background management server may perform analysis and other processing after receiving the request of the terminal device, and feed back the processing result to the terminal device.
[0080] It should be noted that the questionnaire information structured data processing method provided in the embodiment of the present application is generally executed by the server 305 or the terminal devices 301, 302, 303. Accordingly, the questionnaire information structured data processing system is generally set in the server 305 or the terminal devices 301, 302, 303.
[0081] It should be understood that Figure 3 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.
[0082] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system 400 suitable for implementing an electronic device of an embodiment of the present application. Figure 4 The computer system shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0083] like Figure 4 As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0084] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.
[0085] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above-mentioned functions defined in the system of the present application are executed.
[0086] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0087] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0088] The modules involved in the embodiments described in the present application may be implemented by software or hardware. The modules described may also be set in a processor, for example, it may be described as: a processor includes a determination module, an extraction module, a training module and a screening module. The names of these modules do not constitute a limitation on the modules themselves in some cases, for example, the determination module may also be described as a "module for determining a candidate user set".
[0089] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
[0090] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for processing structured data of questionnaire information, It is characterized in that include: Obtaining the answers to the collected questionnaires, performing word segmentation on the answers, and dividing the long text of the answers into a number of semantically consistent word combinations; Count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and filter out the target vocabulary combination based on the frequency of occurrence of each vocabulary combination; Inputting the target vocabulary combination into a pre-trained text classification algorithm model to obtain a topic corresponding to the target vocabulary combination and a sentiment label attached to the topic; The training process of the text classification algorithm model includes: presetting multiple topics, hierarchically splitting each topic to obtain fine-grained topics under each topic, establishing sentiment tags for each topic and each fine-grained topic, marking the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the sentiment tag, and using the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the sentiment tag as the input and output of the text classification algorithm model to train and obtain the text classification algorithm model.
2. The questionnaire information structured data processing method according to claim 1, It is characterized in that The word segmentation processing of the answer content includes: converting the answer content into string type data, and using an NLP natural language processing algorithm to segment the string corresponding to the answer content.
3. The questionnaire information structured data processing method according to claim 1, It is characterized in that After dividing the long text of the answer content into several semantically consistent vocabulary combinations and before counting the number of occurrences of each vocabulary combination, it also includes: receiving new words added by users and / or receiving users' requests to delete vocabulary combinations that have no business meaning in several of the vocabulary combinations, adding the new words to the vocabulary combinations and / or deleting the vocabulary combinations requested to be deleted by the users from the several vocabulary combinations.
4. The questionnaire information structured data processing method according to claim 3, It is characterized in that The emotion tags include positive, neutral and negative, and a mapping relationship is established between the emotion tags and the vocabulary combination. When the user clicks on the emotion tag, the user jumps to the answer content corresponding to the vocabulary combination mapped by the emotion tag.
5. The questionnaire information structured data processing method according to claim 1, It is characterized in that The method of filtering out target vocabulary combinations according to the frequency of occurrence of each vocabulary combination includes: generating a word cloud visualization chart according to the frequency of occurrence of each vocabulary combination, and filtering out target vocabulary combinations based on preset rules according to the word cloud visualization chart, wherein the preset rules are that the vocabulary combinations are ranked in the top N in terms of text size in the word cloud visualization chart, and N is a positive integer.
6. A questionnaire information structured data processing system, It is characterized in that include: A word segmentation module is used to obtain the answer content of the collected questionnaire questions, perform word segmentation processing on the answer content, and divide the long text of the answer content into a number of semantically consistent word combinations; The statistical module is used to count the number of occurrences of each vocabulary combination, calculate the frequency of occurrence of each vocabulary combination, and screen out the target vocabulary combination according to the frequency of occurrence of each vocabulary combination; An analysis and processing module, used for inputting the target vocabulary combination into a pre-trained text classification algorithm model to obtain a topic corresponding to the target vocabulary combination and a sentiment tag attached to the topic; The training process of the text classification algorithm model includes: presetting multiple topics, hierarchically splitting each topic to obtain fine-grained topics under each topic, establishing sentiment tags for each topic and each fine-grained topic, marking the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the sentiment tag, and using the mapping relationship between the vocabulary combination and the topic, the fine-grained topic and the sentiment tag as the input and output of the text classification algorithm model to train and obtain the text classification algorithm model.
7. The questionnaire information structured data processing system according to claim 6, It is characterized in that The word segmentation processing of the answer content includes: converting the answer content into string type data, and using an NLP natural language processing algorithm to segment the string corresponding to the answer content.
8. An electronic device, It is characterized in that The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the questionnaire information structured data processing method according to any one of claims 1 to 5.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the questionnaire information structured data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Internet public opinion analysis method and device and computer readable storage medium
CN108959383A