Data processing method and system for professional approval questionnaire

By using the combination method of logical analysis model and semantic components in the data processing of nurse occupational identity questionnaire, the problem of inaccurate data processing in the existing technology is solved, and high-quality data analysis and scientific decision-making are achieved.

CN120181402AActive Publication Date: 2025-06-20NANJING UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510638043.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The data processing method of the existing nurse occupational identity questionnaire relies on manual verification or simple single-dimensional logic verification rules, making it difficult to effectively identify and eliminate logical errors, resulting in inaccurate data analysis and decision-making bias.

Method used

A data processing method of a professional identity questionnaire is adopted. By receiving the questionnaire result data from the online questionnaire system, the logical analysis model is used to identify logical conflicts, and the semantic components express and rewritten the answer content, further analyze the rewritten content, calculate the logical conflict score, and determine the wrong questionnaire with the mean.

Benefits of technology

This method can accurately identify deep logical errors, improve the accuracy and comprehensiveness of logical error recognition, ensure the authenticity and reliability of data, provide high-quality data for nurse professional identity research, and assist scientific decision-making in the nursing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181402A_ABST
    Figure CN120181402A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a data processing method and system for professional approval questionnaires. The method comprises the steps that all questionnaire questions in all questionnaire result data are recognized, a logic analysis model is used for analyzing first logic conflict scores of all questionnaire questions and other questions one by one, and a plurality of target questionnaire questions with the first logic conflict scores higher than a preset threshold value are obtained through screening; using a semantic component to express and rewrite answer content in a single target questionnaire question to obtain a plurality of rewritten answer content, and using a logic analysis model to analyze second logic conflict scores of the target questionnaire questions replacing the rewritten answer content one by one; and if the mean value of the second logic conflict scores corresponding to any target questionnaire question is higher than a preset threshold value, determining that the questionnaire result data to which the target questionnaire question belongs is an error questionnaire, and clearing the error questionnaire. According to the invention, accurate identification of questionnaires with logic errors can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and more particularly, to a data processing method and system for a professional identity questionnaire. Background Art

[0002] In the field of medical and health, nurses, as the core group directly involved in clinical nursing, patient care and other work, their professional identity is directly related to the quality of nursing services, team stability and patient satisfaction.

[0003] Due to the high daily work intensity and complex and changeable work scenarios of nurses, when participating in a professional identity questionnaire survey, it is easy to fill in the questionnaire hastily due to being busy at work and not understanding the questions properly. For example, in questions related to the length of nursing work and the degree of job burnout, there may be contradictory answers such as a short daily working hours but a high degree of job burnout; in questions about the career development expectations and turnover intention of nursing, there will also be unreasonable logic such as longing for career promotion but clearly expressing a recent plan to leave the job. If these questionnaires with logical errors are mixed into the valid data, it will greatly interfere with the accurate analysis of the current situation of nurses' professional identity, and cause deviations in decisions such as nursing talent cultivation strategies and incentive mechanisms based on the data. The existing data processing methods for nurses' professional identity questionnaires mostly rely on manual verification one by one or setting simple single-dimensional logical verification rules. Manual verification not only consumes a large amount of time and labor costs, but also it is easy for the verification personnel to make omissions due to fatigue; the single-dimensional logical verification rules are difficult to capture the complex logical contradictions formed by multiple factors such as work pressure, professional achievement, and doctor-patient relationship in nurses' professional identity questionnaires, and cannot effectively eliminate the questionnaires with logical errors.

[0004] Therefore, for the professional identity questionnaires of the nurse group, there is an urgent need for a more efficient and intelligent data processing method to accurately identify and eliminate the questionnaires with logical errors, so as to provide reliable data support for scientific decision-making in the nursing industry. Summary of the Invention

[0005] In response to this, the present invention provides a data processing method, system, electronic device, computer storage medium and computer program product for a professional identity questionnaire to solve at least one of the above technical problems.

[0006] In the first aspect of the present invention, a data processing method for a professional identity questionnaire is provided, including the following method steps: Receiving a number of questionnaire result data from an online questionnaire system; Identify each questionnaire question in the questionnaire result data, and use a logical analysis model to analyze the first logical conflict score of each questionnaire question with other questions one by one, and screen out several target questionnaire questions whose first logical conflict score is higher than the preset threshold; the questionnaire questions include questions and answer contents; Use semantic components to rewrite the expression of the answer content in a single target questionnaire question to obtain several rewritten answer contents, and use a logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one; If the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, it is determined that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and it is cleared.

[0007] In the second aspect of the present invention, a data processing system for a professional identity questionnaire is provided. The system includes a receiving unit, an identifying unit, and a clearing unit; The receiving unit is used to receive several questionnaire result data from an online questionnaire system; The identifying unit is used to identify each questionnaire question in the questionnaire result data, and use a logical analysis model to analyze the first logical conflict score of each questionnaire question with other questions one by one, and screen out several target questionnaire questions whose first logical conflict score is higher than the preset threshold; the questionnaire questions include questions and answer contents; And, use semantic components to rewrite the expression of the answer content in a single target questionnaire question to obtain several rewritten answer contents, and use a logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one; The clearing unit is used to, if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, determine that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and clear it.

[0008] In the third aspect of the present invention, an electronic device is provided. The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the computer program is executed by the processor, the method described in any one of the foregoing is implemented.

[0009] In the fourth aspect of the present invention, a computer storage medium is provided. The computer storage medium stores a computer program that can be executed by a processor to implement the method described in any one of the foregoing.

[0010] In the fifth aspect of the present invention, a computer program product is provided. The computer program product includes a computer program that can be executed by a processor to implement the method described in any one of the foregoing.

[0011] This data processing method rewrites the expression of the answer content to the target questionnaire questions, mines logical contradictions from multiple dimensions, calculates scores in combination with a logical analysis model, and determines incorrect questionnaires based on the average value. This method effectively avoids misjudgment caused by expression, accurately identifies deep logical errors, greatly improves the accuracy and comprehensiveness of logical error identification, ensures the authenticity and reliability of the retained data, provides high-quality data for the research on nurses' professional identity, and helps scientific decision-making in the nursing industry. Brief Description of the Drawings

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0013] Figure 1 is a schematic flowchart of a data processing method for a professional identity questionnaire disclosed in an embodiment of the present invention; Figure 2 is a schematic structural diagram of a combined algorithm model disclosed in an embodiment of the present invention; Figure 3 is a schematic structural diagram of a data processing system for a professional identity questionnaire disclosed in an embodiment of the present invention. Detailed Embodiments

[0014] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0015] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0016] As Figure 1 shown, an embodiment of the present invention discloses a data processing method for a professional identity questionnaire, including the following method steps: S10, receive a number of questionnaire result data from an online questionnaire system.

[0017] The questionnaire investigators upload the questionnaire in the online questionnaire system in advance. The nurses to be surveyed log in to the online questionnaire system to answer each question in the questionnaire. For example, the questionnaire includes survey questions on three dimensions: professional benefit, professional identity, and work engagement. Among them, the professional benefit dimension focuses on the material and spiritual benefits that nurses obtain from their occupations, such as salary and welfare, professional achievement, etc.; the professional identity dimension focuses on nurses' recognition of their own occupations, emotional belonging, etc.; the work engagement dimension focuses on the degree of connection between nurses and the work environment, colleagues, organizations, etc.

[0018] The execution subject of the solution of the present invention is a pre-constructed server or terminal, which can receive a number of questionnaire result data transmitted from the online questionnaire system by establishing a connection with the online questionnaire system. Each questionnaire result data contains a number of questions and corresponding answer contents.

[0019] S20. Identify each questionnaire question in each of the questionnaire result data, use a logical analysis model to analyze the first logical conflict score between each questionnaire question and other questions one by one, and screen out a number of target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents.

[0020] Due to the heavy daily workload of nurses and possible understanding deviations when participating in questionnaire filling, the questionnaire answer contents are very likely to have logical contradictions, and the logical errors of a single question will have a negative impact on the validity of the entire questionnaire data.

[0021] It can be understood that it is difficult to ensure the accuracy of the judgment by directly judging the questionnaire questions through simple logical rules. Therefore, the present invention calls a pre-constructed logical analysis model to perform logical conflict analysis on it. The logical analysis model incorporates logical relationship rules summarized based on nursing industry knowledge and a large amount of historical questionnaire data, such as association rules between working hours and work pressure, professional skill mastery level and training experience, etc. These logical relationship rules are obtained by pre-training of the logical analysis model. The logical analysis model is preferably constructed, trained or fine-tuned based on Transformer or large language models, and will not be elaborated here.

[0022] Match the logical relationship between each questionnaire question and all other questions in the questionnaire, and calculate the first logical conflict score (the highest value) according to the above logical relationship rules. If the first logical conflict score is higher than the preset threshold, it indicates that there may be a logical conflict, and at this time, it is screened as a target questionnaire question.

[0023] Through this step, the part that may have logical contradictions can be locked from numerous questionnaire questions, effectively narrowing the scope of subsequent in-depth analysis, and significantly improving the efficiency of logical error identification.

[0024] S30. Use semantic components to rewrite the response content in a single target questionnaire question to obtain several rewritten response contents, and use the logical analysis model to analyze the second logical conflict scores of the target questionnaire questions with the rewritten response contents replaced one by one.

[0025] In actual questionnaire data, some responses may seemingly present logical conflicts due to factors such as expression methods and language habits, but actually have reasonable meanings, or there may be hidden logical errors that were not discovered in the calculation of the first logical conflict scores.

[0026] To avoid misjudgment and discover deep - seated logical problems, for the selected target questionnaire questions, the present invention further uses semantic components based on natural language processing technology to perform expression rewriting operations such as synonym replacement and sentence pattern reorganization on the response content without changing the core semantics of the response. Among them, the semantic components perform multiple expression rewritings on the response content in a single target questionnaire question, thereby obtaining multiple rewritten response contents. Replacing the original response content in the target questionnaire question with these rewritten response contents one by one results in the corresponding number of new target questionnaire questions after rewriting.

[0027] Subsequently, use the logical analysis model again to separately re - analyze the logical relationships of each new target questionnaire question after replacing and rewriting the content, and calculate the corresponding number of second logical conflict scores.

[0028] S40. If the mean of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, then determine that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and clear it.

[0029] For each target questionnaire question, summarize the corresponding multiple second logical conflict scores and calculate the mean of these scores. If the calculated mean is higher than the preset threshold, it indicates that the mean of the logical conflict scores from multiple expression perspectives is on the high side, meaning that no matter from which expression method to examine, there are obvious logical unreasonableness in this questionnaire regarding the relevant questions. Thus, it can be inferred that this questionnaire is probably an incorrect questionnaire. At this time, the server or the terminal clears this questionnaire from the dataset to avoid its interference with subsequent data analysis. This data - processing method, by rewriting the expression of the response content of the target questionnaire question, mines logical contradictions from multiple dimensions, combines the logical analysis model to calculate scores and determines incorrect questionnaires by the mean value. This method effectively avoids misjudgment caused by expression, accurately identifies deep - seated logical errors, greatly improves the accuracy and comprehensiveness of logical error identification, ensures that the retained data is true and reliable, provides high - quality data for nurse professional identity research, and helps the nursing industry make scientific decisions.

[0030] As an example, the use of semantic components to rewrite the response content in a single target questionnaire question to obtain a number of rewritten response contents includes: Determine the number of expression rewrites based on the first logical conflict score, and formulate multiple rewrite intensities corresponding to the number of expression rewrites; wherein, the intensity difference between adjacent rewrite intensities is fixed; Using semantic components to rewrite the response content in a single target questionnaire question respectively with each of the rewrite intensities as a constraint condition to obtain a number of rewritten response contents.

[0031] The first logical conflict scores of different target questionnaire questions reflect the degree differences of their logical contradictions. If the same number and intensity of expression rewrites are uniformly adopted, it may not be possible to fully discover deep logical errors, or over-rewrite the content with reasonable logic itself, resulting in misjudgment. Therefore, the present invention sets to differentially determine the number and intensity of expression rewrites according to the severity of logical conflicts to achieve precise analysis.

[0032] Specifically: Pre-establish a corresponding relationship model between the first logical conflict score and the number of expression rewrites. For example, set the mapping rule between the score interval and the number of rewrites. The higher the first logical conflict score of the target questionnaire question, the more corresponding expression rewrites. At the same time, according to the fixed intensity difference, formulate multiple different levels of rewrite intensities, such as low intensity (only perform simple synonym replacement), medium intensity (perform sentence pattern adjustment and partial semantic expansion), high intensity (greatly reorganize the sentence structure and deeply expand the semantics), to form a standardized rewrite intensity system.

[0033] In this way, the expression rewrite operation can adapt to different degrees of logical conflicts, ensure more comprehensive and in-depth analysis of problems with serious logical contradictions, and perform appropriate analysis on problems with minor logical contradictions, improving the pertinence and efficiency of the analysis.

[0034] Then, for each target questionnaire question, sequentially use the multiple pre-formulated rewrite intensities as constraint conditions, call the semantic components based on natural language processing technology, and rewrite the response content according to the corresponding rewrite intensity. For example, first perform simple synonym replacement with low intensity to generate the first version of the rewritten content; then perform sentence pattern adjustment and semantic expansion with medium intensity to generate the second version of the rewritten content; finally perform deep semantic reorganization with high intensity to generate the third version of the rewritten content, and so on, to obtain rewritten response contents with multiple different expressions.

[0035] Through multi-intensity expression rewrites, the logical relationships of target questionnaire questions can be tested from multiple dimensions, potential logical errors can be maximally discovered, misjudgment or missed judgment caused by expression problems can be reduced, and the accuracy and comprehensiveness of logical error identification can be significantly improved, providing a more reliable data basis for subsequent determination of incorrect questionnaires.

[0036] It should be noted that the present invention does not specifically limit the mapping rules in the above corresponding relationship model and the specific expression forms of different rewriting intensities.

[0037] As an example, with each of the above rewriting intensities as constraints, a semantic component is used to rewrite the answer content in a single target questionnaire question to obtain several rewritten answer contents, including: Call a combined algorithm model that pre - constructs a fusion of a BERT semantic generation sub - model and a generative adversarial network, and convert each of the above rewriting intensities into a first intensity and a second intensity using different normalization functions; The BERT semantic generation sub - model takes each of the above first intensities as a constraint, rewrites the answer content in a single target questionnaire question, and outputs multiple preliminary rewritten answer contents greater than the number of the above - mentioned rewritten expressions; The generative adversarial network performs generative adversarial processing on each of the above preliminary rewritten answer contents, obtains the generative adversarial difficulty, and calculates whether the difference (all differences) between the sorted generative adversarial difficulties is higher than the above - mentioned second intensity; If so, screen out the preliminary rewritten answer contents corresponding to the number of the above - mentioned rewritten expressions as the target rewritten answer contents; if not, control the BERT semantic generation sub - model to perform the rewriting of expressions again.

[0038] As Figure 2 shown, the present invention pre - constructs a combined algorithm model based on a BERT semantic generation sub - model and a generative adversarial network (GAN). The BERT semantic generation sub - model in it can perform multiple rewritings of the answer content in a single target questionnaire question under the constraints of each set rewriting intensity, while the generative adversarial network (GAN) uses its internal generation model and discriminant model to analyze and calculate the generative adversarial difficulty of each preliminary rewritten answer content obtained by rewriting, and then determines whether the difference in the generative adversarial difficulty between these preliminary rewritten answer contents exceeds the above - mentioned determined rewriting intensity. If so, it means that the gap between the preliminary rewritten answer contents generated by the BERT semantic generation sub - model is sufficient and can be directly used; otherwise, it means that the preliminary rewritten answer contents generated by the BERT semantic generation sub - model are too similar, and it is not very meaningful to perform a secondary logical conflict analysis and calculation on the target questionnaire question with the rewritten answer content replaced. Specifically: First, the original rewriting intensities are respectively transformed into parameter forms suitable for the BERT semantic generation sub-model and the generative adversarial network through different normalization functions, obtaining multiple first intensities and a second intensity. The second intensity is used to measure whether the difference in the generative adversarial difficulty is large enough, that is, whether the gap between adjacent preliminary rewritten answer contents in the sorted queue is large enough. For example, the Min-Max normalization can be used to transform each rewriting intensity into the corresponding first intensity, and the Z-Score normalization can be used to transform each rewriting intensity into a third intensity. The average value of the differences between adjacent third intensities after sorting is used as the second intensity.

[0039] The BERT semantic generation sub-model performs in-depth semantic understanding on the input answer content, and uses each first intensity as a reference or constraint condition to perform at least one of operations such as lexical substitution, sentence pattern reorganization, and semantic expansion on the answer content (corresponding to different rewriting intensities), thereby generating multiple preliminary rewritten answer contents. The number of preliminary rewritten answer contents should be significantly higher than the aforementioned determined number of statement rewritings to facilitate selection.

[0040] Then, an evaluation and analysis are carried out on the "gap" situation among the multiple preliminary rewritten answer contents generated by the BERT semantic generation sub-model. If the "gap" is insufficient, it means that these preliminary rewritten answer contents are still too similar, and the BERT semantic generation sub-model has not generated preliminary rewritten answer contents that meet each first intensity, and it is necessary to regenerate and evaluate; otherwise, it means that these preliminary rewritten answer contents already have significantly different expression ways (without changing the original meaning of the answer content), that is, the BERT semantic generation sub-model has generated preliminary rewritten answer contents that meet each first intensity, and selection can be made from them.

[0041] The present invention uses a generative adversarial network to evaluate and analyze the above "gap" situation. Specifically, it calculates the "difficulty" shown during the "generation" and "adversarial" processes of the generative adversarial network for the input preliminary rewritten answer content, and uses this "difficulty" to characterize the easy recognition degree of the core semantics of the preliminary rewritten answer content. Subsequently, by comparing the gaps between the easy recognition degrees of the core semantics of each preliminary rewritten answer content, it indirectly analyzes whether the "gap" among the multiple preliminary rewritten answer contents generated by the BERT semantic generation sub-model is large enough.

[0042] As an example, as Figure 2 shown, the generative adversarial network includes a generative model and a discriminative model. Then, the generative adversarial network performs generative adversarial processing on each of the preliminary rewritten answer contents to obtain a generative adversarial difficulty, including: After receiving the preliminary rewritten response content, the generation model performs secondary packaging on the preliminary rewritten response content according to a preset packaging strategy; the packaging strategy is dynamically generated based on the attention mechanism; The discriminative model extracts features from the preliminarily rewritten response content after secondary packaging, and conducts similarity analysis with the feature vectors of the input original preliminarily rewritten response content. If the obtained similarity is higher than the similarity threshold, it is recorded that the discrimination of the preliminarily rewritten response content after secondary packaging is successful; otherwise, the discrimination is recorded as failed. Based on the above records, the discrimination accuracy rate is calculated, and based on the discrimination accuracy rate, the generation adversarial difficulty is calculated.

[0043] The present invention uses the generation model in the generative adversarial network to perform secondary packaging on the preliminary rewritten response content, and the discriminative model to distinguish the preliminarily rewritten response content after secondary packaging, and obtains the generation adversarial difficulty by statistically calculating the discrimination accuracy rate in this "adversarial" process. Obviously, the higher the discrimination accuracy rate, the lower the generation adversarial difficulty.

[0044] After receiving the preliminary rewritten response content, the generation model calls the packaging strategy dynamically generated based on the attention mechanism. The attention mechanism will perform in-depth semantic analysis on the response content to identify key semantic components and sentence structures. For example, for the response to the question about work pressure in the nurse professional identity questionnaire, it focuses on key semantics such as "overtime hours" and "number of patients".

[0045] Then, from the pre-constructed vocabulary library containing function words, nursing scenario-related words, synonyms, etc., according to the probabilities set for different rewriting intensities (the insertion probability is 10%-20% for low rewriting intensity and 40%-60% for high rewriting intensity, where the rewriting intensity refers to the intensity recorded when the BERT semantic generation sub-model rewrites the response content into the rewritten response content), words are inserted at appropriate grammatical positions. For example, rewrite "The pressure of nursing work is great" into "In the ward with more patients, the pressure of daily nursing work is really not small". Thus, more diverse rewritten versions are generated, making the rewritten content more complex and changeable in terms of semantic level and language expression, providing diverse discrimination samples for the discriminative model, and enhancing the comprehensiveness of the generative adversarial network's evaluation of the logical rationality of the rewritten content.

[0046] The discriminative model determines whether the content after secondary packaging is essentially the same as the original preliminary rewritten response content. Its discrimination accuracy rate can be used to characterize the recognizability of the core semantics of the preliminary rewritten response content, and further indirectly analyze whether the "gap" between multiple preliminary rewritten response contents generated by the BERT semantic generation sub-model is large enough.

[0047] Specifically, the discriminant model uses a deep learning algorithm to extract features from the preliminary rewritten answer content after secondary packaging and converts it into a feature vector. At the same time, the feature vector of the preliminary rewritten answer content of the input before secondary packaging is extracted, for example, by calculating the cosine similarity of the two to perform similarity analysis. For example, when analyzing nurses' answers related to professional identity, the similarity of semantic features such as "professional pride" and "professional value" in the content before and after secondary packaging is compared. If the similarity is higher than the pre-set similarity threshold (such as 0.8), the content after secondary packaging is judged to be similar to the original content, and it is recorded as a successful discrimination; otherwise, it is recorded as a failed discrimination.

[0048] Based on the discrimination records of the discriminant model for multiple times (such as 10 rounds or 20 rounds), the discrimination accuracy is calculated, that is, the ratio of successful discrimination times to the total number of discrimination times. Then, according to the formula such as Generative Adversarial Difficulty = 1-Discrimination Accuracy, the Generative Adversarial Difficulty is calculated.

[0049] It is understandable that the higher the difficulty of generative adversarial training, the more difficult it is for the content packaged by the generative model to be recognized by the discriminant model, and the more complex and obscure the initial rewritten answer content is in terms of logic and semantic expression, and the lower the recognition degree; conversely, it means that the rewritten content is relatively simple and direct, and the recognition degree is high.

[0050] As an example, the method further includes: outputting the erroneous questionnaire for personnel verification, and clearing the erroneous questionnaire after verification and confirmation.

[0051] In order to avoid mistaken deletion, the determined erroneous questionnaire can also be output to relevant personnel for verification. If the relevant personnel also confirm that it is an erroneous questionnaire, it will be deleted.

[0052] like Figure 3 As shown, the embodiment of the present invention further provides a data processing system 100 for a career identity questionnaire, the system comprising a receiving unit 101, an identification unit 102 and a clearing unit 103; The receiving unit 101 is used to receive a number of questionnaire result data from an online questionnaire system; The identification unit 102 is used to identify each questionnaire question in each questionnaire result data, use the logic analysis model to analyze the first logic conflict score of each questionnaire question and other questions one by one, and screen out a number of target questionnaire questions whose first logic conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents; and, using the semantic component to rewrite the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents, and using the logic analysis model to analyze the second logic conflict scores of the target questionnaire questions in which the rewritten answer contents have been replaced one by one; The clearing unit 103 is configured to determine that the questionnaire result data to which a target questionnaire question belongs is an incorrect questionnaire and clear it if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than a preset threshold.

[0053] As an example, the recognition unit 102 is configured to: Determine the number of rewritten expressions according to the first logical conflict score, and formulate a plurality of rewriting intensities corresponding to the number of rewritten expressions; wherein, the intensity difference between adjacent rewriting intensities is fixed; Respectively using each of the rewriting intensities as a constraint condition, use semantic components to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents.

[0054] As an example, the recognition unit 102 is configured to: Call a pre-constructed combined algorithm model of a fused BERT semantic generation sub-model and a generative adversarial network, and convert each of the rewriting intensities into a first intensity and a second intensity by using different normalization functions; The BERT semantic generation sub-model takes each of the first intensities as a constraint condition, rewrites the answer content in a single target questionnaire question, and outputs a plurality of preliminary rewritten answer contents greater than the number of rewritten expressions; The generative adversarial network performs generative adversarial processing on each of the preliminary rewritten answer contents to obtain a generative adversarial difficulty, and calculates whether the difference between the sorted generative adversarial difficulties is higher than the second intensity; If so, screen out the preliminary rewritten answer content corresponding to the number of rewritten expressions as the target rewritten answer content; if not, control the BERT semantic generation sub-model to perform rewriting again.

[0055] As an example, the generative adversarial network includes a generation model and a discriminant model, then the recognition unit 102 is configured to: After receiving the preliminary rewritten answer content, the generation model performs secondary packaging on the preliminary rewritten answer content according to a preset packaging strategy; the packaging strategy is dynamically generated based on an attention mechanism; The discriminant model extracts features from the preliminarily rewritten answer content after secondary packaging, and performs similarity analysis with the feature vector of the original preliminarily rewritten answer content input. If the obtained similarity is higher than the similarity threshold, record that the discrimination of the preliminarily rewritten answer content after secondary packaging is successful, otherwise record the discrimination as failed; Calculate the discrimination accuracy rate based on the above records, and calculate the generative adversarial difficulty based on the discrimination accuracy rate.

[0056] As an example, the clearing unit 103 is further configured to: Output the error questionnaire for personnel verification, and clear the error questionnaire after verification and confirmation.

[0057] An embodiment of the present invention also provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the computer program is executed by the processor, the method described in any one of the foregoing is implemented.

[0058] An embodiment of the present invention also provides a computer storage medium, which stores a computer program executable by a processor to implement the method described in any one of the foregoing.

[0059] An embodiment of the present invention also provides a computer program product, which includes a computer program executable by a processor to implement the method described in any one of the foregoing.

[0060] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0061] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data processing method for a professional identity questionnaire, characterized by: The method comprises the following steps: Receiving a number of questionnaire result data from an online questionnaire system; The questionnaire questions use a logic analysis model to analyze the first logic conflict scores of each questionnaire question and other questions one by one, and select a number of target questionnaire questions whose first logic conflict scores are higher than a preset threshold; the questionnaire questions include questions and answer contents; Using semantic components to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents, and using a logical analysis model to analyze the second logical conflict scores of the target questionnaire questions that have replaced the rewritten answer contents one by one; If the average of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, the occupational identity survey questionnaire to which the target questionnaire question belongs is determined to be an incorrect questionnaire and is cleared.

2. The data processing method of a professional identity questionnaire according to claim 1 is characterized by: The method of using semantic components to express and rewrite the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents includes: Determine the number of expression rewrites according to the first logic conflict score, and formulate a plurality of rewrite strengths corresponding to the number of expression rewrites; wherein the strength difference between adjacent rewrite strengths is fixed; Taking each of the rewriting strengths as constraint conditions, semantic components are used to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents.

3. The data processing method of a professional identity questionnaire according to claim 2 is characterized by: The rewriting strengths are used as constraints respectively, and the answer content in a single target questionnaire question is expressed and rewritten using semantic components to obtain a number of rewritten answer contents, including: Calling a pre-built combined algorithm model that integrates the BERT semantic generation sub-model and the generative adversarial network, and converting each of the rewriting strengths into each first strength and second strength using different normalization functions; The BERT semantic generation sub-model uses each of the first strengths as a constraint condition to rewrite the answer content in a single target questionnaire question, and outputs a plurality of preliminary rewritten answer contents greater than the number of rewritten expressions; The generative adversarial network performs generative adversarial processing on each of the preliminary rewritten answer contents to obtain a generative adversarial difficulty, and calculates whether the difference between the sorted generative adversarial difficulties is higher than the second strength; If so, the preliminary rewritten answer content corresponding to the number of expression rewrites is screened out as the target rewritten answer content; if not, the BERT semantic generation sub-model is controlled to rewrite the expression.

4. The data processing method of a professional identity questionnaire according to claim 3 is characterized by: The generative adversarial network includes a generative model and a discriminative model, and the generative adversarial network performs generative adversarial processing on each of the preliminary rewritten answer contents to obtain a generative adversarial difficulty, including: After receiving the preliminary rewritten answer content, the generative model performs secondary packaging on the preliminary rewritten answer content according to a pre-set packaging strategy; the packaging strategy is dynamically generated based on an attention mechanism; The discriminant model extracts features from the preliminary rewritten answer content after secondary packaging, and performs similarity analysis with the feature vector of the input original preliminary rewritten answer content. If the obtained similarity is higher than the similarity threshold, the discrimination of the preliminary rewritten answer content after secondary packaging is recorded as successful, otherwise it is recorded as failed. The discrimination accuracy is calculated based on the above records, and the generation adversarial difficulty is calculated based on the discrimination accuracy.

5. The data processing method of a professional identity questionnaire according to claim 1 is characterized by: The method further includes: outputting the erroneous questionnaire for personnel to verify, and clearing the erroneous questionnaire after verification.

6. A data processing system for a professional identity questionnaire, characterized in that: The system comprises a receiving unit, an identification unit and a clearing unit; The receiving unit is used to receive a number of questionnaire result data from the online questionnaire system; The identification unit is used to identify each questionnaire question in each questionnaire result data, use the logic analysis model to analyze the first logic conflict score of each questionnaire question and other questions one by one, and screen out a number of target questionnaire questions whose first logic conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents; and, using the semantic component to rewrite the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents, and using the logic analysis model to analyze the second logic conflict scores of the target questionnaire questions in which the rewritten answer contents have been replaced one by one; The clearing unit is used to determine that the questionnaire result data to which any target questionnaire question belongs is an erroneous questionnaire and clear it if the average value of the second logical conflict scores corresponding to the target questionnaire question is higher than a preset threshold.

7. The data processing system for a professional identity questionnaire according to claim 6 is characterized by: The identification unit is used for: Determine the number of expression rewrites according to the first logic conflict score, and formulate a plurality of rewrite strengths corresponding to the number of expression rewrites; wherein the strength difference between adjacent rewrite strengths is fixed; Taking each of the rewriting strengths as constraint conditions, semantic components are used to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.

9. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method according to any one of claims 1 to 5.

10. A computer program product, characterized in that: The computer program product comprises a computer program executable by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Semantic-based image adaptive method by combination of slit cropping and non-homogeneous mapping

    CN101923703A

  • Virtual chatting method and system relieving psychological pressure of adolescents

    CN105206284A

  • System and device for evaluating psychological state based on text semantic vector model

    CN110570941A

  • Copywritting recommendation method and device and electronic equipment

    CN110852793A

  • Data processing method, device and storage medium

    EP4297039A1