Data processing method and system for professional identity questionnaire

The method uses a logic analysis model and semantic rewriting to accurately identify and remove erroneous nurse surveys, improving data quality for informed nursing industry decisions.

CN120181402BActive Publication Date: 2025-07-15NANJING UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510638043.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-15
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing nurse occupational identity questionnaire data processing methods rely on manual verification or single-dimensional logical verification, and cannot effectively identify and eliminate logical errors, resulting in data analysis deviations.

Method used

The method of combining logical analysis model and semantic components is used to identify the logical conflict scores between questionnaire problems and other problems, express rewriting and multi-dimensional logical analysis are carried out to determine the logical errors of the questionnaire.

Benefits of technology

Accurately identify and eliminate logical errors questionnaires, improve data accuracy and comprehensiveness, and provide reliable data support for scientific decision-making in the nursing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181402B_ABST
    Figure CN120181402B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data processing, and provides a data processing method and system for a professional identity questionnaire. The method includes: identifying each questionnaire question in each questionnaire result data, using a logical analysis model to analyze the first logical conflict score between each questionnaire question and other questions one by one, and screening out several target questionnaire questions whose first logical conflict score is higher than a preset threshold; using a semantic component to rewrite the answer content in a single target questionnaire question to obtain several rewritten answer contents, and using the logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one; if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, it is determined that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and is cleared. The present invention can accurately identify questionnaires with logical errors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a data processing method and system for a professional identity questionnaire. Background Art

[0002] In the field of medical and health, nurses, as the core group directly involved in clinical nursing, patient care and other work, their professional identity is directly related to the quality of nursing services, team stability and patient satisfaction.

[0003] Due to the heavy daily workload and complex and changeable work scenarios of nurses, when participating in a professional identity questionnaire survey, it is easy to fill in the questionnaire hastily due to busy work and misunderstand the questions. For example, in the question regarding the working hours of nursing work and the degree of job burnout, there may be contradictory answers such as a short daily working hours but a serious job burnout; in the question about the career development expectations and turnover intention of nursing profession, there will also be an unreasonable logic of longing for career promotion but clearly expressing a recent plan to leave the job. If these questionnaires with logical errors are mixed into the valid data, it will greatly interfere with the accurate analysis of the current situation of nurses' professional identity, and cause deviations in decisions such as nursing talent training strategies and incentive mechanisms based on the data.

[0004] The existing data processing methods for nurses' professional identity questionnaires mostly rely on manual verification one by one or setting simple single-dimensional logical verification rules. Manual verification not only consumes a large amount of time and labor costs, but also it is easy for the verification personnel to make omissions due to fatigue; the single-dimensional logical verification rules are difficult to capture the complex logical contradictions formed by multiple factors such as work pressure, professional achievement, and doctor-patient relationship in the nurses' professional identity questionnaire, and cannot effectively eliminate the questionnaires with logical errors.

[0005] Therefore, for the professional identity questionnaire of nurses, there is an urgent need for a more efficient and intelligent data processing method to accurately identify and eliminate the questionnaires with logical errors, so as to provide reliable data support for scientific decision-making in the nursing industry. Summary of the Invention

[0006] In view of this, the present invention provides a data processing method, system, electronic device, computer storage medium and computer program product for a professional identity questionnaire to solve at least one of the above technical problems.

[0007] In the first aspect of the present invention, a data processing method for a professional identity questionnaire is provided, including the following method steps:

[0008] Receiving a number of questionnaire result data from an online questionnaire system;

[0009] Identify each questionnaire question in the questionnaire result data, and use a logical analysis model to analyze the first logical conflict score of each questionnaire question with other questions one by one, and screen out several target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents;

[0010] Use semantic components to rewrite the expression of the answer content in a single target questionnaire question to obtain several rewritten answer contents, and use a logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one;

[0011] If the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, it is determined that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and it is cleared.

[0012] In the second aspect of the present invention, a data processing system for a professional identity questionnaire is provided, and the system includes a receiving unit, an identifying unit, and a clearing unit;

[0013] The receiving unit is used to receive several questionnaire result data from an online questionnaire system;

[0014] The identifying unit is used to identify each questionnaire question in the questionnaire result data, and use a logical analysis model to analyze the first logical conflict score of each questionnaire question with other questions one by one, and screen out several target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents;

[0015] And, use semantic components to rewrite the expression of the answer content in a single target questionnaire question to obtain several rewritten answer contents, and use a logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one;

[0016] The clearing unit is used to, if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, determine that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and clear it.

[0017] In the third aspect of the present invention, an electronic device is provided, and the electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and when the computer program is executed by the processor, the method described in any one of the foregoing is implemented.

[0018] In the fourth aspect of the present invention, a computer storage medium is provided, and the computer storage medium stores a computer program that can be executed by a processor to implement the method described in any one of the foregoing.

[0019] In a fifth aspect of the present invention, there is provided a computer program product comprising a computer program executable by a processor to implement the method as described in any one of the foregoing.

[0020] This data processing method rewrites the expression of the answer content to the target questionnaire questions, mines logical contradictions from multiple dimensions, combines a logical analysis model to calculate scores, and determines the error questionnaires by the mean value. This method effectively avoids misjudgment caused by expression, accurately identifies deep logical errors, greatly improves the accuracy and comprehensiveness of logical error identification, ensures the authenticity and reliability of the retained data, provides high-quality data for the research on nurses' professional identity, and helps the nursing industry make scientific decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a schematic flowchart of a data processing method for a professional identity questionnaire disclosed in an embodiment of the present invention;

[0023] Figure 2 is a schematic structural diagram of a combined algorithm model disclosed in an embodiment of the present invention;

[0024] Figure 3 is a schematic structural diagram of a data processing system for a professional identity questionnaire disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope protected by the present application.

[0026] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0027] As Figure 1 shown, an embodiment of the present invention discloses a data processing method for a professional identity questionnaire, including the following method steps:

[0028] S10. Receive several questionnaire result data from an online questionnaire system.

[0029] Questionnaire investigators previously upload questionnaires in the online questionnaire system, and the surveyed nurses log in to the online questionnaire system to answer each question in the questionnaire. For example, the questionnaire includes survey questions in three dimensions: professional benefit sense, professional identity, and work embedding. Among them, the professional benefit sense dimension focuses on the material and spiritual benefits that nurses obtain from their occupations, such as salary and welfare, professional achievement, etc.; the professional identity dimension focuses on nurses' recognition of their own occupations, emotional belonging, etc.; the work embedding dimension focuses on the degree of closeness of the connection between nurses and the work environment, colleagues, organizations, etc.

[0030] The execution subject of the solution of the present invention is a pre-constructed server or terminal, which can receive several questionnaire result data transmitted from the online questionnaire system by establishing a connection with the online questionnaire system. Each questionnaire result data contains several questions and corresponding answer contents.

[0031] S20. Identify each questionnaire question in each of the questionnaire result data, use a logical analysis model to analyze the first logical conflict score between each questionnaire question and other questions one by one, and screen out several target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents.

[0032] Due to the heavy daily workload of nurses and possible understanding deviations when participating in questionnaire filling, the questionnaire answer contents are very likely to have logical contradictions, and the logical errors of individual questions will have a negative impact on the validity of the entire questionnaire data.

[0033] It can be understood that it is difficult to ensure the accuracy of the determination by directly judging questionnaire questions through simple logical rules. Therefore, the present invention calls a pre-constructed logical analysis model to perform logical conflict analysis on it. The logical analysis model incorporates logical relationship rules summarized based on nursing industry knowledge and a large amount of historical questionnaire data, such as correlation rules between working hours and work pressure, professional skill mastery level and training experience, etc. These logical relationship rules are obtained by pre-training the logical analysis model. The logical analysis model is preferably constructed, trained, or fine-tuned based on a Transformer or a large language model, and the details are not described herein again.

[0034] Match the logical relationship between each questionnaire question and all other questions in the questionnaire, and calculate the first logical conflict score (the highest value) according to the above logical relationship rules. If the first logical conflict score is higher than the preset threshold, it indicates that there may be a logical conflict, and at this time, it is screened as a target questionnaire question.

[0035] Through this step, it is possible to lock down the parts that may have logical contradictions among numerous questionnaire questions, effectively narrowing down the scope of subsequent in-depth analysis, and significantly improving the efficiency of logical error identification.

[0036] S30. Use semantic components to rewrite the response content in a single target questionnaire question to obtain several rewritten response contents, and use the logical analysis model to analyze the second logical conflict scores of the target questionnaire questions with the rewritten response contents replaced one by one.

[0037] In actual questionnaire data, some responses may seemingly present logical conflicts due to factors such as expression methods and language habits, but actually have reasonable meanings, or there may be hidden logical errors that were not discovered in the calculation of the first logical conflict scores.

[0038] To avoid misjudgment and dig out deep-seated logical problems, for the selected target questionnaire questions, the present invention further uses semantic components based on natural language processing technology to perform expression rewriting operations such as synonym replacement and sentence pattern reorganization on the response content without changing the core semantics of the response. Among them, the semantic components perform multiple expression rewritings on the response content in a single target questionnaire question to obtain multiple rewritten response contents, and replacing the original response content in the target questionnaire question with these rewritten response contents one by one results in the corresponding number of new target questionnaire questions after rewriting.

[0039] Subsequently, use the logical analysis model again to re-analyze the logical relationships of each new target questionnaire question after replacing the rewritten content, and calculate the corresponding number of second logical conflict scores.

[0040] S40. If the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, it is determined that the questionnaire result data to which the target questionnaire question belongs is an incorrect questionnaire and it is cleared.

[0041] For each target questionnaire question, summarize the corresponding multiple second logical conflict scores and calculate the average value of these scores. If the calculated average value is higher than the preset threshold, it indicates that the average value of the logical conflict scores from multiple expression perspectives is on the high side, meaning that no matter from which expression method to examine, there are obvious logical unreasonable points in this questionnaire in the relevant questions. From this, it can be inferred that this questionnaire is probably an incorrect questionnaire. At this time, the server or the terminal clears this questionnaire from the dataset to avoid its interference with subsequent data analysis.

[0042] This data processing method rewrites the answers to the target questionnaire questions, explores logical contradictions from multiple dimensions, calculates scores based on the logical analysis model, and uses the mean to determine incorrect questionnaires. This method effectively avoids misjudgments caused by expressions, accurately identifies deep logical errors, greatly improves the accuracy and comprehensiveness of logical error identification, ensures that the retained data is authentic and reliable, provides high-quality data for nurse professional identity research, and helps the nursing industry make scientific decisions.

[0043] As an example, the use of semantic components to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents includes:

[0044] Determine the number of expression rewrites according to the first logic conflict score, and formulate a plurality of rewrite strengths corresponding to the number of expression rewrites; wherein the strength difference between adjacent rewrite strengths is fixed;

[0045] Taking each of the rewriting strengths as constraint conditions, semantic components are used to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents.

[0046] The first logical conflict scores of different target questionnaire questions reflect the degree of logical contradiction. If the same number and intensity of expression rewriting are used, it may not be possible to fully explore deep logical errors, or excessive rewriting of logically reasonable content may lead to misjudgment. Therefore, the present invention is set to differentiate the number and intensity of expression rewriting according to the severity of the logical conflict to achieve accurate analysis.

[0047] Specifically: a corresponding relationship model between the first logical conflict score and the number of rewritten statements is established in advance, for example, a mapping rule between the score range and the number of rewritten statements is set. The higher the first logical conflict score of the target questionnaire question, the more corresponding statement rewrites are required. At the same time, according to the fixed intensity difference, multiple different levels of rewriting intensity are formulated, such as low intensity (only simple synonym replacement), medium intensity (sentence adjustment and partial semantic expansion), and high intensity (substantially reorganize the sentence structure and deeply expand the semantics), forming a standardized rewriting intensity system.

[0048] In this way, the statement rewriting operation can adapt to different degrees of logical conflicts, ensuring a more comprehensive and in-depth analysis of problems with serious logical contradictions and a moderate analysis of problems with minor logical contradictions, thereby improving the pertinence and efficiency of the analysis.

[0049] Then, for each target questionnaire question, a plurality of pre - formulated rewriting intensities are successively used as constraints, and a semantic component based on natural language processing technology is called to rewrite the answer content according to the corresponding rewriting intensity. For example, first perform simple synonym replacement at a low intensity to generate a first - version rewritten content; then perform sentence - pattern adjustment and semantic expansion at a medium intensity to generate a second - version rewritten content; finally, perform in - depth semantic recombination at a high intensity to generate a third - version rewritten content, and so on, to obtain rewritten answer contents with multiple different expressions.

[0050] Through the rewriting of expressions with multiple intensities, the logical relationship of the target questionnaire question can be tested from multiple dimensions, potential logical errors can be maximally mined, misjudgment or missed judgment caused by expression problems can be reduced, and the accuracy and comprehensiveness of logical error recognition can be significantly improved, providing a more reliable data basis for subsequent determination of incorrect questionnaires.

[0051] It should be noted that the present invention does not specifically limit the mapping rules in the above - mentioned correspondence model and the specific expression forms of different rewriting intensities.

[0052] As an example, taking each of the rewriting intensities as a constraint condition, using a semantic component to rewrite the answer content in a single target questionnaire question to obtain several rewritten answer contents, including:

[0053] Call a combined algorithm model that pre - constructs a fusion of a BERT semantic generation sub - model and a generative adversarial network, and convert each of the rewriting intensities into a first intensity and a second intensity by using different normalization functions;

[0054] The BERT semantic generation sub - model takes each of the first intensities as a constraint condition to rewrite the answer content in a single target questionnaire question, and outputs a plurality of preliminary rewritten answer contents greater than the number of the expression rewritings;

[0055] The generative adversarial network performs generative adversarial processing on each of the preliminary rewritten answer contents to obtain a generative adversarial difficulty, and calculates whether the difference (all differences) between the sorted generative adversarial difficulties is higher than the second intensity;

[0056] If so, screen out the preliminary rewritten answer contents corresponding to the number of the expression rewritings as the target rewritten answer contents; if not, control the BERT semantic generation sub - model to perform expression rewriting again.

[0057] Such as Figure 2As shown, the present invention constructs a combined algorithm model based on a BERT semantic generation sub-model and a generative adversarial network (GAN) in advance. The BERT semantic generation sub-model can rewrite the response content in a single target questionnaire question multiple times under the constraints of various set rewriting intensities. The generative adversarial network (GAN) uses its internal generative model and discriminative model to analyze and calculate the generative adversarial difficulty of each preliminarily rewritten response content obtained by rewriting, and then determines whether the difference in generative adversarial difficulty between these preliminarily rewritten response contents exceeds the aforementioned determined rewriting intensity. If so, it means that the differences between the preliminarily rewritten response contents generated by the BERT semantic generation sub-model are sufficient and can be directly used; otherwise, it means that the preliminarily rewritten response contents generated by the BERT semantic generation sub-model are too similar, and it is not meaningful to perform a secondary logical conflict analysis and calculation on the target questionnaire question with the rewritten response content replaced by using a logical analysis model. Specifically:

[0058] First, the original rewriting intensities are respectively transformed into parameter forms adapted to the BERT semantic generation sub-model and the generative adversarial network through different normalization functions to obtain multiple first intensities and a second intensity. The second intensity is used to measure whether the difference in generative adversarial difficulty is large enough, that is, whether the difference between adjacent preliminarily rewritten response contents in the sorted queue is large enough. For example, the Min-Max normalization can be used to transform each rewriting intensity into the corresponding first intensity, and the Z-Score normalization can be used to transform each rewriting intensity into a third intensity. The average value of the differences between adjacent third intensities after sorting is used as the second intensity.

[0059] The BERT semantic generation sub-model deeply understands the input response content and uses each first intensity as a reference or constraint condition to perform at least one of operations such as lexical substitution, sentence pattern reorganization, and semantic expansion on the response content (corresponding to different rewriting intensities), thereby generating multiple preliminarily rewritten response contents. The number of preliminarily rewritten response contents should be significantly higher than the aforementioned determined number of statement rewritings to facilitate selection.

[0060] Then, an evaluation and analysis are performed on the "difference" situation among the multiple preliminarily rewritten response contents generated by the BERT semantic generation sub-model. If the "difference" is insufficient, it means that these preliminarily rewritten response contents are still too similar, and the BERT semantic generation sub-model has not generated preliminarily rewritten response contents that meet each first intensity, and they need to be regenerated and evaluated; otherwise, it means that these preliminarily rewritten response contents already have significantly different expression forms (without changing the original meaning of the response content), that is, the BERT semantic generation sub-model has generated preliminarily rewritten response contents that meet each first intensity, and selection can be made from them.

[0061] The present invention uses a generative adversarial network to evaluate and analyze the above-mentioned "gap" situation. Specifically, it calculates the "difficulty" shown during the "generation" and "adversarial" processes of the generative adversarial network for the initially rewritten answer content input, and uses this "difficulty" to characterize the recognizability of the core semantics of the initially rewritten answer content. Subsequently, by comparing the gaps between the recognizabilities of the core semantics of each initially rewritten answer content, it indirectly analyzes whether the "gap" between multiple initially rewritten answer contents generated by the BERT semantic generation sub-model is large enough.

[0062] As an example, as Figure 2 shown, the generative adversarial network includes a generative model and a discriminative model. Then, the generative adversarial network performs generative adversarial processing on each of the initially rewritten answer contents to obtain a generative adversarial difficulty, including:

[0063] After receiving the initially rewritten answer content, the generative model performs secondary packaging on the initially rewritten answer content according to a preset packaging strategy; the packaging strategy is dynamically generated based on the attention mechanism;

[0064] The discriminative model extracts features from the secondarily packaged initially rewritten answer content and performs similarity analysis with the feature vector of the original initially rewritten answer content input. If the obtained similarity is higher than the similarity threshold, it records that the discrimination of the secondarily packaged initially rewritten answer content is successful; otherwise, it records discrimination failure.

[0065] Based on the above records, the discrimination accuracy rate is calculated, and based on the discrimination accuracy rate, the generative adversarial difficulty is calculated.

[0066] The present invention uses the secondary packaging of the initially rewritten answer content by the generative model in the generative adversarial network and the discrimination of the secondarily packaged initially rewritten answer content by the discriminative model, and obtains the generative adversarial difficulty by statistically analyzing the discrimination accuracy rate in this "adversarial" process. Obviously, the higher the discrimination accuracy rate, the lower the generative adversarial difficulty.

[0067] After receiving the initially rewritten answer content, the generative model invokes a packaging strategy dynamically generated based on the attention mechanism. The attention mechanism performs in-depth semantic analysis on the answer content, identifies key semantic components and sentence structures. For example, for answers to questions about work pressure in the nurse professional identity questionnaire, it focuses on key semantics such as "overtime hours" and "number of patients".

[0068] Then, from a pre-built vocabulary library containing function words, nursing scenario-related vocabulary, synonyms, etc., according to the probabilities set for different rewriting intensities (the insertion probability is 10%-20% for low rewriting intensity and 40%-60% for high rewriting intensity, where the rewriting intensity refers to the intensity recorded when the BERT semantic generation sub-model rewrites the answer content into the rewritten answer content), insert the vocabulary at appropriate grammatical positions. For example, rewrite "The pressure of nursing work is high" into "In the ward with a large number of patients, the pressure of daily nursing work is really not small". Thus, generate a more diverse set of rewritten versions, making the rewritten content more complex and variable in terms of semantic level and language expression, providing diverse discrimination samples for the discrimination model, and enhancing the comprehensiveness of the evaluation of the logical rationality of the rewritten content by the generative adversarial network.

[0069] The discrimination model determines whether the content after secondary packaging is essentially the same as the original preliminary rewritten answer content. Its discrimination accuracy can be used to characterize the recognizability of the core semantics of the preliminary rewritten answer content, and then indirectly analyze whether the "gap" between multiple preliminary rewritten answer contents generated by the BERT semantic generation sub-model is large enough.

[0070] Specifically, the discrimination model uses a deep learning algorithm to extract features from the preliminary rewritten answer content after secondary packaging and transform it into a feature vector. At the same time, extract the feature vector of the input preliminary rewritten answer content before secondary packaging, for example, conduct a similarity analysis by calculating the cosine similarity between the two. For example, when analyzing the answers related to nurses' professional identity, compare the similarity degree of semantic features such as "professional pride" and "professional value sense" in the content before and after secondary packaging. If the similarity is higher than a pre-set similarity threshold (such as 0.8), it is determined that the content after secondary packaging is similar to the original content, and it is recorded as a successful discrimination; otherwise, it is recorded as a failed discrimination.

[0071] Based on the discrimination records of the discrimination model for multiple times (such as 10 rounds, 20 rounds), calculate the discrimination accuracy, that is, the proportion of the number of successful discriminations to the total number of discriminations. Then, according to the formula such as generative adversarial difficulty = 1 - discrimination accuracy, calculate the generative adversarial difficulty.

[0072] It can be understood that the higher the generative adversarial difficulty, the more difficult it is for the content after being packaged by the generative model to be recognized by the discrimination model, the more complex and implicit the preliminary rewritten answer content is in terms of logic and semantic expression, and the lower the recognizability; on the contrary, it indicates that the rewritten content is relatively simple and direct, and the recognizability is high.

[0073] As an example, the method further includes: outputting the error questionnaire for personnel to verify, and then clearing the error questionnaire after verification and confirmation.

[0074] To avoid accidental deletion, the identified incorrect questionnaires can also be output to relevant personnel for verification. If the relevant personnel also confirm that they are incorrect questionnaires, then they are cleared.

[0075] As Figure 3 shown, an embodiment of the present invention also provides a data processing system 100 for a professional identity questionnaire. The system includes a receiving unit 101, an identifying unit 102, and a clearing unit 103;

[0076] The receiving unit 101 is configured to receive a number of questionnaire result data from an online questionnaire system;

[0077] The identifying unit 102 is configured to identify each questionnaire question in each of the questionnaire result data, analyze the first logical conflict score between each questionnaire question and other questions one by one using a logical analysis model, and screen out a number of target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire questions include questions and answer contents;

[0078] And, use a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents, and analyze the second logical conflict score of the target questionnaire question with the rewritten answer content replaced one by one using a logical analysis model;

[0079] The clearing unit 103 is configured to determine that the questionnaire result data to which a target questionnaire question belongs is an incorrect questionnaire and clear it if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than a preset threshold.

[0080] As an example, the identifying unit 102 is configured to:

[0081] Determine the number of expression rewrites according to the first logical conflict score, and formulate multiple rewrite intensities corresponding to the number of expression rewrites; the intensity difference between adjacent rewrite intensities is fixed;

[0082] Respectively using each of the rewrite intensities as a constraint condition, use a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents.

[0083] As an example, the identifying unit 102 is configured to:

[0084] Call a pre-constructed combined algorithm model that fuses a BERT semantic generation submodel and a generative adversarial network, and convert each of the rewrite intensities into a first intensity and a second intensity using different normalization functions;

[0085] The BERT semantic generation sub-model takes each of the first intensities as a constraint condition, rewrites the response content in a single target questionnaire question, and outputs multiple preliminary rewritten response contents greater than the number of the rewritten expressions;

[0086] The generative adversarial network performs generative adversarial processing on each of the preliminary rewritten response contents to obtain a generative adversarial difficulty, and calculates whether the difference between the sorted generative adversarial difficulties is higher than the second intensity;

[0087] If so, screen out the preliminary rewritten response content corresponding to the number of the rewritten expressions as the target rewritten response content; if not, control the BERT semantic generation sub-model to perform the expression rewriting again.

[0088] As an example, the generative adversarial network includes a generative model and a discriminative model. Then, the recognition unit 102 is used for:

[0089] After receiving the preliminary rewritten response content, the generative model performs secondary packaging on the preliminary rewritten response content according to a preset packaging strategy; the packaging strategy is dynamically generated based on the attention mechanism;

[0090] The discriminative model extracts features from the preliminarily rewritten response content after secondary packaging, and performs similarity analysis with the feature vector of the original preliminarily rewritten response content input. If the obtained similarity is higher than the similarity threshold, record that the discrimination of the preliminarily rewritten response content after secondary packaging is successful, otherwise record the discrimination as failed;

[0091] Calculate the discrimination accuracy rate based on the above records, and calculate the generative adversarial difficulty based on the discrimination accuracy rate.

[0092] As an example, the clearing unit 103 is further used for:

[0093] Output the incorrect questionnaire for personnel to verify, and then perform clearing processing on the incorrect questionnaire after verification and confirmation.

[0094] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the computer program is executed by the processor, the method described in any one of the foregoing items is implemented.

[0095] An embodiment of the present invention further provides a computer storage medium, which stores a computer program that can be executed by a processor to implement the method described in any one of the foregoing items.

[0096] An embodiment of the present invention also provides a computer program product, which includes a computer program that can be executed by a processor to implement the method described in any one of the foregoing items.

[0097] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0098] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data processing method for a professional identity questionnaire, characterized in that: It includes the following method steps: Receive a number of questionnaire result data from an online questionnaire system; Use a logical analysis model to analyze the first logical conflict scores between each questionnaire question and other questions one by one for the questionnaire questions, and screen out a number of target questionnaire questions whose first logical conflict scores are higher than a preset threshold; the questionnaire questions include questions and answer contents; Use a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents, and use a logical analysis model to analyze the second logical conflict scores of the target questionnaire questions after replacing the rewritten answer contents one by one; wherein, the using a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents includes: determining the number of expression rewrites according to the first logical conflict score, and formulating multiple rewrite intensities corresponding to the number of expression rewrites; wherein, the intensity difference between adjacent rewrite intensities is fixed; respectively using each rewrite intensity as a constraint condition, use a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents; wherein, the respectively using each rewrite intensity as a constraint condition, use a semantic component to rewrite the answer content in a single target questionnaire question to obtain a number of rewritten answer contents includes: calling a pre-constructed combined algorithm model integrating a BERT semantic generation sub-model and a generative adversarial network, converting each rewrite intensity into each first intensity and second intensity by using different normalization functions; the BERT semantic generation sub-model takes each first intensity as a constraint condition, rewrites the answer content in a single target questionnaire question, and outputs multiple preliminary rewritten answer contents greater than the number of expression rewrites; the generative adversarial network performs generative adversarial processing on each preliminary rewritten answer content to obtain a generative adversarial difficulty, and calculates whether the difference between the sorted generative adversarial difficulties is higher than the second intensity; if so, screen out the preliminary rewritten answer contents corresponding to the number of expression rewrites as target rewritten answer contents; if not, control the BERT semantic generation sub-model to perform expression rewriting again; If the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than the preset threshold, it is determined that the occupational identity questionnaire to which the target questionnaire question belongs is an incorrect questionnaire and it is cleared.

2. The data processing method of a professional identity questionnaire according to claim 1, wherein: The generative adversarial network includes a generative model and a discriminant model, then the generative adversarial network performs generative adversarial processing on each preliminary rewritten answer content to obtain a generative adversarial difficulty, including: After receiving the preliminary rewritten answer content, the generative model performs secondary packaging on the preliminary rewritten answer content according to a preset packaging strategy; the packaging strategy is dynamically generated based on an attention mechanism; The discriminant model extracts features from the secondary-packaged preliminary rewritten answer content, performs similarity analysis with the feature vector of the input original preliminary rewritten answer content, if the obtained similarity is higher than the similarity threshold, record that the discrimination of the secondary-packaged preliminary rewritten answer content is successful, otherwise record that the discrimination fails; Calculate the discrimination accuracy based on the above records, and calculate the generation adversarial difficulty based on the discrimination accuracy.

3. The data processing method of a professional identity questionnaire according to claim 1, wherein: The method further includes: outputting the error questionnaire for personnel verification, and performing a clearing process on the error questionnaire after verification and confirmation.

4. A data processing system for a professional identity questionnaire, characterized in that, The system includes a receiving unit, an identifying unit, and a clearing unit; The receiving unit is configured to receive a plurality of questionnaire result data from an online questionnaire system; The identifying unit is configured to identify each questionnaire question in each of the questionnaire result data, analyze the first logical conflict score between each questionnaire question and other questions one by one using a logical analysis model, and screen out a plurality of target questionnaire questions whose first logical conflict score is higher than a preset threshold; the questionnaire question includes a question and an answer content; In addition, use a semantic component to rewrite the expression of the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents, and use a logical analysis model to analyze the second logical conflict score of the target questionnaire question after replacing the rewritten answer content one by one; wherein, the use of a semantic component to rewrite the expression of the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents includes: determining the number of expression rewrites according to the first logical conflict score, and formulating a plurality of rewrite intensities corresponding to the number of expression rewrites; wherein, the intensity difference between adjacent rewrite intensities is fixed; respectively using each of the rewrite intensities as a constraint condition, use a semantic component to rewrite the expression of the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents; wherein, the respectively using each of the rewrite intensities as a constraint condition, use a semantic component to rewrite the expression of the answer content in a single target questionnaire question to obtain a plurality of rewritten answer contents includes: calling a pre-constructed combined algorithm model that integrates a BERT semantic generation sub-model and a generative adversarial network, converting each of the rewrite intensities into a first intensity and a second intensity using different normalization functions; the BERT semantic generation sub-model takes each of the first intensities as a constraint condition, rewrites the expression of the answer content in a single target questionnaire question, and outputs a plurality of preliminary rewritten answer contents greater than the number of expression rewrites; the generative adversarial network performs a generative adversarial process on each of the preliminary rewritten answer contents, obtains the generation adversarial difficulty, and calculates whether the difference between the sorted generation adversarial difficulties is higher than the second intensity; if so, screen out the preliminary rewritten answer content corresponding to the number of expression rewrites as the target rewritten answer content; if not, control the BERT semantic generation sub-model to perform expression rewriting again; The clearing unit is configured to determine that the questionnaire result data to which a target questionnaire question belongs is an error questionnaire and clear it if the average value of the second logical conflict scores corresponding to any target questionnaire question is higher than a preset threshold.

5. An electronic device, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and the computer program, when executed by the processor, implements the method according to any one of claims 1-3.

6. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method according to any one of claims 1-3.

7. A computer program product, characterized in that: The computer program product includes a computer program that can be executed by a processor to implement the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Virtual chatting method and system relieving psychological pressure of adolescents

    CN105206284A

  • Data processing method, device and storage medium

    EP4297039A1