A method and system for preventing cheating in NPS surveys
Patent Information
- Application Number
- CN202610809064.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-01
AI Technical Summary
[0009]针对现有技术的不足,本申请提供一种NPS调研反作弊方法及其系统,其目的在于解决或改善现有技术中存在的访问控制技术对人工逻辑矛盾作弊完全无效、行为特征分析技术无法校验内容逻辑且误判率高、以及深度学习模型成本高昂、可解释性差且落地周期长等技术问题
[0058] 1. Achieved accurate identification of highly concealed logical contradictions in cheating: By constructing a "score-cause association rule base" and a "domain sentiment dictionary," and designing a "three-level logical consistency verification pipeline," it systematically achieved the detection of logical contradictions in both the "score-cause" and "score-text" dimensions for the first time in an NPS survey scenario. This breaks through the limitations of traditional access control and behavioral analysis technologies that cannot reach the content logic, and also overcomes the problems of insufficient targeting and blind spots in detection dimensions of complex AI models in this scenario.
Smart Images

Figure CN122674035A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer software and data security technology, and in particular to an NPS survey anti-fraud method and system. Background Technology
[0002] Net Promoter Score (NPS) surveys, as a core tool for measuring customer loyalty and product reputation, are widely used in online survey systems across finance, e-commerce, and SaaS industries, where high data credibility and interpretability are crucial. Its standard data structure includes two core data categories: quantitative rating questions (0-10 points) and structured reason selections plus open-text feedback. The correlation between ratings and user feedback (including selected reasons and open-text) is key to evaluating data validity.
[0003] In real-world survey scenarios, some users, either to obtain rewards or to answer arbitrarily, generate a significant proportion of false data. Typical cheating behaviors include giving high scores (to recommend) but providing detailed negative feedback, or giving low scores (to criticize) but choosing positive reasons. If such questionnaires with "contradictory ratings and content logic" are not effectively identified, they will directly pollute the data pool, mislead companies' judgments of customer loyalty, affect product iteration and service optimization decisions, and increase the manpower costs of subsequent data cleaning.
[0004] Currently, the technical means used in the industry to address this type of data quality issue mainly fall into the following three categories:
[0005] The first category is access control technology, such as IP restrictions, device fingerprinting, and time thresholds for answering questions. Its core objective is to identify and block "machine cheating" initiated by automated scripts. However, this type of technology is completely ineffective against cheating behavior that involves "manually filled out answers but with logical contradictions." This type of cheating cannot be blocked because its behavioral characteristics are consistent with normal answers. Furthermore, it relies too heavily on IP and device restrictions and is prone to mistakenly blocking multiple legitimate responses from the same user (such as employees of a company filling out surveys on different devices).
[0006] The second category is behavioral feature-based technology, which determines whether answers are "random" by monitoring behavioral data such as answer duration, mouse movements, and page interaction frequency. This type of technology can only identify abnormal behavior and cannot judge the logical rationality of the answer content itself. This technology cannot identify cheating such as "normal answer duration but contradictory content" (e.g., carefully filling out a "10 points + extremely poor product" questionnaire in 10 minutes), and it cannot distinguish between "concise and effective feedback" and "random answers," easily misjudging short, effective text feedback as invalid.
[0007] The third category consists of deep learning / NLP (Natural Language Processing) technologies that rely on pre-trained large-scale sentiment analysis models (such as BERT and RoBERTa) to verify the authenticity of text content. While this type of technology can capture complex semantic features, its deployment and maintenance face significant bottlenecks: First, model training requires large-scale labeled corpora and high computational costs, posing a barrier to small and medium-sized survey platforms; second, the model output is a "black box" conclusion, lacking interpretability and failing to meet data compliance and auditing requirements; finally, the model takes more than four weeks from training and parameter tuning to deployment, making it unable to quickly respond to the changing needs of different questionnaire scenarios.
[0008] In summary, existing technologies have significant limitations: they may only defend against machine attacks but cannot identify logical contradictions, or they may only analyze behavioral characteristics without considering the content itself, or while they can process content, they are costly, poorly interpretable, and inflexible. Therefore, the industry urgently needs a solution that can accurately, efficiently, and interpretably identify logically contradictory cheating in NPS surveys, and is easy to deploy. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this application provides an NPS survey anti-fraud method and system, which aims to solve or improve the technical problems existing in the prior art, such as access control technology being completely ineffective against human logical contradiction fraud, behavioral feature analysis technology being unable to verify content logic and having a high false judgment rate, and deep learning models being expensive, having poor interpretability, and having a long implementation cycle.
[0010] To achieve the above objectives, firstly, this application provides an NPS survey anti-fraud method, comprising the following steps:
[0011] S1. Obtain users' NPS survey responses;
[0012] S2. Based on the pre-built score-cause association rule base, perform contradiction detection on the score values contained in the answer sheet and the selected cause options, and output the cause contradiction score;
[0013] S3. Based on a pre-built domain sentiment dictionary, a deterministic text analysis algorithm is used to calculate the sentiment tendency of the score values and open text contained in the answer sheet, and the calculation results are compared with the expected sentiment direction corresponding to the score values to output the text contradiction score.
[0014] S4. Based on the cause-of-fact contradiction score obtained in step S2 and the text-of-fact contradiction score obtained in step S3, determine the suspiciousness score of the answer sheet.
[0015] S5. Determine the validity of the questionnaire based on the suspiciousness score.
[0016] Optionally, the step of pre-building the score-cause association rule base in step S2 includes:
[0017] Establish a logical mapping relationship between NPS score ranges and cause options. The rules include score ranges, required items, prohibited items, option strength levels, and rule priorities.
[0018] Optionally, the logic of intensity grading specifically includes: dividing the cause options into three levels—strong, medium, and weak—based on emotional intensity, with different levels corresponding to different weights, in order to quantify the degree of conflict.
[0019] Optionally, a cold start optimization step may be included before step S2:
[0020] Initial questionnaire samples are collected from the target domain. The co-occurrence frequency of reason options and ratings is analyzed by clustering to generate high-frequency co-occurrence combination suggestions for domain experts to screen and form an initial rule base.
[0021] Optionally, the contradiction detection in step S2 further includes:
[0022] Retrieve the score of the answer sheet and the set of selected reason options;
[0023] Match the corresponding rule entry in the score-reason association rule base based on the score value;
[0024] Verify whether the set of reason options meets the minimum satisfaction condition of the matched rule entry, and whether it contains prohibited items of the matched rule entry;
[0025] Based on the strength of the option that violates the rule and the priority of the rule, the cause of contradiction is quantitatively calculated.
[0026] Optionally, the step of pre-building the domain sentiment lexicon in step S3 includes:
[0027] Manually input single-digit core sentiment keywords for the relevant domain;
[0028] Statistical filtering algorithms were used to select candidate sentiment words from historical valid response corpora;
[0029] Candidate sentiment words undergo final manual review and are then added to the database;
[0030] The domain sentiment dictionary includes a positive lexicon, a negative lexicon, a degree adverb lexicon, and a negation lexicon.
[0031] Optionally, step S3, which involves calculating the sentiment of the open text using a deterministic text analysis algorithm, includes:
[0032] The open text is segmented and divided to distinguish between core sentences and auxiliary sentences, and the core sentences are assigned a higher weight than the auxiliary sentences.
[0033] The frequency of positive and negative words in open texts was statistically analyzed, and the weight coefficients in the degree adverbial lexicon were used for adjustment.
[0034] Identify negative words and, based on the rules of the negative word effective window, reverse the sentiment tendency only for sentiment words within the negative word effective window;
[0035] Based on the adjusted and reversed weighted counts of positive and negative words, the difference between the two is calculated as the sentiment value of the open text.
[0036] Optionally, step S3, which compares the calculated result with the expected sentiment direction corresponding to the score value and outputs a text contradiction score, further includes:
[0037] The calculated text sentiment tendency value is compared with the expected sentiment direction corresponding to the score value;
[0038] If the sentiment value of the text is completely opposite to the expected sentiment direction, it is judged as a logical contradiction;
[0039] Based on the determination of logical contradictions, the open text is matched with predefined contradictory sentence patterns. If the match is successful, the quantitative level corresponding to the logical contradiction is increased.
[0040] Based on the determination of logical contradictions and the quantification level, the text contradiction score is output.
[0041] Optionally, the number of domain sentiment words in the single digit ranges from 3 to 5.
[0042] Optionally, following step S3, the method further includes:
[0043] Perform duplicate and spam detection on open text and generate duplicate and spam tags accordingly;
[0044] The duplication detection is achieved by generating a local sensitive hash value for the open text and comparing its similarity with the hash values of historical responses.
[0045] Spam detection is achieved by matching open text with a pre-defined spam rule base, which includes rules for duplicate characters, meaningless characters, and irrelevant information.
[0046] Optionally, determining the suspiciousness score of the questionnaire in step S4 includes: weighting and summing the cause contradiction score, text contradiction score, duplicate mark and spam mark to obtain the suspiciousness score.
[0047] Optionally, step S5, determining the validity of the questionnaire, further includes performing a grading process:
[0048] If the doubt score is below the first threshold, it is considered a normal answer and submission is allowed;
[0049] If the suspicion score is between the first threshold and the second threshold, it is judged as a medium-risk answer, and a confirmation prompt will pop up for the user;
[0050] If the suspiciousness score is higher than the second threshold, it is judged as a high-risk answer, and an interception prompt will be displayed to the user and the submission will be blocked.
[0051] Secondly, this application provides an NPS survey anti-fraud system, including:
[0052] The data acquisition module is used to obtain users' NPS survey responses;
[0053] The verification module, based on the score-cause association rule base, performs contradiction detection on the score value and the selected cause option, and outputs the cause contradiction score;
[0054] Based on a domain sentiment dictionary, a deterministic text analysis algorithm is used to calculate and compare the sentiment of open texts, and output a text contradiction score.
[0055] The decision engine module is used to determine the suspiciousness score of the questionnaire based on the cause-of-fact contradiction score and the text-of-fact contradiction score, and to determine the validity of the questionnaire based on the suspiciousness score.
[0056] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed by a computer processor, can implement any of the methods described in the first aspect above.
[0057] This application has at least the following beneficial effects:
[0058] 1. Achieved accurate identification of highly concealed logical contradictions in cheating: By constructing a "score-cause association rule base" and a "domain sentiment dictionary," and designing a "three-level logical consistency verification pipeline," it systematically achieved the detection of logical contradictions in both the "score-cause" and "score-text" dimensions for the first time in an NPS survey scenario. This breaks through the limitations of traditional access control and behavioral analysis technologies that cannot reach the content logic, and also overcomes the problems of insufficient targeting and blind spots in detection dimensions of complex AI models in this scenario.
[0059] 2. A lightweight, low-cost deterministic technical path has been constructed, significantly reducing the barriers to deployment and use: the entire process is based on deterministic algorithms such as rule matching, dictionary comparison, and statistical calculations, eliminating the reliance on large-scale labeled corpora, high-performance computing power, and complex deep learning models. This eliminates the need for high training costs and lengthy deployment cycles, making it particularly suitable for rapid application by small and medium-sized research platforms or teams, and solving the core pain points of high cost and difficult implementation of existing AI solutions.
[0060] 3. It provides interpretable and traceable judgment criteria throughout the entire process, meeting data compliance and auditing requirements: All verification steps (such as contradiction judgment and suspiciousness scoring) are based on pre-defined rules, dictionaries, and calculation formulas. Any judgment result can be traced back to a specific rule violation, sentiment word matching, or text pattern. This solves the problem of poor interpretability caused by the "black box" decision-making of deep learning models, and provides reliable and transparent audit logs for data quality control.
[0061] 4. Excellent scenario adaptability and flexible scalability: Through a rule base and dictionary built using a hybrid "human + machine" approach, and dynamically configurable verification weights and tiered thresholds, this solution can quickly adapt to business corpora and evaluation standards in different industries such as finance, e-commerce, and SaaS, enabling rapid cold start in new business areas. The system's modular design also facilitates enabling or adjusting different levels of verification functions as needed.
[0062] 5. While ensuring effective anti-fraud measures, a good user experience was also taken into account: the decision engine's tiered handling strategy based on the degree of suspicion avoided a "one-size-fits-all" approach to blocking. Normal users were allowed to proceed without any notice, while potentially contradictory or low-quality responses were flagged or blocked in a tiered manner, guiding users to provide effective feedback. This achieved an effective balance between improving data quality and maintaining a positive survey experience.
[0063] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of the present invention are given below. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.
[0065] Figure 1 A schematic diagram of the overall process of the NPS survey anti-fraud method provided in the embodiments of this application;
[0066] Figure 2 This is a schematic diagram of the optional execution path for the two-dimensional logic verification in the embodiments of this application;
[0067] Figure 3 This is a schematic diagram of the scoring-open text consistency verification process based on a deterministic text analysis algorithm in an embodiment of this application. Detailed Implementation
[0068] This application will now be described more fully below with reference to the accompanying drawings, in which various embodiments are illustrated. However, this application may be implemented in many different ways and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be exhaustive and complete, and will fully convey the scope of this application to those skilled in the art.
[0069] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting. Unless the context clearly indicates otherwise, the terms “a,” “an,” “the,” and “at least one” as used herein are not intended to indicate a limitation on quantity, but are intended to include both singular and plural forms. “At least one” should not be construed as limited to the quantity “a.” “Or” means “and / or.” The term “and / or” includes any and all combinations of one or more of the associated listed items.
[0070] Unless otherwise specified, all terms used herein, including technical and scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art. Terms defined in commonly used dictionaries shall be interpreted as having the same meaning as in the relevant technical context, and shall not be construed as having a formal meaning in an idealized or overly formal sense unless expressly defined in the specification.
[0071] The terms “comprising” or “including” indicate a nature, quantity, step, operation, or combination thereof, but do not exclude other natures, quantities, steps, operations, or combinations thereof.
[0072] It should be noted that the step numbers (e.g., S1, S2, etc.) mentioned herein are only used to quickly refer to the corresponding operation process and are not intended to limit these steps to being performed strictly in the order of the numbers. Unless it is explicitly stated that a certain step must be performed before or after another step, or the order can be uniquely determined by a person skilled in the art based on technical logic, the steps of this invention may be interchanged, combined, or performed simultaneously in actual implementation.
[0073] The embodiments of the technical solution of this application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of this application more clearly, and are therefore only examples and should not be used to limit the scope of protection of this application.
[0074] like Figure 1 and Figure 2 As shown, the first aspect of this application provides an NPS survey anti-fraud method, including the following steps:
[0075] First, the NPS survey questionnaires submitted by users are obtained. Then, based on a pre-built score-cause association rule base, contradiction detection is performed on the score values in the questionnaire and the reasons selected by the user, outputting a cause-contradiction score. Based on a pre-built domain sentiment dictionary, a deterministic text analysis algorithm is used to calculate the sentiment tendency of the open text in the questionnaire, and the calculation result is compared with the expected sentiment direction corresponding to the score value, outputting a text-contradiction score. Next, based on the obtained cause-contradiction score and text-contradiction score, the suspiciousness score of the questionnaire is determined. Finally, based on this suspiciousness score, the validity of the questionnaire is judged.
[0076] This application first constructs a "score-cause association rule base" and a "domain sentiment dictionary" adapted to NPS scenarios. These are generated through a manually controllable hybrid strategy to balance "recognition accuracy" and "operational efficiency," providing a core basis for subsequent verification.
[0077] The construction of the score-cause association rule base first establishes the logical mapping relationship between score ranges and cause options. Rules can include five core elements: score range, mandatory items, prohibited items, option strength grading, and rule priority. Specifically, the system divides NPS scores into different ranges, such as 0-6 for detractors, 7-8 for neutrals, and 9-10 for recommenders. "Mandatory items" refer to the expected sentiment and specific items among the cause options selected in the questionnaire within the score range. For example, the detractor range might require at least one "strongly negative" or two "moderately negative" cause options. "Prohibited items" refer to contradictory options not allowed among the cause options selected in the questionnaire within the score range. Option strength grading divides cause options into three levels based on sentiment intensity: "strong," "moderate," and "weak," with different levels corresponding to different weights to quantify the degree of contradiction. "Rule priority" is used for adjudication in case of rule conflicts. In addition, it can also include "minimum satisfaction conditions" for determining whether mandatory items are met, such as "one strongly negative or two moderately negative."
[0078] To accelerate initial setup, before performing conflict detection on the score values and selected cause options in the questionnaire based on a pre-built score-cause association rule base, the system can also include a cold start optimization process. First, it collects historical valid questionnaire samples (a certain size of valid samples) from the target domain. Using methods such as cluster analysis, it statistically analyzes the co-occurrence frequency of different score ranges and cause options, automatically generating high-frequency co-occurrence combination suggestions. Subsequently, domain experts can quickly filter, adjust, and confirm these suggestions to generate the initial rule base, significantly shortening the cold start cycle. This rule base is typically stored in a configurable structured format such as JSON for easy dynamic adjustment and maintenance. An example rule is shown below:
[0079] {
[0080] "score_range": [0,6], / / Detractor rating range
[0081] "expected_rule": {
[0082] "must_include": [{"level": "strongly negative", "weight": "high", "items": ["extremely poor", "never need again", "serious fault"]},
[0083] {"level": "Medium-Negative", "weight": "Medium", "items": ["Difficult to use", "Price trap", "Slow response"]}],
[0084] "forbid_include": [{"level": "strongly positive", "weight": "high", "items": ["highly recommended", "irreplaceable"]}],
[0085] "min_condition": "1 strong negative arrow OR 2 medium negative arrows"
[0086] },
[0087] "priority": "high" / / Prioritizes handling conflicts
[0088] }
[0089] Specifically, taking the rule for "detractors" (rating range 0-6 points) as an example, its design follows the basic business logic of NPS: users who give low scores should give feedback that leans towards negative evaluations, and are extremely unlikely to include strong positive evaluations. Therefore, the expected_rule in this rule is configured as follows:
[0090] The `must_include` field sets "strongly negative" reason options (such as "extremely poor" or "never again"). This does not mean that all low-scoring responses must select these options, but rather that they serve as an expected anchor point: if there is a complete lack of negative reasons, the reasonableness of the response will be reduced, and the system will generate a corresponding contradictory score.
[0091] The `forbid_include` option explicitly lists the reasons for "strongly positive" reasons (such as "highly recommended" or "irreplaceable"). Any low-scoring answer that selects these options is considered to have triggered a fundamental logical conflict, resulting in a higher inconsistency score.
[0092] The `min_condition` is set to "1 strongly negative reason OR 2 moderately negative reasons". This condition specifies the minimum emotional intensity threshold that the user's choice from the `must_include` range must reach. It allows for some flexibility (e.g., the user can choose 1 most negative reason or 2 moderately negative reasons), but ensures that the negative intensity of the feedback matches the low score. If the user's choice does not meet this minimum condition, the system will determine that its logical consistency is insufficient.
[0093] This rule design, which is based on the inverse correlation between the rating range and the emotional intensity of the cause option (low rating is negative, high rating is positive) and the positive prohibition (low rating prohibits strong positive, high rating prohibits strong negative), is the key to this application's ability to accurately capture the core cheating pattern of "conflict between rating and cause logic".
[0094] The domain sentiment dictionary employs a hybrid strategy of "human customization, machine initial selection, and human final review" to ensure its domain adaptability and accuracy. First, seed words are initialized: an administrator inputs a single-digit number of core domain sentiment words as seeds, accurately representing the core evaluation dimensions of the business. In this embodiment, the single-digit number of domain sentiment words is preferably 3 to 5. Next, candidate words are mined, using statistical filtering algorithms to process historical valid questionnaire data, selecting a list of candidate words with clear sentiment tendencies. This process does not rely on model training, but only on word frequency and co-occurrence relationships. Finally, candidate words are sorted by relevance, and the administrator selects valid words for inclusion in the database. The system records the review log to ensure human controllability. This domain sentiment dictionary includes not only positive and negative word libraries, but also auxiliary word library rules, namely a degree sub-word library and a negation word library. The degree sub-word library is divided into levels according to the degree of reinforcement, corresponding to different weight coefficients, used to adjust the influence of sentiment words; the negation word library limits the effective influence window, only having a reversal effect on a specified number of subsequent sentiment words, avoiding excessive reversal leading to misjudgment. This provides a fine-grained and reliable computational foundation for subsequent text analysis.
[0095] After completing the construction of the score-cause association rule base and the domain sentiment dictionary, the system has the foundation to perform logical consistency checks. Next, this application processes user-submitted questionnaires through a logical consistency check pipeline.
[0096] The first level is the contradiction detection between ratings and selected reasons (structured data validation): Based on a pre-built score-reason association rule base, the system determines the logical consistency between the user's rating and the selected reasons through rule matching. The technical implementation follows these steps: First, the system extracts the rating values and the set of selected reasons from the user's submitted NPS survey questionnaire. Based on the extracted rating values, the system matches the corresponding rule entries in the score-reason association rule base; for example, 0-6 points match the detractor rule, 7-8 points match the neutral rule, and 9-10 points match the recommender rule. Then, the system proceeds to the validation step, comparing the user's set of reasons with the matched rule entries: on the one hand, it verifies whether the set meets the "minimum satisfaction condition" defined by the rule; on the other hand, it verifies whether the set contains reasons that are explicitly "prohibited" by the rule and have a completely opposite emotional direction, such as whether "strongly recommended" or other strongly positive reasons appear in detractor responses. Finally, based on the above verification results, the degree of contradiction is quantified. If the set of cause options violates the "minimum satisfaction condition" or contains "prohibited items," then the cause contradiction score is calculated based on the emotional intensity level (strong, medium, weak) of the violated option and the priority of the rule it belongs to. The more severe the violation (e.g., violating a high-intensity prohibition) or the higher the priority of the violated rule, the higher the calculated cause contradiction score, thus accurately and quantitatively representing the degree of logical inconsistency between the score and the cause.
[0097] The second level is the consistency verification between the rating and the open text (unstructured data verification): Based on a pre-built domain sentiment dictionary, the sentiment tendency of the open text is calculated using deterministic text analysis algorithms, and this calculation result is compared with the expected sentiment direction corresponding to the user rating to determine the logical consistency between the two. For example... Figure 3As shown, the technical implementation steps are as follows: First, the open text is preprocessed, including word segmentation, stop word removal, and sentence division. Specifically, the first and last sentences are identified as core sentences, while the middle parts are designated as auxiliary sentences. Based on this, the system assigns higher weights to sentiment words in the core sentences than to those in the auxiliary sentences. The sentiment word counts in the core sentences are amplified by their weights, while those in the auxiliary sentences are calculated using their base weights, because the core sentences better reflect the user's true attitude. Subsequently, the system enters the sentiment tendency calculation process (deterministic algorithm): Based on a domain sentiment lexicon, the frequency of positive and negative words in the text is counted; using the weight coefficients corresponding to each adverb in the degree adverb library, the counts of sentiment words modified by them are weighted and adjusted; negative words in the text are identified, and strictly following the preset negative word effective window rules, sentiment tendency reversal is performed only on sentiment words within a specified window after the negative word, while words outside the window are unaffected. After completing the position weight allocation, degree adverb weighting, and negation word reversal processing, the system obtains the weighted total count of positive words and the weighted total count of negative words respectively. By calculating the difference between the two, a definite value can be obtained as the sentiment tendency value of the open text, quantifying the overall sentiment direction of the text.
[0098] This step focuses on detecting "obvious logical contradictions," without pursuing sophisticated sentiment analysis. Complex expressions (such as irony and double negation) can be categorized for subsequent review due to their near-neutral sentiment values. After obtaining the sentiment value, the system executes the contradiction judgment logic: comparing the calculated sentiment value with the expected sentiment direction corresponding to the score (e.g., low score expected to be negative, high score expected to be positive); if the two are completely opposite, a serious contradiction is judged. During this process, the system can also match predefined contradictory sentence patterns (e.g., "Although [positive], but [negative]", "[negative], yet given a high score"). If the text matches such patterns, the contradiction judgment level is strengthened. Finally, the system quantifies and calculates the text contradiction score based on the severity of the contradiction judgment and whether it matches contradictory sentence patterns.
[0099] To make the above solution clearer, a specific and simplified implementation example is provided below. Assume that an e-commerce platform's NPS survey system has pre-set basic rules and a dictionary. A core rule is defined as follows: for ratings between 0-6 points (detractors), the `must_include` must include moderately negative reason options (e.g., "slow logistics"), the `min_condition` is set to "at least one moderately negative reason", and the `forbid_include` prohibits strongly positive reason options (e.g., "great value for money"). The domain sentiment dictionary includes positive words (e.g., "good", "fast"), negative words (e.g., "bad", "slow"), and corresponding rules for handling negative words.
[0100] There is an NPS questionnaire submitted by a user, with the following content: Rating = 2 points, Selected Reason option = [“Great Value for Money”], Open Text = “The product quality is very poor, I will not recommend it.” The system's processing flow for this questionnaire is as follows: First, the data acquisition module acquires the questionnaire data. Then, the verification module performs a two-dimensional logical verification. In the first-level verification, the system matches the rating of 2 points to the detractor rule and finds that the user's selected “Great Value for Money” is a strongly positive item explicitly prohibited by the rule, which seriously violates the `forbid_include`. Therefore, the system quantifies and calculates a high reason contradiction score (e.g., 85 points). In the second-level verification, the system performs deterministic text analysis on the open text “The product quality is very poor, I will not recommend it.” Through dictionary matching, it identifies the negative word “poor”. Combined with the processing rules for the negative word “not,” it calculates an overall sentiment tendency value that is clearly negative (e.g., -3.5). This is completely consistent with the expected sentiment direction of the low score of 2 points, and no logical contradiction is found. Therefore, the system outputs a low text contradiction score (e.g., 5 points). The decision engine module then calculates a suspicion score of 45 points based on the obtained causal contradiction score (85 points) and textual contradiction score (5 points), according to a preset weighting configuration (e.g., 50% each). Assuming the system's preset lower limit for medium risk is 40 points, this score is judged as medium risk, and the questionnaire will be intercepted by the system or marked as highly suspicious, requiring further manual review. This example fully demonstrates the entire process of identifying the fundamental contradiction between the score and the causal options through two-dimensional logical verification and making a judgment accordingly.
[0101] Furthermore, after completing the consistency verification between the score and the open text, this application may include a third-level auxiliary verification step. Its core function is to combine locality-sensitive hashing (LSH) technology with rule matching to quickly filter out batches of duplicate cheating content and meaningless spam text from the open text. The specific implementation is as follows:
[0102] In duplicate detection, the system generates a local sensitive hash value for the open text and calculates the similarity between this hash value and the hash values of recently submitted historical answers. If the similarity is higher than a set threshold, the current text is determined to be highly duplicated with historical text, and a duplicate tag is generated accordingly.
[0103] In spam detection, the system matches open text against a pre-defined spam rule base. This rule base contains various pre-defined rule types, primarily including: repeating character rules, used to detect text consisting of consecutive repeating characters reaching a certain length; meaningless character rules, used to detect text composed of meaningless symbols or garbled characters; and irrelevant information rules, used to detect text containing advertising links, contact information, or other content unrelated to the survey. If the text matches any spam rule, a spam tag is generated accordingly.
[0104] The system ultimately outputs duplicate and garbage tags for subsequent decision calculations.
[0105] After completing the aforementioned three-level logical consistency verification pipeline and obtaining the cause-contradiction score, text-contradiction score, duplicate marker, and spam marker, this application uses a decision engine module to perform final judgment and processing on the responses. This module calculates a suspicion score based on the quantitative indicators output from the three-level verification, classifies responses according to their validity, and executes corresponding processing strategies, balancing "accurate identification" and "user experience." In this embodiment, the suspicion score can be a comprehensive suspicion score. Its specific implementation includes the following steps:
[0106] First, the system configures weights. The decision engine assigns configurable weight coefficients to each verification indicator. Based on the information contribution and scenario adaptability of each verification dimension, an objective weighting method is used to determine the general weights, with a total weight sum of 1. The core principle of weight configuration is: logical contradiction indicators (causal contradiction score, textual contradiction score) are assigned higher weights, while auxiliary verification indicators (duplicate marking, spam marking) are assigned appropriate weights, which can be dynamically adjusted according to business scenarios. The configuration method is: weight parameters are stored through a JOSN configuration file, allowing administrators to flexibly adjust them according to specific business needs without modifying the core code.
[0107] Next, the system calculates the overall suspicion score. The decision engine receives the output from the verification pipeline, namely the cause-contradiction score, text-contradiction score, duplicate marker, and spam marker, and obtains the suspicion score by weighted summation of the cause-contradiction score, text-contradiction score, duplicate marker, and spam marker. The specific formula is:
[0108] The overall suspiciousness score is calculated as follows: Cause contradiction score × W1 + Text contradiction score × W2 + Duplicate marker × W3 + Spam marker × W4. Here, W1, W2, W3, and W4 are the corresponding weight coefficients, and following the aforementioned principle, the weights of W1 and W2 are greater than those of W3 and W4. Duplicate markers and spam markers are binary variables (0 for normal, 1 for abnormal) and directly participate in the weighted calculation.
[0109] Finally, the system executes a tiered processing strategy. Based on the calculated comprehensive suspiciousness score, the decision engine divides the questionnaires into three risk levels according to two preset thresholds (the first threshold being low and the second threshold being high), and executes different processing strategies accordingly:
[0110] If the overall suspiciousness score is lower than the first threshold, it is judged as a normal answer, and the system allows the answer sheet to be submitted directly without any processing.
[0111] If the overall suspicion score is between the first and second thresholds, it is judged as a medium-risk answer. The system will pop up a prompt window to remind the user to confirm that their answer is relevant to the question and contains no meaningless content. The user can only submit after confirming.
[0112] If the overall suspiciousness score is higher than the second threshold, it is judged as a high-risk answer. The system will pop up an interception window for the user, prompting the user to fill in the answer according to the question stem and preventing the invalid submission.
[0113] The grading thresholds are determined based on the distribution characteristics of a large amount of actual survey data and combined with statistical principles, aiming to balance the pass rate of normal questionnaires and the detection rate of cheating questionnaires. The verification indicators, comprehensive suspiciousness scores, and final processing results of all questionnaires are anonymized and stored to form a complete audit log for subsequent rule calibration, effect analysis, and compliance traceability.
[0114] To enable those skilled in the art to better understand this application, the method of this application is described below through a complete embodiment. It should be noted that the numerical values, rules, and results in this example are merely illustrative and do not constitute any limitation on the scope of protection.
[0115] Scenario and Process Description: Assume an online education platform is conducting an NPS survey for its "Live Courses" product. The administrator has completed system initialization: First, a score-cause association rule library was constructed. For example, for the detractor rule in the score range [0,6], it is stipulated that it must contain at least one "strongly negative" or two "moderately negative" reasons, prohibiting the inclusion of strong positive reasons such as "strongly recommended," and setting a high priority. At the same time, a domain sentiment dictionary containing positive / negative word libraries, degree adverbial word libraries, and negation word libraries was constructed, and the effective window for negation words was set to the last two words. The decision engine weight configuration is as follows: cause contradiction score weight W1=0.4, text contradiction score weight W2=0.4, and duplicate and spam label weights are each 0.1; the grading thresholds are set as follows: low risk <30 points, medium risk 30-70 points, and high risk >70 points.
[0116] User A submitted a questionnaire with a score of 2. The selected reason was "["Great gains"], and the open text was "The teacher's lecture was excellent, and the content was extremely practical. I will definitely recommend it to my classmates!". The system processed the questionnaire as follows: After the data acquisition module obtained the questionnaire, the first-level verification matched the score 2 against the detractor rule. It found that the user's selected reason "Great gains" was a strongly positive item prohibited by the rule, seriously violating the prohibited inclusion item and not meeting the minimum condition. Therefore, a high reason contradiction score of 80 was calculated. The second-level verification performed deterministic text analysis on the open text: After word segmentation, the core sentence and auxiliary sentence were distinguished and assigned higher weights. The sentiment words "good," "practical," and "recommended," as well as the degree adverbs "very" and "extremely," were identified. After positional weighting and degree adverb adjustment, the sentiment tendency value was calculated to be significantly +6.75. This positive sentiment tendency is completely opposite to the negative sentiment expected by score 2. Therefore, a logical contradiction was determined, and a text contradiction score of 75 was output. The third-level verification performed local sensitive hash duplicate detection and spam rule matching on the text. No anomalies were found, so the duplicate mark was 0 and the spam mark was 0. The decision engine then calculates the overall suspiciousness score: 80*0.4 + 75*0.4 + 0*0.1 + 0*0.1 = 62. This score falls within the medium-risk range of 30-70. The system classifies this as a medium-risk answer and displays a pop-up window to the user, requiring them to confirm the consistency of their answer before submission. All verification indicators, calculated scores, and judgment results are recorded in the audit log.
[0117] The above embodiments are merely specific examples to more intuitively illustrate the technical solutions of this application. Those skilled in the art should understand that in practical applications, the scoring interval division, rule entries, sentiment dictionary content, weight parameters, threshold settings, etc., can all be adjusted and configured according to different business scenarios. Any solution that adopts the core idea of this application, "deterministic logic verification based on a pre-built rule base and dictionary," to achieve anti-fraud measures in NPS surveys falls within the protection scope of this application.
[0118] Secondly, this application also provides an NPS survey anti-fraud system. This system includes: a data acquisition module, a verification module, and a decision engine module.
[0119] The data acquisition module is used to implement the steps of obtaining user NPS survey responses in the aforementioned method.
[0120] The verification module, connected to the data acquisition module, is used to implement the aforementioned logical consistency verification. It is configured to perform the following processing:
[0121] Based on a pre-built score-cause association rule base, the system performs contradiction detection on the score value in the answer sheet and the selected cause option, and outputs the cause contradiction score.
[0122] Based on a pre-built domain sentiment dictionary, a deterministic text analysis algorithm is used to calculate the sentiment tendency of open text in the questionnaire, and the calculation results are compared with the expected sentiment direction corresponding to the score to output the text contradiction score.
[0123] The decision engine module, connected to the verification module, is used to implement the aforementioned decision output steps. It is configured to: determine the suspiciousness score of the questionnaire based on the cause-of-fact contradiction score and the text-of-fact contradiction score output by the verification module, and make a final judgment on the validity of the questionnaire based on the suspiciousness score.
[0124] The data acquisition module, verification module, and decision engine module in the above system are divided according to functional logic. In specific implementation, they can be embodied as software program units, hardware circuits, or a combination of both. For example, the system of this invention can be deployed entirely on a backend server or cloud platform; it can also adopt a distributed deployment approach, in which some data acquisition and preliminary verification functions are executed on user terminal devices (such as smartphones or personal computers), while the core verification and decision engine functions are executed on the server side.
[0125] Thirdly, this application also provides a computer-readable storage medium. The storage medium can be an internal storage unit of a terminal or server, such as a hard disk or memory, or it can be an external storage device, such as a plug-in hard disk, a smart memory card, or a flash memory card. The storage medium stores a computer program (i.e., instructions), which, when executed by a processor, can implement the NPS survey anti-fraud method as described in any of the foregoing embodiments of this application.
[0126] Specifically, when the program is executed by the processor, it can implement a method including the following steps: obtaining the user's NPS survey questionnaire; performing contradiction detection on the score value and the selected reason option based on a pre-built score-cause association rule base; performing consistency verification on the score value and the open text using a deterministic text analysis algorithm based on a pre-built domain sentiment dictionary; determining the suspiciousness score of the questionnaire based on the output contradiction score; and determining the validity of the questionnaire based on the suspiciousness score. The beneficial effects brought by the storage medium are the same as those in the corresponding method embodiment, and will not be repeated here.
[0127] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium.
[0128] This application provides an anti-fraud method and system for NPS survey scenarios. The method constructs a score-cause association rule base and a domain sentiment dictionary, and designs a three-level logical consistency verification pipeline to achieve deterministic and quantifiable detection of logical contradictions in the two dimensions of "rating-cause" and "rating-text". This method abandons the reliance on complex deep learning models and features low deployment cost, strong interpretability, and flexible scenario adaptability. It can effectively identify human-induced logical contradiction fraud, significantly improve the quality and credibility of NPS survey data, and provide a reliable data foundation for subsequent customer loyalty analysis and product decision-making.
[0129] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for preventing cheating in NPS surveys, characterized in that, Includes the following steps: S1. Obtain users' NPS survey responses; S2. Based on the pre-built score-cause association rule base, perform contradiction detection on the score values contained in the questionnaire and the selected cause options, and output the cause contradiction score; S3. Based on a pre-built domain sentiment dictionary, a deterministic text analysis algorithm is used to calculate the sentiment tendency of the score values and open text contained in the questionnaire, and the calculation results are compared with the expected sentiment direction corresponding to the score values to output a text contradiction score. S4. Based on the cause contradiction score obtained in step S2 and the text contradiction score obtained in step S3, determine the suspiciousness score of the questionnaire; S5. Based on the doubt score, determine the validity of the questionnaire.
2. The method according to claim 1, characterized in that, The step of pre-building the score-cause association rule base in step S2 includes: Establish a logical mapping relationship between NPS score ranges and cause options. The rules include score ranges, required items, prohibited items, option strength levels, and rule priorities.
3. The method according to claim 2, characterized in that, The logic of the intensity grading specifically includes: dividing the cause options into three levels—strong, medium, and weak—based on emotional intensity, with different levels corresponding to different weights, in order to quantify the degree of conflict.
4. The method according to claim 2, characterized in that, A cold start optimization step is included before step S2: Initial questionnaire samples are collected from the target domain. The co-occurrence frequency of reason options and ratings is analyzed by clustering to generate high-frequency co-occurrence combination suggestions for screening and forming an initial rule base.
5. The method according to claim 1, characterized in that, The contradiction detection in step S2 further includes: Obtain the score of the answer sheet and the set of selected reason options; Match the corresponding rule entry in the score-reason association rule base according to the score value; Verify whether the set of reason options meets the minimum satisfaction condition of the matched rule entry, and whether it contains prohibited items of the matched rule entry; The conflict score of the cause is quantitatively calculated based on the strength of the option that violates the rule and the priority of the rule.
6. The method according to claim 1, characterized in that, The step of pre-constructing the domain sentiment dictionary in step S3 includes: Manually input single-digit core sentiment keywords for the relevant domain; Statistical filtering algorithms were used to select candidate sentiment words from historical valid response corpora; The candidate sentiment words are subject to final manual review and then added to the database. The domain sentiment dictionary includes a positive lexicon, a negative lexicon, a degree adverb lexicon, and a negation lexicon.
7. The method according to any one of claims 1 or 6, characterized in that, Step S3, which involves calculating the sentiment of the open text using a deterministic text analysis algorithm, includes: The open text is segmented and divided to distinguish between core sentences and auxiliary sentences, and the core sentences are assigned a higher weight than the auxiliary sentences. The frequency of positive and negative words in the open text is statistically analyzed, and the weight coefficients in the degree adverbial lexicon are used for adjustment. Identify negative words, and based on the rules of the negative word effective window, reverse the sentiment tendency only for sentiment words within the negative word effective window; Based on the adjusted and reversed weighted counts of positive and negative words, the difference between the two is calculated as the sentiment value of the open text.
8. The method according to claim 7, characterized in that, Step S3, which compares the calculation result with the expected sentiment direction corresponding to the score value and outputs a text contradiction score, further includes: The calculated text sentiment tendency value is compared with the expected sentiment direction corresponding to the score value; If the text sentiment tendency value is completely opposite to the expected sentiment direction, it is determined to be a logical contradiction; Based on the determination that there is a logical contradiction, the open text is matched with a predefined contradictory sentence pattern. If the match is successful, the quantification level corresponding to the logical contradiction is increased. Based on the determination and quantification level of the logical contradiction, a text contradiction score is output.
9. The method according to claim 6, characterized in that, The number of domain sentiment words in the single digit ranges from 3 to 5.
10. The method according to claim 1, characterized in that, Following step S3, the following is also included: The open text is subjected to duplicate detection and spam detection, and duplicate tags and spam tags are generated accordingly; The duplication detection is achieved by generating a local sensitive hash value for the open text and comparing its similarity with the hash values of historical responses. The spam detection is achieved by matching the open text with a preset spam rule base, which includes rules for repeated characters, meaningless characters, and irrelevant information.
11. The method according to claim 10, characterized in that, The step S4 of determining the suspiciousness score of the questionnaire includes: weighting and summing the cause contradiction score, text contradiction score, duplicate mark and spam mark to obtain the suspiciousness score.
12. The method according to any one of claims 1 or 11, characterized in that, Step S5, determining the validity of the questionnaire, further includes performing a grading process: If the suspiciousness score is lower than the first threshold, it is determined to be a normal answer and submission is allowed; If the suspicion score is between the first threshold and the second threshold, it is determined to be a medium-risk answer, and a confirmation prompt will be displayed to the user. If the suspiciousness score is higher than the second threshold, it is determined to be a high-risk answer, and an interception prompt will be displayed to the user to prevent submission.
13. An NPS survey anti-fraud system, characterized in that, include: The data acquisition module is used to obtain users' NPS survey responses; The verification module, based on the score-cause association rule base, performs contradiction detection on the score value and the selected cause option, and outputs the cause contradiction score; Based on the aforementioned domain sentiment dictionary, the open text is analyzed using a deterministic text analysis algorithm to calculate and compare sentiment tendencies, and the resulting text contradiction score is output. The decision engine module is used to determine the suspiciousness score of the questionnaire based on the cause contradiction score and the text contradiction score, and to determine the validity of the questionnaire based on the suspiciousness score.
14. A computer-readable storage medium, characterized in that, The medium stores instructions that, when executed by a computer processor, can implement any of the methods described in claims 1 to 12.