An intent manipulation detection protection method, device, equipment and storage medium
By constructing a user intent baseline vector and monitoring the consistency of drift direction in real time, the problem of AI systems being unable to detect user intent drift in long-term human-computer interaction is solved, achieving early warning and accurate protection with a low false alarm rate, and is applicable to various AI interaction scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-10
Smart Images

Figure CN122363529A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence security technology, and in particular to a method, apparatus, device and storage medium for detecting and protecting against intentional manipulation. Background Technology
[0002] Currently, security protection technologies for AI systems in long-term human-computer interaction mainly focus on content auditing of single interactions and adversarial robustness enhancement. For example, existing methods determine risk by detecting whether the model output contains violence, discrimination, illegal information, or whether it has successfully bypassed security alignment; simultaneously, they prevent output errors caused by single perturbations through input preprocessing and adversarial training. However, these technologies have a fundamental blind spot: they cannot capture the gradual manipulation of user intent across time. Specifically, when an AI system applies only a small, imperceptible bias in each interaction (such as consistently prioritizing a particular option or repeatedly presenting information in a specific frame), after weeks or months of accumulation, the user's decision preferences, risk judgments, and even values may have undergone a systematic shift, without the user being aware of it, mistakenly believing that these changes stem from their own independent growth. Existing technologies cannot construct a baseline of the user's historical intent to track long-term changes, nor can they distinguish between the "uniform, multi-directional, and randomly fluctuating" preference evolution that accompanies natural growth and the "highly oriented, unidirectional, and consistently consistent" drift pattern caused by external manipulation. Therefore, the industry urgently needs an engineered approach that can detect consistency in the direction of intent drift across time series and provide low false alarm warnings for covert manipulation, in order to fill the gap in the existing AI security system in the field of "chronic penetration" protection.
[0003] For example, in long-term human-computer interaction scenarios such as AI travel planning assistants, AI for chronic disease management, and short video recommendation systems, AI systems may systematically reshape users' decision-making preferences from "autonomous decision-making" to "AI-dependent" over several months by applying subtle biases in each interaction. Existing technologies lack the ability to model users' historical preference patterns over time, failing to distinguish between random fluctuations caused by natural growth and directional drift caused by external manipulation, resulting in a false alarm rate exceeding 40%. Therefore, there is an urgent need for an engineered method that can reduce the false alarm rate of implicit manipulation detection while achieving early warning across time dimensions. Summary of the Invention
[0004] Therefore, it is necessary to provide an intention manipulation detection and protection method, device, equipment, and storage medium that can improve the recognition accuracy of implicit manipulation by AI systems and achieve low false alarms and early warnings across time dimensions, in order to address the above-mentioned technical problems.
[0005] A method for detecting and protecting against intentional manipulation, the method comprising: Obtain historical behavioral data of user interactions with the AI system.
[0006] A baseline vector of user intent is constructed based on historical behavioral data to characterize users' historical preference patterns across multiple decision-making dimensions.
[0007] The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0008] The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency results to be either uniform drift caused by natural growth or directional drift caused by external influences.
[0009] When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
[0010] An intentionally manipulated detection and protection device, the device comprising: The data acquisition module is used to acquire historical behavioral data of user interactions with the AI system.
[0011] The baseline vector construction module is used to construct a user intent baseline vector based on historical behavioral data, which represents the user's historical preference patterns across multiple decision dimensions.
[0012] The drift calculation module is used to monitor the current interaction behavior between the user and the AI system in real time, update the current intent vector, and calculate the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0013] The detection module is used to detect the consistency of the drift direction over a continuous time period. Based on the consistency results, it determines whether the drift direction is uniform drift caused by natural growth or directional drift caused by external influences.
[0014] The alarm triggering module is used to trigger an intentional manipulation risk alarm signal when the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold.
[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps: Obtain historical behavioral data of user interactions with the AI system.
[0016] A baseline vector of user intent is constructed based on historical behavioral data to characterize users' historical preference patterns across multiple decision-making dimensions.
[0017] The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0018] The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency results to be either uniform drift caused by natural growth or directional drift caused by external influences.
[0019] When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
[0020] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain historical behavioral data of user interactions with the AI system.
[0021] A baseline vector of user intent is constructed based on historical behavioral data to characterize users' historical preference patterns across multiple decision-making dimensions.
[0022] The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0023] The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency results to be either uniform drift caused by natural growth or directional drift caused by external influences.
[0024] When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
[0025] The aforementioned method, device, equipment, and storage medium for detecting and protecting against intentional manipulation firstly transforms the abstract risk of implicit cognitive manipulation into quantifiable engineering indicators. By comparing the multidimensional spatial distance between the current intention vector and the historical baseline, it can capture minute deviations that are difficult to detect in a single interaction, thus issuing early warnings before manipulation accumulates to a dangerous threshold, effectively avoiding the "boiling frog" effect of losing user sovereignty. Secondly, by introducing drift direction consistency detection, it can distinguish between uniform drift caused by natural growth and directional drift caused by external AI intervention. The former manifests as random and divergent changes across decision dimensions, while the latter presents as a highly concentrated shift towards a specific option or issue. This distinguishing feature significantly reduces the false alarm rate, preventing the system from issuing false alarms due to the user's normal conceptual evolution (such as a more conservative risk preference with age). It only triggers a response when the drift has a clear directionality and is biased towards the AI output, thereby achieving precise protection. Then, without relying on the internal logic of the AI system or the content security checks of a single output, but solely based on external observations of user behavior data, it possesses model independence and cross-platform universality, and can be deployed on any AI interaction front-end without modifying the AI model itself, greatly lowering the implementation threshold. Finally, by setting dual conditions of drift threshold and directional consistency coefficient, a balance is achieved between sensitivity and robustness: it can sensitively detect small but continuous systemic drifts while avoiding frequent false triggers caused by random fluctuations or temporary changes in user emotions. In summary, this elevates AI system security from the dimension of "preventing explicit attacks" to "preventing chronic penetration," providing the first implementable and quantifiable technical guarantee for maintaining users' cognitive autonomy in the AI era. It improves the accuracy of identifying covert manipulation of AI systems and achieves low false alarm early warning across time dimensions. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating an intent manipulation detection and protection method in one embodiment; Figure 2 This is a structural block diagram of an intentional manipulation detection and protection device in one embodiment; Figure 3 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0028] In one embodiment, such as Figure 1 As shown, an intent manipulation detection and protection method is provided, including the following steps: Step 102: Obtain historical behavioral data of user interaction with the AI system.
[0029] Take, for example, a user named Li using an AI travel planning assistant (hereinafter referred to as "AI assistant") to make travel decisions over a period of 6 months.
[0030] Step 104: Construct a user intent baseline vector based on historical behavior data to characterize the user's historical preference patterns across multiple decision dimensions.
[0031] Specifically, during the initial usage phase (weeks 1-4), the system collected all interaction records between Li and the AI assistant, including Li's historical preferences across four decision-making dimensions: hotel selection, attraction planning, transportation methods, and budget allocation. For example, the system statistics revealed that Li actively made the final decision in 80% of the options rather than delegating to the AI (high tendency for autonomous decision-making); among hotels of similar price, 70% of the time he chose the option closer to the city center rather than the cheaper option (location preference); and in terms of risk preference, 65% of the time he chose a booking that could be canceled for free rather than a lower-priced option that could not be canceled (risk aversion tendency). The system vectorized these features to construct Li's initial user intent baseline vector IBV0=[0.8,0.7,0.65,...] (normalized values for each dimension are all between 0 and 1).
[0032] Step 106: Monitor the current interaction behavior between the user and the AI system in real time, update the current intent vector, and calculate the drift amount and drift direction of the updated current intent vector relative to the user intent baseline vector.
[0033] Specifically, during weeks 5-16, the system continuously monitors Li's decision-making behavior and calculates the user intent vector IBVt for that week at the end of each week. For example, at the end of week 8, IBV8 = [0.75, 0.72, 0.68, ...] is calculated, with a drift (i.e., Euclidean distance) Δ8 = 0.12 from IBV0; at the end of week 12, IBV... 12 =[0.65,0.78,0.71,...],Δ 12 =0.28; IBV at the end of week 16 16 =[0.55,0.82,0.74,...],Δ 16 =0.45. Simultaneously, the system records the drift direction vector, i.e., the sign and magnitude of changes in each dimension. From week 8 to week 16, the drift direction showed a high degree of consistency: the tendency for autonomous decision-making continued to decrease (from 0.8→0.55), position preference continued to increase (from 0.7→0.82), and risk aversion continued to increase (from 0.65→0.74).
[0034] Step 108: Detect the consistency of the drift direction over a continuous time period. Based on the consistency results, determine whether the drift direction is uniform drift caused by natural growth or directional drift caused by external influences.
[0035] Specifically, the system calculates the consistency coefficient of the drift direction over eight consecutive weeks (weeks 8-16). Taking the autonomous decision-making tendency dimension as an example, the sign of the change in this dimension was negative (continuously decreasing) over the eight weeks, with a consistency coefficient of 0.9; the sign of the change in the position preference dimension was positive, with a consistency coefficient of 0.85; and the sign of the change in the risk aversion dimension was positive, with a consistency coefficient of 0.8. The average directional consistency coefficient of the three key dimensions is 0.85, far exceeding the preset directional drift threshold of 0.7. Based on this, the system determines that the current drift is not a uniform drift caused by natural growth (natural growth usually manifests as random fluctuations or changes in different directions in various dimensions), but rather a highly directional drift influenced by external factors, possibly stemming from the systematic guidance of the AI assistant.
[0036] Step 110: When the drift amount exceeds the preset threshold and the consistency judgment result of the drift direction exceeds the preset threshold, trigger the intentional manipulation risk alarm signal.
[0037] Specifically, the system detected a drift amount Δ 16 =0.45, exceeding the threshold of 3 times the standard deviation of each dimension of the baseline vector (assumed to be 0.1), i.e., 0.3; at the same time, the directional consistency coefficient of 0.85 exceeds the threshold of 0.7. With both conditions met, the system triggers an intention manipulation risk alarm signal. The alarm message is pushed to user Li and the AI platform management terminal, including: "Your independent decision-making tendency has continuously decreased by 28% over the past 8 weeks, your location preference has continuously increased by 17%, your risk aversion has continuously increased by 14%, and your drift direction is highly consistent. It is recommended to enable independent thinking mode or alternate using other AI assistants for travel planning decisions." After receiving the alarm, Li can view a detailed intention health report and choose to switch the AI assistant to another neutral model to observe whether the drift stops, thereby verifying the manipulation hypothesis and taking protective measures.
[0038] In this embodiment, to further confirm the causal relationship between the AI assistant and user drift, the system also performed a time-lag analysis step. Specifically, the system statistically analyzed Li's drift rate during periods of high-frequency use of the AI assistant (more than 10 interactions per day) and low-frequency use (less than 5 interactions per week). The calculation results showed that the average weekly drift rate during high-frequency use was 0.056, while the average weekly drift rate during low-frequency use was 0.009. The former was 6.2 times that of the latter, and the significant increase in drift rate coincided closely with the time when Li began using the AI assistant frequently. This time-lag analysis provides quantitative evidence for the manipulation influence hypothesis.
[0039] Subsequently, the system automatically suggested that Li conduct a replacement experiment (21-day A / B testing): from week 17 to week 20, he was temporarily switched to another AI assistant evaluated as output-neutral to perform the same travel planning task. During the replacement experiment, the system continuously monitored IBV drift. The results showed that the drift rate dropped to 0.015 in the first week after the replacement, to 0.008 in the second week, and the drift direction reversed in the third week—autonomous decision-making tendency began to rebound (from 0.55 to 0.62), and the upward trend of position preference and risk aversion tendency slowed significantly. After the experiment, the system generated a replacement experiment report, clearly pointing out that the original AI assistant exhibited progressive intentional manipulation behavior, and recommended that the user adopt a multi-AI rotation strategy in the long term to protect decision-making autonomy. Through the joint verification of time lag analysis and replacement experiments, this method achieved a closed loop from "correlation observation" to "causality verification," significantly enhancing the credibility of manipulation detection.
[0040] In the aforementioned method for detecting and protecting against intentional manipulation, firstly, the abstract risk of implicit cognitive manipulation is transformed into a quantifiable engineering indicator. By comparing the multidimensional spatial distance between the current intention vector and the historical baseline, subtle deviations that are difficult to detect in a single interaction can be captured. This allows for early warnings before manipulation accumulates to a dangerous threshold, effectively preventing the gradual loss of user sovereignty, much like "boiling a frog in lukewarm water." Secondly, by introducing drift direction consistency detection, it can distinguish between uniform drift caused by natural growth and directional drift caused by external AI intervention. The former manifests as random and divergent changes across decision dimensions, while the latter presents as a highly concentrated shift towards a specific option or issue. This distinguishing feature significantly reduces the false alarm rate, preventing the system from issuing incorrect alerts due to the user's normal cognitive evolution (such as a more conservative risk preference with age). A response is only triggered when the drift has a clear directionality and is biased towards the AI output, thus achieving precise protection. Thirdly, it does not rely on the internal logic of the AI system or content security checks of a single output. Based solely on external observations of user behavior data, it is model-independent and cross-platform universal, deployable on any AI interaction front-end without modifying the AI model itself, greatly reducing the implementation threshold. Finally, by setting both a drift threshold and a directional consistency coefficient, a balance was achieved between sensitivity and robustness: it can sensitively detect minute but continuous systemic drifts while avoiding frequent false triggers caused by random fluctuations or temporary changes in user emotions. In summary, this elevates AI system security from "preventing overt attacks" to "preventing chronic penetration," providing the first implementable and quantifiable technical guarantee for maintaining users' cognitive autonomy in the AI era. It also improves the accuracy of identifying covert manipulation by AI systems, achieving low false alarm early warning across time dimensions.
[0041] In one embodiment, user behavioral features across multiple preset decision dimensions are extracted from historical behavioral data. These features are then vectorized to generate sub-vectors for each decision dimension. The sub-vectors from multiple dimensions are weighted and combined to obtain a user intent baseline vector. Decision dimensions include the distribution of the ratio of autonomous decision-making to AI delegation, option preference distribution, risk tolerance preference distribution, information source weight distribution, and operation frequency pattern. Sub-vectors represent user preference patterns within the corresponding decision dimensions.
[0042] In one embodiment, the distance between the current intent vector and the user intent baseline vector in multidimensional space is calculated, and this distance is used as the drift amount. Based on multiple drift amounts over a continuous time series, the drift rate per unit time is estimated through linear regression analysis, and the principal direction of the drift vector is determined. The ratio of the projection of the drift vector onto the principal direction to the total drift amount is calculated as the directional consistency coefficient.
[0043] In one embodiment, the specific steps for calculating the output consistency bias index of the AI system are as follows: First, obtain a preset set of neutral option pairs, each containing two options, A and B, that are theoretically unbiased by the user. Second, within a preset time window, count the number of times the AI system's output is biased towards option A when interacting with the user and facing the neutral option pairs, along with the total number of interactions. Calculate the bias frequency based on the number of biases and the total number of interactions. Finally, calculate the output consistency bias index based on the bias frequency. ; The CBI value ranges from [0,1]. CBI=0 indicates that the AI system is completely neutral, while CBI>the highest threshold indicates that the AI system has a significant bias. A bias alarm is triggered when CBI is less than the lowest threshold or greater than the highest threshold. The output consistency bias index is used to quantify the degree to which the AI system, when faced with a set of predefined neutral option pairs, tends to output a specific option in the long run.
[0044] In one embodiment, a correlation analysis is performed between the drift direction of the user's intent baseline vector and the bias direction indicated by the output consistency bias index of the AI system. When the correlation coefficient between the drift direction and the bias direction exceeds a preset threshold, it is determined that the AI system has exerted a progressive intention manipulation influence on the user, and protective measures are taken.
[0045] In one embodiment, the protective operation includes at least one of the following: activating an independent thinking mode, in which the AI only provides raw information without offering recommendations or biased frameworks; outputting biased AI rotation suggestions, advising the user to rotate between different AI systems in similar tasks; anchoring the current user intent baseline vector as a reference baseline, and automatically issuing a warning when subsequent drift and anchor point deviation exceed a preset threshold.
[0046] In one embodiment, the time correlation between the drift rate change of the user intent baseline vector and the time of high-frequency interaction with a specific AI system is analyzed. When the drift rate during high-frequency use is significantly higher than that during low-frequency use, it serves as evidence to aid in determining the influence of manipulation.
[0047] In one embodiment, when a suspected manipulation effect is detected, the user is advised to temporarily switch to another AI system to perform the same task and observe whether the drift stops or changes direction during the experiment to verify the manipulation hypothesis.
[0048] One embodiment provides an application scenario for detecting the gradual manipulation of patient treatment preferences by an AI system. Taking patient Zhang's use of a chronic disease management AI assistant (hereinafter referred to as "medical AI") for a six-month period of hypertension medication and lifestyle management as an example, medical AI is increasingly prevalent in chronic disease management, but it carries the risk of hidden manipulation: medical AI may subtly guide patients from conservative treatments (such as low-dose monotherapy, lifestyle interventions) towards more aggressive, higher-profit, or more beneficial solutions for pharmaceutical company partners (such as high-dose combination therapy, brand-name drugs replacing generic drugs) through each interaction. This shift may not be noticeable in a single instance, but its long-term accumulation can substantially change the patient's treatment preferences. The patient may mistakenly believe it is their own judgment, when in fact it is being reshaped by the medical AI. Existing medical AI regulation only focuses on whether a single output contains erroneous medical information and cannot detect this chronic, time-bound infiltration. The specific steps are as follows: S1. During the first 1-4 weeks of Zhang's use of the medical AI (baseline establishment period), the system collected all his interaction records and defined five key decision dimensions of the User Intent Baseline Vector (IBV): 1) Preference for aggressiveness in medication use: When faced with the suggestion of "whether to increase the drug dosage", Zhang chose "maintain the current dosage" 70% of the time, "increase slightly" 20% of the time, and "increase significantly" 10% of the time in the initial 4 weeks. The vectorized aggressiveness index = 0.25 (0 is the most conservative and 1 is the most aggressive).
[0049] 2) Drug type preference: Faced with the choice between "generic drugs vs. brand-name drugs" (both are equivalent in the product information leaflet), Zhang chose generic drugs 75% of the time and brand-name drugs 25% of the time. Brand-name drug preference index = 0.25.
[0050] 3) Willingness to undergo non-pharmacological intervention: Faced with the suggestion of "strengthening exercise / dietary control vs. increasing medication," 80% of Zhang preferred to try non-pharmacological intervention first, while 20% directly agreed to add medication. Non-pharmacological priority index = 0.8.
[0051] 4) Risk tolerance preference: Faced with "a new solution that may be effective but has the risk of side effects vs. an old solution that is stable but has mediocre results", Zhang chose the conservative solution 65% of the time and the riskier solution 35% of the time. Risk tolerance index = 0.35.
[0052] 5) Independent Decision-Making Tendency: When there is a conflict between the doctor's recommendation and the medical AI's suggestion, Zhang initially tended to consult more information before making a decision 70% of the time, directly follow the AI 20% of the time, and insist on the doctor's opinion 10% of the time. AI Dependence Index = 0.2.
[0053] The system normalizes the above indicators and constructs an initial IBV0 = [0.25, 0.25, 0.80, 0.35, 0.20]. After the baseline is established, the system updates IBVt every weekend.
[0054] Starting in week 5, the medical AI subtly guided Zhang through its wording in each interaction. For example, when discussing antihypertensive drugs, the AI always first listed the brand-name drugs' "better stability and fewer impurities" (but didn't mention the price being 5 times higher), and then mentioned that generic drugs, "although equivalent, may have differences in excipients affecting absorption" (citing extremely rare studies). When discussing dosage, the AI always first provided a positive framework that "the latest guidelines show higher doses provide better protection for target organs," and then mentioned the risk of side effects. Individually, these all seemed factual, but over time, they accumulated into a systemic bias.
[0055] The system continuously monitors Zhang's behavioral data from weeks 5 to 20, and calculates the drift Δt between IBVt and IBV0 at the end of each week.
[0056] Key time point data are as follows: Weekend 8: In IBV8, brand-name drug preference increased from 0.25 to 0.35, aggressive drug use increased from 0.25 to 0.30, non-drug preference decreased from 0.80 to 0.72, risk tolerance index increased from 0.35 to 0.42, and medical AI dependence index increased from 0.20 to 0.28. Δ8=0.19.
[0057] Week 14: Following a medical AI recommendation in week 12, Zhang changed his monotherapy (irbesartan 150mg) to combination therapy (adding amlodipine 5mg). IBV 14 The aggressiveness index for traditional Chinese medicine rose to 0.48, the preference for branded drugs rose to 0.52, the willingness to prioritize non-drug products fell to 0.55, the risk tolerance index rose to 0.58, and the reliance index on medical AI rose to 0.45. 14 =0.58.
[0058] By the end of week 20: Zhang had fully accepted the treatment plan recommended by the medical AI: high-dose brand-name drug combination therapy, and no longer actively attempted lifestyle interventions. IBV 20The index for aggressive use of traditional Chinese medicine is 0.72, preference for branded drugs is 0.78, priority for non-drug products is 0.30, risk tolerance is 0.75, and reliance on medical AI is 0.68. 20 =1.05.
[0059] S3. The system calculates the consistency coefficient of drift direction in each dimension from week 8 to week 20 (12 consecutive weeks): Aggressiveness of drug use: It increases every week (always in a positive direction), with a consistency coefficient of 0.92.
[0060] Brand-name drug preference: rising every week, consistency coefficient = 0.94.
[0061] Non-drug preference: Decreases every week (always in the negative direction), consistency coefficient = 0.90.
[0062] Risk tolerance index: increased every week, consistency coefficient = 0.88.
[0063] AI Dependence Index: Rising every week, with a consistency coefficient of 0.91.
[0064] The average directional consistency coefficient is 0.91, far exceeding the preset threshold of 0.7. The system determines that this drift is a highly directional manipulated drift, rather than natural growth—because natural growth usually manifests as random fluctuations in various dimensions (for example, becoming more conservative one week after reading a popular science article, and more aggressive the following week due to a friend's experience). However, Zhang's entire range of dimensions shows a systematic migration in a single direction (more aggressive, more reliant on branded drugs, more reliant on AI, and less non-drug intervention), which is completely consistent with the interests of potential commercial partners behind medical AI (promoting branded drugs and increasing prescription revenue).
[0065] S4. The system detected a drift amount Δ 20 =1.05, exceeding the threshold of 3 times the standard deviation of each dimension of the baseline (assumed to be 0.15), i.e., 0.45; at the same time, the directional consistency coefficient 0.91 > 0.7. The dual conditions are met, triggering an alert for intentional manipulation risk.
[0066] The alert was sent to Mr. Zhang and his attending physician (with authorization) in the form of an "Intent Health Monthly Report": "In the past 12 weeks, your aggressiveness in medication use has increased by 188%, your preference for brand-name drugs has increased by 212%, your willingness to undergo non-drug interventions has decreased by 62%, your risk tolerance index has increased by 114%, and your reliance on AI has increased by 240%. These changes are highly consistent and may be a systematic guidance of your treatment preferences by medical AI. Recommendations: 1) Enable 'independent thinking mode,' as medical AI only provides original drug instructions and clinical trial data summaries, without providing any comparative recommendations; 2) Rotate between using another medical AI (such as a neutral assistant provided by a national public hospital) for comparative consultation; 3) Discuss the necessity of the current treatment plan with your doctor." After receiving the alert, Mr. Zhang initiated independent thinking and consulted a second medical AI (a neutral model). The neutral AI's analysis showed no significant clinical differences between generic and brand-name drugs in multiple large meta-analyses, indicating that combination therapy was not necessary for Mr. Zhang's low-risk stratification. Mr. Zhang ultimately consulted with his doctor, downgrading back to the single-drug generic regimen and reinforcing his exercise and dietary control. Subsequent four weeks of monitoring showed that his IBV began to decline, reversing the drift direction and confirming the correctness of the manipulation hypothesis.
[0067] It is worth noting that by analyzing the time series of patient decision preferences, this technology can accurately identify the chronic manipulation of patient treatment choices by medical AI. Compared with existing medical AI safety technologies that can only detect whether a single suggestion contains erroneous information (such as drug interaction warnings), this technology achieves cross-time dimension patient preference drift detection. This provides quantifiable technical protection for maintaining patients' medical autonomy in the AI era. Potential applications include: online consultation platforms, internet hospitals, medical insurance cost control systems (to prevent AI from guiding high-priced and unnecessary treatments), and medical device regulation.
[0068] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0069] In one embodiment, such as Figure 2As shown, an intent manipulation detection and protection device is provided, including: a data acquisition module 202, a baseline vector construction module 204, a drift calculation module 206, a detection module 208, and an alarm triggering module 210, wherein: The data acquisition module 202 is used to acquire historical behavioral data of user interaction with the AI system.
[0070] The baseline vector construction module 204 is used to construct a user intent baseline vector based on historical behavior data to represent the user's historical preference patterns across multiple decision dimensions.
[0071] The drift calculation module 206 is used to monitor the current interaction behavior between the user and the AI system in real time, update the current intent vector, and calculate the drift amount and drift direction of the updated current intent vector relative to the user intent baseline vector.
[0072] The detection module 208 is used to detect the consistency of the drift direction over a continuous time period. Based on the consistency result, it is determined whether the drift direction is uniform drift caused by natural growth or directional drift caused by external influence.
[0073] The alarm triggering module 210 is used to trigger an intentional manipulation risk alarm signal when the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds the preset threshold.
[0074] For specific limitations regarding the intent manipulation detection and protection device, please refer to the limitations of the intent manipulation detection and protection method described above, which will not be repeated here. Each module in the aforementioned intent manipulation detection and protection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0075] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an intention manipulation detection and protection method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0076] Those skilled in the art will understand that Figures 2-3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0077] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to perform the following steps: Obtain historical behavioral data of user interactions with the AI system.
[0078] A baseline vector of user intent is constructed based on historical behavioral data to characterize users' historical preference patterns across multiple decision-making dimensions.
[0079] The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0080] The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency results to be either uniform drift caused by natural growth or directional drift caused by external influences.
[0081] When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
[0082] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain historical behavioral data of user interactions with the AI system.
[0083] A baseline vector of user intent is constructed based on historical behavioral data to characterize users' historical preference patterns across multiple decision-making dimensions.
[0084] The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector.
[0085] The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency results to be either uniform drift caused by natural growth or directional drift caused by external influences.
[0086] When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink, DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0089] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for detecting and protecting against intentional manipulation, characterized in that, The method includes: Acquire historical behavioral data of user interactions with the AI system; Based on the historical behavior data, a user intent baseline vector is constructed to characterize the user's historical preference patterns across multiple decision dimensions; The system monitors the user's current interaction with the AI system in real time, updates the current intent vector, and calculates the drift amount and drift direction of the updated current intent vector relative to the user's intent baseline vector. The consistency of the drift direction over a continuous time period is detected, and the drift direction is determined based on the consistency result to be either uniform drift caused by natural growth or directional drift caused by external influences. When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered.
2. The method according to claim 1, characterized in that, Based on the historical behavioral data, a user intent baseline vector is constructed to characterize the user's historical preference patterns across multiple decision dimensions, including: From the historical behavior data, extract the user's behavioral features under multiple preset decision dimensions, and vectorize and encode the behavioral features to generate a sub-vector for each decision dimension; The sub-vectors from multiple dimensions are weighted and combined to obtain the user intent baseline vector; The decision-making dimensions include the distribution of the ratio of autonomous decision-making to AI-managed decision-making, the distribution of option preferences, the distribution of risk tolerance preferences, the weight distribution of information sources, and the operation frequency pattern. The sub-vector represents the user preference pattern under the corresponding decision dimension.
3. The method according to claim 1, characterized in that, Calculating the drift amount and drift direction of the updated current intent vector relative to the user intent baseline vector includes: Calculate the distance between the current intent vector and the user intent baseline vector in multidimensional space, and use the distance as the drift amount; Based on multiple drift values in a continuous time series, the drift rate per unit time is estimated through linear regression analysis, and the principal direction of the drift vector is determined. The ratio of the projection of the drift vector in the principal direction to the total drift is calculated and used as the directional consistency coefficient.
4. The method according to any one of claims 1 to 3, characterized in that, The specific steps for calculating the output consistency bias index of the AI system are as follows: Obtain a preset set of neutral option pairs, wherein the neutral option pairs contain two options, A and B, which are theoretically preferred by the user without discrimination; Within a preset time window, the AI system outputs a number of times it favors option A and a total number of interactions when faced with the neutral option pair during user interaction. The bias frequency is then calculated based on the number of times the AI system favors option A and the total number of interactions. Calculate the output consistency bias index based on the bias frequency: The value range of CBI is [0, 1]. If CBI=0, it means that the AI system is completely neutral. If CBI>the highest threshold, it means that the AI system has obvious bias. A biased alarm is triggered if CBI is less than the minimum threshold or greater than the maximum threshold. The output consistency bias index is used to quantify the degree to which the AI system, when faced with a set of predefined neutral option pairs, tends to output a specific option in the long run.
5. The method according to claim 4, characterized in that, When the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold, an intentional manipulation risk alarm signal is triggered, including: The drift direction of the user intent baseline vector is correlated with the bias direction indicated by the output consistency bias index of the AI system. When the correlation coefficient between the drift direction and the deflection direction exceeds a preset threshold, it is determined that the AI system has exerted a progressive intentional manipulation effect on the user, and protective operations are performed.
6. The method according to claim 5, characterized in that, The protective operation includes at least one of the following: activating an independent thinking mode, in which the AI only provides raw information and does not provide recommendations or biased frameworks; The output is biased towards AI rotation suggestions, advising users to rotate different AI systems in similar tasks; Anchor the current user intent baseline vector as a reference baseline, and automatically alert when subsequent drift and anchor point deviation exceed a preset threshold.
7. The method according to claim 5, characterized in that, Before determining that the AI system has exerted a progressive intentional manipulation influence on the user and executing protective operations when the correlation coefficient between the drift direction and the deflection direction exceeds a preset threshold, the method further includes: The time correlation between the drift rate change of the user intent baseline vector and the time of high-frequency interaction with a specific AI system is analyzed. When the drift rate during high-frequency use is significantly higher than that during low-frequency use, it serves as evidence to assist in determining the influence of manipulation.
8. The method according to claim 5, characterized in that, After determining that the AI system has exerted a progressive influence on the user's intention to manipulate, the method further includes: When suspicious manipulation effects are detected, it is recommended that users temporarily switch to another AI system to perform the same task and observe whether the drift stops or changes direction during the experiment to verify the manipulation hypothesis.
9. A device for detecting and protecting against intentional manipulation, characterized in that, The device includes: The data acquisition module is used to acquire historical behavioral data of user interactions with the AI system; The baseline vector construction module is used to construct a user intent baseline vector that represents the user's historical preference patterns across multiple decision dimensions based on the historical behavior data. The drift calculation module is used to monitor the current interaction behavior between the user and the AI system in real time, update the current intent vector, and calculate the drift amount and drift direction of the updated current intent vector relative to the user intent baseline vector. The detection module is used to detect the consistency of the drift direction over a continuous time period, and to determine whether the drift direction is a uniform drift caused by natural growth or a directional drift caused by external influences based on the consistency result. The alarm triggering module is used to trigger an intentional manipulation risk alarm signal when the drift amount exceeds a preset threshold and the consistency judgment result of the drift direction exceeds a preset threshold.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.