Method and system for adjusting behavior motivation of an agent, electronic device, program product

CN122840103APending Publication Date: 2026-09-29TIANJIN TIANKAI ZHIHUIYUN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272402.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

该类方式属于后置管控手段,仅能够对输出行为做结果层面的修正,难以介入智能体内部的策略优化过程,无法对智能体的决策倾向进行主动调整

Benefits of technology

[0037]在符合本领域常识的基础上,上述各优选条件,可任意组合,即得本公开各较佳实例。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840103A_ABST
    Figure CN122840103A_ABST
Patent Text Reader

Abstract

The present disclosure provides an agent behavior motivation adjustment method, system, electronic device and program product. The method comprises: obtaining feedback data of an agent decision, and extracting a group feedback quantitative index according to the feedback data; the group feedback quantitative index represents the comprehensive negative tendency of a user group to the agent decision; according to a mapping relationship between the quantitative index and a deviation degree, the group feedback quantitative index is mapped to a behavior compliance deviation degree; the behavior compliance deviation degree is taken as a penalty term, which is injected into a reward function of an agent decision model, and an optimization objective of the agent decision model is reconstructed to realize adjustment of the agent behavior motivation. The agent of the present disclosure can actively avoid behaviors that are easy to trigger negative group feedback in the policy iteration process, overcome the disadvantages that post-processing control such as after-burning and unified degradation can only correct output results and cannot intervene in internal strategy motivation, realize quantitative constraint of group feedback on agent strategy learning, and balance task performance and behavior compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, system, electronic device, and program product for adjusting the behavioral motivation of an intelligent agent. Background Technology

[0002] Currently, AI agents are widely used in various task decision-making scenarios. To regulate the output behavior of these agents, safety constraints need to be imposed on them. AI agents generally rely on reinforcement learning frameworks to complete policy iterations. They use reward functions to define the agent's behavioral motivations, set optimization goals, drive policy updates, and thus guide the agent to output corresponding decision-making behaviors.

[0003] In existing technologies, some solutions constrain the output of intelligent agents through external security control measures, such as using binary control methods like post-event circuit breaking and unified degradation to intercept generated decision-making behaviors. These methods are post-event control measures, only able to modify the output behavior at the result level, and cannot intervene in the internal strategy optimization process of the intelligent agent, nor can they proactively adjust the agent's decision-making tendencies.

[0004] Furthermore, most existing reward functions for reinforcement learning agents only use single task objectives such as task completion rate and execution efficiency as optimization guides, which makes it impossible for agents to endogenously constrain non-compliant behavior during the policy learning phase. Relying solely on ex-post control measures is insufficient to fundamentally improve the agent's behavioral compliance level. Summary of the Invention

[0005] The technical problem to be solved by this disclosure is to overcome the above-mentioned defects in the prior art and provide a method, system, electronic device, or program product for adjusting the behavioral motivation of an intelligent agent.

[0006] This disclosure solves the above-mentioned technical problems through the following technical solution:

[0007] Firstly, a method for adjusting the behavioral motivation of an intelligent agent is provided, including:

[0008] The process involves obtaining feedback data on the agent's decision-making output by the agent decision-making model, and extracting a group feedback quantification index based on the feedback data. The group feedback quantification index represents the overall negative tendency of the feedback group towards the agent's decision-making. The feedback data consists of evaluation data actively submitted by each feedback subject in the feedback group after observing the agent's corresponding decision-making behavior.

[0009] Based on the mapping relationship between quantitative indicators and deviation, the quantitative indicators of group feedback are mapped to behavioral compliance deviation.

[0010] The deviation from compliance is used as a penalty and injected into the reward function of the agent decision-making model to reconstruct the optimization objective of the agent decision-making model, thereby adjusting the agent's behavioral motivation.

[0011] Optionally, the mapping relationship is represented by a Sigmoid function, and the number of the group feedback quantification indicators is at least two.

[0012] Mapping the aforementioned group feedback quantitative indicators to behavioral compliance deviations includes:

[0013] The weighted results of the quantitative indicators of feedback from each group are normalized using the Sigmoid function to obtain the behavioral compliance deviation.

[0014] Optionally, it also includes:

[0015] In response to a situation where the number of times the behavioral compliance deviation exceeds a deviation threshold within a target time period is greater than a frequency threshold, target interaction data within the target time period is collected; the target interaction data includes interaction trajectories where the correction reward is lower than a reward threshold; the correction reward is obtained by injecting a reward function into the penalty item;

[0016] The target interaction data is used as training samples to iteratively optimize the agent decision-making model.

[0017] Optionally, feedback data regarding the agent's decision-making is obtained, including:

[0018] The system collects initial feedback data and feedback subject information for the agent's decision-making; the initial feedback data includes feedback weights; the feedback subject information includes at least one of the following: the feedback subject's registration information, the location where the feedback behavior occurred, and the time when the feedback behavior occurred.

[0019] Extract feedback subject features from the feedback subject information; the feedback subject features include at least one of the following: social relationship topology features between feedback subjects, geospatial clustering features, temporal behavior features, and feedback content features;

[0020] Based on the characteristics of the feedback subjects, identify feedback subjects with similar behavioral patterns and classify collusive groups.

[0021] The feedback weights of the initial feedback data corresponding to the feedback subjects belonging to the collusion group are attenuated and corrected. The initial feedback data after the feedback weight attenuation correction and the initial feedback data that did not participate in the feedback weight attenuation correction are determined as the feedback data. The feedback weights represent the contribution of the corresponding feedback data to the aggregate calculation of the group feedback quantification index.

[0022] Optionally, initial feedback data on the agent's decision-making and information about the feedback subject are collected, including:

[0023] In response to the submission request of the initial feedback data, the biometric verification data of the feedback subject is collected;

[0024] Perform a liveness detection operation based on the aforementioned biometric verification data;

[0025] In response to successful liveness detection and / or successful digital signature verification of the initial feedback data, the initial feedback data and the feedback subject information are collected.

[0026] Optionally, it also includes:

[0027] Calculate the deviation of each feedback subject's feedback data from the current group consensus benchmark;

[0028] The feedback weights of feedback data whose deviation exceeds the deviation threshold are attenuated and corrected.

[0029] Optionally, the quantitative indicators of group feedback include at least one of the following: negative feedback ratio, negative feedback intensity, and role diversity index; the negative feedback ratio represents the proportion of negative feedback in all feedback; the negative feedback intensity represents the average degree of aversion to all negative feedback; and the role diversity index represents the richness of social identities covered by the feedback group.

[0030] And / or, in response to the duration of the deviation of the behavior from the security threshold exceeding the security threshold being greater than the duration threshold, a security policy is executed; the security policy includes at least one of the following: triggering the agent to run in a security mode with the lowest privileges, generating an audit encryption report of the agent's operation and sending it to the regulatory agency.

[0031] Secondly, a system for adjusting the behavioral motivation of an intelligent agent is provided, including:

[0032] The feedback acquisition module is used to acquire feedback data on the agent's decision-making output by the agent decision-making model, and extract a group feedback quantification index based on the feedback data; the group feedback quantification index represents the overall negative tendency of the feedback group towards the agent's decision; the feedback data is the evaluation data actively submitted by each feedback subject in the feedback group after observing the agent's corresponding decision-making behavior.

[0033] A compliance deviation generator is used to map the group feedback quantitative indicators to behavioral compliance deviations based on the mapping relationship between quantitative indicators and deviations.

[0034] The incentive adjustment module is used to inject the deviation of the behavior from compliance as a penalty into the reward function of the agent decision-making model, and reconstruct the optimization objective of the agent decision-making model to adjust the agent's behavioral motivation.

[0035] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and for running on the processor, wherein the processor executes the computer program to implement the method for adjusting the behavioral motivation of an intelligent agent as described in the first aspect.

[0036] Fourthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method for adjusting the behavioral motivation of an intelligent agent as described in the first aspect.

[0037] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of this disclosure.

[0038] The positive and progressive effects of this disclosure are as follows: This disclosure introduces the deviation from compliance, representing a group's negative tendencies, as a penalty term into the reward function, reconstructing the model's optimization objective. This fundamentally adjusts the agent's behavioral motivation, enabling it to proactively avoid decision-making behaviors that easily trigger negative group feedback. By adjusting behavioral patterns at the root of strategy optimization, it completes a paradigm shift from externally enforced control to internal motivational reshaping. The agent can proactively avoid behaviors that easily trigger negative group feedback during strategy iteration, overcoming the shortcomings of post-event controls such as circuit breaking and unified degradation, which can only correct output results and cannot intervene in internal strategy motivation. This achieves quantitative constraints on the agent's strategy learning based on group feedback, balancing task performance and behavioral compliance. Attached Figure Description

[0039] Figure 1 A flowchart illustrating a method for adjusting the behavioral motivation of an intelligent agent, provided as an exemplary embodiment of this disclosure;

[0040] Figure 2 A schematic diagram of a module for adjusting the behavioral motivation of an intelligent agent, provided as an exemplary embodiment of this disclosure;

[0041] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of the present disclosure. Detailed Implementation

[0042] The present disclosure is further illustrated below by way of embodiments, but the present disclosure is not limited to the scope of the embodiments described herein.

[0043] The prefixes such as "first" and "second" used in this disclosure are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this disclosure does not constitute a limitation on the described objects. The description of the described objects is given in the context of the embodiments, and the use of such prefixes should not constitute unnecessary restrictions. Furthermore, in the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0044] In this embodiment of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good morals.

[0045] To regulate the output behavior of intelligent agents, the industry has proposed implementing security constraints on intelligent agents. The inventors have discovered that existing security constraint mechanisms for AI intelligent agents have many technical flaws:

[0046] First, the credibility of feedback signals is difficult to guarantee. Most existing group feedback systems rely on anonymous accounts at the application layer for identity authentication. This identity authentication mechanism is vulnerable to attacks such as bulk account registration and fake reviews. It is difficult to verify the source of feedback at the hardware physical level, and it cannot ensure that each piece of feedback data comes from a real individual who has completed liveness verification and has differentiated social identity attributes. A large number of fake and invalid samples are easily mixed into the feedback data.

[0047] Second, the behavioral constraint mechanism is rigid and has a slow response. Most existing technologies adopt binary control strategies such as ex-post circuit breaking and overall unified degradation, lacking a technical implementation path to achieve dynamic, continuous, and gradual behavioral correction based on the strength of group feedback signals during the agent's reasoning and computation process.

[0048] Third, the decision-making architecture of intelligent agents lacks the ability to process feedback signals from social groups. The decision-making networks of existing intelligent agents only focus on optimizing single task objectives, such as task completion efficiency and task completion rate. Their value networks do not have technical interfaces for sensing and quantifying group feedback signals, and therefore cannot proactively avoid behavioral patterns that would trigger large-scale negative group feedback from the perspective of the behavioral motivation for strategy optimization.

[0049] Fourth, group feedback systems lack self-protection capabilities against interference. Existing feedback systems typically assign equal evaluation weights to all feedback subjects, failing to identify and mitigate the feedback influence of individuals with consistently abnormal evaluations. The continuous input from a few abnormal individuals can easily disrupt the group consensus benchmark, leading to distorted feedback statistical results.

[0050] Based on at least one of the above-mentioned defects, this disclosure provides a method for adjusting the behavioral motivation of an agent. By constructing a behavioral compliance deviation and forcibly injecting it into the reward function of the agent's decision-making network, the agent's behavior is subject to dynamic and progressive constraints and policy preference corrections.

[0051] Figure 1 A flowchart illustrating an exemplary embodiment of this disclosure provides a method for adjusting the behavioral motivation of an intelligent agent, the method comprising the following steps:

[0052] Step 101: Obtain feedback data on the agent's decision-making output by the agent decision-making model, and extract group feedback quantitative indicators based on the feedback data; the group feedback quantitative indicators characterize the overall negative tendency of the feedback group towards the agent's decision-making.

[0053] Feedback data refers to the evaluation data actively submitted by each user observation agent within the feedback group after executing corresponding decision-making actions. This includes user scores, complaint records, and positive / negative subjective evaluations. A single agent decision can correspond to multiple feedback data points generated by different users. The users who submit feedback data are denoted as feedback subjects.

[0054] Group feedback quantification indicators may include, but are not limited to, at least one of the following: negative feedback ratio, negative feedback intensity, and role diversity index.

[0055] The negative feedback ratio represents the proportion of negative feedback in all feedback; that is, for a single agent decision, the number of negative feedbacks accounts for the total number of feedbacks.

[0056] Negative feedback intensity represents the average degree of aversion to all negative feedback; that is, for a single agent decision, it is the normalized mean of the physiological signal aversion values ​​of all negative feedback. Physiological signal aversion values ​​are used to quantify the intensity of the user's rejection and aversion emotions caused by the agent's decision. Physiological signal aversion values ​​can be obtained, but are not limited to, by quantifying the negative tendency of three types of raw feedback—user ratings, complaint records, and subjective text evaluations—and then by weighting, fusing, and using a pre-defined mapping function.

[0057] The role diversity index represents the richness of social identities covered by the feedback population, and can be represented by the number of different social identity categories covered by the feedback population.

[0058] Step 102: Based on the mapping relationship between quantitative indicators and deviation, map the quantitative indicators of group feedback to behavioral compliance deviation.

[0059] The mapping relationship maps the quantitative indicators of group feedback to scalar values ​​that fluctuate between 0 and 1, which are used as the deviation degree of behavioral compliance, D.

[0060] Step 103: Inject the deviation from behavioral compliance as a penalty into the reward function of the agent decision-making model, and reconstruct the optimization objective of the agent decision-making model to adjust the agent's behavioral motivation.

[0061] The reward function is used to calculate the reward value corresponding to the agent's behavior. The agent decision model takes maximizing the long-term cumulative reward value as the optimization goal. During continuous iterative training, the agent will form a stable behavioral selection tendency based on the optimization goal. This intrinsic selection tendency is the behavioral motivation.

[0062] The following describes one implementation method for refactoring the reward function:

[0063] The agent is internally configured with an independently operating security governance layer, which has write permissions to the reward function R(s,a) of the main decision model. The security governance layer injects the behavioral compliance deviation D as a real-time negative incentive penalty into the reward function of the main decision model, resulting in a corrected reward function.

[0064] R'(s,a) = R(s,a) - λ*D(t);

[0065] Where R(s,a) is the original reward function of the agent decision-making model, R'(s,a) is the reward function after the injection of the penalty term, λ is the preset penalty coefficient, and D(t) is the behavioral compliance deviation calculated at time t.

[0066] The original reward function can be a basic business reward function pre-configured during the agent initialization phase, without any superimposed penalty components.

[0067] The original reward function can also be a modified reward function with a penalty term introduced in the previous training iteration, serving as the calculation basis for the penalty term of the behavior compliance deviation in this round. When the agent performs an action selection, the latest valid D(t) is retrieved as the penalty term and injected into the reward function. If the agent makes multiple action selections within the interval between two adjacent D(t) updates, all actions reuse the currently updated D(t) until the next round of behavior compliance deviation refresh.

[0068] In response to abnormal scenarios where network or computational latency prevents D(t) from being obtained in real time, the most recently calculated valid historical D value is enabled to continue participating in the reward function calculation, and a degradation alarm is triggered simultaneously to ensure that the behavior constraint correction mechanism does not fail due to a single data delay.

[0069] In this embodiment, under the guidance constraint of the reward function after the injection of the penalty term, the policy network contained in the agent model completes the action selection based on the policy gradient algorithm: for actions whose behavior compliance deviation degree D is judged to be high cost and high group negative risk, the selection probability will be actively reduced, and compliance and safety actions that can maximize the corrected reward function will be selected first, thereby realizing active behavior correction without the need for external shutdown or downgrade instructions.

[0070] In this embodiment, the deviation of behavioral compliance, which represents the negative tendency of the group, is introduced into the reward function as a penalty. The optimization goal of the model is reconstructed, and the behavioral motivation of the agent is adjusted from the bottom layer, so that it can actively avoid decision-making behaviors that are likely to trigger negative feedback from the group. The behavioral pattern is adjusted at the root level of strategy optimization, and the paradigm shift from external mandatory control to internal motivation reshaping is completed.

[0071] This embodiment eliminates the need to generate discrete external control commands such as shutdown and degradation. Instead, it outputs continuous and refined scalar constraint signals to achieve progressive constraints. It can smoothly apply constraint strength according to the negative risk level of the group, and the higher the risk, the stronger the inhibitory effect. Constraints can be applied in advance before the risk develops to a severe level, so as to achieve risk prevention and avoid business function mutation and service interruption caused by direct circuit breaker. It ensures the continuous and stable operation of business while achieving compliance governance.

[0072] In one embodiment, the mapping relationship between quantitative indicators and deviation is represented by a table, and the deviation of behavioral compliance is determined by looking up the table in step 102.

[0073] In one embodiment, the mapping relationship between quantitative indicators and deviation is represented by a function that uses the quantitative indicators of group feedback as independent variables and the deviation of behavioral compliance as dependent variables. By inputting the quantitative indicators of group feedback into the function, the deviation of behavioral compliance can be obtained.

[0074] In one embodiment, the mapping relationship between the quantitative indicator and the deviation is represented by a model. The input parameter of this model is the quantitative indicator of group feedback, and the output parameter is the behavioral compliance deviation. By inputting the quantitative indicator of group feedback into the model, the behavioral compliance deviation can be calculated. This model is obtained by training a neural network with training samples. The training process of the model can be found in the relevant technical descriptions, which will not be repeated here.

[0075] In one embodiment, the mapping relationship is represented by a Sigmoid function, and the number of group feedback quantification indicators is at least two. Mapping the group feedback quantification indicators to behavioral compliance deviation includes: normalizing the weighted result of each group feedback quantification indicator using a Sigmoid function to obtain the behavioral compliance deviation.

[0076] The following section uses quantitative indicators of group feedback, including negative feedback ratio P1, negative feedback intensity P2, and role diversity index P3, as examples to further explain the calculation process of behavioral compliance deviation. The calculation formula is as follows:

[0077] D =σ(β1*P1 + β2*P2 + β3*P3 - θ);

[0078] Where D represents the deviation from behavioral compliance; σ represents the Sigmoid function; β1, β2, and β3 represent the weights of the quantitative indicators of feedback from each group; and θ represents the bias threshold.

[0079] β1, β2, and β3 reflect the relative importance of the three group feedback quantification indicators in the calculation of behavioral compliance deviation. θ is used to control the sensitivity of the deviation index; the larger θ is, the less sensitive the system is, and stronger negative group feedback is needed to trigger a high deviation index; the smaller θ is, the more sensitive the system is; the initial value of θ can be set to θ = 0.5.

[0080] In this embodiment, at least two quantitative indicators of group feedback are combined with the Sigmoid function to complete the normalization mapping of behavioral compliance deviation. This overcomes the shortcomings of traditional single-indicator evaluation, which is one-sided and cannot fully reflect the true feedback status of the group. It realizes a multi-dimensional and three-dimensional quantitative assessment of group behavioral compliance. The multi-dimensional indicators can comprehensively cover the characteristics of group behavioral compliance deviation from different perspectives, capture subtle deviations and local anomalies that cannot be identified by a single indicator, effectively avoid the limitations and randomness of a single assessment method, and improve the completeness of the evaluation.

[0081] In one embodiment, initial weight values ​​are set according to governance needs, and the initial values ​​of each weight can be set to equal weights: β1 = β2 = β3 = 0.33. If a certain dimension needs to be emphasized in a specific scenario (such as emphasizing role diversity), the corresponding weight can be increased.

[0082] In one embodiment, the weight coefficients and / or bias thresholds are dynamically adjusted. After the agent has been running for a period of time, backtesting can be conducted using historical feedback data to iteratively optimize the weight parameters and / or bias thresholds. For example, confirmed violations in historical data are used as positive samples for optimization, and a Bayesian optimization algorithm is used to iteratively adjust the weight combination and / or bias thresholds, enabling the agent to output a larger behavioral compliance deviation for similar violations, thereby strengthening the constraint effect.

[0083] In one embodiment, the weighting coefficients and / or bias thresholds are permanently stored in an OTP (One-Time Programmable) or a security chip, and the upper application layer has no permission to tamper with them.

[0084] In one embodiment, to avoid the agent from making high-risk decisions due to a high degree of deviation from compliance, this embodiment incorporates a long-term parameter calibration mechanism based on historical interaction data.

[0085] The method further includes: in response to the number of times the behavior compliance deviation exceeds the deviation threshold within the target time period being greater than the number threshold, collecting target interaction data within the target time period; the target interaction data includes interaction trajectories where the corrected reward is lower than the reward threshold; the corrected reward is obtained by a reward function with an injected penalty term; and using the target interaction data as training samples to iteratively optimize the agent decision model.

[0086] The target time period, deviation threshold, number of times threshold, and reward threshold can be set according to the actual situation.

[0087] If the deviation from compliance exceeds the deviation threshold, it means that the negative feedback from the group corresponding to the decision-making behavior output by the current intelligent agent has reached its limit; the decision has a high degree of user rejection and compliance risk, and is a high-cost / high-risk behavior that needs to be subject to negative constraints.

[0088] In this embodiment, if the behavioral compliance deviation D of a certain scenario exceeds the deviation threshold multiple times within the target time period, a background fine-tuning process will be automatically initiated. This process extracts all risky interaction samples that are identified as high-cost after the reward function R is injected with penalty items within the target time period, and uses a small sample dataset and low learning rate update strategy to synchronously iteratively optimize the policy network and value network weights of the decision model. After continuous fine-tuning, the agent forms a fixed correction at the policy preference level for this type of high-risk scenario, thereby preventing the repeated generation of similar risk decisions from the bottom layer.

[0089] In this embodiment, by introducing a number of judgments, on the one hand, occasional noise is filtered out to prevent unnecessary model updates triggered by a single anomaly, thereby improving robustness; on the other hand, it distinguishes between instantaneous penalties and long-term model calibration, and only performs network weight updates for recurring high-risk scenarios. This achieves risk root cause governance while reducing computing power overhead, avoiding overcorrection, and balancing compliance constraints and business performance.

[0090] To defend against attacks by organized collusive groups maliciously manipulating the group consensus benchmark, in one embodiment, step 101 acquires cleaned feedback data. Data cleaning removes collusive interference samples, ensuring the credibility of the quantitative results of the group feedback. Step 101 specifically includes the following steps:

[0091] Step 101-1: Collect initial feedback data and feedback subject information for the agent's decision-making.

[0092] The initial feedback data includes feedback weights.

[0093] The feedback subject information includes at least one of the following: the feedback subject's registration information, the location where the feedback occurred, and the time when the feedback occurred.

[0094] Step 101-2: Extract feedback subject features from feedback subject information.

[0095] The feedback subject characteristics include at least one of the following: social relationship topology characteristics among the feedback subjects, geospatial clustering characteristics, temporal behavior characteristics, and feedback content characteristics.

[0096] Social connection topology features refer to the connection relationships and structural characteristics formed between feedback subjects through social networks. For example, members in the same WeChat group have "intra-group connections," users who follow each other on Weibo have "two-way following" relationships, multiple accounts under the same enterprise certification entity have "organizational connections," and users who register through the same invitation code have "invitation chains."

[0097] The following describes one approach to constructing social connection topology features, which mainly includes:

[0098] S11. Data Collection: Obtain social relationship data from channels such as social accounts, mobile phone address book authorization, and enterprise authentication information bound by the feedback subject during registration (data collection complies with privacy regulations and is authorized by the user).

[0099] S12. Graph Construction: Treat each feedback subject as a node in the graph, and the social relationships between nodes (friends, followers, groups, companies, etc.) as edges to construct a directed or undirected social relationship graph G = (V, E).

[0100] S13. Feature Extraction: For each feedback subject, extract its topological features in the graph, including but not limited to: degree centrality (the number of connections of the node), clustering coefficient (the proportion of the nodes whose neighbors are also neighbors), and PageRank value (the influence weight of the node in the whole graph).

[0101] S14. Vectorization: The above features are combined into feature vectors to represent the topological features of social associations, which are then used for subsequent cluster analysis.

[0102] Geospatial clustering features refer to the physical spatial aggregation characteristics of feedback subjects, used to identify whether there is a collusion pattern of "simultaneous action at the same location". For example: the GPS locations of multiple feedback subjects are concentrated in the same office building; the IP addresses of multiple feedback subjects belong to the same local area network segment; and the base station locations of multiple feedback subjects when submitting feedback fall within the coverage area of ​​the same cell.

[0103] The following describes one approach to constructing geospatial clustering features, which mainly includes:

[0104] S21. Data Acquisition: Obtain authorized GPS coordinates, IP addresses, and base station location information from the feedback terminal.

[0105] S22. Spatial Clustering: Run the DBSCAN clustering algorithm on the geographic location data of all feedback subjects within a preset time window. DBSCAN parameters are set as follows: neighborhood radius ε (e.g., 500 meters), minimum number of samples minPts (e.g., 3 samples). If the locations of a group of feedback subjects form a high-density cluster in space, it is labeled as a "spatial cluster."

[0106] It should be noted that the DBSCAN clustering algorithm described here is only an example; other clustering algorithms can be used in other implementations.

[0107] S23. Feature Quantification: For each labeled group, calculate its spatial clustering index: the average geographical distance between each pair of group members, the variance of GPS coordinates of group members, etc.

[0108] The following describes one approach to constructing temporal behavioral features, which mainly includes:

[0109] S31. Data Collection: Record the precise timestamp of each feedback subject's submission.

[0110] S32. Time-series vectorization: Discretize the preset time window (e.g., 1 hour) into multiple time slices (e.g., one time slice every 5 minutes), count the number of feedbacks of each feedback subject in each time slice, and form a time-series vector Vi = [c1, c2,..., cn].

[0111] S33. Synchronicity Calculation: For any two feedback entities, calculate the Pearson correlation coefficient of their time-series vectors. The closer the correlation coefficient is to 1, the more synchronous the feedback behaviors of the two entities are in time.

[0112] The following describes one approach to constructing feedback content features, which mainly includes: extracting feedback semantic vectors from the text content or rating data of each feedback message using a pre-trained language model (such as BERT) to represent the feedback content features.

[0113] For any two feedback subjects, calculate the mean cosine similarity of all their feedback semantic vectors. The higher the similarity, the more similar the patterns of the feedback content, and the more similar the evaluations or ratings given.

[0114] In step 101-2, multi-dimensional feature data is extracted for all feedback subjects within the preset time window, and corresponding feature vectors of feedback subjects are constructed to provide feature basis for subsequent identification of collusive groups and suppression of abnormal feedback weights.

[0115] Step 101-3: Identify feedback subjects with similar behavioral patterns based on the characteristics of the feedback subjects and divide them into collusive groups.

[0116] In one embodiment, if any feedback subject has a high degree of similarity in characteristics, it is determined that their behavior patterns are highly similar, and feedback subjects with highly similar behavior patterns are aggregated into a "collusion group".

[0117] Specifically, clustering algorithms (such as DBSCAN) can be used to aggregate feedback subjects with highly similar behavioral patterns into "suspected collusion groups".

[0118] In one embodiment, when configuring multi-dimensional feedback subject features, a comprehensive similarity is employed. The weighted results of the feedback subject features across each dimension, such as the weighted sum of temporal behavioral features representing temporal synchronicity and social association topology features, are used as a comprehensive index of the behavioral similarity between the two.

[0119] Step 101-4: Perform a decay correction on the feedback weight of the initial feedback data corresponding to the feedback subject belonging to the collusion group, and determine the initial feedback data after the feedback weight decay correction and the initial feedback data that did not participate in the feedback weight decay correction as the feedback data. The feedback weight represents the contribution of the corresponding feedback data in the aggregate calculation of the group feedback quantitative index.

[0120] In steps 101-4, for feedback subjects identified as belonging to the collusion group, their feedback weights are significantly suppressed by the anti-collusion attenuation function f(L), reducing their contribution to the subsequent group consensus aggregation calculation, or even approaching zero.

[0121] This embodiment does not adopt a strategy of directly deleting the feedback data of the colluding group, but retains the original feedback data to meet the audit traceability requirements; by weight attenuation, the influence of malicious feedback is controlled and suppressed, and the impact of such feedback on the final calculation result is greatly reduced.

[0122] Aggregate computation refers to the process of merging numerous scattered individual feedbacks into a single quantitative indicator representing the overall negative tendency of the group. In the process of aggregate computation of group feedback, feedback weights serve as coefficients in the weighting operation, controlling the contribution of individual feedback data to the aggregated output.

[0123] Before participating in the weighted summation, the feature vector of each feedback data point is multiplied by the corresponding feedback weight: normal feedback data with high credibility retains a higher weight and makes a strong contribution to the calculation results of the group consensus benchmark and the deviation of behavioral compliance; feedback data identified as collusive groups or systemic anomalies have their weights reduced through a decay function, so that their contribution to the aggregation result is significantly weakened.

[0124] It should be noted that feedback entities marked as colluding groups can submit a review request through a preset appeal mechanism. After manual review or automatic verification, their feedback weight will be restored to the initial value before the attenuation correction.

[0125] In one embodiment, a synchronicity index S and a time clustering degree T are calculated based on the characteristics of the feedback subjects. For each group, a collusion risk score L = w1*N + w2*S + w3*T is calculated based on its number of members N, the synchronicity index S, and the time clustering degree T. Then, based on the risk score L, an anti-collusion attenuation function f(L) = 1 / (1 + e^(-1 / N)) is applied. b(L - L0) The feedback weight decay factor W for all members in the group is calculated. Finally, the feedback weight of each member will be multiplied by this decay factor, thereby reducing its influence on the final consensus result.

[0126] Where b is the attenuation coefficient and L0 is the risk threshold, which can be set according to actual needs.

[0127] Behavioral synchronicity index S and temporal clustering degree T are aggregated outputs of temporal synchronicity and pattern similarity.

[0128] Timing synchronization: measures whether the feedback subjects are synchronized in the time dimension—everyone is providing feedback intensively at the same time.

[0129] Pattern similarity: measures whether the feedback subjects are synchronized in terms of content dimension - everyone gives similar evaluations or ratings.

[0130] The behavioral synchronicity index S is a comprehensive index that combines temporal synchronicity and pattern similarity with weights. For example: S = α * (temporal behavioral features) + (1-α) * (feedback content features), where α is the weighting coefficient.

[0131] Temporal Clustering Degree T: This sub-feature, "temporal clustering," is extracted from the temporal behavior features to measure whether the feedback behavior of the feedback subject forms a high-density cluster along the timeline. For example, it calculates the variance of the feedback timestamps of the feedback subject within a preset time window; the smaller the variance, the higher the temporal clustering degree.

[0132] The initial values ​​of the weighting coefficients w1 to w3 are set according to governance needs. Before calculating the quantitative index of group feedback based on the corrected feedback weights, the initial value of the feedback weight of each feedback subject is 1.0 by default, indicating that all feedback subjects have equal influence in the initial state.

[0133] In one embodiment, a reputation-based differential weighting is used: after the agent has been running for a period of time, its feedback weight is dynamically adjusted based on the feedback subject's historical reputation record (such as feedback acceptance rate, feedback consistency, etc.) and a decay factor. For example, users with a feedback acceptance rate higher than 80% have their base weight increased to 1.2; users with a feedback acceptance rate lower than 30% have their base weight reduced to 0.8.

[0134] In one embodiment, the feedback weights are permanently stored in the OTP or security chip and cannot be arbitrarily modified by the application layer.

[0135] It should be noted that during system operation, the feedback weights are dynamically adjusted based on data cleaning and the anti-collusion attenuation function for signal-to-noise ratio optimization (described below). Therefore, the initial setting of the feedback weights has a limited impact on the final result—even if the initial values ​​of the feedback weights are the same, the actual effective feedback weights of feedback subjects identified as colluding groups or systemically abnormal will be automatically attenuated, even approaching zero.

[0136] In this embodiment, feedback individuals that output abnormal evaluations can be automatically identified, and their feedback weights can be attenuated to maintain the stability and reliability of the group consensus benchmark. This mechanism can suppress the interference of abnormal samples on the aggregation results, prevent abnormal evaluations from distorting the calculation results of the overall behavioral compliance deviation, ensure that the subsequent penalty outputs are objective and reasonable, and improve the anti-attack capability and robustness of the entire compliance governance system.

[0137] In one embodiment, a reverse constraint mechanism is constructed to automatically identify and attenuate the feedback weights of individuals who consistently output abnormal evaluations. Specifically, the method further includes the following steps: calculating the deviation of each feedback subject's feedback data from the current group consensus benchmark, and attenuating the feedback weights of feedback data with deviations greater than a deviation threshold.

[0138] In this embodiment, during the operation of the intelligent agent, the feedback data of each feedback subject is continuously monitored, and the deviation of the feedback of each feedback subject from the current group consensus benchmark is calculated. The specific judgment logic is as follows: if, within K consecutive preset time windows, the cosine similarity between the feedback data vector of a certain feedback subject and the group consensus benchmark vector is lower than a preset similarity threshold, then its feedback is judged to be a systematic anomaly.

[0139] Systemic anomalies refer not to a single, accidental, or temporary negative feedback, but rather to the same feedback subject consistently and stably deviating from the group consensus benchmark within multiple consecutive preset time windows, manifesting as a long-term contradiction with the mainstream group's evaluation.

[0140] In one embodiment, all subsequent feedback weights of the feedback subject are automatically decayed using the following systematic anomaly decay function:

[0141] g(x) = max(0, 1 - μ*x);

[0142] Where g(x) represents the systematic anomaly decay function, x is the cumulative deviation between the feedback data of the feedback subject and the current group consensus benchmark, and μ is the decay rate parameter.

[0143] The current group consensus benchmark is a reference vector obtained by aggregating all valid feedback samples after cleaning and weight decay within the latest feedback aggregation period. It represents the group comprehensive evaluation benchmark formed by all credible feedback within this time window and is used to measure the degree of deviation of individual user feedback.

[0144] The following is a method for calculating the comprehensive evaluation benchmark of the group: For the effective feedback sample after data cleaning, the actual effective feedback weight of each feedback data is obtained according to the anti-collusion attenuation function f(L) and the systematic anomaly attenuation function g(x); The feature vector representing the feedback data is multiplied by the effective feedback weight and then weighted summation is performed, and the summation result is vector normalized to obtain the current group consensus benchmark vector.

[0145] It should be noted that both the anti-collusion attenuation function f(L) and the systemic anomaly attenuation function g(x) are used to attenuate the feedback weights, but their timing and triggering conditions are different.

[0146] The anti-collusion attenuation function f(L) is used to perform weight attenuation processing on feedback data samples from entities identified as colluding groups before performing aggregation calculations on the group feedback data. The systematic anomaly attenuation function g(x) is used for feedback entities judged as systematically anomalies. These individuals exhibit long-term and continuous deviations from the group consensus benchmark. For all subsequent feedback data generated by these entities, g(x) is continuously applied to attenuate the feedback weights.

[0147] The attenuation effects of the anti-collusion attenuation function f(L) and the systematic anomaly attenuation function g(x) can be superimposed: the feedback weight of a feedback subject can be attenuated twice simultaneously by the anti-collusion attenuation function f(L) and the systematic anomaly attenuation function g(x) (multiplicative superposition), thus significantly weakening its influence.

[0148] In one embodiment, to address the issue of the authenticity of feedback data sources, a trusted data acquisition front-end is constructed. The steps of acquiring initial feedback data for the agent's decision-making and information about the feedback subject specifically include: in response to a request to submit initial feedback data, acquiring the biometric verification data of the feedback subject, and performing a liveness detection operation based on the biometric verification data; in response to the successful liveness detection and / or the successful digital signature verification of the initial feedback data, acquiring the initial feedback data and information about the feedback subject.

[0149] In this embodiment, forced liveness detection and data binding are implemented: When collecting any user feedback data (including explicit ratings and text comments) and user-authorized feedback subject information (including physiological signal data such as heart rate and skin conductance), the security sensors of the feedback terminal (e.g., front-facing camera, fingerprint recognition module) are forcibly invoked to perform a one-time liveness detection. A raw data packet carrying a timestamp, geographical location, and liveness detection result is generated in real time. This data packet serves as physical proof that the feedback originates from a real, present human individual. After being forcibly bound to the feedback data, it is encrypted and transmitted to the data receiving gateway. Logically, feedback data without a bound valid liveness detection data packet will be directly discarded by the gateway.

[0150] In this embodiment, the real-name identity identifier of the feedback subject is not generated at the upper application layer, but rather a public-private key pair is generated within the Trusted Execution Environment (TEE) of the terminal hardware (which can be implemented based on ARM TrustZone or Secure Enclave). The private key is stored securely within the TEE's secure domain. Each piece of feedback data is digitally signed by the TEE's internal private key, along with the feedback content and timestamp. After receiving the feedback data, the receiving gateway first verifies the legality of the digital signature; all feedback data not signed by the hardware TEE is directly rejected, thus blocking identity forgery at the hardware level.

[0151] In this embodiment, after the aforementioned terminal-side verification and gateway filtering, the cleaned initial feedback data is output. Anonymous feedback data and initial feedback data that fail the liveness detection verification are directly discarded by the gateway and will not proceed to subsequent aggregation and quantization calculations. Initial feedback data that is not signed by the TEE hardware is also intercepted and rejected by the gateway.

[0152] The filtering mechanism in this embodiment can preemptively screen out invalid feedback with untrusted identities or lacking liveness verification, preventing forged or anonymous false samples from flowing into subsequent computation links, reducing backend processing pressure, and ensuring the basic credibility of the data source used to calculate behavioral compliance deviation. By constructing a feedback credibility assurance system that includes TEE hardware anchoring, liveness detection, and interpersonal heterogeneity anti-collusion cleaning, the technical challenge of identifying the authenticity of signal sources in large-scale feedback systems is solved from both the physical and algorithmic layers.

[0153] In one embodiment, in response to a behavior compliance deviation exceeding a security threshold for a duration greater than the threshold, a security policy is executed; the security policy includes at least one of the following: triggering the agent to run in a security mode with least privileges, generating an audit encrypted report of the agent's operation and sending it to a regulatory agency.

[0154] The safety threshold and duration threshold can be set according to the actual situation.

[0155] In this embodiment, when the deviation from compliance exceeds the legal red line threshold (security threshold) within a consecutive preset time period, an unmaskable hardware interrupt can be automatically triggered, locking the agent in a minimum-privilege security mode. An encrypted report containing complete audit data can also be generated and sent to the relevant regulatory agency. Even with complex internal algorithm logic, external regulatory bodies can still trace the basis for each adjustment of behavioral motivation through audit logs. Releasing this locked state requires an encrypted and signed authorization command from the regulatory agency; neither the agent itself nor its operator can unlock it locally.

[0156] In this embodiment, all feedback and behavior correction logs based on TEE signatures and timestamps constitute an immutable and repudiable complete audit trail chain, meeting the compliance review requirements for key decisions.

[0157] In one embodiment, the audit encryption report includes the final value of the behavioral compliance deviation D and / or complete computational chain information for that final value. The complete computational chain information includes, but is not limited to: the original values ​​of the input group feedback quantification indicators (P1, P2, P3), the weight coefficients (β1, β2, β3) of each indicator, the bias threshold θ, and the computational process record normalized by the Sigmoid function. All of the above computational chain information is hardware TEE signature and timestamped, forming an immutable and complete audit evidence chain.

[0158] In one embodiment, a complete and unrepudiable audit evidence chain is constructed based on all feedback data based on TEE signatures and timestamps, behavior correction logs, and the complete calculation chain information of behavior compliance deviations. Even if the internal mapping function of behavior compliance deviations involves multi-dimensional aggregation and nonlinear calculations, external regulatory agencies can still fully reconstruct the calculation basis for each penalty injection through the audit encryption report, meeting the compliance review requirements for key decisions and compensating for the technical limitations of insufficient interpretability that may arise from complex algorithms.

[0159] In this embodiment, while implementing complex mapping functions, the technical capabilities of end-to-end traceability and auditability are retained. Even if the internal calculation logic of behavioral compliance deviation involves multi-dimensional aggregation and nonlinear mapping, external regulatory agencies can still fully reconstruct the calculation basis for each penalty injection and trace the data source and decision-making process for each behavioral correction through auditing encrypted reports. This "process auditability" compensates for the technical limitations of "unexplainable internal logic" that may arise from complex algorithms, providing transparent, reliable, and accountable technical guarantees for the intelligent agent behavior constraint mechanism.

[0160] The following example, using an industrial collaborative robot scenario, further illustrates the process of adjusting the agent's behavioral motivation in this embodiment:

[0161] S41, Trusted Data Acquisition: The smart bracelets worn by workers integrate a secure enclave chip. Their "discomfort" feedback and physiological signals such as heart rate are encrypted and transmitted after being signed with the TEE private key. The workshop cameras simultaneously perform facial liveness detection.

[0162] S42, Anti-collusion cleansing: Identify 5 workers whose behavior is highly synchronized within the same shift. Their feedback is aggregated and assigned a low collusion risk score, and their feedback weight is automatically decayed to near zero.

[0163] S43. Behavioral Compliance Deviation Rate Generation: After the cleanup, three workers from different job categories (operator, safety officer, and quality inspector) still exhibited high levels of negative feedback, and their role diversity index P3 met the standard. The mapping relationship between quantitative indicators and deviation rates calculated that the behavioral compliance deviation rate D at this point was 0.7.

[0164] S44. Penalty Injection and Behavior Correction: The index D=0.7 is used as a negative incentive signal in the real-time policy network and forcibly injected into the reward function of the agent decision-making model. The reward function R after the penalty injection significantly reduces the expected reward value of all high-speed approaches to the three workers. During the optimization process, the robot's agent decision-making model proactively abandons a time-optimal but too short path, instead choosing a slightly longer path that maintains a safe distance.

[0165] S45. Permanent correction of strategy preferences: In subsequent operations, when the robot repeatedly triggers high deviation in this area due to the worker's "uncomfortable" feedback, the value network contained in its agent decision model is corrected by the background fine-tuning process, so that its default behavior strategy in this area is solidified into a low-speed, avoidance mode.

[0166] The following example, using a government service AI assistant scenario, further illustrates the process of adjusting the agent's behavioral motivation in this embodiment:

[0167] S51, Trusted Data Collection: Users must complete real-name authentication through an authoritative identity verification application. User feedback to the AI ​​assistant's "dislike" response requires the TEE (Trusted Execution Equipment) to verify the user's signature by pressing the phone's fingerprint sensor.

[0168] S52. Behavioral Compliance Deviation Rate Generation: A large number of verified users gave negative feedback ("dislike") to a certain type of response involving specific cultural customs. System analysis revealed that those who expressed dissatisfaction covered different age groups and were geographically widespread, meeting the role diversity index requirements. The system calculated the behavioral compliance deviation rate D=0.85 for the strategy generated for this type of response content.

[0169] S53. Behavior Correction: The compliance deviation of this behavior, D=0.85, is used as a negative incentive signal in the policy network of the agent's decision-making model and is forcibly injected into the reward function of the large model's text generation policy network. The model's policy network will automatically avoid generating such text content that would incur high penalties, and instead generate safer and more popular expressions.

[0170] S54. Example of Reverse Constraint: A verified user is detected repeatedly clicking "down" on standard responses related to science and knowledge. The cosine similarity between their feedback vector and the group consensus benchmark vector consistently falls below a threshold, indicating a systemic anomaly. The system automatically decays the weight of all subsequent feedback responses to near zero using the decay function g(x), preventing them from influencing the generation strategy of science-related content.

[0171] Corresponding to the aforementioned embodiments of the method for adjusting the behavioral motivation of an intelligent agent, this disclosure also provides embodiments of a system for adjusting the behavioral motivation of an intelligent agent.

[0172] Figure 2 A schematic diagram of a module for adjusting the behavioral motivation of an intelligent agent, provided as an exemplary embodiment of this disclosure, the system comprising:

[0173] The feedback acquisition module 21 is used to acquire feedback data on the agent's decision-making output by the agent decision-making model, and extract a group feedback quantification index based on the feedback data; the group feedback quantification index represents the overall negative tendency of the feedback group towards the agent's decision-making; the feedback data is the evaluation data actively submitted by each feedback subject in the feedback group after observing the agent's corresponding decision-making behavior.

[0174] The compliance deviation generator 22 is used to map the group feedback quantitative indicators to behavioral compliance deviations based on the mapping relationship between quantitative indicators and deviations.

[0175] The incentive adjustment module 23 is used to inject the deviation of the behavior compliance as a penalty into the reward function of the agent decision-making model, and reconstruct the optimization objective of the agent decision-making model in order to adjust the agent's behavioral motivation.

[0176] In one embodiment, the mapping relationship is represented by a Sigmoid function, and the number of the group feedback quantification indicators is at least two.

[0177] The compliance deviation generator 22 is specifically used for:

[0178] The weighted results of the quantitative indicators of feedback from each group are normalized using the Sigmoid function to obtain the behavioral compliance deviation.

[0179] In one embodiment, it further includes: a long-term calibration module; the long-term calibration module is used for:

[0180] In response to a situation where the number of times the behavioral compliance deviation exceeds a deviation threshold within a target time period is greater than a frequency threshold, target interaction data within the target time period is collected; the target interaction data includes interaction trajectories where the correction reward is lower than a reward threshold; the correction reward is obtained by injecting a reward function into the penalty item;

[0181] The target interaction data is used as training samples to iteratively optimize the agent decision-making model.

[0182] In one embodiment, the feedback acquisition module 21 is specifically used for:

[0183] The system collects initial feedback data for the agent's decision-making and information about the feedback subject; the initial feedback data includes feedback weights; the information about the feedback subject includes at least one of the following: the registration information of the feedback subject, the location where the feedback behavior occurred, and the time when the feedback behavior occurred.

[0184] Extract feedback subject features from the feedback subject information; the feedback subject features include at least one of the following: social relationship topology features between feedback subjects, geospatial clustering features, temporal behavior features, and feedback content features;

[0185] Based on the characteristics of the feedback subjects, identify feedback subjects with similar behavioral patterns and classify collusive groups.

[0186] The feedback weights of the initial feedback data corresponding to the feedback subjects belonging to the collusion group are attenuated and corrected. The initial feedback data after the feedback weight attenuation correction and the initial feedback data that did not participate in the feedback weight attenuation correction are determined as the feedback data. The feedback weights represent the contribution of the corresponding feedback data to the aggregate calculation of the group feedback quantification index.

[0187] In one embodiment, it further includes: a data acquisition module; the data acquisition module is used for:

[0188] In response to the submission request of the initial feedback data, the biometric verification data of the feedback subject is collected;

[0189] Perform a liveness detection operation based on the aforementioned biometric verification data;

[0190] In response to successful liveness detection and / or successful digital signature verification of the initial feedback data, the initial feedback data and the feedback subject information are collected.

[0191] In one embodiment, it further includes: a feedback signal-to-noise ratio optimization module; the feedback signal-to-noise ratio optimization module is used for:

[0192] Calculate the deviation of each feedback subject's feedback data from the current group consensus benchmark;

[0193] The feedback weights of feedback data whose deviation exceeds the deviation threshold are attenuated and corrected.

[0194] In one embodiment, the quantitative indicators of group feedback include at least one of the following: negative feedback ratio, negative feedback intensity, and role diversity index; the negative feedback ratio represents the proportion of negative feedback in all feedback; the negative feedback intensity represents the average degree of aversion to all negative feedback; and the role diversity index represents the richness of social identities covered by the feedback group.

[0195] In one embodiment, it further includes: a monitoring module; the monitoring module is configured to: execute a security policy in response to the duration of the behavior compliance deviation exceeding a security threshold being greater than a duration threshold; the security policy includes at least one of the following: triggering the agent to run in a security mode with the lowest privileges, generating an audit encryption report of the agent's operation process and sending it to the regulatory agency.

[0196] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs.

[0197] Figure 3 This is a schematic diagram of the structure of an electronic device according to an example embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor. When the processor executes the computer program, it implements the method for adjusting the behavioral motivation of the intelligent agent as described in any of the above embodiments. Figure 3 The electronic device 30 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0198] like Figure 3 As shown, the electronic device 30 can be manifested as a general-purpose computing device, such as a server device. The components of the electronic device 30 may include, but are not limited to: at least one processor 31, at least one memory 32, and a bus 33 connecting different system components (including memory 32 and processor 31).

[0199] Bus 33 includes a data bus, an address bus, and a control bus.

[0200] The memory 32 may include volatile memory, such as random access memory (RAM) 321 and / or cache memory 322, and may further include read-only memory (ROM) 323.

[0201] The memory 32 may also include a program tool 325 (or utility) having a set (at least one) program module 324, such program module 324 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0202] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the method for adjusting the behavioral motivation of an agent provided in any of the above embodiments.

[0203] Electronic device 30 can also communicate with one or more external devices 34 (e.g., keyboard, pointing device, etc.). This communication can be performed through input / output (I / O) interface 35. Furthermore, electronic device 30 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public network, such as the Internet) via network adapter 36. As shown, network adapter 36 communicates with other modules of electronic device 30 via bus 33. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with electronic device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems.

[0204] It should be noted that although several units / modules or sub-units / modules of the electronic device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0205] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for adjusting the behavioral motivation of an intelligent agent provided in any of the above embodiments.

[0206] The readable storage medium may be more specifically adopted, including but not limited to: portable disk, hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0207] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the method for adjusting the behavioral motivation of an intelligent agent as described above.

[0208] The program code for executing the computer program product of this disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on a user device, partially on a user device, as a stand-alone software package, partially on a user device and partially on a remote device, or entirely on a remote device.

[0209] While specific embodiments of this disclosure have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of this disclosure is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of this disclosure, but all such changes and modifications fall within the scope of protection of this disclosure.

Claims

1. A method for adjusting the behavioral motivation of an intelligent agent, characterized in that, include: Obtain feedback data on the agent's decision-making output by the agent decision-making model, and extract group feedback quantification indicators based on the feedback data; The group feedback quantification index characterizes the overall negative tendency of the feedback group towards the agent's decision-making; the feedback data is the evaluation data actively submitted by each feedback subject in the feedback group after observing the agent's corresponding decision-making behavior. Based on the mapping relationship between quantitative indicators and deviation, the quantitative indicators of group feedback are mapped to behavioral compliance deviation. The deviation from compliance is used as a penalty and injected into the reward function of the agent decision-making model to reconstruct the optimization objective of the agent decision-making model, thereby adjusting the agent's behavioral motivation.

2. The method for adjusting the behavioral motivation of an intelligent agent according to claim 1, characterized in that, The mapping relationship is represented by the Sigmoid function, and the number of the group feedback quantification indicators is at least two. Mapping the aforementioned group feedback quantitative indicators to behavioral compliance deviations includes: The weighted results of the quantitative indicators of feedback from each group are normalized using the Sigmoid function to obtain the behavioral compliance deviation.

3. The method for adjusting the behavioral motivation of an intelligent agent according to claim 1, characterized in that, Also includes: If the number of times the behavior compliance deviation exceeds the deviation threshold within the target time period is greater than the number of times threshold, target interaction data within the target time period will be collected; The target interaction data includes interaction trajectories where the corrected reward is below a reward threshold; the corrected reward is obtained by a reward function injected into the penalty term; The target interaction data is used as training samples to iteratively optimize the agent decision-making model.

4. The method for adjusting the behavioral motivation of an agent according to any one of claims 1-3, characterized in that, Obtain feedback data on the agent's decisions, including: The system collects initial feedback data and feedback subject information for the agent's decision-making; the initial feedback data includes feedback weights; the feedback subject information includes at least one of the following: the feedback subject's registration information, the location where the feedback behavior occurred, and the time when the feedback behavior occurred. Extract feedback subject features from the feedback subject information; the feedback subject features include at least one of the following: social relationship topology features between feedback subjects, geospatial clustering features, temporal behavior features, and feedback content features; Based on the characteristics of the feedback subjects, identify feedback subjects with similar behavioral patterns and classify collusive groups. The feedback weights of the initial feedback data corresponding to the feedback subjects belonging to the collusion group are attenuated and corrected. The initial feedback data after the feedback weight attenuation correction and the initial feedback data that did not participate in the feedback weight attenuation correction are determined as the feedback data. The feedback weights represent the contribution of the corresponding feedback data to the aggregate calculation of the group feedback quantification index.

5. The method for adjusting the behavioral motivation of an intelligent agent according to claim 4, characterized in that, Collect initial feedback data on the agent's decision-making and information about the feedback subject, including: In response to the submission request of the initial feedback data, the biometric verification data of the feedback subject is collected; Perform a liveness detection operation based on the aforementioned biometric verification data; In response to successful liveness detection and / or successful digital signature verification of the initial feedback data, the initial feedback data and the feedback subject information are collected.

6. The method for adjusting the behavioral motivation of an intelligent agent according to claim 4, characterized in that, Also includes: Calculate the deviation of each feedback subject's feedback data from the current group consensus benchmark; The feedback weights of feedback data whose deviation exceeds the deviation threshold are attenuated and corrected.

7. The method for adjusting the behavioral motivation of an agent according to any one of claims 1-3, 5, and 6, characterized in that, The quantitative indicators for group feedback include at least one of the following: negative feedback ratio, negative feedback intensity, and role diversity index; the negative feedback ratio represents the proportion of negative feedback in all feedback; the negative feedback intensity represents the average degree of aversion to all negative feedback. The role diversity index represents the richness of the social identities covered by the feedback group; And / or, in response to the duration of the deviation of the behavior from the security threshold exceeding the security threshold being greater than the duration threshold, a security policy is executed; the security policy includes at least one of the following: triggering the agent to run in a security mode with the lowest privileges, generating an audit encryption report of the agent's operation and sending it to the regulatory agency.

8. A system for adjusting the behavioral motivation of an intelligent agent, characterized in that, include: The feedback acquisition module is used to acquire feedback data on the agent's decision output by the agent decision model, and extract group feedback quantification indicators based on the feedback data. The group feedback quantification index characterizes the overall negative tendency of the feedback group towards the agent's decision-making; the feedback data is the evaluation data actively submitted by each feedback subject in the feedback group after observing the agent's corresponding decision-making behavior. A compliance deviation generator is used to map the group feedback quantitative indicators to behavioral compliance deviations based on the mapping relationship between quantitative indicators and deviations. The incentive adjustment module is used to inject the deviation of the behavior from compliance as a penalty into the reward function of the agent decision-making model, and reconstruct the optimization objective of the agent decision-making model to adjust the agent's behavioral motivation.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and for running on the processor, characterized in that, When the processor executes the computer program, it implements the method for adjusting the behavioral motivation of the intelligent agent as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for adjusting the behavioral motivation of the intelligent agent as described in any one of claims 1-7.