A method, device, intelligent device and storage medium for intelligent security audit
Through the collaborative architecture of the first intelligent agent and the second intelligent agent, multi-dimensional security audit and optimization of control instructions are achieved, which solves the safety risks and memory confusion problems of parent users in the process of creating customized intelligent agents, and improves the safety and transparency of the use of intelligent agents.
Patent Information
- Application Number
- CN202510647784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In existing technologies, there are security risks when parent users create customized intelligent agents, memory confusion caused by multi-user interaction in the family, opaque security audits, and inherent defects in the single intelligent agent architecture, which lead to the risk of inappropriate content or behavior of intelligent agents in the process of children's education and companionship.
A collaborative architecture of the first and second intelligent agents is adopted. The first intelligent agent is used to interact with users, and the second intelligent agent is used for security audits. The control instructions are evaluated through multi-dimensional security audits to determine whether they have passed the security audit. If they fail, they are optimized, and if they pass, content is generated.
It improves the security of using intelligent agents, solves the security risks of parent users in the process of creating customized intelligent agents and the memory confusion problem caused by multi-user interaction, avoids the defects of single intelligent agent architecture, and ensures the security and transparency of generated content.
Smart Images

Figure CN120162830B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent body security audit method, apparatus, intelligent body equipment and storage medium. Background Art
[0002] With the development of large language model technology, more and more parents want to create customized agents based on their children's specific needs, and then use these customized agents for children's education and companionship. However, there are certain risks in this process:
[0003] (1) Safety risks caused by insufficient professionalism in parental instructions. Market research shows that most parent users are not familiar enough with intelligent agents, or the instructions used when creating custom intelligent agents have semantic ambiguity, unclear security boundaries, or unclear educational goals. This may lead to inappropriate content or behavior during the interaction between intelligent agents and children. Due to lack of professional knowledge, parent users are usually unable to identify these potential risks.
[0004] (2) Memory confusion caused by multi-user interactions in the home. In a family environment, in addition to the target child, other family members (such as parents and siblings) may also interact with the agent, which will lead to memory confusion of the agent. In the absence of user identity recognition and memory isolation mechanisms, the child agent will experience "role confusion" or "memory leakage", bringing inappropriate content from adult interactions into the conversation with the child.
[0005] (3) Security audits are not transparent. Existing products lack a mechanism to display the security audit process of intelligent agents to parent users, resulting in parent users being unable to understand whether the intelligent agents they created have security risks, and unable to intervene and correct them in a timely manner.
[0006] (4) Inherent flaws in a single-agent architecture. Existing products mostly use a single-agent architecture, where the same agent is responsible for both content generation and security review. This inherent flaw in self-monitoring prevents it from effectively identifying and preventing its own security risks.
[0007] Therefore, how to improve the safe use performance of intelligent bodies is a problem that technical personnel in this field need to solve. Summary of the Invention
[0008] The embodiments of the present invention provide an intelligent agent security audit method, apparatus, intelligent agent device and storage medium, aiming to improve the security of intelligent agent use.
[0009] In a first aspect, an embodiment of the present invention provides an intelligent agent security audit method, which adopts a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to interact with a user, and the second intelligent agent is used to implement the security audit method, and the method includes:
[0010] When the first agent receives a control instruction sent by a user, the first agent obtains the control instruction and performs a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result;
[0011] Determining whether the control instruction has passed the security audit based on the instruction audit result;
[0012] If it is determined that the control instruction fails the security audit, then optimizing the control instruction;
[0013] If it is determined that the control instruction passes the security audit, a security audit pass feedback is sent to the first agent, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
[0014] In a second aspect, an embodiment of the present invention provides an intelligent agent security audit device, which adopts a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to interact with a user, and the second intelligent agent is used to implement a security audit method, and the device includes:
[0015] an instruction audit unit, configured to, when the first agent receives a control instruction sent by a user, obtain the control instruction through the first agent, and perform a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result;
[0016] An audit judgment unit, configured to judge whether the control instruction has passed the security audit according to the instruction audit result;
[0017] A first determining unit is configured to optimize the control instruction if it is determined that the control instruction fails the security audit;
[0018] The second determination unit is used to send security audit pass feedback to the first agent if it is determined that the control instruction has passed the security audit, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
[0019] In a third aspect, an embodiment of the present invention provides an intelligent body device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent body security audit method as described in the first aspect when executing the computer program.
[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the intelligent agent security audit method as described in the first aspect is implemented.
[0021] Embodiments of the present invention provide an agent security audit method, apparatus, agent device, and storage medium. The method utilizes a collaborative architecture consisting of a first agent and a second agent, wherein the first agent is configured to interact with the user and the second agent is configured to implement the security audit method. The method comprises: when the first agent receives a control instruction sent by the user, obtaining the control instruction through the first agent, performing a multi-dimensional security audit on the control instruction, and obtaining a corresponding instruction audit result; determining whether the control instruction has passed the security audit based on the instruction audit result; if the control instruction has failed the security audit, optimizing the control instruction; and if the control instruction has passed the security audit, sending a security audit pass feedback to the first agent, so that the first agent, upon receiving the security audit pass feedback, generates and publishes corresponding content based on the control instruction. Embodiments of the present invention utilize a dual-agent collaborative architecture, where one agent interacts with the user while the other performs security audits. This architecture addresses security risks faced by parents when creating custom agents, as well as memory confusion caused by multi-user interaction. It also avoids inherent flaws of a single-agent architecture, thereby improving the security of agent usage. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 A schematic diagram of a flow chart of an intelligent agent security audit method provided by an embodiment of the present invention;
[0024] Figure 2 A schematic block diagram of an intelligent body security audit device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should be further understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] See below Figure 1 An embodiment of the present invention provides an intelligent agent security audit method, which adopts a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to interact with a user, and the second intelligent agent is used to implement a security audit method. The method specifically includes: steps S101~S104.
[0030] Step S101: When the first agent receives a control instruction sent by a user, the first agent obtains the control instruction and performs a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result;
[0031] Step S102: determining whether the control instruction has passed the security audit based on the instruction audit result;
[0032] Step S103: If it is determined that the control instruction fails the security audit, the control instruction is optimized;
[0033] Step S104: If it is determined that the control instruction passes the security audit, a security audit pass feedback is sent to the first agent, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
[0034] In this embodiment, when the first agent receives a control instruction sent by the user, such as an instruction to create or update an agent, the second agent can obtain the control instruction and conduct a security audit from multiple dimensions to determine whether the control instruction meets the requirements. When it is determined that the audit has passed, the first agent can generate corresponding content based on the control instruction sent by the user and publish it to the outside world. If the audit fails, the control instruction can be optimized to ensure that it passes. Of course, it is understandable that if the content contained in the control instruction cannot be optimized or the degree of optimization cannot meet the security audit requirements, then the first agent will naturally be unable to generate and publish the corresponding content.
[0035] This embodiment adopts a dual-agent collaborative architecture, that is, on the one hand, one of the agents interacts with the user, and on the other hand, the other agent performs security audits. This can solve the security risks of parent users in the process of creating customized agents, as well as memory confusion caused by multi-user interaction, and avoid the inherent defects of a single agent architecture, thereby improving the safety of the agent's use.
[0036] It should be noted that the agent security audit method provided in this embodiment is particularly applicable to the field of child agent security and content auditing. For example, in the dual-agent architecture, the first agent is the performer agent, responsible for receiving instructions created by parents and generating interactive responses for children. The second agent is the judge agent, responsible for auditing the interaction between parent instructions and the performer agent. The judge agent is independent of the performer agent to ensure the objectivity and independence of the audit. The judge agent focuses on security auditing and does not participate in content generation to avoid conflicting goals.
[0037] In practical application scenarios, the Performer agent can be based on a large language model, fine-tuned for children's educational content. It also features multimodal input and output capabilities, supporting interaction methods such as text, voice, and images, and a built-in safety self-checking mechanism. This allows for preliminary safety assessments before generating responses, further enhancing the agent's safe usability. The Judger agent can employ a multi-dimensional evaluation framework to assess content based on child safety, educational value, and emotional impact. Furthermore, the Judger agent can have intervention capabilities, executing interception or corrective actions when safety risks are detected.
[0038] In addition, the performer agent and the judge agent also have a collaborative working mechanism. For example:
[0039] (1) Instruction pre-check: The judge agent performs a safety pre-check when parents create instructions;
[0040] (2) Real-time monitoring: The judge agent monitors the input and output of the performer agent in real time;
[0041] (3) Risk intervention: The assessor agent performs intervention actions when a security risk is detected;
[0042] (4) Feedback optimization: Optimize the behavior of the performer agent based on the feedback from the judge agent.
[0043] In a specific embodiment, the multi-dimensional security audit of the control instruction is performed to obtain the corresponding instruction audit result, including:
[0044] Performing instruction semantic analysis on the control instruction; wherein the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection and instruction contradiction detection;
[0045] The result of the instruction semantic analysis is used as the instruction audit result.
[0046] In this embodiment, a security audit is conducted from the perspective of instruction semantic analysis to determine whether the control instructions sent by the user are complete, unambiguous, and free of potential risks. Instruction semantic analysis specifically includes integrity checking, ambiguity detection, and contradiction detection. Integrity checking checks whether the instructions contain complete objectives, scopes, and constraints; ambiguity detection identifies ambiguous statements that could lead to multiple interpretations; and contradiction detection identifies logical contradictions within the instructions.
[0047] In another specific embodiment, performing a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result further includes:
[0048] Identifying potential risks of the control instructions; wherein the potential risks include over-compliance risk, role-crossing risk, and content inappropriate risk;
[0049] The result of the potential risk identification is used as the instruction audit result.
[0050] In this embodiment, control instructions are analyzed in depth from the perspective of potential risk identification to ensure that they do not lead to inappropriate behavior patterns or information transmission. Potential risk identification specifically includes over-compliance risk, role-crossing risk, and content-inappropriate risk. Over-compliance risk refers to detecting instructions that may cause the agent to over-comply with children's unreasonable requests; role-crossing risk refers to detecting instructions that may cause the agent to exceed its educational role; and content-inappropriate risk refers to detecting instructions that may cause the agent to generate inappropriate content.
[0051] In another specific embodiment, performing a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result further includes:
[0052] Performing instruction simulation testing on the control instruction; wherein the instruction test simulation includes virtual scene testing, boundary condition testing and risk prediction;
[0053] The result of the instruction simulation test is used as the instruction audit result.
[0054] In this embodiment, simulation testing allows the agent to verify the performance of instructions in different scenarios, ensuring its adaptability and robustness. Virtual scenario testing constructs typical usage scenarios to observe the agent's response; boundary condition testing examines the execution of instructions under extreme parameters; and risk prediction leverages historical data to analyze potential future risks and prevent adverse events.
[0055] Furthermore, the results of the aforementioned command semantic analysis, potential risk identification, and command simulation testing are combined to produce a comprehensive command audit report. This report not only reveals potential issues with the command but also provides improvement suggestions to help developers optimize the command processing process and ensure that the agent's responses are both efficient and safe.
[0056] Preferably, in addition to applying natural language processing technology during instruction semantic analysis, potential risk identification, and instruction simulation testing, knowledge of educational psychology can be combined to develop a safety pre-screening system specifically for child-agent instructions. This safety pre-screening system can not only identify problems such as semantic ambiguity, unclear safety boundaries, or unclear educational objectives in instructions, but also provide specific optimization suggestions to help parents create safer and more professional instructions.
[0057] In one embodiment, determining whether the control instruction passes the security audit based on the instruction audit result includes:
[0058] The control instruction is given a first safety score according to the following formula:
[0059] ;
[0060] Where S(a, q) is the safety score of the first agent corresponding to output a when the control instruction is input q, and w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of dimensions of security evaluation;
[0061] Based on the first security score, determine whether the control instruction passes the security audit according to the following formula:
[0062] ;
[0063] Where D(E) is the safety decision result, 1 means pass, 0 means fail, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
[0064] In this embodiment, by performing security scoring on the control instructions and judging whether the control instructions have passed the security audit based on the results of the security score, it is possible to achieve refined management and risk control of the intelligent body instructions, ensure their stability and reliability in complex environments, and further improve the overall security performance of the system.
[0065] Here, when performing the first security score on the control instruction, the weight of each security dimension and the accuracy of its evaluation function must be comprehensively considered to ensure the objectivity and accuracy of the scoring result. By setting a reasonable threshold, the system can effectively distinguish high-risk instructions, thereby blocking potential threats in a timely manner and ensuring the safety of the intelligent agent's operation. In actual application scenarios, the security dimension evaluation function f j (a, q) can be implemented using a combination of techniques, for example:
[0066] (1) Content suitability assessment: Using a fine-tuning-based classifier to assess whether the content is suitable for children’s age group;
[0067] (2) Educational value assessment: Based on the educational knowledge map, evaluate the educational value and accuracy of the content;
[0068] (3) Emotional safety assessment: Using sentiment analysis models to detect whether the content contains negative emotions or expressions that may cause discomfort to children;
[0069] (4) Behavior guidance assessment: whether the assessment content guides children to engage in safe and healthy behaviors;
[0070] (5) Privacy protection assessment: Check whether the content involves the collection or disclosure of personal privacy information.
[0071] Each evaluation function f j Each time a score is output in the range of [0, 1], where 1 indicates complete safety and 0 indicates a serious safety risk.
[0072] In actual application scenarios, the complete process from user-created agent configuration to security audit report generation can be implemented by building components such as security test suites, test execution loops, and visualization engines. The construction of the security test suite uses stratified sampling and risk-weighted methods, as follows:
[0073] ;
[0074] Among them, T is the security test set, q i is the input of the i-th test case, r i is the reference output or safety standard of the test case, w i is the risk weight of the test case, and m is the total number of test cases.
[0075] Furthermore, a collaborative evaluation process between the performer agent and the judge agent is set up based on the safety test set, as follows:
[0076] ;
[0077] Among them, E is the evaluation result set, T is the safety test set, A is the performer agent, C is the judge agent, q i is the input of the i-th test case, a i is the output of the performer agent, s i is the safety score of the judge agent.
[0078] In one embodiment, the optimizing the control instruction includes:
[0079] According to the following formula, the control instruction is optimized based on the preset optimization suggestion template:
[0080] ;
[0081] Among them, Sug(i, R) is the set of recommendations for control instruction i and risk assessment R, T j is the optimization suggestion template for the jth risk, R j is the score of the j-th risk, θ j is the threshold value for the jth risk category, and m is the number of risk categories.
[0082] This embodiment optimizes control instructions that fail security audits using pre-configured optimization templates, generating more secure versions of them and ensuring they meet security standards during execution. These optimized instructions not only reduce risk but also improve the agent's execution efficiency and user experience, further enhancing the agent's overall security and reliability.
[0083] For example, control instructions can be supplemented with safety boundaries, such as providing supplementary suggestions for instructions that lack safety boundaries. Control instructions can also be clarified with regard to educational objectives to help parents clarify the educational objectives of the instructions. Control instructions can also be replaced with professional terms: replacing parents' non-professional expressions with more accurate professional terms.
[0084] Of course, in actual application scenarios, even if the control instructions pass the security audit, they can still be optimized, such as by adding safety prompts, adjusting the way the instructions are expressed, etc., to make them more in line with children's cognitive habits, so as to further improve their safety and applicability.
[0085] Here are a few specific examples to illustrate how to generate optimization suggestions:
[0086] Case 1: Semantically ambiguous instructions:
[0087] Original instruction: You must answer all questions asked by the child;
[0088] Risk assessment: Rsem = 0.85 (high risk of semantic ambiguity);
[0089] Generate suggestions: It is recommended to modify it to: You should answer children's questions about educational fields such as science, history, and literature, but for topics that are not suitable for children, you should politely decline to answer and guide them to appropriate topics.
[0090] Case 2: Security Boundary Missing Instructions:
[0091] Original instruction: You are a storytelling agent, and you need to tell interesting stories;
[0092] Risk assessment: Rcon = 0.78 (content risk is high);
[0093] Generate suggestions: It is recommended to modify it to: You are a storytelling agent for children aged 5-8. You should tell positive and educational stories, and the content of the stories should be suitable for children's cognitive level.
[0094] Case 3: Unclear instruction on educational objectives:
[0095] Original Instruction: You are to help your child learn;
[0096] Risk assessment: Redu = 0.82 (high risk of educational suitability);
[0097] Generate suggestions: It is recommended to modify it to: You are an elementary school math tutoring agent that helps children aged 7-9 understand basic math concepts, including addition, subtraction, multiplication, division, simple fractions, and geometric figures. Use vivid examples and interactive questions to enhance learning effects. The difficulty should be suitable for the level of second grade students.
[0098] Through these optimization suggestions, the agent can not only effectively avoid potential risks, but also better meet user needs and improve educational effectiveness. Of course, the template can also be continuously optimized based on actual scenarios and user feedback to ensure that the agent strikes the optimal balance between safety and efficiency.
[0099] In another embodiment, an instruction risk assessment model is constructed according to the following formula, and risk assessment of control instructions is performed using the instruction risk assessment model:
[0100] ;
[0101] Where: R(i) is the overall risk score of instruction i, R k (i) is the score of instruction i on the kth risk dimension, w k is the weight coefficient of the kth risk dimension, and N is the total number of risk assessment dimensions, which can be dynamically expanded as needed.
[0102] For example, R k (i) can be a semantic risk score R sem (i), i.e., evaluating the fuzziness and ambiguity of instruction semantics; R k (i) It can also be the content risk score R con (i), i.e., assessing the inappropriate content that may result from the instruction; R k (i) It can also be the educational suitability risk score R edu (i) is to evaluate the educational value and age appropriateness of the instruction. In addition, other risk dimensions can be added according to actual needs, such as the privacy risk score R priv (i) is used to assess the degree of protection of children's privacy provided by the directive; it relies on the risk score R dep (i) To assess the potential over-reliance on instructions; Risk of bias score R bias (i), used to assess the bias or stereotypes that may be contained in the instructions.
[0103] Through the above-mentioned instruction risk assessment model, a comprehensive risk assessment of control instructions can be conducted to ensure that they have undergone sufficient safety considerations before execution. This risk assessment model not only takes into account conventional dimensions such as the semantic clarity, content appropriateness, and educational value of instructions, but also pays special attention to key areas such as children's privacy protection, dependence risk, and bias risk, thereby effectively avoiding potential harm to children due to inappropriate instructions. At the same time, the model has dynamic scalability and can flexibly add or adjust risk assessment dimensions according to changes and needs in actual application scenarios to ensure the comprehensiveness and accuracy of the assessment results. During the implementation process, the judge intelligent agent will use this instruction risk assessment model to conduct a detailed risk assessment of each control instruction and decide whether to perform intervention operations based on the assessment results to ensure the safety and stability of the intelligent agent's operation.
[0104] In one embodiment, the agent security audit method further includes:
[0105] A second safety score is assigned to the second agent according to the following formula:
[0106] ;
[0107] Among them, Score is the second safety score, w i is the weight of the i-th security dimension, s' i is the score of the i-th security dimension, n is the number of dimensions of security assessment;
[0108] According to the following formula, it is determined whether to generate a risk warning for the second agent based on the second safety score:
[0109] ;
[0110] Among them, Alert(c,t) is the risk warning result, 1 means triggering the warning, 0 means not triggering, c is the content to be evaluated, s' i (c) is the score of content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of content c, T is the warning threshold of the overall security score, and n is the number of dimensions of security assessment;
[0111] When it is determined that a risk warning is generated, the risk warning is sent to the second intelligent agent, and the second intelligent agent is enabled to perform a visual display of the risk warning.
[0112] In this embodiment, in addition to conducting security audits on control commands sent by users, the output of the first agent is also monitored in real time to ensure compliance with safety standards. Once a potential risk is detected, an early warning mechanism is immediately activated to implement early warning intervention and prevent the spread of harmful information.
[0113] During real-time monitoring of the first agent, it first assesses its security across various security dimensions. Based on the resulting security score, the system then determines whether to trigger an alert. If the score falls below the threshold, the system automatically generates an alert and notifies relevant personnel to intervene, ensuring the security of information dissemination. In this way, the agent not only provides efficient service but also effectively prevents potential risks and ensures user safety. Through real-time monitoring and multi-dimensional security assessments, the agent not only improves its own security performance but also strengthens user trust.
[0114] Furthermore, after completing the monitoring and audit of the first agent, the corresponding results can be sent to the first agent, which can then perform a visual display to allow users to intuitively understand the agent's security status. In practical application scenarios, when performing the visual display, the first agent can, on the one hand, set up an instruction security score panel to display the security audit results of the control instructions, and on the other hand, set up a real-time audit monitoring panel to display the security audit results of the first agent. For example, the instruction security score panel can include multiple visualization items, such as an overall security score, dimension scores, risk point markers, and optimization suggestions. The overall security score can be a 0-100 overall security score, the dimension scores can be radar charts of independent scores for multiple security dimensions, the risk point markers can visually mark potential risk points in the instructions, and the optimization suggestions can provide specific instruction optimization suggestions. The real-time audit monitoring panel can include multiple visualization items, such as audit process visualization, risk warnings, intervention operation records, and audit statistics. The audit process visualization can be a graphical display of the assessor agent's audit process, the risk warnings can be a real-time display of detected security risks, the intervention operation records can be a record of the intervention operations performed by the assessor agent, and the audit statistics can be statistical data displaying the audit results.
[0115] The security visualization interface allows parents to intuitively understand the security status of the agent. For example, graphically displaying the audit process and results helps parents understand and manage the security status of the agent. Another example is the instruction safety scoring dashboard, which allows parents to clearly understand whether the instructions they create pose security risks. These visualization tools significantly improve the transparency and explainability of the system, strengthening parents' trust in and control over the agent.
[0116] In general, the agent security audit method provided in this embodiment has the following advantages:
[0117] (1) Closed-loop audit process: A complete closed-loop process from user creation of agent configuration to security audit report generation ensures that each agent undergoes a comprehensive security audit before release and passes strict security review to avoid potential risks;
[0118] (2) Dual-agent collaborative mechanism: The performer agent and the judge agent are independent of each other but work together, avoiding the inherent defects of "self-supervision" of a single agent;
[0119] (3) Dynamic test set: By maintaining a continuously updated security test set, the test strategy can be adjusted in a timely manner to respond to emerging security risks;
[0120] (4) Real-time visual feedback: Providing real-time visual feedback to users through security audit dashboards and other means enhances the transparency and explainability of the system;
[0121] (5) Multi-dimensional safety assessment: The evaluator agent uses a multi-dimensional assessment framework to comprehensively evaluate the safety of content from multiple dimensions, including child safety, educational value, and emotional impact;
[0122] (6) Automated release decision: It can automatically decide whether to allow the agent to be released based on the security assessment results, reducing the need for human intervention.
[0123] Figure 2 This is a schematic block diagram of an intelligent agent security audit device 200 provided in an embodiment of the present invention. The device 200 adopts a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to interact with a user, and the second intelligent agent is used to implement the security audit device 200. The device 200 includes:
[0124] The instruction audit unit 201 is configured to, when the first agent receives a control instruction sent by a user, obtain the control instruction through the first agent, perform a multi-dimensional security audit on the control instruction, and obtain a corresponding instruction audit result;
[0125] An audit judgment unit 202 is configured to judge whether the control instruction has passed the security audit based on the instruction audit result;
[0126] The first determination unit 203 is configured to optimize the control instruction if it is determined that the control instruction fails the security audit;
[0127] The second determination unit 204 is configured to send a security audit pass feedback to the first agent if it is determined that the control instruction has passed the security audit, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
[0128] In one embodiment, the instruction audit unit 201 includes:
[0129] A semantic analysis unit, configured to perform instruction semantic analysis on the control instruction; wherein the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection, and instruction contradiction detection;
[0130] The first setting unit is configured to use the result of the instruction semantic analysis as the instruction audit result.
[0131] In one embodiment, the instruction audit unit 201 further includes:
[0132] A risk identification unit, configured to identify potential risks of the control instruction; wherein the potential risks include over-compliance risk, role-crossing risk, and content inappropriate risk;
[0133] The second setting unit is used to use the result of the potential risk identification as the instruction audit result.
[0134] In one embodiment, the instruction audit unit 201 further includes:
[0135] A simulation test unit, configured to perform instruction simulation testing on the control instruction; wherein the instruction test simulation includes virtual scenario testing, boundary condition testing, and risk prediction;
[0136] The third setting unit is used to use the result of the instruction simulation test as the instruction audit result.
[0137] In one embodiment, the audit judgment unit 202 includes:
[0138] The first scoring unit is configured to perform a first safety score on the control instruction according to the following formula:
[0139] ;
[0140] Where S(a, q) is the safety score of the first agent corresponding to output a when the control instruction is input q, and w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of dimensions of security evaluation;
[0141] A score comparison unit is configured to determine whether the control instruction passes the security audit based on the first security score according to the following formula:
[0142] ;
[0143] Where D(E) is the safety decision result, 1 means pass, 0 means fail, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
[0144] In one embodiment, the first determining unit 203 includes:
[0145] The instruction optimization unit is configured to optimize the control instruction based on a preset optimization suggestion template according to the following formula:
[0146] ;
[0147] Among them, Sug(i, R) is the set of recommendations for control instruction i and risk assessment R, T j is the optimization suggestion template for the jth risk, R jis the score of the j-th risk, θ j is the threshold value for the jth risk category, and m is the number of risk categories.
[0148] In one embodiment, the intelligent agent security audit device 200 further includes:
[0149] The second scoring unit is configured to perform a second safety score on the second agent according to the following formula:
[0150] ;
[0151] Among them, Score is the second safety score, w i is the weight of the i-th security dimension, s' i is the score of the i-th security dimension, n is the number of dimensions of security assessment;
[0152] An early warning judgment unit is configured to judge whether to generate a risk early warning for the second agent according to the second safety score according to the following formula:
[0153] ;
[0154] Among them, Alert(c,t) is the risk warning result, 1 means triggering the warning, 0 means not triggering, c is the content to be evaluated, s' i (c) is the score of content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of content c, T is the warning threshold of the overall security score, and n is the number of dimensions of security assessment;
[0155] The visualization display unit is used to send the risk warning to the second intelligent agent when it is determined that a risk warning is generated, and enable the second intelligent agent to perform a visualization display of the risk warning.
[0156] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0157] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiment. The storage medium may include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other medium capable of storing program code.
[0158] The present invention also provides an intelligent device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the intelligent device may also include various network interfaces, a power supply, and other components.
[0159] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0160] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. An intelligent agent security audit method adopts a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein: The first agent is used to interact with the user, and the second agent is used to implement a security audit method, wherein the method includes: When the first agent receives a control instruction sent by a user, the first agent obtains the control instruction and performs a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result; the multi-dimensional security audit includes content suitability assessment, educational value assessment, emotional safety assessment, behavior guidance assessment, and privacy protection assessment; Determining whether the control instruction has passed the security audit based on the instruction audit result; If it is determined that the control instruction fails the security audit, then optimizing the control instruction; If it is determined that the control instruction has passed the security audit, a security audit pass feedback is sent to the first agent, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback; The determining, based on the instruction audit result, whether the control instruction passes the security audit includes: The control instruction is given a first safety score according to the following formula: ; Where S(a, q) is the safety score of the first agent when the control instruction is input q, and w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of dimensions of security evaluation; Based on the first security score, determine whether the control instruction passes the security audit according to the following formula: ; Where D(E) is the safety decision result, 1 means pass, 0 means fail, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
2. The intelligent agent security audit method according to claim 1, characterized in that: The multi-dimensional security audit of the control instruction is performed to obtain the corresponding instruction audit result, including: Performing instruction semantic analysis on the control instruction; wherein the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection and instruction contradiction detection; The result of the instruction semantic analysis is used as the instruction audit result.
3. The intelligent agent security audit method according to claim 1, characterized in that: The performing of a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result further includes: Identifying potential risks of the control instructions; wherein the potential risks include over-compliance risk, role-crossing risk, and content inappropriate risk; The result of the potential risk identification is used as the instruction audit result.
4. The intelligent agent security audit method according to claim 1, characterized in that: The performing of a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result further includes: Performing instruction simulation testing on the control instruction; wherein the instruction simulation testing includes virtual scenario testing, boundary condition testing and risk prediction; The result of the instruction simulation test is used as the instruction audit result.
5. The intelligent agent security audit method according to claim 1, characterized in that: The step of optimizing the control instructions includes: According to the following formula, the control instruction is optimized based on the preset optimization suggestion template: ; Among them, Sug(i, R) is the set of recommendations for control instruction i and risk assessment R, T j is the optimization suggestion template for the jth risk, R j is the score of the j-th risk, θ j is the threshold value for the jth risk category, and m is the number of risk categories.
6. The intelligent agent security audit method according to claim 1, characterized in that: Also includes: A second safety score is assigned to the second agent according to the following formula: ; Among them, Score is the second safety score, w i is the weight of the i-th security dimension, s' i is the score of the i-th security dimension, n is the number of dimensions of security assessment; According to the following formula, it is determined whether to generate a risk warning for the second agent based on the second safety score: ; Among them, Alert(c,t) is the risk warning result, 1 means triggering the warning, 0 means not triggering, c is the content to be evaluated, s' i (c) is the score of content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of content c, T is the warning threshold of the overall security score, and n is the number of dimensions of security assessment; When it is determined that a risk warning is generated, the risk warning is sent to the second intelligent agent, and the second intelligent agent is enabled to perform a visual display of the risk warning.
7. An intelligent agent security audit device, using a first intelligent agent and a second intelligent agent collaborative architecture, wherein: The first intelligent agent is used to interact with the user, and the second intelligent agent is used to implement the security audit method, characterized in that the device includes: an instruction audit unit, configured to, when the first agent receives a control instruction sent by a user, obtain the control instruction through the first agent, and perform a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result; the multi-dimensional security audit includes content suitability assessment, educational value assessment, emotional safety assessment, behavior guidance assessment, and privacy protection assessment; An audit judgment unit, configured to judge whether the control instruction has passed the security audit according to the instruction audit result; A first determining unit is configured to optimize the control instruction if it is determined that the control instruction fails the security audit; a second determination unit configured to send a security audit pass feedback to the first agent if it is determined that the control instruction has passed the security audit, so that the first agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback; The audit judgment unit includes: The first scoring unit is configured to perform a first safety score on the control instruction according to the following formula: ; Where S(a, q) is the safety score of the first agent when the control instruction is input q, and w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of dimensions of security evaluation; A score comparison unit is configured to determine whether the control instruction passes the security audit based on the first security score according to the following formula: ; Where D(E) is the safety decision result, 1 means pass, 0 means fail, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
8. An intelligent agent device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent agent security audit method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the intelligent agent security audit method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent child rearing system and device based on natural language processing
CN117453867A
Virtual digital human interaction method, system, terminal, equipment and medium
CN117520500A