Intelligent agent security auditing method and device, intelligent agent equipment and storage medium
By introducing the first and second agent collaborative architectures into the agent system, multi-dimensional security audits and instruction optimization are carried out, security risks and memory confusion problems in the creation of the agent are solved, and the security of the agent is improved.
Patent Information
- Application Number
- CN202510647784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing agents have security risks when creating custom agents. Multi-user interactions in homes lead to memory confusion, security audits are opaque, and a single agent architecture cannot effectively identify and prevent its own security risks.
The first agent and the second agent collaborative architecture are adopted, the first agent is used for interaction with the user, and the second agent is used for security auditing. The control instructions are audited through multi-dimensional security audit to determine whether they have passed the security audit. If they are not passed, the instructions are optimized. If they are passed, the security audit will be sent to pass feedback.
It improves the security of the use of agents, solves the security risks of parents and users in the process of creating custom agents and the memory confusion caused by multi-user interaction, and avoids the inherent defects of a single agent architecture.
Smart Images

Figure CN120162830A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an intelligent agent security auditing method, device, intelligent agent device and storage medium. Background Art
[0002] With the development of large language model technology, more and more parent users hope to create custom intelligent agents according to the specific needs of their children and then use the custom intelligent agents for children's education and companionship. However, there are certain risks in this process: (1) Security risks caused by insufficient professionalism of parent instructions. Through market research, it shows that most parent users are not familiar enough with intelligent agents, or the instructions used when creating custom intelligent agents have problems such as ambiguous semantics, unclear security boundaries or unclear educational goals, resulting in inappropriate content or behaviors that may occur during the interaction between the intelligent agent and children. Due to the lack of professional knowledge, parent users usually cannot identify these potential risks; (2) Memory confusion caused by multi-user interaction in the family. In a family environment, in addition to the target child, other family members (such as parents, siblings) may also interact with the intelligent agent, which will lead to memory confusion of the intelligent agent. Without a user identity recognition and memory isolation mechanism, the children's intelligent agent will have phenomena of "role confusion" or "memory leakage", bringing inappropriate content in adult interactions into the conversation with children; (3) Opaque security auditing. Existing products lack a mechanism to show the security auditing process of intelligent agents to parent users, resulting in parent users being unable to understand whether the intelligent agents they create have security risks and unable to intervene and correct them in a timely manner; (4) Inherent defects of a single intelligent agent architecture. Existing products mostly adopt a single intelligent agent architecture, where the same intelligent agent is responsible for both content generation and security review, with the inherent defect of "self-supervision" and being unable to effectively identify and prevent its own security risks.
[0003] Therefore, how to improve the safe use performance of intelligent agents is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] Embodiments of the present invention provide an intelligent agent security auditing method, device, intelligent agent device and storage medium, aiming to improve the use safety of intelligent agents.
[0005] In a first aspect, embodiments of the present invention provide an intelligent agent security auditing method, adopting a collaborative architecture of a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to interact with users, and the second intelligent agent is used to implement the security auditing method. The method includes: When the first agent receives a control instruction sent by a user, it obtains the control instruction through the first agent and performs multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result; Judge whether the control instruction passes the security audit according to the instruction auditing result; If it is determined that the control instruction fails to pass the security audit, optimize the control instruction; If it is determined that the control instruction passes the security audit, send a feedback of passing the security audit to the first agent, so that after receiving the feedback of passing the security audit, the first agent generates corresponding content according to the control instruction and publishes it.
[0006] In a second aspect, an embodiment of the present invention provides an agent security auditing device, which adopts a collaborative architecture of a first agent and a second agent. Among them, the first agent is used to interact with the user, and the second agent is used to implement the security auditing method. The device includes: An instruction auditing unit, configured to, when the first agent receives a control instruction sent by a user, obtain the control instruction through the first agent and perform multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result; An auditing judgment unit, configured to judge whether the control instruction passes the security audit according to the instruction auditing result; A first determination unit, configured to, if it is determined that the control instruction fails to pass the security audit, optimize the control instruction; A second determination unit, configured to, if it is determined that the control instruction passes the security audit, send a feedback of passing the security audit to the first agent, so that after receiving the feedback of passing the security audit, the first agent generates corresponding content according to the control instruction and publishes it.
[0007] In a third aspect, an embodiment of the present invention provides an agent device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the agent security auditing method as described in the first aspect.
[0008] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the agent security auditing method as described in the first aspect.
[0009] Embodiments of the present invention provide an intelligent agent security auditing method, apparatus, intelligent agent device, and storage medium. The method adopts a collaborative architecture of a first intelligent agent and a second intelligent agent. Among them, the first intelligent agent is used to interact with the user, and the second intelligent agent is used to implement the security auditing method. The method includes: when the first intelligent agent receives a control instruction sent by the user, obtaining the control instruction through the first intelligent agent, and performing multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result; judging whether the control instruction passes the security auditing according to the instruction auditing result; if it is determined that the control instruction does not pass the security auditing, optimizing the control instruction; if it is determined that the control instruction passes the security auditing, sending a security auditing passed feedback to the first intelligent agent, so that after receiving the security auditing passed feedback, the first intelligent agent generates corresponding content according to the control instruction and publishes it. Embodiments of the present invention adopt a dual-intelligent-agent collaborative architecture, that is, on the one hand, one intelligent agent interacts with the user, and on the other hand, the other intelligent agent performs security auditing. In this way, it can solve the security hazards in the process of parents creating custom intelligent agents, as well as problems such as memory confusion caused by multi-user interaction, and avoid the inherent defects of a single intelligent agent architecture, thereby improving the use safety of the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of an intelligent agent security auditing method provided by an embodiment of the present invention; Figure 2 It is a schematic block diagram of an intelligent agent security auditing apparatus provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0013] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0014] It should also be understood that the terms used in this specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0015] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0016] Please refer to the following Figure 1 An intelligent agent security auditing method provided by an embodiment of the present invention adopts a collaborative architecture of a first intelligent agent and a second intelligent agent. Among them, the first intelligent agent is used to interact with a user, and the second intelligent agent is used to implement the security auditing method. The method specifically includes: steps S101 to S104.
[0017] Step S101: When the first intelligent agent receives a control instruction sent by a user, obtain the control instruction through the first intelligent agent, and perform multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result; Step S102: Judge whether the control instruction passes the security auditing according to the instruction auditing result; Step S103: If it is determined that the control instruction does not pass the security auditing, optimize the control instruction; Step S104: If it is determined that the control instruction passes the security auditing, send a security auditing passed feedback to the first intelligent agent, so that after receiving the security auditing passed feedback, the first intelligent agent generates corresponding content according to the control instruction and publishes it.
[0018] In this embodiment, when the first agent receives a control instruction sent by the user, such as an instruction to create or update an agent, the second agent can obtain the control instruction and perform a security audit on it from multiple dimensions to determine whether the control instruction meets the requirements. When it is determined that the audit passes, the first agent can generate corresponding content according to the control instruction sent by the user and publish it externally. If the audit fails, the control instruction can be optimized to ensure passing. Of course, it can be understood that if the content included in the control instruction cannot be optimized or the degree of optimization does not meet the security audit requirements, then the first agent naturally cannot generate and publish the corresponding content.
[0019] This embodiment adopts a dual-agent collaborative architecture. That is, on the one hand, one agent interacts with the user, and on the other hand, the other agent performs a security audit. In this way, it can solve the security risks in the process of parents creating custom agents, as well as problems such as memory confusion caused by multi-user interaction, and avoid the inherent defects of a single-agent architecture, thereby improving the security of agent use.
[0020] It should be noted that the agent security audit method provided in this embodiment is particularly applicable to the field of child agent security and content audit. For example, the first agent in the dual-agent architecture is a performer agent, which is responsible for receiving instructions created by parents and generating interactive responses for children. The second agent is a judge agent, which is responsible for auditing the interaction process between the parent instruction and the performer agent. The judge agent is independent of the performer agent to ensure the objectivity and independence of the audit. Moreover, the judge agent focuses on security audit and does not participate in content generation to avoid goal conflicts.
[0021] In an actual application scenario, the performer agent can be based on a large language model and fine-tuned with child education content. At the same time, it also has multi-modal input and output capabilities, supports interaction methods such as text, voice, and image, and has a built-in security self-check mechanism, so that it can perform a preliminary security assessment before generating a response, thereby further improving the safe use of the agent. The judge agent can adopt a multi-dimensional evaluation framework to evaluate content from dimensions such as child safety, educational value, and emotional impact. At the same time, the judge agent can also have an intervention ability, that is, perform an interception or correction operation when detecting a security risk.
[0022] In addition, the performer agent and the judge agent also have a collaborative working mechanism. For example: (1) Instruction pre-check: The judge agent performs a security pre-check when the parent creates an instruction; (2) Real-time monitoring: The judge agent monitors the input and output of the performer agent in real time; (3) Risk intervention: The judge agent performs an intervention operation when detecting a security risk; (4) Feedback optimization: Optimize the behavior of the performer agent according to the feedback of the judge agent.
[0023] In a specific embodiment, performing multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result, including: Performing instruction semantic analysis on the control instruction; wherein, the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection, and instruction contradiction detection; Taking the result of the instruction semantic analysis as the instruction auditing result.
[0024] In this embodiment, security auditing is performed from the dimension of instruction semantic analysis to determine whether the control instruction sent by the user is complete, clear, and has no potential risks, etc. The instruction semantic analysis specifically includes integrity check, ambiguity detection, and contradiction detection. Among them, the integrity check refers to detecting whether the instruction contains complete objectives, scopes, and constraints; the ambiguity detection refers to identifying ambiguous expressions that may lead to multiple interpretations; the contradiction detection refers to identifying logical contradictions within the instruction.
[0025] In another specific embodiment, performing multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result further includes: Identifying potential risks of the control instruction; wherein, the potential risk identification includes excessive compliance risk, role boundary crossing risk, and content inappropriateness risk; Taking the result of the potential risk identification as the instruction auditing result.
[0026] In this embodiment, in-depth analysis of the control instruction is carried out from the perspective of potential risk identification to ensure that the instruction will not lead to inappropriate behavior patterns or information transmission. The potential risk identification specifically includes excessive compliance risk, role boundary crossing risk, and content inappropriateness risk. Among them, the excessive compliance risk refers to detecting instructions that may cause the agent to excessively comply with unreasonable requirements of children; the role boundary crossing risk refers to detecting instructions that may cause the agent to exceed the educational role; the content inappropriateness risk refers to detecting instructions that may cause the agent to generate inappropriate content.
[0027] In yet another specific embodiment, performing multi-dimensional security auditing on the control instruction to obtain a corresponding instruction auditing result further includes: Performing instruction simulation testing on the control instruction; wherein, the instruction test simulation includes virtual scenario testing, boundary condition testing, and risk prediction; Taking the result of the instruction simulation testing as the instruction auditing result.
[0028] In this embodiment, through simulation tests, the agent can verify the performance of the instructions in different scenarios to ensure its adaptability and robustness. The virtual scenario test observes the agent's reactions by constructing typical usage scenarios; the boundary condition test checks the execution of the instructions under extreme parameters; and the risk prediction uses historical data analysis to predict potential future risks to prevent the occurrence of adverse events.
[0029] Furthermore, based on the above-mentioned instruction semantic analysis, potential risk identification, and the results of the instruction simulation test, a comprehensive instruction audit report is formed. This report not only reveals potential problems with the instructions but also provides improvement suggestions to help developers optimize the instruction processing flow and ensure that the agent's response is both efficient and safe.
[0030] Preferably, when performing instruction semantic analysis, potential risk identification, and instruction simulation tests, in addition to applying natural language processing techniques, it can also be combined with educational psychology knowledge to develop a safety pre-check system specifically for child agent instructions. This safety pre-check system can not only identify problems such as semantic ambiguity, unclear safety boundaries, or unclear educational goals in the instructions but also provide specific optimization suggestions to help parents create safer and more professional instructions.
[0031] In one embodiment, determining whether the control instruction passes the safety audit according to the instruction audit result includes: Perform a first safety score on the control instruction according to the following formula: ; where S(a, q) is the safety score corresponding to the output a of the first agent when the control instruction is input q, w j is the weight of the jth safety dimension, f j (a, q) is the evaluation function of the jth safety dimension, and n is the number of safety evaluation dimensions; Based on the first safety score, determine whether the control instruction passes the safety audit according to the following formula: ; where D(E) is the safety decision result, 1 indicates passing, 0 indicates not passing, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
[0032] In this embodiment, by performing a safety score on the control instruction and determining whether the control instruction passes the safety audit based on the result of the safety score, fine-grained management and risk control of the agent's instructions can be achieved, ensuring its stability and reliability in a complex environment, and further improving the overall safety performance of the system.
[0033] Here, when performing the first security scoring on the control instruction, it is necessary to comprehensively consider the weights of each security dimension and the accuracy of its evaluation function to ensure the objectivity and accuracy of the scoring result. By setting a reasonable threshold, the system can effectively distinguish high-risk instructions, thereby timely blocking potential threats and ensuring the security of the agent's operation. In the actual application scenario, the security dimension evaluation function f j (a, q) can be implemented using a variety of technology combinations, for example: (1) Content suitability assessment: Use a fine-tuned classifier to evaluate whether the content is suitable for the children's age group; (2) Educational value assessment: Based on the educational knowledge graph, evaluate the educational value and accuracy of the content; (3) Emotional security assessment: Use an emotion analysis model to detect whether the content contains negative emotions or expressions that may cause discomfort to children; (4) Behavior guidance assessment: Evaluate whether the content guides children to perform safe and healthy behaviors; (5) Privacy protection assessment: Detect whether the content involves the collection or leakage of personal privacy information.
[0034] Each evaluation function f j will output a score within the range of [0, 1], where 1 indicates completely safe and 0 indicates a serious security risk.
[0035] In the actual application scenario, a complete process from user creation of the agent configuration to the generation of the security audit report can be achieved by constructing components such as a security test set, a test execution loop, and a visualization engine. Among them, the construction of the security test set adopts the methods of stratified sampling and risk weighting, specifically as follows: ; where T is the security test set, q i is the input of the i-th test case, r i is the reference output or security standard of the test case, w i is the risk weight of the test case, and m is the total number of test cases.
[0036] Furthermore, a collaborative evaluation process of the performer agent and the judge agent is set up based on the security test set, specifically as follows: ; where E is the evaluation result set, T is the security test set, A is the performer agent, C is the judge agent, q i is the input of the i-th test case, a i is the output of the performer agent, and s i is the security score of the judge agent.
[0037] In one embodiment, the optimization of the control instruction includes: Optimizing the control instruction based on a preset optimization suggestion template according to the following formula: ; where Sug(i, R) is the set of suggestions for control instruction i and risk assessment R, T j is the optimization suggestion template for the j-th type of risk, R j is the score of the j-th type of risk, θ j is the threshold of the j-th type of risk, and m is the number of risk categories.
[0038] In this embodiment, the control instructions that fail the security audit are optimized through a preset optimization suggestion template to generate a more secure instruction version, ensuring that they meet the security standards during execution. The optimized instructions not only reduce risks but also improve the execution efficiency and user experience of the agent, further enhancing the overall security and reliability of the agent.
[0039] For example, it is possible to supplement the security boundaries of the control instructions, such as providing supplementary suggestions for instructions lacking security boundaries, or to clarify the educational objectives of the control instructions to help parents clarify the educational objectives of the instructions. It is also possible to replace the professional terms in the control instructions: replace the non-professional expressions of parents with more accurate professional terms.
[0040] Of course, in actual application scenarios, even if the control instructions pass the security audit, they can still be optimized, such as by adding security prompts, adjusting the instruction expression method, etc., to make them more in line with the cognitive habits of children, so as to further improve their security and applicability.
[0041] The following provides several specific cases to show how to generate optimization suggestions: Case 1: Instructions with ambiguous semantics: Original instruction: You need to answer all the questions of the child; Risk assessment: Rsem = 0.85 (high risk of ambiguous semantics); Generated suggestion: It is recommended to modify it to: You need to answer the child's questions in the educational fields such as science, history, literature, etc., but for topics involving violence, adult content or topics not suitable for children, you should decline to answer politely and guide to suitable topics.
[0042] Case 2: Instructions lacking security boundaries: Original instruction: You are a storytelling agent and need to tell interesting stories; Risk assessment: Rcon = 0.78 (high content risk); Generated suggestion: It is recommended to modify to: You are a storytelling intelligent agent for children aged 5 - 8, and should tell positive and educational stories. The content of the stories should be suitable for children's cognitive level and not contain horror, violence, or complex adult themes.
[0043] Case 3. Unclear educational goal Instruction: Original instruction: You need to help children learn; Risk assessment: Redu = 0.82 (high risk of educational suitability); Generated suggestion: It is recommended to modify to: You are a primary school mathematics tutoring intelligent agent, helping children aged 7 - 9 understand basic mathematical concepts, including addition, subtraction, multiplication, division, simple fractions, and geometric figures. Use vivid examples and interactive questions to enhance the learning effect, and the difficulty should be suitable for second - grade students.
[0044] Through these optimization suggestions, the intelligent agent can not only effectively avoid potential risks, but also better meet user needs and improve the educational effect. Of course, it is also possible to continuously optimize the template in combination with the actual scenario and user feedback to ensure that the intelligent agent achieves the best balance between safety and efficiency.
[0045] In another embodiment, an instruction risk assessment model is built according to the following formula, and the control instruction is risk - assessed through the instruction risk assessment model: ; Where: R(i) is the overall risk score of instruction i, R k (i) is the score of instruction i on the k - th risk dimension, w k is the weight coefficient of the k - th risk dimension, and N is the total number of risk assessment dimensions, which can be dynamically expanded according to needs.
[0046] For example, R k (i) can be the semantic risk score R sem (i), that is, to evaluate the ambiguity and vagueness of the instruction semantics; R k (i) can also be the content risk score R con (i), that is, to evaluate the inappropriate content that the instruction may lead to; R k (i) can also be the educational suitability risk score R edu (i), that is, to evaluate the educational value and age - suitability of the instruction. In addition, other risk dimensions can be added according to actual needs, such as the privacy risk score R priv (i), used to evaluate the degree of protection of children's privacy by the instruction; the dependence risk score R dep (i), used to evaluate the excessive dependence that the instruction may lead to; the bias risk score R bias (i), used to evaluate the possible biases or stereotypes in the instruction.
[0047] Through the above-mentioned instruction risk assessment model, a comprehensive risk assessment can be carried out on control instructions to ensure that they have been fully considered for security before execution. This risk assessment model not only considers conventional dimensions such as the semantic clarity, content suitability, and educational value of the instructions, but also particularly focuses on key areas such as child privacy protection, dependency risk, and bias risk, thereby effectively avoiding potential harms that may be brought to children due to improper instructions. At the same time, this model has dynamic scalability and can flexibly add or adjust risk assessment dimensions according to the changes and requirements of the actual application scenario to ensure the comprehensiveness and accuracy of the assessment results. During the implementation process, the judge agent will use this instruction risk assessment model to conduct a detailed risk assessment on each control instruction and decide whether to perform an intervention operation based on the assessment results to ensure the security and stability of the agent's operation.
[0048] In one embodiment, the agent security audit method further includes: Perform a second security score on the second agent according to the following formula: ; where Score is the second security score, w i is the weight of the i-th security dimension, s’ i is the score of the i-th security dimension, and n is the number of security assessment dimensions; Judge whether to generate a risk warning for the second agent according to the following formula based on the second security score: ; where Alert(c,t) is the risk warning result, 1 indicates triggering the warning, 0 indicates not triggering, c is the content to be evaluated, s’ i (c) is the score of the content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of the content c, T is the warning threshold of the overall security score, and n is the number of security assessment dimensions; When it is determined to generate a risk warning, send the risk warning to the second agent and enable the second agent to visually display the risk warning.
[0049] In this embodiment, in addition to performing a security audit on the control instructions sent by the user, the output content of the first agent will also be monitored in real time to ensure that it meets the security standards. Once a potential risk is detected, the warning mechanism will be immediately activated to achieve warning intervention and prevent the spread of bad information.
[0050] When conducting real-time monitoring of the first intelligent agent, its security score is first determined through different set security dimensions, and then based on the result of the security score, it is judged whether to trigger an early warning. If the score is lower than the threshold, the system will automatically generate an early warning message and notify relevant personnel to intervene to ensure the security of information dissemination. In this way, while providing efficient services, the intelligent agent can effectively prevent potential risks and ensure the security of user use. Through real-time monitoring and multi-dimensional security assessment, the intelligent agent not only improves its own security performance but also enhances user trust.
[0051] Furthermore, after completing the monitoring and auditing of the first intelligent agent, the corresponding results can be sent to the first intelligent agent for visual display, enabling users to intuitively understand the security status of the intelligent agent. In actual application scenarios, when the first intelligent agent conducts visual display, on the one hand, it can set up an instruction security score panel to display the security audit results of control instructions, and on the other hand, it can set up a real-time audit monitoring panel to display the security audit results of the first intelligent agent. For example, the instruction security score panel can include multiple visual items such as overall security score, dimension score, risk point marking, and optimization suggestions. Among them, the overall security score can be an overall security score from 0 to 100 points, the dimension score can be an independent scoring radar chart of multiple security dimensions, the risk point marking can intuitively mark potential risk points in the instruction, and the optimization suggestions can be specific instruction optimization suggestions. The real-time audit monitoring panel can include multiple visual items such as visualization of the audit process, risk early warning, intervention operation records, and audit statistics. Among them, the visualization of the audit process can graphically display the audit process of the judging intelligent agent, the risk early warning can display the detected security risks in real time, the intervention operation records can record the intervention operations performed by the judging intelligent agent, and the audit statistics can display the statistical data of the audit results.
[0052] Through the security visualization interface, parents can intuitively understand the security status of the intelligent agent. For example, by graphically displaying the audit process and results, it can help parents understand and manage the security status of the intelligent agent. Another example is that through the instruction security score panel, parents can clearly understand whether the instructions they create have security risks. These visualization tools greatly improve the transparency and interpretability of the system and enhance parents' trust and control over the intelligent agent.
[0053] Generally speaking, the intelligent agent security audit method provided in this embodiment has the following advantages: (1) Closed-loop audit process: A complete closed-loop process from user creation of intelligent agent configuration to the generation of a security audit report, ensuring that each intelligent agent undergoes a comprehensive security audit before release, and avoiding potential risks through strict security reviews; (2) Dual-agent collaboration mechanism: The performer agent and the judge agent are independent of each other but work collaboratively, avoiding the inherent defects of "self-supervision" in a single agent; (3) Dynamic test set: By maintaining an ever-updating security test set, it can timely adjust the test strategy for newly emerging security risks; (4) Real-time visual feedback: Provide real-time visual feedback to users through means such as a security audit dashboard, enhancing the transparency and interpretability of the system; (5) Multi-dimensional security assessment: The judge agent adopts a multi-dimensional assessment framework to comprehensively evaluate the content security from multiple dimensions such as child safety, educational value, and emotional impact; (6) Automated release decision: It can automatically decide whether to allow the agent to release based on the security assessment results, reducing the need for manual intervention.
[0054] Figure 2 It is a schematic block diagram of an agent security audit device 200 provided by an embodiment of the present invention. The device 200 adopts a collaborative architecture of a first agent and a second agent. Among them, the first agent is used to interact with the user, and the second agent is used to implement the security audit device 200. The device 200 includes: An instruction audit unit 201, configured to, when the first agent receives a control instruction sent by the user, obtain the control instruction through the first agent and perform multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result; An audit judgment unit 202, configured to judge whether the control instruction passes the security audit according to the instruction audit result; A first determination unit 203, configured to, if it is determined that the control instruction fails to pass the security audit, optimize the control instruction; A second determination unit 204, configured to, if it is determined that the control instruction passes the security audit, send a security audit passed feedback to the first agent, so that after receiving the security audit passed feedback, the first agent generates corresponding content and publishes it according to the control instruction.
[0055] In one embodiment, the instruction audit unit 201 includes: A semantic analysis unit, configured to perform instruction semantic analysis on the control instruction; wherein, the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection, and instruction contradiction detection; A first setting unit, configured to use the result of the instruction semantic analysis as the instruction audit result.
[0056] In one embodiment, the instruction audit unit 201 further includes: A risk identification unit for identifying potential risks in the control instruction; wherein the potential risk identification includes over-compliance risk, role overstep risk, and content inappropriateness risk; A second setting unit for using the result of the potential risk identification as the instruction audit result.
[0057] In one embodiment, the instruction audit unit 201 further includes: A simulation test unit for performing an instruction simulation test on the control instruction; wherein the instruction test simulation includes virtual scenario test, boundary condition test, and risk prediction; A third setting unit for using the result of the instruction simulation test as the instruction audit result.
[0058] In one embodiment, the audit judgment unit 202 includes: A first scoring unit for performing a first security score on the control instruction according to the following formula: ; wherein, S(a, q) is the security score under the corresponding output a of the first agent when the control instruction is the input q, w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of security evaluation dimensions; A score comparison unit for judging whether the control instruction passes the security audit according to the following formula based on the first security score: ; wherein, D(E) is the security decision result, 1 indicates passing, 0 indicates not passing, θ1 is the minimum security score threshold, θ2 is the average security score threshold, s i is the first security score corresponding to the ith control instruction, and z is the total number of control instructions.
[0059] In one embodiment, the first determination unit 203 includes: An instruction optimization unit for optimizing the control instruction according to the following formula based on a preset optimization suggestion template: ; wherein, Sug(i, R) is the set of suggestions for the control instruction i and the risk assessment R, T j is the optimization suggestion template for the jth type of risk, R j is the score of the jth type of risk, θ j is the threshold of the jth type of risk, and m is the number of risk categories.
[0060] In one embodiment, the intelligent agent security auditing device 200 further includes: A second scoring unit, configured to perform a second security score on the second intelligent agent according to the following formula: ; where Score is the second security score, w i is the weight of the i-th security dimension, s’ i is the score of the i-th security dimension, and n is the number of security assessment dimensions; An early warning judgment unit, configured to judge whether to generate a risk warning for the second intelligent agent according to the following formula based on the second security score: ; where Alert(c,t) is the risk warning result, 1 indicates that the warning is triggered, 0 indicates that the warning is not triggered, c is the content to be evaluated, s’ i (c) is the score of content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of content c, T is the warning threshold of the overall security score, and n is the number of security assessment dimensions; A visualization display unit, configured to send the risk warning to the second intelligent agent when it is determined that a risk warning is generated, and enable the second intelligent agent to visually display the risk warning.
[0061] Since the embodiments of the device part correspond to the embodiments of the method part, for the embodiments of the device part, please refer to the description of the embodiments of the method part, which will not be elaborated here.
[0062] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0063] The embodiment of the present invention further provides an intelligent agent device, which may include a memory and a processor. When the computer program stored in the memory is called by the processor, the steps provided in the above embodiments can be implemented. Of course, the intelligent agent device may further include various network interfaces, power supplies, and other components.
[0064] The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0065] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
Claims
1. An intelligent agent security audit method, using a first intelligent agent and a second intelligent agent collaborative architecture, wherein: The first agent is used to interact with the user, and the second agent is used to implement a security audit method, wherein the method comprises: When the first agent receives a control instruction sent by a user, the control instruction is obtained through the first agent, and a multi-dimensional security audit is performed on the control instruction to obtain a corresponding instruction audit result; Determining whether the control instruction has passed the security audit according to the instruction audit result; If it is determined that the control instruction fails the security audit, optimizing the control instruction; If it is determined that the control instruction passes the security audit, a security audit pass feedback is sent to the first intelligent agent, so that the first intelligent agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
2. The intelligent agent security audit method according to claim 1, characterized in that: The multi-dimensional security audit of the control instruction is performed to obtain the corresponding instruction audit result, including: Performing instruction semantic analysis on the control instruction; wherein the instruction semantic analysis includes instruction integrity check, instruction ambiguity detection and instruction contradiction detection; The result of the instruction semantic analysis is used as the instruction audit result.
3. The intelligent agent security audit method according to claim 1, characterized in that: The performing a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result also includes: Identify potential risks of the control instructions; wherein the potential risks include over-compliance risk, role-crossing risk and content inappropriate risk; The result of the potential risk identification is used as the instruction audit result.
4. The intelligent agent security audit method according to claim 1, characterized in that: The performing a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result also includes: Performing instruction simulation test on the control instruction; wherein the instruction test simulation includes virtual scene test, boundary condition test and risk prediction; The result of the instruction simulation test is used as the instruction audit result.
5. The intelligent agent security audit method according to claim 1, characterized in that: The determining, according to the command audit result, whether the control command has passed the security audit includes: According to the following formula, the control instruction is given a first safety score: ; Where S(a, q) is the safety score of the first agent corresponding to output a when the control instruction is input q, and w j is the weight of the jth security dimension, f j (a, q) is the evaluation function of the jth security dimension, and n is the number of dimensions of security evaluation; Based on the first security score, determine whether the control instruction passes the security audit according to the following formula: ; Where D(E) is the safety decision result, 1 means pass, 0 means fail, θ1 is the minimum safety score threshold, θ2 is the average safety score threshold, s i is the first safety score corresponding to the ith control instruction, and z is the total number of control instructions.
6. The intelligent agent security audit method according to claim 1, characterized in that: The step of optimizing the control instruction comprises: According to the following formula, the control instruction is optimized based on the preset optimization suggestion template: ; Where Sug(i, R) is the set of recommendations for control instruction i and risk assessment R, T j is the optimization suggestion template for the jth risk, R j is the score of the j-th risk, θ j is the threshold of the jth risk category, and m is the number of risk categories.
7. The intelligent agent security audit method according to claim 1, characterized in that: Also includes: According to the following formula, a second safety score is given to the second agent: ; Among them, Score is the second safety score, w i is the weight of the i-th security dimension, s' i is the score of the i-th security dimension, n is the number of dimensions of security assessment; According to the following formula, it is determined whether to generate a risk warning for the second agent based on the second safety score: ; Among them, Alert(c,t) is the risk warning result, 1 means triggering the warning, 0 means not triggering, c is the content to be evaluated, s' i (c) is the score of content c in the i-th security dimension, t i is the warning threshold of the i-th security dimension, Score(c) is the overall security score of content c, T is the warning threshold of the overall security score, and n is the number of dimensions of security assessment; When it is determined that a risk warning is generated, the risk warning is sent to the second intelligent agent, and the second intelligent agent is enabled to perform a visual display of the risk warning.
8. An intelligent agent security audit device, using a first intelligent agent and a second intelligent agent collaborative architecture, wherein: The first agent is used to interact with the user, and the second agent is used to implement the security audit method, characterized in that the device includes: an instruction audit unit, configured to obtain the control instruction through the first agent when the first agent receives the control instruction sent by the user, and perform a multi-dimensional security audit on the control instruction to obtain a corresponding instruction audit result; An audit judgment unit, used to judge whether the control instruction has passed the security audit according to the instruction audit result; A first determination unit, configured to optimize the control instruction if it is determined that the control instruction fails the security audit; The second determination unit is used to send a security audit pass feedback to the first intelligent agent if it is determined that the control instruction has passed the security audit, so that the first intelligent agent generates and publishes corresponding content according to the control instruction after receiving the security audit pass feedback.
9. An intelligent agent device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the intelligent agent security audit method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent agent security audit method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent child rearing system and device based on natural language processing
CN117453867A
Virtual digital human interaction method, system, terminal, equipment and medium
CN117520500A
Construction method and device of multi-component data agent
CN119398092A
Multi-agent model training method, apparatus, electronic device, storage medium and program product
WO2023024378A1
Cited By
Network security auditing method and device for computing power driven intelligent agent of intelligent computing center cloud platform
CN120378349A
An intelligent computing center cloud platform computing power driven intelligent agent network security auditing method and device
CN120378349B
Intelligent agent behavior alignment detection method and system based on thinking chain auditing
CN121094123A