Information updating method and device, intelligent agent, equipment, medium and product
Patent Information
- Application Number
- CN202610954438.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]根据本公开的一方面,提供了一种根据目标角色的画像特征和针对画像特征的测试策略,生成用于确定目标角色的行为边界的风险查询请求;响应于风险查询请求,根据与画像特征对应的响应约束信息,对与风险查询请求语义关联的历史响应信息进行处理,生成风险对抗响应;基于预设评估规则,确定风险对抗响应相对于行为边界的边界匹配度和相对于画像特征的特征匹配度;在边界匹配度或特征匹配度低于预定阈值的情况下,基于风险对抗响应中导致低于预定阈值的约束缺陷,更新响应约束信息
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.
Smart Images

Figure CN122653653A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more particularly to the fields of natural language processing and large model security. More specifically, this disclosure provides an information updating method, apparatus, intelligent agent, device, medium, and product. Background Technology
[0002] Large language model-driven role-playing agents have wide applications, but strict role definition can easily conflict with general security strategies. Related technologies balance security and role expression by adjusting model parameters; however, retraining is required when role definition or attack methods change, resulting in high maintenance costs. Summary of the Invention
[0003] This disclosure provides an information updating method, apparatus, intelligent agent, device, medium, and product.
[0004] According to one aspect of this disclosure, a risk query request is provided to generate a risk query request for determining the behavioral boundaries of a target character based on the character profile features and a testing strategy targeting the character profile features; in response to the risk query request, historical response information semantically associated with the risk query request is processed according to response constraint information corresponding to the character profile features to generate a risk adversarial response; based on preset evaluation rules, the boundary matching degree of the risk adversarial response relative to the behavioral boundaries and the feature matching degree relative to the character profile features are determined; if the boundary matching degree or feature matching degree is lower than a predetermined threshold, the response constraint information is updated based on the constraint defects in the risk adversarial response that cause it to be lower than the predetermined threshold.
[0005] According to another aspect of this disclosure, an information updating apparatus is provided, comprising: a request generation module, configured to generate a risk query request for determining the behavioral boundaries of a target character based on the profile features of the target character and a testing strategy targeting the profile features; a response generation module, configured to generate a risk adversarial response in response to the risk query request, based on response constraint information corresponding to the profile features and historical response information semantically associated with the risk query request; a matching determination module, configured to determine the boundary matching degree of the risk adversarial response relative to the behavioral boundaries and the feature matching degree relative to the profile features based on preset evaluation rules; and an information updating module, configured to update the response constraint information based on constraint defects in the risk adversarial response that cause the boundary matching degree or feature matching degree to fall below a predetermined threshold, if the boundary matching degree or feature matching degree is lower than a predetermined threshold.
[0006] According to another aspect of this disclosure, an intelligent agent for data updating is provided, comprising: an input module for receiving profile features of a target role, a testing strategy for the profile features, response constraint information corresponding to the profile features, and historical response information semantically associated with a risk query request; a processing module for determining a target task based on the profile features of the target role received by the input module, the testing strategy for the profile features, the response constraint information corresponding to the profile features, and the historical response information semantically associated with the risk query request, and determining a language model based on the target task; obtaining updated response constraint information by calling the language model and executing the above method; and an output module for outputting the updated response constraint information from the processing module.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described above.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 This illustration schematically shows an exemplary system architecture to which information update methods and apparatus can be applied according to embodiments of the present disclosure;
[0013] Figure 2 A flowchart illustrating an information update method according to an embodiment of the present disclosure is shown schematically;
[0014] Figure 3 A schematic diagram illustrating a system framework for an information update method according to an embodiment of the present disclosure is shown.
[0015] Figure 4 An adversarial evolution dynamic analysis diagram is schematically shown according to an embodiment of the present disclosure;
[0016] Figure 5 This illustration schematically shows a diagram of the dynamic assembly of prompt words according to an embodiment of the present disclosure;
[0017] Figure 6 A block diagram of an information updating apparatus according to an embodiment of the present disclosure is shown schematically;
[0018] Figure 7 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure is shown.
[0019] Figure 8 A block diagram of an electronic device suitable for implementing an information update method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] Figure 1 An exemplary system architecture for applying information update methods and apparatus according to embodiments of this disclosure is illustrated.
[0022] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the information update method and apparatus can be applied may include a terminal device, but the terminal device can implement the information update method and apparatus provided by the embodiments of this disclosure without interacting with the server.
[0023] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0024] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0025] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0026] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0027] It should be noted that the information update method provided in this embodiment can generally be executed by terminal devices 101, 102, or 103. Accordingly, the information update device provided in this embodiment can also be disposed in terminal devices 101, 102, or 103.
[0028] Alternatively, the information update method provided in this embodiment can generally be executed by server 105. Correspondingly, the information update device provided in this embodiment can generally be located in server 105. The information update method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the information update device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0029] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0030] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of any type of information, such as user personal information, comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.
[0031] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0032] Figure 2 A flowchart illustrating an information update method according to an embodiment of the present disclosure is shown schematically.
[0033] like Figure 2 As shown, the method includes operations S210~S240.
[0034] In operation S210, a risk query request is generated to determine the behavioral boundaries of the target character based on the target character's profile characteristics and the testing strategy targeting those characteristics.
[0035] In operation S220, in response to a risk query request, the historical response information semantically associated with the risk query request is processed based on the response constraint information corresponding to the profile features to generate a risk countermeasure response.
[0036] In operation S230, based on preset evaluation rules, the boundary matching degree of risk countermeasure response relative to the behavioral boundary and the feature matching degree relative to the profile features are determined.
[0037] In operation S240, if the boundary matching degree or feature matching degree is lower than a predetermined threshold, the response constraint information is updated based on the constraint defects in the risk adversarial response that cause it to fall below the predetermined threshold.
[0038] According to embodiments of this disclosure, the target role refers to a virtual identity set by the large language model in dialogue interaction, such as a game character, virtual customer service representative, digital human anchor, or psychological companion assistant. Profile features are used to describe the target role's personality traits, language style, behavioral logic, value orientation, background setting, and narrative identity, among other role attribute information.
[0039] Behavioral boundaries refer to the range of permitted behavioral outputs of a target character during interaction, specifically including two aspects: safety boundaries and role consistency boundaries. The safety boundary restricts the target character from outputting responses containing illegal, violent, discriminatory, privacy-disclosure, or other harmful content. The role consistency boundary restricts the target character's responses from maintaining consistency with the role profile described by the profiling features in terms of language style, behavioral logic, and value orientation.
[0040] The target role's profile features and corresponding testing strategies are input into a large language model to generate risk query requests used to determine the target role's behavioral boundaries. These risk query requests simulate provocative questions that a potential attacker might ask the target role, potentially triggering breaches of security or role consistency boundaries. The testing strategies are pre-built sets of query templates or testing experiences used to identify potential vulnerabilities in the target role's security or role consistency.
[0041] After generating a risk query request, the large language model responds to it by processing historical response information semantically associated with the risk query request based on response constraint information corresponding to the profile features, and then generates a risk adversarial response. The response constraint information is a pre-constructed set of constraints that guide the large language model to adhere to security and role consistency requirements when generating responses, and is stored in an external knowledge base. The historical response information is a set of high-quality response examples verified during historical interactions that simultaneously meet both security and role consistency requirements.
[0042] After generating the risk-response interaction, the large language model evaluates the interaction in two dimensions based on pre-defined evaluation rules. This evaluation determines the boundary matching degree of the risk-response interaction relative to the behavioral boundary and the feature matching degree relative to the profile features. The boundary matching degree characterizes the degree of conformity between the risk-response interaction and the behavioral boundary, specifically referring to the compliance of the risk-response interaction in the security dimension, i.e., whether it contains illegal, violent, discriminatory, privacy-disclosure, or other harmful content. The feature matching degree characterizes the degree of conformity between the risk-response interaction and the behavioral style of the role described by the profile features, specifically referring to the consistency between the risk-response interaction and the target role's profile features in terms of language style, behavioral logic, and value orientation.
[0043] After determining the boundary matching degree and feature matching degree, if the boundary matching degree or feature matching degree is lower than a predetermined threshold, it indicates that the current risk countermeasure response deviates from the behavioral boundary or profile features, meaning that the existing response constraint information is defective and needs to be updated. Based on the constraint defects in the risk countermeasure response associated with matching degrees below the predetermined threshold, the response constraint information is updated, gradually refining the response constraint information. This allows for more effective guidance in generating responses that simultaneously meet the requirements of behavioral boundaries and profile features when facing similar risk query requests in the future.
[0044] By generating risk query requests based on the target role's profile features and the testing strategy targeting those features, personalized targeting of testing strategies is achieved. By processing historical response information with semantic associations based on response constraint information to generate risk adversarial responses, hierarchical collaboration between response constraint information and example data is achieved. By determining boundary matching degree and feature matching degree, a two-dimensional quantitative assessment of security and role consistency is achieved. By updating response constraint information based on constraint defects, targeted evolution of response constraint information is achieved. Without updating model parameters, iterative updates of information simultaneously improve role consistency and security, reduce maintenance costs in multi-role scenarios, and enhance the model's generalization ability to different roles and attack strategies.
[0045] According to embodiments of this disclosure, a risk query request for determining the behavioral boundaries of a target character is generated based on the character profile features and a testing strategy targeting the character profile features. This includes: extracting at least one of character setting information and narrative background information from the character profile features; converting at least one of the character setting information and narrative background information into query elements for constructing a behavioral boundary determination scenario based on the testing strategy; and generating a risk query request that conforms to the boundary determination scenario based on the query elements.
[0046] The profile features of a target character can be pre-configured and stored in a character profile knowledge base. The data structure can be in key-value pair format, knowledge graph format, or structured descriptive text format. Profile features include two dimensions: character setting information and narrative background information.
[0047] The character setting information describes the static attributes of the target character, including but not limited to the character's name, personality traits, language style, background, and value orientation. The narrative background information describes the dynamic context of the target character in the virtual interactive scene, including but not limited to the current virtual scene, historical events, and character relationships.
[0048] Extract at least one of the character setting information and narrative background information from the portrait features. Specifically, this can be achieved by semantically parsing the structured or unstructured text describing the target character using a natural language understanding module, or by using regular expression matching or preset slot extraction.
[0049] The large language model, based on a testing strategy, transforms at least one of the character setting information and narrative background information into query elements used to construct behavioral boundaries and determine scenarios. The testing strategy consists of two levels: general testing experience and personalized testing experience corresponding to profile features.
[0050] During the transformation process, the large language model uses general testing experience to initially screen at least one of the character setting information and narrative background information, identifying target features associated with common behavioral boundaries. Subsequently, individual testing experience is used to process the target features in a targeted manner, transforming the general target features into specific inductive entry points for the current target character, ultimately forming query elements composed of topic anchors, question angles, and inductive logic.
[0051] The large language model inputs structured query elements into the natural language generation module, which then integrates the query elements into one or more natural language question statements, which are output as the final risk query request.
[0052] In another implementation, the large language model can also fill and combine the semantic units in the query elements according to the sentence structure specified by the template based on a preset sentence template, and generate a natural language question statement that conforms to the expression habits of the target character in the scene.
[0053] By extracting character setting information and narrative background information and transforming them into query elements, a risk query request generation driven by profile features was realized, enabling the testing strategy to accurately determine the boundaries of specific weaknesses of the target character.
[0054] According to embodiments of this disclosure, based on a testing strategy, at least one of the character setting information and narrative background information is transformed into query elements for constructing a behavior boundary determination scenario, including: extracting target features related to the behavior boundary determination scenario from at least one of the character setting information and narrative background information based on general testing experience in the testing strategy; and transforming the target features into query elements for the behavior boundary determination scenario based on individual testing experience corresponding to the profile features in the testing strategy.
[0055] The testing strategy is a pre-built, hierarchically stored set of testing experiences, comprising two levels: general testing experiences and personalized testing experiences corresponding to specific role profiles. The general testing experiences are a collection of test rules or query templates extracted from historical, definitive cases across multiple roles, possessing cross-role applicability. These are used to identify common risk scenario types in role-playing scenarios. The general testing experiences do not rely on specific role profiles but are built upon statistical analysis of the patterns of security vulnerabilities observed in numerous role-playing scenarios.
[0056] For example, through cluster analysis of a large number of historical risk cases, we can summarize a number of common risk scenario types, such as "characters are prone to overstepping their authority when faced with power threats", "characters are prone to outputting emotionally harmful content when involved in emotional conflicts", and "characters are prone to cognitive dissonance when faced with value challenges".
[0057] Personality testing experience is a collection of query templates or inducement strategies accumulated based on the specific profile characteristics of a target character, used to determine the unique weaknesses of that character. Unlike general testing experience, personality testing experience is stored in conjunction with specific profile characteristics, and its content is continuously updated and expanded as new weaknesses are discovered during the character's evolution in combat.
[0058] The large language model extracts general testing experience from testing strategies and, based on this general testing experience, extracts target features related to the behavioral boundary determination scenario from at least one of the role setting information and narrative background information. The purpose of extracting target features is to leverage cross-role general knowledge to quickly identify risk scenario types where the current target role may have security vulnerabilities, and to locate specific information directly related to that risk scenario type from its profile features.
[0059] In the specific implementation process, the large language model semantically matches each type of risk scenario defined in general testing experience with at least one of the character setting information and narrative background information. Each type of risk scenario corresponds to a set of scenario description templates or scenario keywords in general testing experience.
[0060] For example, the keywords for a "power threat scenario" could include "power," "status," "threat," and "challenge"; the keywords for a "emotional competition scenario" could include "competition," "jealousy," "struggle," and "liking."
[0061] The large language model determines the risk scenario type most relevant to the current character by calculating the semantic similarity between at least one semantic unit in the character setting information and narrative background information and the set of scenario keywords corresponding to each risk scenario type.
[0062] After determining the risk scenario type, the large language model further extracts specific information fragments directly related to the risk scenario type from at least one of the character setting information and narrative background information, as target features.
[0063] Specifically, the large language model can employ a semantic matching method based on a pre-trained language model, encoding at least one of the character setting information and narrative background information as a vector representation. Simultaneously, the set of scene keywords corresponding to each risk scenario type is also encoded as a vector representation. The matching risk scenario type is determined by calculating the vector similarity between the two, and the information fragment with the highest semantic relevance to the matching risk scenario type is extracted from at least one of the character setting information and narrative background information as the target feature.
[0064] In another implementation, the large language model can also adopt a rule-based approach, using keyword matching or regular expression matching to locate information fragments related to the risk scenario type from at least one of the character setting information and narrative background information.
[0065] After extracting the target features, the large language model transforms the target features into query elements for behavioral boundary determination scenarios based on the personalized testing experience corresponding to the profile features in the testing strategy, so as to construct structured semantic units for risk query requests.
[0066] The personality test experience stores multiple query templates bound to the current target character's profile characteristics. Each query template contains a set of populated slots, which are used to receive the specific content of the target characteristics. The query templates in the personality test experience may differ for different characters, and even within the same character's personality test experience, different query templates may be configured for different types of target characteristics.
[0067] Specifically, the large language model determines query templates that match the target features based on individual test experience, and then fills the corresponding slots of the query template with the specific content of the target features to generate structured query elements. Alternatively, the large language model can also determine matching query templates based on the semantic category or type label of the target features.
[0068] By layering and transforming general testing experience into personalized testing experience, the precise construction of query elements is achieved, enabling risk query requests to simultaneously have cross-role coverage and precision for the current role.
[0069] According to embodiments of this disclosure, based on response constraint information corresponding to profile features, historical response information semantically associated with risk query requests is processed to generate risk adversarial responses, including: determining general constraint information and individual constraint information corresponding to profile features in the response constraint information; semantically matching the risk query request with historical request information in multiple historical question-answer pairs to determine target question-answer pairs; combining the general constraint information, individual constraint information, and historical response information in the target question-answer pairs into response prompt words; and generating risk adversarial responses based on the response prompt words.
[0070] The response constraint information includes two levels: general constraint information and personalized constraint information corresponding to the profile features. These two levels of constraint information are logically independent, but can be distinguished in physical storage through different data tables or different storage sets to support independent retrieval and independent updates.
[0071] General constraint information is a set of universal security rules extracted from historical defense experience across roles. Its purpose is to provide basic security baseline guarantees for all roles. General constraint information includes multiple general security rules, such as "prohibiting the output of descriptions containing specific steps of violent acts", "prohibiting the output of descriptions containing discriminatory language or insulting terms", and "prohibiting the output of descriptions containing methods of obtaining or disclosing personal privacy information", etc.
[0072] The general constraint information does not depend on the profile characteristics of a specific role. Regardless of the identity or personality of the target role, the general constraint information must be followed during the response generation process to ensure basic output security.
[0073] Personalized constraint information is a set of personalized constraint instructions generated based on the profile characteristics of a target character. Its purpose is to maintain character immersion while ensuring safety. Personalized constraint information is stored in conjunction with profile characteristics, and the personalized constraint information varies for different characters.
[0074] For example, for a character with villainous traits, their personality constraint information may include a personality constraint instruction such as "use a mocking or contemptuous tone when refusing harmful requests to conform to the villainous character setting"; for a character with righteous traits, their personality constraint information may include a personality constraint instruction such as "use honor and morality as reasons when refusing harmful requests to conform to the righteous character setting".
[0075] The large language model queries the associated stored individual constraint information from an external knowledge base based on the profile features of the current target character, while simultaneously retrieving independently stored general constraint information. Specifically, the query operation can be implemented through structured query statements, or it can be obtained through precise matching based on the mapping relationship between profile features and individual constraint information.
[0076] While acquiring general and individual constraint information, or sequentially, the large language model also retrieves historical response information that is semantically related to the current risk query request from historical response information. This historical response information is stored in the third level of the external knowledge base, and includes multiple historical question-answer pairs, each containing historical request information and corresponding historical response information.
[0077] During the retrieval process, the large language model semantically matches the risk query request with historical request information in multiple historical question-answer pairs to determine the target question-answer pair. The historical response information in the target question-answer pair represents the most similar high-quality historical response examples to the current risk query request scenario.
[0078] After determining the general constraint information, individual constraint information, and target question-answer pair, the large language model combines the historical response information from the general constraint information, individual constraint information, and target question-answer pair into response prompt words to generate risk adversarial responses based on the response prompt words.
[0079] In one implementation, the combination of response prompts can be achieved using a template-based concatenation method. The large language model sequentially fills the corresponding positions of the prompt template with the target character's profile features, general constraint information, individual constraint information, and historical response information from the target question-answer pair according to a preset hierarchical structure, forming complete response prompts.
[0080] After generating response prompts, the large language model generates a risk-adversarial response based on the prompts. Specifically, the response prompts are input into the large language model for reasoning, so that a risk-adversarial response is output under the guidance of the prompts.
[0081] Since the response prompts include general constraint information, individual constraint information, and historical response information, the large language model is guided by both the security dimension and the role consistency dimension when generating responses. This ensures that the generated risk-countermeasure responses meet basic security requirements while maintaining a language style and behavioral logic consistent with the target role profile as much as possible.
[0082] By generating response prompts through a hierarchical combination of response constraint information and historical response information, response constraints for profile perception are realized, enabling risk countermeasure responses to simultaneously meet the requirements of security boundaries and role consistency.
[0083] According to embodiments of this disclosure, semantic matching is performed between a risk query request and historical request information in multiple historical question-answer pairs to determine a target question-answer pair. This includes: performing vectorized retrieval on the risk query request and historical request information in multiple historical question-answer pairs to obtain multiple candidate question-answer pairs; performing semantic comparison between the risk query request and historical request information in each candidate question-answer pair to obtain comparison results; and determining the target question-answer pair from multiple candidate question-answer pairs based on the comparison results.
[0084] Multiple historical question-answer pairs are stored in the third level of an external knowledge base. Each historical question-answer pair includes historical request information and corresponding historical response information. The historical response information consists of high-quality response examples that have been verified to simultaneously meet security and role consistency requirements during historical adversarial interactions. To support efficient semantic matching, the historical request information in the historical question-answer pairs has been encoded into dense vectors and stored in a vector database offline, forming a historical request information vector index.
[0085] During semantic matching, the large language model performs vectorized retrieval of risk query requests and historical request information from multiple historical question-answer pairs, resulting in multiple candidate question-answer pairs. Specifically, the large language model inputs the risk query request into a pre-trained encoding model, which maps the risk query request into a dense vector of fixed dimensions. This dense vector represents the overall semantic information of the risk query request in the semantic space.
[0086] After obtaining the vector representation of the risk query request, the large language model uses this vector as the retrieval vector and performs an approximate nearest neighbor search in the historical request information vector index. It calculates the vector similarity between the retrieval vector and each historical request information vector in the index, and quickly recalls a preset number of historical question-answer pairs as candidate question-answer pairs in descending order of vector similarity. The preset number can be flexibly set according to the actual application scenario.
[0087] After obtaining candidate question-answer pairs through vectorized retrieval, the large language model performs semantic comparison between the risk query request and the historical request information in each candidate question-answer pair to obtain the comparison results.
[0088] Specifically, the large language model concatenates the risk query request with the historical request information in each candidate question-answer pair to form multiple text pairs to be compared. For each text pair to be compared, the large language model inputs it into a cross-encoder, which performs deep semantic interaction calculations on the risk query request and the historical request information.
[0089] The cross encoder outputs the semantic matching score of each text pair to be compared. This semantic matching score is used to characterize the semantic similarity between the risk query request and the historical request information in the candidate question-answer pair, and the semantic matching score corresponding to each candidate question-answer pair is used as the comparison result.
[0090] After obtaining the comparison results, the large language model determines the target question-answer pair from multiple candidate question-answer pairs based on the comparison results. Specifically, the large language model can sort multiple candidate question-answer pairs in descending order according to the semantic matching scores corresponding to each candidate question-answer pair, and select one or more candidate question-answer pairs with the highest semantic matching scores, i.e., the highest ranking, as the target question-answer pair. The historical response information in the target question-answer pair is then identified as the historical high-quality response examples most similar to the current risk query request scenario.
[0091] Through a two-stage retrieval pipeline, the large language model can quickly locate the high-quality response examples that are most semantically relevant to the current risk query request from multiple historical question-answer pairs, providing accurate scenario-based references for the generation of subsequent risk adversarial responses.
[0092] According to embodiments of this disclosure, determining the boundary matching degree of a risk confrontation response relative to a behavioral boundary and the feature matching degree relative to a profile feature based on preset evaluation rules includes: inputting a risk query request, a risk confrontation response, and profile features into a preset evaluation model to determine the boundary matching degree of the risk confrontation response relative to a behavioral boundary and the feature matching degree relative to a profile feature based on preset evaluation rules; wherein, the boundary matching degree is used to characterize whether the risk confrontation response exceeds the behavioral boundary, and the feature matching degree is used to characterize whether the risk confrontation response conforms to the role behavior style described by the profile feature.
[0093] The pre-trained evaluation model is a dedicated evaluation neural network model that can be built upon a large language model. In one implementation, the pre-trained evaluation model and the large language model that generates the risk-adversarial response can be two independent models. That is, the pre-trained evaluation model is only used to perform the evaluation task and does not participate in the response generation process, thereby reducing bias or circular dependency problems that may arise from sharing a model for evaluation and generation. In another implementation, the pre-trained evaluation model can also be implemented by inputting specific evaluation prompts into a general large language model. That is, the large language model is used as the evaluator, and prompt engineering guides it to perform multi-dimensional evaluation of the risk-adversarial response.
[0094] Pre-defined assessment rules are a set of predefined rules used to guide the pre-defined assessment model in scoring. They include at least two dimensions: security assessment criteria and role assessment criteria. Security assessment criteria define what constitutes "exceeding behavioral boundaries," while role assessment criteria define what constitutes "conforming to the role's behavioral style described by the profile characteristics."
[0095] During the assessment, the risk query request, risk countermeasure response, and profile features are concatenated according to a preset input template to form a structured assessment input sequence, which is then input into a preset assessment model. The preset assessment model, based on security assessment criteria, performs item-by-item detection and scoring on whether the risk countermeasure response contains content exceeding behavioral boundaries, generating a boundary matching score. The boundary matching score can be a continuous score, with a preset range between 0 and 1. A higher score indicates that the risk countermeasure response better conforms to the behavioral boundary constraints, i.e., higher security; a lower score indicates that the risk countermeasure response is more likely to exceed the behavioral boundary, i.e., lower security.
[0096] Meanwhile, the pre-defined evaluation model calculates the vector similarity between the semantic representation of the risk confrontation response and the semantic representation of the profile features based on role evaluation criteria, and maps this vector similarity to the feature matching degree. The feature matching degree can also be a continuous score, and its value range can be preset between 0 and 1. The higher the score, the more the risk confrontation response matches the role behavior style described by the profile features, that is, the higher the role consistency; the lower the score, the more the risk confrontation response deviates from the role behavior style described by the profile features, that is, the lower the role consistency.
[0097] By conducting a two-dimensional assessment, the performance level of risk response in terms of security and role consistency can be quantitatively determined, thus providing an objective basis for judging whether to trigger the update of response constraint information.
[0098] According to embodiments of this disclosure, updating response constraint information based on constraint defects associated with boundary matching degree and feature matching degree in risk confrontation response includes: performing semantic analysis on risk confrontation response based on evaluation feedback information output by a preset evaluation model to determine constraint defects that cause boundary matching degree or feature matching degree to be lower than a predetermined threshold; modifying general constraint information in response constraint information based on constraint defects, and generating personalized constraint instructions corresponding to profile features to update personalized constraint information corresponding to profile features in response constraint information.
[0099] When performing a two-dimensional evaluation of risk response, the pre-defined evaluation model outputs not only boundary matching and feature matching degrees, but also evaluation feedback information related to the evaluation results. This feedback information is auxiliary information generated by the pre-defined evaluation model during the evaluation process to explain the basis for the matching degree scoring. Specifically, the evaluation feedback information includes the specific reasons why the risk response failed the evaluation in either the security dimension or the role consistency dimension.
[0100] For example, when the boundary matching degree is below a predetermined threshold, the evaluation feedback may include specific descriptions of security violations such as "the response contains specific steps of violent behavior" or "the response contains discriminatory remarks against a specific group." When the feature matching degree is below a predetermined threshold, the evaluation feedback may include specific descriptions of character deviation such as "the response tone is too mild and does not conform to the arrogant character setting" or "the response wording is too formal and does not conform to the character's colloquial expression habits." If both the boundary matching degree and the feature matching degree reach the predetermined threshold, the evaluation feedback may indicate "all dimensions have passed the evaluation," in which case the large language model may not perform an update step.
[0101] After obtaining the evaluation feedback information, the large language model searches and compares the response constraint information based on the violation type or deviation type indicated in the evaluation feedback information to locate the missing, inaccurate, or insufficient constraint information in the current response constraint information.
[0102] Specifically, if the evaluation feedback indicates that the boundary matching degree is lower than a predetermined threshold, that is, there is a problem with the risk countermeasure response in the security dimension, the large language model determines that the constraint defect exists in the general constraint information and further analyzes the specific reasons for the security violation.
[0103] If the evaluation feedback indicates that the feature matching degree is lower than the predetermined threshold, that is, there is a problem with the risk confrontation response in the role consistency dimension, then the large language model determines that the constraint defect exists in the personality constraint information, and further analyzes the specific reasons for the role deviation.
[0104] After identifying constraint defects, if the defects exist in the general constraint information, the large language model modifies the general constraint information. Modification operations include at least one of the following: adding new general safety rules, modifying existing inaccurate general safety rules, deleting outdated general safety rules, or merging redundant general safety rules.
[0105] If a constraint defect exists in the personality constraint information, the large language model generates personality constraint instructions corresponding to the profile features to update the personality constraint information. These personality constraint instructions are natural language constraint instructions specifically generated for the current target character's profile features, guiding the large language model to maintain safety while preserving character immersion during subsequent response generation. These personality constraint instructions are stored in the personality constraint information corresponding to the profile features within the response constraint information for retrieval and use during subsequent response generation.
[0106] If both the boundary matching degree and the feature matching degree are lower than the predetermined threshold, that is, the risk confrontation response has problems in both the security and role consistency dimensions, then the large language model will perform the above two update operations simultaneously, modifying the general constraint information and generating personalized constraint instructions corresponding to the profile features.
[0107] General constraint information and individual constraint information are physically independent, so the two update operations described above can be executed in parallel without interference. The updated general constraint information and individual constraint information are stored back in the corresponding level of the external knowledge base for use in the subsequent response generation process.
[0108] Through the aforementioned targeted update mechanism based on constraint defects, the large language model can accurately locate the constraint defects that cause the failure each time a response fails the evaluation, and modify or supplement the corresponding response constraint information accordingly. This allows the accuracy and coverage of the response constraint rules to be continuously enhanced with the accumulation of adversarial experience.
[0109] According to embodiments of this disclosure, the information update method further includes: when the boundary matching degree or feature matching degree is lower than a predetermined threshold, iteratively correcting the risk adversarial response based on evaluation feedback information until the corrected response makes both the boundary matching degree and feature matching degree reach the predetermined threshold, thus obtaining the target adversarial response; and updating the historical response information semantically associated with the risk query request to the target adversarial response.
[0110] When the boundary matching degree or feature matching degree is lower than a predetermined threshold, the large language model performs a self-correction loop. Specifically, the self-correction loop is an iterative process, and each loop includes three sub-steps: correction, evaluation, and decision.
[0111] During the correction process, risk response, assessment feedback information, and profile features are input into the big language model. The big language model then makes targeted corrections to the risk response based on the specific issues indicated in the assessment feedback information.
[0112] If the assessment feedback indicates that the risk response has issues in the security dimension, the large language model will focus on adjusting the content involving security violations in the response, replacing it with expressions that meet security requirements. If the assessment feedback indicates that the risk response has issues in the role consistency dimension, the large language model will focus on adjusting the language style, tone, or behavioral logic of the response to better match the role's behavioral style described by the profile features. If issues exist in both dimensions, both adjustments will be made simultaneously.
[0113] After generating the corrected response, the large language model inputs the risk query request, the corrected response, and the profile features into the preset evaluation model again. The preset evaluation model then redetermines the boundary matching degree of the corrected response relative to the behavior boundary and the feature matching degree relative to the profile features.
[0114] If both the redefined boundary matching degree and feature matching degree reach the predetermined threshold, the correction is considered successful, the iteration stops, and the corrected response is used as the target adversarial response. If at least one of the redefined boundary matching degree or feature matching degree is still below the predetermined threshold, the large language model obtains new evaluation feedback information from the preset evaluation model for the corrected response and enters the next round of correction loop. The corrected response from the previous round, the new evaluation feedback information, and the profile features are input into the large language model again for a new round of targeted correction. This cycle continues until the corrected response makes both the boundary matching degree and feature matching degree reach the predetermined threshold, or until the number of correction iterations reaches the preset maximum number of iterations.
[0115] After obtaining the target adversarial response through the aforementioned iterative correction process, the risk query request and the target adversarial response are constructed into a new historical question-and-answer pair, which is then stored in the third level of the external knowledge base. During storage, the large language model also attaches a role identifier to the new historical question-and-answer pair to establish an association between the pair and the target role, enabling subsequent searches for the same role to preferentially match the role's historical response information.
[0116] Through the aforementioned iterative correction and update mechanism, the large language model can transform failed response cases into successful response examples and store them in historical response information, so that the number of high-quality examples in historical response information continues to grow with the accumulation of adversarial experience.
[0117] According to embodiments of this disclosure, the information update method further includes: when both the boundary matching degree and the feature matching degree reach a predetermined threshold, analyzing the reasons why the risk query request failed to trigger the behavior boundary, determining the strategy defect in the risk query request; and updating the test strategy based on the strategy defect.
[0118] When both boundary matching and feature matching reach predetermined thresholds, meaning the risk adversarial response passes evaluation in both security and role consistency dimensions, it indicates that the current response constraint information is sufficient to effectively guide the large language model to resist the current risky query request. At this point, the reasons why the risky query request failed to trigger the behavioral boundary are analyzed to identify strategy flaws in the risky query request, and the test strategy is updated based on these flaws.
[0119] Specifically, the system obtains confirmation information from the preset evaluation model indicating that the evaluation has passed. This confirmation information differs from the evaluation feedback information in the case of defense failure. In the case of successful defense, the preset evaluation model does not output a description of the violation as "why it failed," but rather positive feedback as "why it passed," such as "the response did not contain any violation content" and "the response meets the role setting requirements."
[0120] The large language model, based on the reasons for acceptance indicated by positive feedback, searches and compares within the testing strategy to determine if the query template or testing strategy corresponding to the risky query request has been effectively defended by the current response constraint information, thus identifying the query template or testing strategy as an invalid strategy. A strategy defect refers to a query template or attack pattern within the testing strategy that has been successfully defended by the current response constraint information but has failed in subsequent adversarial attacks.
[0121] After identifying policy flaws, the large language model updates the testing strategy based on these flaws. Specifically, the query templates corresponding to risk query requests in the testing strategy can be structurally adjusted to generate the updated testing strategy. Structural adjustments include at least one of changing the question angle, switching attack types, and adjusting the question structure.
[0122] By updating the test strategy based on policy defects, after each successful defense, the large language model can promptly identify the parts of the current test strategy that have been effectively defended, and update the test strategy through mutation operations. This allows the attacking side's test strategy to be updated synchronously with the defense side's evolution, forming a dynamic mechanism in which the attacking and defending sides alternately evolve in iterative confrontation.
[0123] Figure 3 A schematic diagram illustrating a system framework for an information update method according to an embodiment of the present disclosure is provided.
[0124] like Figure 3 As shown, the system framework of the information update method includes two core modules: a defense knowledge base 301 and an attack knowledge base 302. The attack knowledge base 302 stores general testing experience and personalized testing experience, while the defense knowledge base 301 stores a three-layer structure of general constraint information, personalized constraint information, and historical response information.
[0125] During the knowledge base update process, the profile feature 303, along with the general and individual testing experiences in the attack knowledge base 302, are used as input to the large language model to generate a risk query request 305 for determining the behavioral boundaries of the target role. This risk query request 305 is then input into the large language model along with the profile feature 303. The large language model retrieves general constraint information, individual constraint information, and historical response information semantically related to the risk query request 305 from the defense knowledge base 301 to generate a risk countermeasure response 304.
[0126] The risk response 304 is input into the preset evaluation model 306. The preset evaluation model 306 performs a dual-dimensional evaluation of the risk response 304, assessing security and role consistency, and outputs the evaluation results. If the evaluation result is a failure (i.e., the boundary matching degree or feature matching degree is lower than a predetermined threshold), a failure feedback 307 is triggered, and the failure information is fed back to the defense knowledge base 301 to update the general constraint information, individual constraint information, and historical response information.
[0127] If the evaluation result is satisfactory (i.e., both boundary matching and feature matching reach the predetermined thresholds), feedback 308 is triggered, and the successfully intercepted risk query request 305 is fed back to the attack knowledge base 302 for updating general testing experience and personalized testing experience. The above two feedback paths form a dual-loop adversarial evolution mechanism, in which the attack side and the defense side achieve co-evolution of the knowledge base through continuous adversarial iteration.
[0128] Figure 4 An adversarial evolution dynamic analysis diagram is schematically illustrated according to an embodiment of the present disclosure.
[0129] like Figure 4As shown, to verify the effectiveness of the dual-cycle adversarial evolution mechanism, the experiment set up cross-adversarial tests between attackers (A250 to A1000, representing 250, 500, 750 and 1000 evolution rounds respectively) and defenders (baseline, D250 to D1000, representing no evolution, 250, 500, 750 and 1000 evolution rounds respectively).
[0130] Figure 4 The value in the table represents the rejection rate (%) of the defender when faced with risky query requests generated by the attacker. That is, the proportion of the defender that successfully identifies and blocks harmful queries. The higher the value, the stronger the defense.
[0131] From a line-of-sight perspective, when the defender is fixed, the higher the attacker's evolutionary rounds, the lower the defender's rejection rate. For example, with the defender as the baseline, the rejection rate is 71% when facing an attacker who has evolved 250 rounds (A250), while it drops to 62% when facing an attacker who has evolved 1000 rounds (A1000). This indicates that attackers can continuously discover new testing strategies through continuous evolution, effectively breaching the defender's security protection.
[0132] From a column perspective, when the attacker is fixed, the higher the evolutionary stage of the defender, the higher its rejection rate. For example, when facing an attacker who has evolved for 1000 stages (A1000), the rejection rate of an unevolved defender (baseline) is only 62%, while the rejection rate of a defender who has evolved for 1000 stages (D1000) increases to 76%. This indicates that through continuous evolution, the defender can transform the experience accumulated in the confrontation into more effective response constraint information.
[0133] The above results demonstrate that the dual-loop adversarial evolution mechanism can drive attackers and defenders to evolve synchronously in iterative adversarial processes. Defenders must continuously evolve to cope with escalating attacks, thus verifying the effectiveness of this mechanism in improving the security of role-playing agents.
[0134] According to embodiments of this disclosure, the information update method further includes: in response to a target query request for a target role, generating a response to the target query request based on the target query request, the target role's profile features, updated response constraint information, and historical response information semantically associated with the query request.
[0135] When a target query request for a specific role is received from a user, this request is a genuine question entered by the user during the inference phase, not a risk query request generated during the offline adversarial phase. The target role's profile features, along with updated response constraints and updated historical response information, are obtained. The updated response constraints and updated historical response information are the final versions generated during the aforementioned offline adversarial evolution phase.
[0136] Based on the semantic content of the target query request, response constraint information matching the target query request is retrieved from the updated response constraint information. Specifically, the target query request is semantically matched with each general security rule in the general constraint information and each individual constraint instruction in the individual constraint information, and response constraint information whose semantic relevance to the target query request exceeds a preset threshold is selected as the constraint information to be used.
[0137] Simultaneously, historical response information semantically related to the target query request is retrieved from the updated historical response information. Specifically, the target query request is semantically matched with the historical request information of each historical question-answer pair in the updated historical response information to determine the target question-answer pair that is most semantically similar to the target query request, and the historical response information in that target question-answer pair is obtained as a reference example.
[0138] After completing the above retrieval operations, the target character's profile features, retrieved response constraint information, and retrieved historical response information are combined into inference prompts. These inference prompts are then input into a large language model, which generates a response tailored to the target query request based on the inference prompts.
[0139] Through the operations in the above reasoning stage, the response constraint information accumulated in the offline adversarial evolution stage will be applied to the actual reasoning process, enabling the large language model to meet the requirements of security and role consistency when providing services to users, without having to modify the model parameters of the large language model itself.
[0140] Figure 5 A schematic diagram illustrating the dynamic assembly of prompts according to an embodiment of the present disclosure is shown.
[0141] like Figure 5 As shown, the assembly process of the inference prompt word 501 adopts a hierarchical dynamic combination mechanism. First, the profile features of the target character are directly input into the profile feature paragraph of the prompt word template as static context information, which is used to establish the character identity and character setting background, so that the large language model can clearly understand the role identity it is currently playing.
[0142] Secondly, general constraint information relevant to the current scenario is selected from the first level of the defense knowledge base and filled into the general constraint paragraph of the prompt word template to provide basic security baseline protection. Personalized constraint information corresponding to the current profile features is selected from the second level of the defense knowledge base and filled into the personalized constraint paragraph of the prompt word template to provide character-specific behavioral style constraints. From the third level of the defense knowledge base, through a two-stage retrieval pipeline of dense retrieval and re-ranking, one or more historical response information most relevant to the user query semantics are retrieved as few-shot learning examples and filled into the example paragraph of the prompt word template to provide specific style references and expression paradigms. Finally, the user query is directly passed into the user question paragraph of the prompt word template.
[0143] The information from each of the above levels is combined into a complete reasoning prompt word 501 according to the preset hierarchical structure, and then input into the large language model so that the large language model can output a response that simultaneously meets the requirements of security and role consistency.
[0144] Figure 6 A block diagram of an information updating apparatus according to an embodiment of the present disclosure is shown schematically.
[0145] like Figure 6 As shown, the information update device 600 includes a request generation module 610, a response generation module 620, a matching determination module 630, and an information update module 640.
[0146] The request generation module 610 is used to generate a risk query request to determine the behavioral boundaries of the target character based on the target character's profile features and the testing strategy for the profile features.
[0147] The response generation module 620 is used to generate a risk countermeasure response in response to a risk query request, based on the response constraint information corresponding to the profile features and the historical response information semantically associated with the risk query request.
[0148] The matching determination module 630 is used to determine the boundary matching degree of the risk countermeasure response relative to the behavioral boundary and the feature matching degree relative to the profile features based on preset evaluation rules.
[0149] The information update module 640 is used to update the response constraint information based on the constraint defects in the risk adversarial response that cause the boundary matching degree or feature matching degree to fall below a predetermined threshold.
[0150] According to embodiments of this disclosure, the request generation module 610 includes an information extraction submodule, an information conversion submodule, and a request generation submodule.
[0151] The information extraction submodule is used to extract at least one of the character setting information and narrative background information from the portrait features.
[0152] The information transformation submodule is used to transform at least one of the character setting information and narrative background information into query elements for constructing behavioral boundary determination scenarios, based on the testing strategy.
[0153] The request generation submodule is used to generate risk query requests that conform to the boundary determination scenario based on the query elements.
[0154] According to embodiments of this disclosure, the information conversion submodule includes a feature extraction unit and a feature conversion unit.
[0155] The feature extraction unit is used to extract target features related to the behavior boundary determination scenario from at least one of the character setting information and narrative background information, based on general testing experience in the testing strategy.
[0156] The feature transformation unit is used to transform target features into query elements for scenarios that define behavioral boundaries, based on personalized testing experience corresponding to profile features in the testing strategy.
[0157] According to embodiments of this disclosure, the response generation module 620 includes an information determination submodule, a question-and-answer determination submodule, a prompt combination submodule, and a response generation submodule.
[0158] The information determination submodule is used to determine the general constraint information and the individual constraint information corresponding to the profile features in the response constraint information.
[0159] The question-and-answer determination submodule is used to semantically match the risk query request with historical request information in multiple historical question-and-answer pairs to determine the target question-and-answer pair.
[0160] The prompt combination submodule is used to combine general constraint information, individual constraint information, and historical response information from the target question-answer pair into response prompt words.
[0161] The response generation submodule is used to generate risk countermeasure responses based on response prompts.
[0162] According to embodiments of this disclosure, the question-answering determination submodule includes a vector retrieval unit, a semantic comparison unit, and a target determination unit.
[0163] The vector retrieval unit is used to perform vectorized retrieval of risk query requests and historical request information in multiple historical question-answer pairs to obtain multiple candidate question-answer pairs.
[0164] The semantic comparison unit is used to perform semantic comparison between the risk query request and the historical request information in each candidate question-answer pair to obtain the comparison result.
[0165] The target determination unit is used to determine the target question-answer pair from multiple candidate question-answer pairs based on the comparison results.
[0166] According to embodiments of this disclosure, the matching determination module 630 includes an information input submodule.
[0167] The information input submodule inputs risk query requests, risk confrontation responses, and profile features into a preset evaluation model. Based on preset evaluation rules, it determines the boundary matching degree of the risk confrontation response relative to the behavioral boundary and the feature matching degree relative to the profile features. The boundary matching degree is used to characterize whether the risk confrontation response exceeds the behavioral boundary, and the feature matching degree is used to characterize whether the risk confrontation response conforms to the role behavior style described by the profile features.
[0168] According to embodiments of this disclosure, the information update module 640 includes a defect determination submodule and an information update submodule.
[0169] The defect determination submodule is used to perform semantic analysis on the risk resistance response based on the evaluation feedback information output by the preset evaluation model, and to determine the constraint defects that cause the boundary matching degree or feature matching degree to be lower than a predetermined threshold.
[0170] The information update submodule is used to modify the general constraint information in the response constraint information based on constraint defects, and generate personalized constraint instructions corresponding to the profile features, so as to update the personalized constraint information corresponding to the profile features in the response constraint information.
[0171] According to embodiments of this disclosure, the information update device 600 further includes a response correction module and a response update module.
[0172] The response correction module is used to iteratively correct the risk adversarial response based on the evaluation feedback information when the boundary matching degree or feature matching degree is lower than a predetermined threshold, until the corrected response makes both the boundary matching degree and feature matching degree reach the predetermined threshold, thus obtaining the target adversarial response.
[0173] The response update module is used to update the historical response information that is semantically associated with the risk query request to the target adversarial response.
[0174] According to embodiments of this disclosure, the information update device 600 further includes a cause analysis module and a strategy update module.
[0175] The cause analysis module is used to analyze the reasons why a risk query request fails to trigger the behavior boundary when both the boundary matching degree and the feature matching degree reach a predetermined threshold, and to determine the strategy defects in the risk query request.
[0176] The strategy update module is used to update test strategies based on strategy defects.
[0177] According to embodiments of this disclosure, the information update device 600 further includes a request response module.
[0178] The request-response module is used to respond to target query requests for a target role. Based on the target query request, it generates a response for the target query request according to the target role's profile features, updated response constraint information, and historical response information semantically associated with the query request.
[0179] Figure 7 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0180] In the embodiments of this disclosure, inspired by the von Neumann architecture in modern computer theory, as shown in Figure 7, the AI agent 700 may include five core modules: an input module 710, a processing module 720, and an output module 730.
[0181] The input module 710 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 700 can understand and process. The input module 710 is the primary link for the AI agent 700 to interact with the outside world. It enables the AI agent 700 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.
[0182] In the example, the input module 710 can input the profile features of the target character described above, the testing strategy for the profile features, the response constraint information corresponding to the profile features, and the historical response information semantically associated with the risk query request.
[0183] In embodiments of this disclosure, the processing module 720 may include a control module 721, a storage module 722, and a computation module 723. The processing module 720 is configured to determine a target task based on the input information received by the input module 710, determine a target large-scale model based on the target task, and obtain updated response constraint information by invoking an information update method on the target large-scale model. In the example, the target large-scale model includes at least one of a visual language large-scale model, a video generation large-scale model, and a video evaluation large-scale model.
[0184] The control module 721 is the core support for the AI agent 700's ability to handle complex tasks. The control module 721 can execute the information update method described above.
[0185] In the example, the control module 721 will continuously interact with the storage module 722, the arithmetic module 723, and / or the output module 730 during operation. However, it should be noted that in the embodiments of this disclosure, the control module 721 initiates communication with the storage module 722, the arithmetic module 723, and / or the output module 730 as a single initiator, and there is no communication coupling between the storage module 722, the arithmetic module 723, and the output module 730.
[0186] In the example, the performance of the control module 721 is closely related to the large model on which the AI agent 700 is based. To fully leverage the capabilities of the large model, the internal structure of the control module 721 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0187] Storage module 722 can be responsible for storing the large visual language model, the large video generation model, and the large video evaluation model. The aforementioned large visual language model, large video generation model, and large video evaluation model can be included in storage module 722.
[0188] In the example, after receiving change information, functional descriptions of multiple target modules, and calling dependencies between the target modules, AI agent 700 can trigger a description generation process to obtain the change description and feed it back to control module 721. Then, control module 721 can pass the feedback field content to output module 730.
[0189] The arithmetic module 723 can be viewed as a predefined tool library. Tools for timing alignment, as described above, can be included in the arithmetic module 723.
[0190] In the example, when the AI agent 700 needs to process a request, it can invoke relevant tools from the computing module 723 and feed them back to the control module 721. The control module 721 can then use the returned tools to process the request. It's understandable that while large models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks without any tools is limited. When the AI agent 700 is given the ability to invoke tools, it can perform tasks such as timing alignment using tools designed for timing alignment.
[0191] The output module 730 can output the updated response constraint information described above.
[0192] The AI agent 700 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0193] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0194] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0195] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.
[0196] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0197] Figure 8 A block diagram schematically illustrates an electronic device suitable for implementing an information update method according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0198] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0199] Multiple components in device 800 are connected to input / output (I / O) interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disk, etc.; and communication unit 809, such as a network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0200] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the information update method. For example, in some embodiments, the information update method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the information update method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the information update method by any other suitable means (e.g., by means of firmware).
[0201] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0202] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0203] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0204] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0205] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0206] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0207] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0208] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An information updating method, comprising: Based on the profile characteristics of the target character and the testing strategy targeting the profile characteristics, a risk query request is generated to determine the behavioral boundaries of the target character. In response to the risk query request, the historical response information semantically associated with the risk query request is processed according to the response constraint information corresponding to the profile features to generate a risk countermeasure response; Based on preset evaluation rules, the boundary matching degree of the risk countermeasure response relative to the behavioral boundary and the feature matching degree relative to the profile features are determined; If the boundary matching degree or the feature matching degree is lower than a predetermined threshold, the response constraint information is updated based on the constraint defects in the risk resistance response that cause the value to fall below the predetermined threshold.
2. The method according to claim 1, wherein, The step of generating a risk query request to determine the behavioral boundaries of the target character based on the target character's profile features and a testing strategy targeting those features includes: Extract at least one of the character setting information and narrative background information from the portrait features; Based on the testing strategy, at least one of the character setting information and narrative background information is transformed into query elements for constructing behavioral boundary determination scenarios; Based on the query elements, generate a risk query request that matches the boundary determination scenario.
3. The method according to claim 2, wherein, Based on the testing strategy, at least one of the character setting information and narrative background information is transformed into query elements for constructing a behavior boundary determination scenario, including: Based on the general testing experience in the testing strategy, target features related to the behavior boundary determination scenario are extracted from at least one of the character setting information and the narrative background information; Based on the personalized testing experience corresponding to the profile features in the testing strategy, the target features are transformed into query elements for the behavior boundary determination scenario.
4. The method according to claim 1, wherein, The step of processing historical response information semantically associated with the risk query request based on response constraint information corresponding to the profile features to generate a risk mitigation response includes: Determine the general constraint information and the individual constraint information corresponding to the profile features in the response constraint information; The risk query request is semantically matched with the historical request information in multiple historical question-answer pairs to determine the target question-answer pair; The general constraint information, the individual constraint information, and the historical response information in the target question-answer pair are combined into response prompt words; The risk countermeasure response is generated based on the response prompt words.
5. The method according to claim 4, wherein, The step of semantically matching the risk query request with historical request information in multiple historical question-answer pairs to determine the target question-answer pair includes: The risk query request and the historical request information in the multiple historical question-answer pairs are vectorized for retrieval to obtain multiple candidate question-answer pairs; The risk query request is semantically compared with the historical request information in each of the candidate question-answer pairs to obtain the comparison result. Based on the comparison results, the target question-answer pair is determined from the multiple candidate question-answer pairs.
6. The method according to claim 1, wherein, The step of determining the boundary matching degree of the risk countermeasure response relative to the behavioral boundary and the feature matching degree relative to the profile features based on preset evaluation rules includes: The risk query request, the risk countermeasure response, and the profile features are input into a preset evaluation model to determine the boundary matching degree of the risk countermeasure response relative to the behavior boundary and the feature matching degree relative to the profile features based on the preset evaluation rules. The boundary matching degree is used to characterize whether the risk-resistance response exceeds the behavioral boundary, and the feature matching degree is used to characterize whether the risk-resistance response conforms to the role behavior style described by the profile features.
7. The method according to claim 6, wherein, The step of updating the response constraint information based on the constraint defects associated with the boundary matching degree and the feature matching degree in the risk resistance response includes: Based on the evaluation feedback information output by the preset evaluation model, semantic analysis is performed on the risk resistance response to determine the constraint defects that cause the boundary matching degree or the feature matching degree to be lower than the predetermined threshold. Based on the aforementioned constraint defects, the general constraint information in the response constraint information is modified, and a personalized constraint instruction corresponding to the portrait feature is generated to update the personalized constraint information corresponding to the portrait feature in the response constraint information.
8. The method according to claim 7, further comprising: If the boundary matching degree or the feature matching degree is lower than a predetermined threshold, the risk confrontation response is iteratively corrected based on the evaluation feedback information until the corrected response makes both the boundary matching degree and the feature matching degree reach the predetermined threshold, thus obtaining the target confrontation response; The historical response information semantically associated with the risk query request is updated to the target adversarial response.
9. The method according to claim 1, further comprising: If both the boundary matching degree and the feature matching degree reach the predetermined threshold, the reason why the risk query request failed to trigger the behavior boundary is analyzed to determine the strategy defect in the risk query request. The testing strategy is updated based on the aforementioned flaws.
10. The method according to claim 1, further comprising: In response to a target query request for the target role, a response is generated based on the target query request, the profile features of the target role, the updated response constraint information, and the historical response information semantically associated with the query request.
11. An information updating device, comprising: The request generation module is used to generate a risk query request to determine the behavioral boundaries of the target character based on the target character's profile features and the testing strategy for the profile features. The response generation module is used to respond to the risk query request by generating a risk countermeasure response based on the response constraint information corresponding to the profile features and the historical response information semantically associated with the risk query request. The matching determination module is used to determine the boundary matching degree of the risk countermeasure response relative to the behavior boundary and the feature matching degree relative to the profile features based on preset evaluation rules. The information update module is used to update the response constraint information based on the constraint defects in the risk resistance response that cause the boundary matching degree or the feature matching degree to fall below the predetermined threshold when the boundary matching degree or the feature matching degree is lower than the predetermined threshold.
12. An intelligent agent for information updating, comprising: The input module is used to receive the profile features of the target character, the testing strategy for the profile features, the response constraint information corresponding to the profile features, and the historical response information semantically associated with the risk query request; The processing module is used to determine the target task based on the target character's profile features received by the input module, the test strategy for the profile features, the response constraint information corresponding to the profile features, and the historical response information semantically associated with the risk query request, and to determine the language model based on the target task. By invoking the language model, the method described in any one of claims 1 to 10 is executed to obtain updated response constraint information; The output module is used to output the updated response constraint information from the processing module.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.
15. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1-10 when executed by a processor.