Dynamic desensitization method and system for agent-oriented data interaction based on sandbox

CN122818345APending Publication Date: 2026-09-25WUHAN YISIJIE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610898880.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

若为了消除这种关联风险而采用过度严苛的阻断或大面积掩码策略,又会导致交互文本中的上下文逻辑严重缺失,破坏了文本的语义保真度,导致后续交互决策失败

Benefits of technology

本发明基于隔离沙箱截获会话并构建会话逻辑流图,利用逆向数据流方程动态计算各节点的活跃隐私语义属性,从而在时间轴和逻辑链条上建立了数据流的溯源映射,能够将当前拟发送的碎片信息与历史已发送的科室、药品等无害信息进行关联回溯,从而在关键泄露发生前定量评估出累积的隐私泄露趋势。同时构建关联路径拓扑并计算最短关联路径距离,将交叉碰撞风险转化为拓扑空间中的几何距离度量,能够模拟外部智能体利用公开排班表或结算碎片数据逆向推理身份的链路,当距离低于安全阈值时标记关联阻断节点,实现靶向性的风险干预,避免了传统方法盲目掩码导致的语义破坏。此外,依据偏差值计算语义稀释量,并利用级联概念树匹配泛化阶数对阻断节点进行语义平滑稀释补偿,将时空特征转化为具有一定模糊度但仍具业务含义的泛化表达,在物理上拉长了碰撞关联路径,既阻断了高概率的身份重标识,又最大程度地保留了理赔审查所需的上下文逻辑。最后通过多维语义保真度校验,定量评估脱敏前后文本在语义空间中的贴合度,确保放行文本在满足严密隐私保护的前提下仍能被准确理解,解决了安全防护与业务可用性之间的物理冲突。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818345A_ABST
    Figure CN122818345A_ABST
Patent Text Reader

Abstract

The application provides a sandbox-based dynamic desensitization method and system for agent data interaction, which comprises the following steps: intercepting agent interaction messages in an isolated sandbox and generating a session logic flow graph; extracting privacy feature attributes and calculating active privacy semantic attributes by using a reverse data flow equation; constructing an associated path topology and calculating the shortest associated path distance to an identity identifier; marking the corresponding node as an associated blocking node when the distance is lower than a security threshold; calculating the semantic dilution amount according to the deviation value, performing semantic smoothing dilution on the blocking node by matching the generalization order of the cascade concept tree, and generating a desensitization candidate text; and releasing the desensitization candidate text when the semantic fidelity is satisfied. The application effectively solves the problem of associated privacy leakage caused by cross-collision of fragmented data in multi-round asynchronous claim interaction, and guarantees the usability of the interaction business.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data desensitization technology, specifically relating to a dynamic desensitization method and system for intelligent agent data interaction based on a sandbox. Background Technology

[0002] In the field of intelligent healthcare, entrusted claims assistance agents can provide necessary supporting materials to external underwriting agents through multi-round, asynchronous data interactions. To protect patient privacy, standard security procedures typically involve removing direct identifying identifiers such as names and ID numbers. However, the medical claims process involves fragmented medical information across multiple dimensions, including the department visited, discharge medications, and consultation time. This information is highly contextually related over time. Traditional static anonymization methods often only filter explicit privacy features in a single message using rules. However, in multi-round asynchronous interactions, sensitive information is often not presented directly as a single high-risk term, but rather as a concatenation of multiple weak privacy attributes through logical connections.

[0003] For example, in the initial interactions, the department information provided by the agent is considered safe due to its lack of specificity. Similarly, in subsequent interactions, the specific names of common medications, without the constraint of a time dimension, are also considered safe medical data. However, when the interaction progresses to later rounds and outputs specific consultation times and other related spatiotemporal features, the external recipient can perform multi-dimensional cross-referencing of the aforementioned department information, specific drug combinations, and precise consultation dates. This can be further matched using external channels such as publicly available hospital expert schedules or fragmented regional medical insurance settlement data, thereby inferring the patient's true identity with a high probability. If overly stringent blocking or large-scale masking strategies are adopted to eliminate this association risk, it can lead to a severe loss of contextual logic in the interactive text, compromising semantic fidelity and causing subsequent interaction decisions to fail. Therefore, how to prevent privacy leaks due to associations in multi-round asynchronous interactions while ensuring the business usability of multi-round interactive texts is a technical problem that needs to be solved. Summary of the Invention

[0004] This invention provides a sandbox-based dynamic desensitization method and system for intelligent agent data interaction to solve the above-mentioned technical problems.

[0005] In a first aspect, the present invention provides a dynamic desensitization method for agent-oriented data interaction based on a sandbox, the method comprising the following steps: Intercept the agent's external interaction messages in the isolation sandbox, extract the session identifier, timestamp and interaction text in the interaction messages, and generate a session logic flow graph by combining the session identifier and timestamp and based on the interaction text; Extract privacy feature attributes and identity identifiers from interactive text, reverse traverse the session logic flow graph based on privacy feature attributes, and use the reverse data flow equation to calculate the active privacy semantic attributes of the current interactive node in the session logic flow graph. Construct an association path topology with active privacy semantic attributes and identity identifiers as topology nodes, and use the association path vector metric algorithm to calculate the shortest association path distance from each active privacy semantic attribute to the identity identifier in the association path topology; When there is a target associated path distance in the shortest associated path distance that is lower than the preset safety threshold, the topology node corresponding to the target associated path distance in the associated path topology is marked as an associated blocking node; The required semantic dilution is calculated based on the deviation between the target association path distance and the security threshold. The generalization order is matched according to the semantic dilution and the cascaded concept tree constructed based on the preset category hierarchy relationship based on privacy feature attributes. The semantic smoothing dilution compensation is performed on the association blocking nodes based on the generalization order to generate desensitized candidate text. Calculate the semantic fidelity of the interactive text and the desensitized candidate text in the multidimensional semantic space, and release the desensitized candidate text when the semantic fidelity meets the preset usability threshold.

[0006] Optionally, the steps of extracting privacy feature attributes and identity identifiers from the interactive text, reverse traversing the session logic flow graph based on the privacy feature attributes, and calculating the active privacy semantic attributes of the current interactive node in the session logic flow graph using the reverse data flow equation include the following steps: The text entity recognition module in the isolation sandbox is invoked to perform entity tagging on the interactive text in order to extract privacy feature attributes and identity identifiers; The reverse control flow analysis algorithm is used to reverse traverse the session logic flow graph to determine the attribute lifecycle of privacy feature attributes in the current interaction node of the session logic flow graph. Configure the initial active state for privacy feature attributes within the attribute lifecycle and generate an initial active variable set; The initial set of active variables is used as the input parameter of the inverse data flow equation in the inverse control flow analysis algorithm, and the active privacy semantic attributes existing in the current interaction node are obtained by solving the inverse data flow equation.

[0007] Optionally, obtaining the active privacy semantic attributes existing in the current interaction node by solving the inverse data flow equation includes the following steps: The conversation logic flow graph is divided into basic interactive semantic blocks consisting of single question-and-answer interactions; Extract the privacy feature entities of the new input in each basic interactive semantic block as a local generated attribute set, and mark the privacy feature attributes of the privacy feature entities in the local generated attribute set as active states; Based on the locally generated attribute set after attribute tagging, the input active variables and output active variables of each basic interactive semantic block are calculated using the reverse survival propagation algorithm. The reverse data flow equation is established by combining the input active variables and output active variables. During the reverse traversal of the session logic flow graph, privacy feature attributes that have exceeded their lifespan are deregistered according to the attribute lifespan endpoint of the attribute lifespan, and the active state corresponding to each basic interaction semantic block is updated. Iteratively solve the inverse data flow equation until the equation converges, and output the active privacy semantic attributes after convergence.

[0008] Optionally, the step of generating a local attribute set based on attribute tagging, calculating the input and output active variables of each basic interaction semantic block using the reverse survival propagation algorithm, and establishing the reverse data flow equation by combining the input and output active variables includes the following steps: Identify the predecessor and successor semantic nodes of each basic interaction semantic block in the session logic flow graph; Using the locally generated attribute set after attribute labeling as the local generation source, an information transmission survival matrix based on predecessor semantic nodes and successor semantic nodes is constructed. The information transmission survival matrix represents the flow relationship of the reverse survival propagation algorithm between different basic interactive semantic blocks. Based on the reverse survival propagation algorithm, a reverse data flow equation is established that uses the input active variables of the successor semantic node as the output active variables of the predecessor semantic node. Perform iterative calculations on the inverse data flow equations under a preset maximum iteration threshold constraint; When the change between two adjacent iterations of the iterative calculation is lower than the preset convergence accuracy, the iterative calculation stops, the inverse data flow equation is solved, and the active privacy semantic attributes are output.

[0009] Optionally, the step of calculating the shortest path distance from each active privacy semantic attribute to the identity identifier in the association path topology using the association path vector metric algorithm includes the following steps: Retrieve the probability of direct association between active privacy semantic attributes and identity identifiers from a pre-built prior knowledge relationship graph; The direct association probability is converted into the edge association weight between topological nodes in the associated path topology; Initialize a distance update vector table pointing to the topology node corresponding to the identity identifier for each active privacy semantic attribute; The shortest associated path distance in the updated vector table is updated by using the associated path vector metric algorithm and updating the distance through information exchange between neighboring topological nodes. Based on the distance-updated vector table, and using the path convergence function, the shortest associated path distance from the topology node corresponding to each active privacy semantic attribute to the topology node corresponding to the identity identifier is calculated in the current associated path topology state.

[0010] Optionally, the step of using the associated path vector metric algorithm and updating the shortest associated path distance in the distance update vector table through information interaction between neighboring topological nodes includes the following steps: Simulate and construct topology status messages between various topology nodes in the associated path topology; Collect the latest edge association weights between topological nodes, and construct a single-hop distance matrix that reflects the local proximity relationships in the topology of associated paths based on the latest edge association weights; Control each active privacy semantic attribute corresponding to the topology node to broadcast the current distance update vector table to neighboring topology nodes through topology status messages; The loop elimination operation blocks the cyclic announcement of topology status messages in the closed path of the associated path topology. When a sudden change in the shortest associated path distance is detected in the distance update vector table, the risk reversal mechanism is triggered and the infinite distance value representing the path blockage is written as the updated shortest associated path distance into the topology status message; Based on the received topology status message and combined with the single-hop distance matrix, the distance update vector table of each topology node is updated using the path distance algorithm to complete the update calculation of the shortest associated path distance.

[0011] Optionally, the step of matching the generalization order based on the semantic dilution amount and the cascaded concept tree constructed based on the preset category hierarchy relationship of privacy feature attributes, and performing semantic smoothing dilution compensation on the associated blocking nodes based on the generalization order to generate desensitized candidate text includes the following steps: Load the preset category hierarchy of privacy feature attributes and construct a cascaded concept tree based on the category hierarchy; Establish the correspondence between semantic dilution and the depth of the cascaded concept tree hierarchy, and match to obtain the target generalization order; Based on the target generalization order, a progressively upward generalization mapping algorithm is performed on the associated blocking nodes in the cascaded concept tree to obtain the target generalized entity; Extract the syntactic skeleton of the interactive text containing the associated blocking nodes, and replace the associated blocking nodes in the interactive text with placeholder indicator tags; The target generalized entity is filled into the syntactic skeleton corresponding to the placeholder indicator label to achieve semantic smoothing dilution compensation and generate desensitized candidate text.

[0012] Optionally, the step of performing a progressively upward generalization mapping algorithm on the associated blocking nodes in the cascaded concept tree according to the target generalization order to obtain the target generalized entity includes the following steps: Initialize the associated blocking node in the cascaded concept tree based on the target generalization order and calculate the initial semantic heat. Calculate the semantic similarity loss between the node to be evaluated and its parent node at the next higher level; The generalized cross-layer probability is calculated by combining the initial semantic heat and semantic similarity loss. If the randomly generated value is lower than the generalized cross-layer probability, the parent node of the previous level is taken as the new node to be evaluated; otherwise, the current node to be evaluated is determined as a temporary entity. The initial semantic heat is iteratively reduced using a preset heat decay coefficient until the initial semantic heat is lower than a preset termination threshold, at which point the temporary entity is used as the target generalized entity.

[0013] In a second aspect, the present invention also provides a sandbox-based dynamic desensitization system for agent-oriented data interaction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the sandbox-based dynamic desensitization method for agent-oriented data interaction as described in any one of the first aspects.

[0014] Thirdly, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform a sandbox-based dynamic desensitization method for agent-oriented data interaction according to any one of the first aspects.

[0015] The beneficial effects of this invention are: This invention intercepts sessions using an isolated sandbox and constructs a session logic flow graph. It dynamically calculates the active privacy semantic attributes of each node using inverse data flow equations, establishing a source mapping of data flow along the timeline and logical chain. This allows for the association and backtracking of currently intended fragmented information with previously sent harmless information such as departments and medications, enabling a quantitative assessment of accumulated privacy leakage trends before critical leaks occur. Simultaneously, it constructs an association path topology and calculates the shortest association path distance, transforming the risk of cross-collision into a geometric distance metric in the topological space. This simulates the link where external agents use public schedules or settlement fragmented data to infer identities. When the distance is below a safety threshold, association blocking nodes are marked, achieving targeted risk intervention and avoiding semantic destruction caused by blind masking in traditional methods. Furthermore, it calculates semantic dilution based on deviation values ​​and uses cascaded concept trees to match the generalization order for semantic smoothing dilution compensation of blocking nodes. This transforms spatiotemporal features into generalized expressions with a certain degree of ambiguity but still business meaning, physically lengthening the collision association path. This not only blocks high-probability identity re-identification but also preserves the contextual logic required for claims review to the greatest extent possible. Finally, through multi-dimensional semantic fidelity verification, the degree of fit between the text before and after de-identification in the semantic space is quantitatively evaluated, ensuring that the released text can still be accurately understood under the premise of strict privacy protection, thus resolving the physical conflict between security protection and business availability. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a dynamic desensitization method for agent-oriented data interaction based on a sandbox, as described in one embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0019] Figure 1 This is a flowchart illustrating a sandbox-based dynamic desensitization method for agent-oriented data interaction in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps. For example Figure 1 As shown, the dynamic desensitization method for agent-oriented data interaction based on a sandbox disclosed in this invention specifically includes the following steps: S101. Intercept the agent's external interaction messages in the isolation sandbox, extract the session identifier, timestamp and interaction text in the interaction messages, and generate a session logic flow graph by combining the session identifier and timestamp and based on the interaction text.

[0020] The process involves intercepting interaction messages sent from an agent operating within an isolated sandbox to an external network using a network protocol parser. Protocol stripping is performed on these messages to extract the session identifier, timestamp, and interaction text from the message header. Multiple interaction messages are then grouped based on their session identifiers, with interaction texts sharing the same identifier grouped into a single session set. The interaction texts within each session set are then linearly arranged according to their timestamps. Semantic dependency analysis and contextual relevance calculation are used to identify the semantic connections and intent transitions between interaction texts with adjacent timestamps. Interaction texts are treated as topology nodes, and semantic connections as directed edges. Connecting adjacent directed edges to these topology nodes automatically generates a session logic flow graph that reflects the evolution of the session. In practice, by analyzing the multi-turn question and answer under the session identifier, the output of the previous question and answer is associated with the input of the subsequent question and answer. If the subsequent text is found to refer to the entity of the previous text or to logically progress, a directional directed arc is established between the corresponding nodes. Finally, the session identifier, timestamp and interactive text are transformed into a session directed acyclic graph with a topological structure to intuitively present the evolution of the session interaction in the time and logical dimensions.

[0021] S102. Extract privacy feature attributes and identity identifiers from the interactive text, reverse traverse the session logic flow graph based on the privacy feature attributes, and use the reverse data flow equation to calculate the active privacy semantic attributes of the current interactive node in the session logic flow graph.

[0022] This process utilizes an entity naming recognition model from natural language processing to scan interactive text, identifying and extracting identity identifiers such as names, ID numbers, and phone numbers, as well as privacy-related attributes such as geographical location, salary range, and interests. Following the directed edges in the session logic flow graph, starting from the current, most recent, or last interactive node, a depth-first or breadth-first traversal is performed upstream, moving backwards along the timestamp and logical flow. For each traversed interactive node, a reverse data flow equation based on reverse control flow is established. The privacy semantic state of subsequent nodes is used as input to the current node, and the active privacy semantic attributes at the current interactive node are calculated using a data flow aggregation operator. Specifically, the active variable set is defined as privacy-related attributes referenced after the current node and not redefined or overridden during this period. By solving the reverse data flow equation, it is calculated which privacy-related semantic attributes still have the potential to be associated with identity identifiers in the current session state. Through iterative solving of the reverse data flow equation, it is possible to identify which privacy-related attributes are still semantically active at the current session interactive node, providing data support for subsequent inference path analysis.

[0023] S103. Construct an association path topology with active privacy semantic attributes and identity identifiers as topology nodes, and use the association path vector metric algorithm to calculate the shortest association path distance from each active privacy semantic attribute to the identity identifier in the association path topology.

[0024] The process involves identifying active privacy semantic attributes and identity identifiers as vertices in the association path topology. Potential associations between these vertices are retrieved from a pre-constructed social relationship or common-sense probability network, establishing edges between them to complete the association path topology construction. An association path vector metric algorithm is employed, assigning a distance vector table pointing to the topology node representing the identity identifier to each topology node representing an active privacy semantic attribute. Each topology node exchanges distance vector information with its neighbors, continuously updating the shortest distance to the identity identifier based on the association weights on the edges. Specifically, the association path vector metric algorithm runs within the topology network, using the inverse of the co-occurrence probability or semantic association distance between privacy feature attributes as edge weights. Each node maintains a state vector and continuously performs distributed iterative computation. After changes in the topology or information exchange, the algorithm outputs the shortest association path distance from any active privacy semantic attribute node to the identity identifier node. This vector-based iterative computation method allows for the determination of the link length and inference difficulty for attribution reasoning from different dimensions of privacy feature attributes to the identity identifier within the current knowledge association network.

[0025] S104. When there is a target associated path distance in the shortest associated path distance that is lower than the preset safety threshold, mark the topology node corresponding to the target associated path distance in the associated path topology as an associated blocking node.

[0026] The process involves setting a security threshold to represent the degree of security isolation. The shortest path distance from each active privacy semantic attribute to the identity identifier is calculated and compared with this threshold. If a shortest path distance is found to be less than the threshold, it indicates an overly strong association between the active privacy semantic attribute and the identity identifier, posing a high risk of leaking the true identity through association inference. In this case, the topology node with a distance below the security threshold is marked as an association blocking node in the association path topology. In practice, the security threshold is statically configured or dynamically adjusted based on the specific application scenario and security level. The comparison process iterates through all shortest path distances. Once a target association path distance meeting the security threshold condition is detected, the corresponding active privacy semantic attribute name is extracted, and the node is highlighted or its attribute is set in the topology, establishing it as a target for subsequent association severing or obfuscation. This threshold-based determination identifies which specific locations or attribute dimensions of privacy information in the conversational text need to be blocked, thereby preventing the channel for locking user identity through multi-attribute joint inference.

[0027] S105. Calculate the required semantic dilution amount based on the deviation between the target association path distance and the security threshold. Match the generalization order based on the semantic dilution amount and the cascaded concept tree constructed based on the preset category hierarchy relationship of privacy feature attributes. Perform semantic smoothing dilution compensation on the association blocking nodes based on the generalization order to generate desensitized candidate text.

[0028] The process involves calculating the numerical difference between the security threshold and the distance to the target association path as the deviation value. The required semantic dilution is obtained through forward mapping based on the magnitude of the deviation value; a larger deviation value indicates a greater required semantic dilution. A cascaded concept tree constructed based on a privacy feature attribute classification system is loaded. Based on the magnitude of the semantic dilution, the corresponding tree level depth in the cascaded concept tree is matched as the generalization order. According to the generalization order, upward concept generalization mapping is performed on the specific privacy-sensitive words marked as association blocking nodes to obtain a more abstract target generalized entity. In specific implementations, if the association blocking node is a specific drug name and the deviation value is large, upward generalization based on the cascaded concept tree replaces the drug name with a drug type. The syntax tree structure of the original interactive text is extracted, and the positions of the association blocking nodes are replaced with placeholder indicators. Then, the generated target generalized entities are filled into the positions of the placeholder indicators, thereby semantically diluting and smoothing the transition of sensitive parts in the original interactive text, ultimately generating desensitized candidate text to replace the original text. Through this concept tree-based hierarchical generalization method, semantic coherence is preserved while reducing the possibility of association reasoning.

[0029] S106. Calculate the semantic fidelity of the interactive text and the desensitized candidate text in the multidimensional semantic space. When the semantic fidelity meets the preset usability threshold, the desensitized candidate text is allowed.

[0030] The process involves using a high-dimensional vector representation model to project both the original interactive text and the newly generated anonymized candidate text into a multi-dimensional semantic space, resulting in two corresponding semantic feature vectors. The cosine similarity or distance between these two semantic feature vectors is used to measure the semantic fidelity of the anonymized candidate text relative to the original interactive text. This semantic fidelity is compared to a preset usability threshold. If the semantic fidelity is greater than or equal to the usability threshold, it indicates that the anonymized candidate text has filtered out sensitive related information while still preserving the semantics and usability of the original text. In this case, the anonymized candidate text is allowed to be sent out in place of the original interactive text. In practice, if the semantic fidelity fails to reach the usability threshold, the semantic loss of the current anonymized text is deemed too large, triggering a mechanism to lower the generalization order or reselect a generalized entity until the generated candidate text meets both the safety distance requirement and passes the semantic fidelity check. This dual-loop constraint mechanism enables proactive, secure, and highly semantically valuable dynamic anonymization and release of data sent out by the intelligent system at the sandbox boundary.

[0031] In one implementation, the steps of extracting privacy feature attributes and identity identifiers from the interactive text, reverse traversing the session logic flow graph based on the privacy feature attributes, and calculating the active privacy semantic attributes of the current interactive node in the session logic flow graph using the reverse data flow equation include the following: The text entity recognition module in the isolation sandbox is invoked to perform entity tagging on the interactive text in order to extract privacy feature attributes and identity identifiers; The reverse control flow analysis algorithm is used to reverse traverse the session logic flow graph to determine the attribute lifecycle of privacy feature attributes in the current interaction node of the session logic flow graph. Configure the initial active state for privacy feature attributes within the attribute lifecycle and generate an initial active variable set; The initial set of active variables is used as the input parameter of the inverse data flow equation in the inverse control flow analysis algorithm, and the active privacy semantic attributes existing in the current interaction node are obtained by solving the inverse data flow equation.

[0032] In this implementation, a text entity recognition module deployed within the isolation sandbox is invoked to perform word segmentation and part-of-speech tagging on the acquired interactive text, thereby identifying and extracting privacy-related attributes and identity identifiers from the text. The text entity recognition module employs a pre-trained bidirectional gated recurrent unit combined with a conditional random field model to perform sequence labeling on the input interactive text. For each character in the interactive text, the probability distribution of its belonging to a specific entity such as a person's name, organization name, place name, phone number, or account is calculated, and the sequence of labels with the highest probability is generated. Through this labeling operation, identity identifiers in the interactive text are extracted and categorized into a set. It extracts and categorizes privacy-related attributes such as geographical location, occupation, family status, and hobbies into a set. In specific implementation, it is assumed that the interactive text is represented as follows: The text entity recognition module establishes a mapping relationship. ,in Indicates the identified Individual identifier, Indicates the identified Each privacy feature attribute. Because the isolation sandbox logically isolates the external network, the above entity recognition and tagging process runs securely within the internal environment of the isolation sandbox, effectively eliminating the risk of leakage of interactive text before it has been securely filtered, and ensuring the reliable operation of subsequent security checks.

[0033] Using a reverse control flow analysis algorithm, the algorithm traverses the conversation logic flow graph, which is constructed from multiple interactive text nodes according to their logical relationships, to determine the lifecycle of each privacy feature attribute within the graph. The algorithm starts from the most recently generated interactive node and visits its predecessor nodes sequentially along the reverse direction of the directed edges. The attribute lifecycle refers to the entire logical span from when a privacy feature attribute is first introduced into the conversation to when it is last referenced and participates in semantic reasoning. In practice, for each extracted privacy feature attribute... The lifespan of the attribute is defined as an interval. ,in This indicates the introduction of privacy feature attributes. The starting interaction node, This indicates the last time a privacy feature attribute was detected. Terminating interaction nodes with semantic relevance. The reverse control flow analysis algorithm analyzes the content of nodes along the traversal path, examines the referential relationships and semantic omissions between different interactive texts, and determines privacy-related attributes. Does the contextual relationship still exist among these predecessor nodes? Once at an earlier predecessor node, are privacy-preserving attributes... If the semantics no longer generate a correlation, it is determined that the boundary of the attribute lifecycle has been reached, and the scope of dynamic desensitization analysis is locked.

[0034] After determining the lifetime of each privacy feature attribute, an initial active state is configured for each privacy feature attribute within its corresponding lifetime interval, thereby generating an initial active variable set. The active state represents the reasoning potential of a specific privacy feature attribute at a given session interaction node, which can be associated with an identity identifier. In specific implementation, for any privacy feature attribute... and the interaction nodes in the session logic flow graph If interactive nodes Within the attribute lifecycle range Within this, there are privacy feature attributes. Configure an initial state value representing the active state and set privacy feature attributes. Add interactive nodes The corresponding initial active variable set. The initial active state configuration function is established as follows: A state value of 1 indicates that the privacy feature attribute is active, while a state value of 0 indicates that it is inactive. All privacy feature attributes with a state value of 1 are collected and combined to form an interaction node. Corresponding initial active variable set By configuring a dedicated set of initial active variables for different interaction nodes, the initial value range of the subsequent reverse data flow analysis is limited, meaningless global redundant calculations are avoided, and the efficiency of state analysis is improved.

[0035] The generated initial set of active variables is used as the input parameter for the reverse data flow equation in the reverse control flow analysis algorithm. By establishing and iteratively solving the reverse data flow equation, the active privacy semantic attributes of the current interaction node in the session logic flow graph are calculated. The reverse control flow analysis algorithm calculates the active variable states at the input and output ends of each interaction node through the reverse propagation of the data flow. In specific implementation, for the first interaction node in the session logic flow graph... Interactive nodes Based on the initial active variable set The corresponding reverse data flow equation is established as follows: in, Indicates the flow into the interaction node The set of active input variables, Indicates the outflow of interaction nodes The output active variable set, and It is an interactive node It is obtained by aggregating the set of active input variables of all successor nodes. Indicates at the interaction node The set of newly generated active variables locally, Indicates at the interaction node The set of cancelled variables. Substituting the initial set of active variables into the inverse data flow equation, and performing iterative calculations, convergence is achieved when the state no longer changes, and the output is... This refers to the active privacy semantic attributes that ultimately exist at the current interaction node, accurately defining the input boundaries of sensitive reasoning association analysis.

[0036] In one implementation, obtaining the active privacy semantic attributes existing at the current interaction node by solving the inverse data flow equation includes the following steps: The conversation logic flow graph is divided into basic interactive semantic blocks consisting of single question-and-answer interactions; Extract the privacy feature entities of the new input in each basic interactive semantic block as a local generated attribute set, and mark the privacy feature attributes of the privacy feature entities in the local generated attribute set as active states; Based on the locally generated attribute set after attribute tagging, the input active variables and output active variables of each basic interactive semantic block are calculated using the reverse survival propagation algorithm. The reverse data flow equation is established by combining the input active variables and output active variables. During the reverse traversal of the session logic flow graph, privacy feature attributes that have exceeded their lifespan are deregistered according to the attribute lifespan endpoint of the attribute lifespan, and the active state corresponding to each basic interaction semantic block is updated. Iteratively solve the inverse data flow equation until the equation converges, and output the active privacy semantic attributes after convergence.

[0037] In this embodiment, the conversation logic flow graph is refined and segmented according to the conversation structure and semantic boundaries, into multiple basic interactive semantic blocks composed of single question-and-answer interactions. The conversation logic flow graph typically contains multi-turn, continuous contextual information. To reduce the computational overhead of reverse data flow analysis and control the anonymization granularity, the interactive text nodes in the conversation logic flow graph need to be merged and divided. Specifically, the user-inputted single prompt text and the agent's single response text to that prompt are merged into an indivisible minimum execution unit, i.e., a basic interactive semantic block. Let the conversation logic flow graph be represented as a graphical structure. It contains multiple interactive nodes that represent interactive text. Using partitioning rules to The interactive nodes in the code are divided into a set of non-overlapping basic interactive semantic blocks. ,in This represents the total number of basic interaction semantic blocks. For any given basic interaction semantic block... It contains a question-and-answer pair, that is ,in Indicates the first The user input interaction node of the wheel, This represents the corresponding agent response interaction node, thus avoiding semantic fragmentation caused by analyzing input or response separately. By performing this block-like partitioning of the session in the sandbox, the number of edges in the control flow graph can be significantly reduced, resulting in a substantial reduction in the size of the subsequent data flow transmission matrix, thereby reducing memory resource consumption while ensuring analysis quality.

[0038] The text content within each basic interactive semantic block is analyzed in depth. Newly input privacy-preserving entities are extracted and used as a set of locally generated attributes for the current block. The active status of the privacy-preserving attributes contained in this set is then marked as active. The set of locally generated attributes refers to privacy-preserving entities newly generated in the current question-and-answer interaction that were not previously defined in the current session flow. In specific implementation, this is done for each basic interactive semantic block. The text entity recognition module is invoked to extract the set of privacy-preserving features. The entity is identified by removing privacy entities that have already appeared in the preceding basic interaction semantic block, thereby determining the set of locally generated attributes that are introduced only for the first time in the current block. The formula for extracting locally generated attribute sets is defined as follows: Subsequently, the active status of all privacy-preserving attributes within the set is forcibly set, thus transitioning them to an active state. This method of locally marking only newly introduced privacy-preserving attributes enables dynamic tracking of incremental privacy information, avoiding interference from historical redundant information in the security judgment of the current node. The set attributes act as the source of active variables in the subsequent reverse data flow, laying the foundation for data flow input for subsequent survival analysis and narrowing the scope of analysis.

[0039] Based on the labeled locally generated attribute set, the inverse survival propagation algorithm is used to calculate the input and output activity variables of each basic interaction semantic block in the session logic flow graph. These two variables are then combined to establish an inverse data flow equation reflecting the transition state of privacy attributes. The input activity variable represents the set of privacy feature attributes that are already active before entering the current semantic block, and the output activity variable represents the set of privacy feature attributes that remain active after leaving the current semantic block. In specific implementation, for any basic interaction semantic block... The reverse survival propagation algorithm first collects the input active variables of all subsequent semantic blocks of the current semantic block, and then calculates the output active variable set of the current semantic block by taking the union of the input variables. Based on this, the set of locally generated attributes first introduced in the current semantic block is combined. and the set of dead variables that have been cancelled in the current semantic block. The inverse data flow equation used to describe the transmission relationship between input and output active variables is established as follows: in, This represents the current basic interactive semantic block obtained through computation. The set of active input variables. By establishing such an algebraic equation containing the relationship between input and output, it is possible to clearly characterize the flow and survival status of privacy information at the semantic block boundary using only local set operations, without needing to analyze global semantics, thus greatly improving the agility of privacy exposure risk assessment.

[0040] During the reverse traversal of the session logic flow graph, the lifecycles of the currently traversed basic interaction semantic blocks and each privacy feature attribute are compared in real time. Based on the attribute's lifecycle endpoint, privacy feature attributes that have exceeded their lifecycle are deregistered, thus dynamically updating the active state corresponding to each basic interaction semantic block. In reverse control flow analysis, privacy attributes are generated from the moment they are introduced and disappear at the end of their lifecycle. Since the reverse traversal is performed from back to front, once a privacy feature attribute is passed... The starting interaction node Moving further upstream, this attribute should no longer remain active. In practice, if the current basic interaction semantic block is detected during reverse traversal... In attribute life cycle In addition, if it indicates that the attribute has died upstream in the current block, it will be added to the deregistered variable set. The update formula using the cancellation condition is expressed as: By unregistering attributes belonging to the set and setting their active state value from 1 to 0, their propagation is prevented during reverse traversal. This dynamic unregistration and update mechanism ensures the validity of active variables during data flow propagation, avoids the erroneous propagation of outdated privacy information that is no longer relevant, thereby significantly eliminating the phenomenon of state over-approximation and significantly improving the efficiency and accuracy of subsequent inference path analysis.

[0041] The established inverse data flow equations are solved iteratively until the sets of active input and output variables for each basic interactive semantic block in the equations no longer change. This indicates that the equations have converged, and the converged active privacy semantic attributes are output. Because context jumps and logical branches may exist in the session logic flow graph, the propagation relationship of privacy attributes between different semantic blocks requires multiple iterations to reach stability. Specifically, the set of active input variables is initialized in the 0th iteration, and in each subsequent iteration... In this process, using the node states obtained from the previous iteration, the set of active input variables for the current round is iteratively calculated through the inverse data flow equation. The iterative calculation formula is: in, Represents basic interactive semantic blocks The set of all directly succeeding semantic blocks. When the convergence condition is satisfied... Corresponds to all If all conditions are met, stop the loop and set the input active variable set corresponding to the current interactive semantic block obtained at this time. This step serves as the final output of the active privacy semantic attribute. It enables a global solution for the information transmission state across the entire graph, ensuring the completeness and accuracy of the active attribute analysis results.

[0042] In one implementation, a locally generated attribute set is generated based on attribute tagging, and the input and output active variables of each basic interaction semantic block are calculated using the reverse survival propagation algorithm. The inverse data flow equation is established by combining the input and output active variables, including the following steps: Identify the predecessor and successor semantic nodes of each basic interaction semantic block in the session logic flow graph; Using the locally generated attribute set after attribute labeling as the local generation source, an information transmission survival matrix based on predecessor semantic nodes and successor semantic nodes is constructed. The information transmission survival matrix represents the flow relationship of the reverse survival propagation algorithm between different basic interactive semantic blocks. Based on the reverse survival propagation algorithm, a reverse data flow equation is established that uses the input active variables of the successor semantic node as the output active variables of the predecessor semantic node. Perform iterative calculations on the inverse data flow equations under a preset maximum iteration threshold constraint; When the change between two adjacent iterations of the iterative calculation is lower than the preset convergence accuracy, the iterative calculation stops, the inverse data flow equation is solved, and the active privacy semantic attributes are output.

[0043] In this embodiment, The topological connection structure of the session logic flow graph is analyzed to identify and determine the direct predecessor and successor semantic nodes corresponding to each basic interaction semantic block within the flow graph. The session logic flow graph is a directed acyclic graph (DAG) consisting of multiple basic interaction semantic blocks as topological vertices, connected by directed edges expressing session context transition logic. By traversing the directed edges in the DAG, the temporal sequence and logical causal flow relationships between the basic interaction semantic blocks can be determined. In specific implementation, for any basic interaction semantic block... By retrieving the basic interactive semantic block Given all directed edges, determine the set of direct predecessor semantic nodes. ; By retrieving basic interactive semantic blocks Given all directed edges, determine the set of direct successor semantic nodes. To enable efficient topological traversal and calculation in subsequent program implementations, the aforementioned adjacency relationships are statically stored in memory using an adjacency list or adjacency matrix, allowing subsequent reverse lookup operations to be implemented via direct array indexing. The formula for the above topological adjacency relationships is defined as follows: In the formula, Indicates the predecessor semantic node. Let E represent the set of directed edges, and let E represent the set of successor semantic nodes. By clearly defining the topological boundaries of predecessors and successors, not only is a logical path established between local semantic blocks and global control flow at the topological level, but also a precise topological neighborhood range is provided for the active state of variables in the subsequent computational data flow. This ensures the integrity of the reverse data transmission analysis from a structural perspective and prevents dangling nodes or missed paths.

[0044] The locally generated attribute set after attribute labeling is used as the local data stream generation source. An information transmission survival matrix based on predecessor and successor semantic nodes is constructed. This matrix establishes an algebraic mapping relationship to characterize the state transition relationship between different basic interactive semantic blocks in the reverse survival propagation algorithm. The information transmission survival matrix is ​​a multidimensional Boolean matrix derived from the adjacency connection relationship of basic interactive semantic blocks, describing the connectivity channels when privacy attributes are reverse-propagated between different semantic blocks. In specific implementation, assuming the total number of basic interactive semantic blocks is y, a matrix of size y is constructed based on the set of predecessor and successor semantic nodes. The information transmission survival matrix M. Elements in the information transmission survival matrix. The definition is as follows: , in the formula, Indicates from the successor semantic node Forward semantic node The connected state that enables reverse information flow transmission. Indicates the successor semantic node The set of direct predecessor semantic nodes. This is achieved by generating a set of attributes from locally generated nodes after labeling. As a local generation source, and combined with the established information transmission survival matrix, it can not only replace the tedious graph depth traversal with matrix multiplication operations, but also intuitively show the evolution and disappearance path of privacy semantic attributes in the session logic flow graph. This provides standardized and normalized data support for the matrix-based rapid iterative solution of subsequent multidimensional data flow equations, and significantly shortens the update time of survival state.

[0045] A reverse data flow equation is established based on the reverse survival propagation algorithm. The input active variables of successor semantic nodes are used as the output active variables of predecessor semantic nodes, realizing the reverse flow of active variables between nodes. During the reverse survival propagation process, the output state of a semantic block is jointly determined by the input states of all directly successor semantic blocks, meaning that active privacy attribute information flows in reverse along the opposite direction of the control flow in the graph structure. Specifically, for any predecessor semantic node... Collect the corresponding set of direct successor semantic nodes. And each successor semantic node in the set of direct successor semantic nodes input active variable set The merge is performed, and the merged set is defined as the predecessor semantic node. Output active variable set The above mapping relationship establishes the following inverse data flow equation: In the formula, Represents the predecessor semantic node The set of active output variables, Indicates the successor semantic node The set of active input variables, Represents the predecessor semantic node The set of direct successor semantic nodes, This represents the set of variables that were cancelled in the predecessor semantic node. This represents the set of locally generated attributes newly generated at the predecessor semantic node. By establishing cascaded inverse data flow equations, the discrete local semantic node states can be chained together into a global data flow chain, ensuring that the propagation path of sensitive privacy information during session transitions can be fully captured, effectively eliminating the risk of analysis gaps during cross-question-answering interactions.

[0046] Under a preset maximum iteration threshold constraint, the established reverse data flow equations are iteratively calculated until a specific convergence criterion is met or the preset maximum number of calculations is reached. Because the session control flow may have multi-branch or loop-flow structures, the equations cannot be solved directly through a single sequential calculation; multiple iterations are necessary to gradually approximate the true values ​​of the active variables at each node. In specific implementation, the maximum iteration threshold is defined as... The iteration counter is d, and d is initialized to 1 before the iteration begins. During each iteration, the active input variable state calculated in the previous round is substituted into the inverse data flow equation to calculate the latest variable state for the current round. The calculation formula for the d-th iteration is defined as follows: In the formula, This represents the predecessor semantic node obtained from the d-th iteration. The set of active input variables, This represents the successor semantic node obtained from the (d-1)th iteration. The set of active input variables. During the execution of the loop iterative calculation, the current iteration number d is monitored in real time at the loop entry point to see if it exceeds the maximum iteration threshold. If the time limit is exceeded, the iterative calculation will be forcibly stopped, thereby avoiding computational dead loops caused by abnormal loop topologies in the data flow graph. This ensures that the overall time consumption for data desensitization and filtering is within a safe and controllable time range, preventing the exhaustion of runtime resources.

[0047] When the change between two adjacent iterations in the iterative calculation is lower than the preset convergence precision, the iterative calculation stops, the inverse data flow equation is solved, and the active privacy semantic attributes are output. The convergence precision criterion measures the stability of the active variable set during the data flow equation solution process. When the active variable sets calculated in two adjacent iterations are completely identical or the change is extremely small, it indicates that the computational state has reached equilibrium. In specific implementation, the preset convergence precision is defined as the critical change threshold. After completing the d-th iteration calculation, the change in adjacent iterations is obtained by calculating the sum of the number of elements in the symmetric difference set of the active variable sets of all semantic blocks between the d-th and d-1-th iterations. The formula for calculating the change in adjacent iterations is defined as follows: In the formula, Let represent the j-th basic interaction semantic block, and y represent the total number of basic interaction semantic blocks. When the decision condition is met... When the current loop iteration process is completed, the latest active input variable set corresponding to each semantic block at the convergence state is obtained. The active privacy semantic attribute obtained from the final solution is output. By introducing a convergence accuracy criterion, the convergence timing of the equation system solution can be captured automatically and sensitively. This ensures the completeness of the calculation results while avoiding meaningless and redundant calculations, greatly ensuring the policy calculation response speed of the dynamic desensitization process.

[0048] In one implementation, calculating the shortest path distance from each active privacy semantic attribute to the identity identifier in the association path topology using the association path vector metric algorithm includes the following steps: Retrieve the probability of direct association between active privacy semantic attributes and identity identifiers from a pre-built prior knowledge relationship graph; The direct association probability is converted into the edge association weight between topological nodes in the associated path topology; Initialize a distance update vector table pointing to the topology node corresponding to the identity identifier for each active privacy semantic attribute; The shortest associated path distance in the updated vector table is updated by using the associated path vector metric algorithm and updating the distance through information exchange between neighboring topological nodes. Based on the distance-updated vector table, and using the path convergence function, the shortest associated path distance from the topology node corresponding to each active privacy semantic attribute to the topology node corresponding to the identity identifier is calculated in the current associated path topology state.

[0049] In this embodiment, the direct association probability between active privacy semantic attributes and identity identifiers is retrieved by accessing a priori knowledge relationship graph pre-stored in local storage. The priori knowledge relationship graph is a semantic network with multi-dimensional entities as nodes and relationships between entities as edges, storing statistical probabilities between different privacy attributes and identity leakage risks. Specifically, for any detected active privacy semantic attribute... and the identified identity tokens Perform a graph query within the prior knowledge relationship graph to check for the existence of connections. and The direct association edges are identified. If direct association edges exist, the co-occurrence probability or inferred association value stored on the direct association edges is read as the direct association probability. The direct association probability obtained from the retrieval is defined as... , represented as This value ranges from 0 to 1. If the prior knowledge graph does not contain a direct edge connecting active privacy semantic attributes and identity identifiers, the default direct association probability value is 0. By introducing retrieval computation based on the prior knowledge base, discrete text symbols can be transformed into quantifiable security indicators with probabilistic measures. This provides a reasonable data source for subsequently quantifying the leakage path risk of privacy inference in the topological network, enhancing the reliability of privacy quantification assessment at the multi-source knowledge association level.

[0050] The retrieved direct association probabilities are converted into edge association weights representing the topological nodes of each entity in the association path topology. Edge association weights are used to measure the distance between two entities in the topological network; the lower the edge association weight, the stronger the association between the two topological nodes, meaning it's easier to infer one topological node from another. In practice, a transformation function is established to map direct association probabilities to edge association weights, smoothing the nonlinear probabilities (originally in the zero-to-one interval) into a non-negative real number space through a mathematical mapping. This is defined from the active privacy semantic attributes. To identity verification The edge association weights between corresponding topological nodes are The conversion formula is as follows: In the formula, This indicates the probability of finding a direct association. This represents a preset constant fine-tuning factor used to prevent logarithmic overflow anomalies when the direct association probability is 0. Through an inverse logarithmic proportional transformation, high association probabilities are converted into low edge association weights, while low association probabilities are converted into extremely high or even infinite weights. Because of this mathematical logarithmic mapping transformation, probability multiplication can be converted into path weight addition, logically simplifying the subsequent shortest path accumulation algorithm and significantly improving the efficiency of strategy processing for associated path network topology metrics.

[0051] For each active privacy semantic attribute corresponding to a topology node, initialize a distance update vector table pointing to the topology node corresponding to the identity identifier. The distance update vector table records and dynamically maintains the currently known shortest estimated distance required to reach the destination identity identifier node from the current active privacy semantic attribute node. Specifically, for each node representing an active privacy semantic attribute in the associated path topology... For each node, a dedicated array structure is allocated as a distance update vector table. Define the first... The distance update vector table corresponding to each active privacy semantic attribute node is as follows: The table contains information on reaching each identity identifier. The initial value of the distance. The initialization rules are as follows: if the privacy semantic attribute is active... With identity markers If there are directly connected edges, then the initial distance in the distance update vector table is set to the corresponding edge correlation weight. Otherwise, set the initial distance to represent an unreachable infinity. Through standardized initialization operations, each local privacy node in the graph structure has an independent starting point for path exploration and state measurement, laying the foundation for subsequent distributed neighbor node information interaction and iterative convergence of the global shortest path, and ensuring the efficient advancement of the algorithm.

[0052] The association path vector metric algorithm updates the shortest association path distance recorded in the distance update vector table by periodically exchanging information between neighboring topological nodes. Neighboring topological nodes are adjacent entity nodes with direct edges in the association path topology. In practice, each topological node sends its distance update vector table to its neighboring topological nodes. (When representing active privacy semantic attributes...) The node receives neighboring topology nodes Distance update vector table sent Then, it will combine the edge correlation weights from the node representing the active privacy semantic attribute to the neighboring topological nodes. The shortest path distance in the table is updated and calculated. The defined update formula is as follows: In the formula, This represents the shortest path distance from the current node to the identity identifier. This represents the path distance from neighboring topology nodes to the identity identifier. This represents the edge correlation weight between adjacent nodes. Through the neighbor information interaction mechanism, local privacy features are continuously propagated forward in the topology, enabling each node to acquire and perceive dynamic changes in the topology structure further away. This gradually corrects the distance assessment bias caused by insufficient information, ultimately yielding an assessment value that closely approximates the true degree of correlation.

[0053] Based on the updated distance update vector table and combined with the path convergence function, the shortest associated path distance from the topology node corresponding to each active privacy semantic attribute to the topology node corresponding to the identity identifier is calculated under the specific topology state of the current associated path topology. The path convergence function is used to determine whether the vector update process has stabilized and to extract the final stable distance value. In specific implementation, the path convergence function is defined as follows: When the values ​​in the distance update vector table no longer fluctuate after multiple rounds of continuous information interaction, the convergence function is triggered. The shortest path distance is defined as... The expression for calculating the path convergence function is as follows: In the formula, Indicates the first After each information interaction, the distance stored in the updated vector table is used. By solving the path convergence function, the final shortest association path distance can be determined from the dynamically interacting vector table. The above calculation steps ensure that the shortest association path distance obtained in the current multi-round interaction context reflects the most realistic privacy inference path, eliminates the computational uncertainty caused by dynamic changes in the topology, and provides a high-confidence data benchmark for subsequent association blocking decisions.

[0054] In one implementation, the shortest association path distance from the topology node corresponding to each active privacy semantic attribute to the topology node corresponding to the identity identifier in the current association path topology state is calculated based on the distance update vector table and using the path convergence function, including the following steps: Extract all candidate path topology routes between the topology node corresponding to each active privacy semantic attribute and the topology node corresponding to the identity identifier; Based on the node overlap between candidate path topologies, calculate the path independence coefficient for each candidate path topology. Obtain the timestamps of the privacy feature attributes corresponding to each topology node in the candidate path topology, and calculate the temporal correlation decay weight between nodes based on the timestamps; The inference path multiplier is calculated by combining the time-related decay weight and the path independence coefficient. The initial path distance calculated based on the distance update vector table is multiplicatively corrected using the inference path multiplier, and the corrected shortest associated path distance is output.

[0055] In this implementation, by performing graph traversal and path search on the associated path topology, all acyclic paths between the topology nodes corresponding to each active privacy semantic attribute and the topology nodes corresponding to the identity identifier are extracted as candidate path topologies. An acyclic path refers to a path sequence from the starting node to the target node that does not contain duplicate nodes. In specific implementations, for a particular active privacy semantic attribute... To specific identity identifier A depth-first path search algorithm is used to perform a full path search, traversing all possible directed or undirected connected edges, filtering out invalid paths containing closed cycles, thereby establishing a candidate set of acyclic paths. Assume that for the ... The active privacy semantic attribute node and the first There are 10 identity nodes, and the search yields a total of 10 candidate acyclic paths. A set of candidate path topological paths is defined as follows: Each candidate pathway topology path Each path consists of a series of intermediate topological nodes and connecting edges arranged in sequence. By comprehensively extracting these potential candidate paths, the limitations of single-path evaluation are overcome, extending the assessment of privacy exposure risks from single-chain analysis to the space of multi-path concurrent reasoning. This provides a complete and detailed topological path basis for in-depth analysis of the risk boundary of leaking true identity through joint reasoning of multiple features.

[0056] Based on the degree of node overlap between different candidate path topologies, the path independence coefficient for each candidate path is calculated. If two paths contain a large number of identical intermediate topological nodes, it indicates that the reasoning processes of the two paths rely on similar information sources and lack high independence, thus requiring a reduction in their corresponding reasoning confidence. In specific implementation, for a given path in the candidate path topology set... Extraction path The set of nodes contained . Path Intersecting the union of the set with the union of all other paths in the set yields the set of overlapping nodes. Define the first... The path independence coefficient of the candidate path topology is The corresponding calculation formula is as follows: In the formula, Representing a path The set of topological nodes on the topology, Indicates excluding path The set of topological nodes on the remaining candidate path topological paths, excluding This indicates the number of elements within the computation set. By calculating the path independence coefficient, the information redundancy between different inference links can be quantified, suppressing the artificially inflated risk of privacy leakage caused by high information overlap, and making the final assessment of multi-path joint inference risk results more consistent with the actual logical relationship.

[0057] The system obtains the timestamps of the privacy feature attributes corresponding to each topological node in the candidate path topology and calculates the temporal correlation decay weight between nodes based on the span of the timestamps in the time dimension. The time intervals generated by different session interactions affect the degree of correlation between privacy attributes; the longer the time span, the lower the confidence of joint inference between different privacy feature attributes. In specific implementation, for the candidate path topology... The timestamps generated by each topology node along the retrieval path in the session logic flow graph are calculated. The difference between the timestamps of the first and last nodes of the path is calculated to obtain the total time span corresponding to the path, which is defined as the time difference. Introduce a preset time decay coefficient. Calculate the first Time-related decay weight of each path The corresponding calculation formula is as follows: . In the formula, This represents the timestamp difference between the first and last topological nodes of the path. This represents the preset time decay constant. This represents a natural constant. By using time-related decay weights to correct the decay of inference links on a time scale, it reflects the natural loss of the correlation strength of scattered privacy information in different sessions during inference over time, thus eliminating false privacy inference paths that span extremely long session periods and have no actual correlation.

[0058] By combining the temporal correlation attenuation weight and the path independence coefficient, a comprehensive calculation is performed to obtain the inference path multiplier, which characterizes the reliability of inference for a specific path. As a dimensionless correction factor, the inference path multiplier can simultaneously reflect path independence in spatial topology and temporal correlation in the temporal dimension, thereby achieving dynamic multidimensional adjustment of static path distances. In specific implementation, for the... The candidate path topology is analyzed, and the calculated path independence coefficient is obtained. Corresponding time-related decay weight Multiplying the two factors yields the inference path multiplier, representing the topological path of the corresponding candidate pathway. The calculation formula is as follows: . In the formula, Indicates the first The path independence coefficient of the candidate path topology. Indicates the first The temporal correlation decay weight of each candidate path topology is used. By combining multiplication, it is ensured that if either the path independence or temporal correlation dimension performs poorly, the inference path multiplier will be significantly reduced, thereby controlling the credibility of the inference path and improving the decision rationality of the dynamic desensitization strategy in the face of conversational scenarios.

[0059] Using the calculated inference path multiplier, the initial path distance calculated based on the distance update vector table is multiplicatively corrected, ultimately outputting the corrected shortest associated path distance. This multiplicative correction operation aims to combine the physical characteristics of both time and space dimensions to correct the purely static distance initially calculated solely based on the inverse of probability, making the measured privacy distance more consistent with real-world logic. In specific implementation, for the... The active privacy semantic attribute and the first The first identification node reads the initial path distance for each path from the distance update vector table. Define the first... The initial path distance corresponding to each path is By using the initial path distance Divide by the inference path multiplier of the corresponding path The corrected path distance is calculated, and the minimum value among all candidate paths is selected as the final output corrected shortest associated path distance. The corresponding correction calculation formula is as follows: In the formula, This represents the initial path distance of the corresponding candidate path topology calculated based on the distance update vector table. This represents the inference path multiplier for the corresponding path. By using a smaller inference path multiplier as the denominator, the distances corresponding to paths with weak correlation are multiplicatively amplified, thereby filtering out false alarms caused by accidental co-occurrence and improving the accuracy of the desensitization scheme.

[0060] In one implementation, updating the shortest associated path distance in the distance update vector table using an associated path vector metric algorithm and through information exchange between neighboring topological nodes includes the following steps: Simulate and construct topology status messages between various topology nodes in the associated path topology; Collect the latest edge association weights between topological nodes, and construct a single-hop distance matrix that reflects the local proximity relationships in the topology of associated paths based on the latest edge association weights; Control each active privacy semantic attribute corresponding to the topology node to broadcast the current distance update vector table to neighboring topology nodes through topology status messages; The loop elimination operation blocks the cyclic announcement of topology status messages in the closed path of the associated path topology. When a sudden change in the shortest associated path distance is detected in the distance update vector table, the risk reversal mechanism is triggered and the infinite distance value representing the path blockage is written as the updated shortest associated path distance into the topology status message; Based on the received topology status message and combined with the single-hop distance matrix, the distance update vector table of each topology node is updated using the path distance algorithm to complete the update calculation of the shortest associated path distance.

[0061] In this embodiment, within the constructed associated path topology, a structured data entity is created for each connected topology node as a topology state message to transmit path change information within the topology network. The topology state message includes the identity of the sending source node, its version sequence number, and the path cost value recorded in the distance update vector table. Specifically, for topology nodes representing active privacy semantic attributes... Define the generated topology state message as a triple. ,in, Represents topology nodes Unique identifier, This represents an incrementing sequence number used to prevent interference from replay and outdated messages. Indicates the current topology node An internally stored and maintained distance update vector table pointing to each identity identifier is used. Topology state messages are packaged in memory with a specific data format and written to a buffer after the sequence number is incremented. By periodically assembling this type of data payload in memory, topology state messages can transform the link connectivity state of local nodes into a structured information stream that can be directly parsed and read by neighboring nodes. This ensures that local state updates can be smoothly propagated to neighboring nodes in a stepwise manner, improving the response agility of local topology change detection and avoiding uncontrollable computational delays during data flow blocking phases.

[0062] The system collects the latest edge association weights between adjacent nodes in the associated path topology in real time. Based on this, a two-dimensional array of double-precision floating-point numbers is allocated in memory to construct a single-hop distance matrix reflecting the local direct proximity connectivity state in the associated path topology. The latest edge association weights change dynamically with session updates and fluctuations in prior knowledge base retrieval results. In specific implementation, it is assumed that there are... Each active privacy semantic attribute node defines a single-hop distance matrix of size equal to... Adjacency cost matrix The first in the matrix Line 1 Column elements The corresponding value calculation rules are as follows: In the formula, This represents the latest edge correlation weights between adjacent nodes. By constructing a single-hop distance matrix, the originally chaotic graph topology edge connections and weight costs are transformed into a standardized and regular two-dimensional numerical representation. This facilitates the parallel and fast querying of single-hop path costs using matrix-vector operations, significantly reducing addressing latency when retrieving local proximity connectivity costs and laying the foundation for fast path state convergence across the entire network. Each topology node representing active privacy semantic attributes, upon detecting a change in its own distance or a timer expiration, calls a socket interface to broadcast its latest distance update vector table to neighboring topology nodes via topology status messages. This broadcast mechanism operates securely in directed or undirected topology networks, ensuring that the latest local metric for path distances quickly spreads to its neighbors.

[0063] In specific implementation, for active privacy semantic attribute nodes First, obtain the set of adjacent topological nodes through the adjacency list. Then, the distance recorded by the current node is used to update the vector table. Encapsulated into topology status messages And through parallel memory message distribution channels or network sending functions, it sends messages to the set of adjacent topology nodes. Each adjacent topological node in The topology status message is delivered asynchronously, with the following expression: .in, Includes topology nodes The latest distance update vector table is used. By deploying this broadcast mechanism, each local privacy node can promptly transmit its own risk status changes to surrounding related nodes, promoting the diffusion of status changes to the entire network at a remote end. This ensures that the path assessment of the entire network can remain consistent, significantly improving the agility of tracking dynamic security risk exposures.

[0064] By incorporating targeted loop elimination operations into the state announcement and vector calculation processes, the generation of infinite loop announcements in the closed-loop paths of the associated path topology is prevented, thus avoiding routing loops and infinite counting problems. Specifically, a complete node path visit sequence is assigned to each path distance record to trace the order of nodes traversed along the path. For topology nodes... In preparation for sending to adjacent nodes When sending a topology status message, first retrieve and check the arriving identity. Path access sequence If adjacent nodes are found If the path is already included in the access sequence, it indicates that a potential loop has been formed. Therefore, the blocking mechanism is immediately activated to prevent the path from reaching the identity identifier. The routing information is announced to neighboring nodes. The logical judgment rules for the above loop checking and elimination are defined as follows: in, Indicates from topology node Arrival Identification The set of identifiers for all topological nodes traversed along the path. By implementing the above filtering restrictions, the closure of potential loop paths is directly cut off at the message sending end, effectively preventing endless loop calculations and the generation of false short-distance values, and ensuring the absolute stability of path distance convergence in multi-round interaction scenarios.

[0065] During operation, the distance update vector table is dynamically monitored. When an abnormal change in the shortest path distance in the distance update vector table is detected, a risk reversal mechanism is triggered. The infinite distance value, representing path blocking, is proactively written into the topology state message as the updated shortest path distance to prevent erroneous routes from propagating further in the network. Shortest path distance mutations typically occur when topology nodes disconnect due to session context transfer, expiration of lifecycle, or execution of security blocking operations. In specific implementation, the shortest path distance values ​​from two adjacent time periods are compared in real time. A preset mutation judgment fluctuation difference is set to... When the following mutation criteria are met: When the path is deemed severely degraded, the current distance value is immediately forcibly set to infinity, i.e., set... The updated shortest path status is then directly written into the topology status message sent to all adjacent nodes. This forced setting operation, which reverses risk, ensures that the link interruption notification is disseminated throughout the network, blocking continuous inference based on the damaged link and thus guaranteeing the leak prevention security of the entire desensitization scheme in the first instance.

[0066] Based on the received topology status message and combined with the single-hop distance matrix already established in memory, the path distance algorithm is used to globally update the distance update vector table of each topology node in the associated path topology, thus completing the update calculation of the shortest associated path distance. The core of the path distance algorithm is to continuously adjust the sum of the shortest weights to the destination node in a distributed manner. In specific implementation, when the topology node... Receive neighboring topology nodes Topology status messages sent Next, the vector table of neighboring topology nodes is extracted from the received message by unpacking. Subsequently, the edge correlation weight between these two adjacent nodes is directly read from the single-hop distance matrix. The path distance update equation is called as follows: Based on the aforementioned calculation equation, the path distance values ​​corresponding to each identity identifier of the topology node representing active privacy semantic attributes in the distance update vector table are scanned, compared, and overwritten one by one. This real-time recalculation mechanism, which combines received neighboring topology messages with the locally stored single-hop distance matrix depth, ensures that any change in link topology overhead can quickly reconstruct the latest shortest distance state. This essentially eliminates privacy leakage assessment errors caused by lag in graph structure information, achieving highly sensitive dynamic filtering defense.

[0067] In one implementation, the process of generating desensitized candidate text by matching the generalization order based on the semantic dilution amount and a cascaded concept tree constructed based on a preset category hierarchy relationship of privacy feature attributes, and by performing semantic smoothing dilution compensation on associated blocking nodes based on the generalization order, includes the following steps: Load the preset category hierarchy of privacy feature attributes and construct a cascaded concept tree based on the category hierarchy; Establish the correspondence between semantic dilution and the depth of the cascaded concept tree hierarchy, and match to obtain the target generalization order; Based on the target generalization order, a progressively upward generalization mapping algorithm is performed on the associated blocking nodes in the cascaded concept tree to obtain the target generalized entity; Extract the syntactic skeleton of the interactive text containing the associated blocking nodes, and replace the associated blocking nodes in the interactive text with placeholder indicator tags; The target generalized entity is filled into the syntactic skeleton corresponding to the placeholder indicator label to achieve semantic smoothing dilution compensation and generate desensitized candidate text.

[0068] In this implementation, a pre-defined hierarchy of privacy feature attributes is loaded from a local security repository. This hierarchy is stored in a structured data file, defining hierarchical concepts of different privacy attribute entities from concrete to abstract. During loading, lexical analysis is performed on the data file to extract entity nodes representing entity categories and to identify the nested hierarchical relationships between nodes, thereby determining the parent pointers between nodes and ultimately completing the bottom-up memory construction of the cascaded concept tree. The cascaded concept tree is a multi-branch tree with concrete privacy entities as leaf nodes, intermediate abstract concepts as internal nodes, and top-level generalized categories as root nodes. In specific implementations, the cascaded concept tree is defined as... In this process Indicates inclusion A collection of nodes at different conceptual levels. This represents the set of directed edges connecting parent and child nodes. If a node... It is a node If the direct parent node is the one above the child node, then a mapping edge is established. The tree's level depth function is defined as follows: The leaf nodes have a depth of 0, and the depth increases by 1 with each upward extension until the root node. By reconstructing discrete privacy entities into a cascaded concept tree according to their hierarchical relationships, a clear hierarchical concept progression between attributes can be established. This provides a standard, hierarchical topological carrier for adaptively adjusting the semantic generalization range based on the degree of security risk deviation.

[0069] A correspondence is established between semantic dilution, which characterizes the difficulty of blocking, and the depth of the cascaded concept tree hierarchy. This allows for quantitative numerical calculations to determine the target generalization order. The semantic dilution depends on the deviation level of the calculated shortest association path distance from a safety threshold; the closer the distance, the greater the exposure risk, and the stronger the required concept dilution. In practice, a preset safety threshold is set as follows: The calculated target path distance corresponding to the currently associated blocking node is Define the target generalization order as... The following formula is established to convert the deviation value into the corresponding generalization order: In the formula, This indicates the preset safety threshold. Indicates the distance of the target associated path. This represents the preset dilution ratio conversion constant. This represents the upper limit of the maximum level depth of the cascading concept tree. This represents semantic dilution. Through this correspondence formula, the abstract security distance deviation can be linearly mapped to a discrete tree-level span. After calculating the target generalization order, this value is directly used as a limitation on the number of iterations for subsequent generalization mapping algorithms. Logically, this translates the physical security distance metric directly into abstract operational instructions at the conceptual level. This mapping mechanism can adaptively allocate generalization depth according to different levels of privacy exposure severity, ensuring large-span generalization when security protection is insufficient, while only slight fuzziness is applied when the risk is low, thus balancing data security and usability.

[0070] Based on the target generalization order obtained from the matching, in the cascaded concept tree, starting from the specific leaf node corresponding to the association blocking node, a progressively upward generalization mapping algorithm is executed to ultimately locate and extract the corresponding superior target generalized entity. The execution of the generalization mapping algorithm is essentially a multi-step upward addressing process guided by the parent node pointer in the cascaded concept tree; each step forward signifies an increase in the level of abstraction of the entity concept. In specific implementation, the initial entity node corresponding to the association blocking node in the cascaded concept tree is defined as... The generalized state transition mapping equation is defined as follows: . In the formula, Indicates the initial entity node to be evaluated. This indicates the operation of finding the direct parent node of the current node. This represents the current iteration mapping step, with values ​​increasing from 1 to the target generalization order. By iteratively applying the above mapping equation... Next, the final node obtained This is the desired target generalized entity. Since this generalization mapping algorithm is strictly limited by the classification logic of the cascaded concept tree, the output superordinate entity must have a natural category similarity with the original sensitive words, thus enabling a fitting concept replacement. While blocking the association reasoning of sensitive features, it preserves the industry attributes or category semantics contained in the original entity to the greatest extent.

[0071] The syntactic skeleton of the interactive text containing association blocking nodes is extracted using a parser, and then directly replaced in memory with specific placeholder labels to represent the text intervals containing the corresponding association blocking nodes. Syntactic skeleton extraction uses a dependency parsing model to identify the core predicates, subject-object frames, and sentence modification relationships in the text, thereby stripping away specific feature words while preserving the sentence's skeletal structure. In specific implementation, let the original input interactive text be represented as... The identified associated blocking node text is Define the placeholder indicator label as follows: By using a string replacement function, a syntactic skeleton text with sensitive words removed is constructed. The corresponding expression is as follows: During the replacement process, the analyzer not only erases specific sensitive and private words, but also uses syntax tree analysis to identify placeholder tags. The part-of-speech and component position within the grammatical structure. This positional locking ensures that subsequent filling of generalized entities does not disrupt the overall grammatical coherence of the sentence. This extraction and replacement mechanism based on syntactic structure analysis physically isolates semantic structure from sensitive content, preserving the narrative structure and tone of the original sentence while protecting data privacy, and providing specific positional and syntactic constraint information for subsequent semantic compensation.

[0072] The target generalized entities obtained in the previous steps are directly filled into the blank positions corresponding to the placeholder indicator tags in the grammatical skeleton. This entity replacement achieves smooth dilution compensation for privacy semantics in the sentence, and finally dynamically assembles in memory to generate desensitized candidate text that can be safely released externally. Smooth dilution compensation, under syntactic dependency constraints, replaces low-order concrete words with high-order abstract words, eliminating inference links while achieving natural coherence of the contextual meaning. In specific implementation, the grammatical skeleton text is extracted. Placeholder label in Obtain the target generalized entity output by the cascaded concept tree generalization mapping algorithm in the previous steps. The desensitization generation equation is constructed as follows: . In the formula, This represents the final generated de-identified candidate text. This represents the extracted grammatical skeleton text. This indicates a placeholder label. This represents the target generalized entity. The desensitization generation equation is used to complete the entity backfilling and text concatenation of the placeholders. This smooth replacement compensation under syntactic skeleton constraints ensures that the newly generated desensitized text not only achieves risk dilution at the security level but also maintains a high degree of naturalness at the reading level, avoiding the decline in interactive experience caused by direct and rigid desensitization.

[0073] In one implementation, the process of performing a progressively upward generalization mapping algorithm on the associated blocking nodes in the cascaded concept tree according to the target generalization order to obtain the target generalized entity includes the following steps: Initialize the associated blocking node in the cascaded concept tree based on the target generalization order and calculate the initial semantic heat. Calculate the semantic similarity loss between the node to be evaluated and its parent node at the next higher level; The generalized cross-layer probability is calculated by combining the initial semantic heat and semantic similarity loss. If the randomly generated value is lower than the generalized cross-layer probability, the parent node of the previous level is taken as the new node to be evaluated; otherwise, the current node to be evaluated is determined as a temporary entity. The initial semantic heat is iteratively reduced using a preset heat decay coefficient until the initial semantic heat is lower than a preset termination threshold, at which point the temporary entity is used as the target generalized entity.

[0074] In this embodiment, the target generalization order obtained from the previous steps is loaded. This determines the optimization depth and logical starting point of the iterative computation. Physical pointers are used to locate the specific entities marked as association blocking nodes in the cascaded concept tree. The specific leaf nodes corresponding to the association blocking nodes are taken as the initial nodes to be evaluated. Based on this, a pre-defined mathematical mapping model is used to calculate the initial semantic heat used to constrain the randomness of subsequent generalization across layers. The node to be evaluated serves as the current focus position for the algorithm's upward traversal and generalization in the cascaded concept tree. The initial state directly determines the convergence speed and stability of subsequent multi-round path exploration. In specific implementation, it is assumed that the leaf node corresponding to the association blocking node in the cascaded concept tree is... Define the current node to be evaluated as And the nodes to be evaluated The memory pointer is directly assigned and initialized to point to a leaf node. The corresponding memory address. After pointer initialization, the initial semantic heat calculation formula is established as follows: . In the formula, This represents the initial semantic heat obtained from the calculation. Indicates the order of generalization of the target. This represents a pre-configured constant thermal gain factor. Initial semantic heat, as a core control indicator measuring the generalization and diffusion energy of the current node in the cascaded concept tree, is directly proportional to the target generalization order. This indicates that the further the deviation from the safety threshold, the stronger the optimization and evolution energy granted in the initial stage, better supporting subsequent adaptive probabilistic jumps between large-span tree levels. This breaks the single-path defect of deterministic generalization at the underlying level, ensuring that the desensitized text maintains high fidelity while possessing a high degree of resistance to inference intrusion.

[0075] After identifying the node to be evaluated, a high-dimensional vector representation model is invoked to calculate the semantic similarity loss between the currently active node and its parent node at the next higher level. Semantic similarity loss quantifies the amount of original semantic information lost when a concept jumps one level up in a cascading concept tree. The hierarchical ascent of the concept tree is essentially a process of continuous abstraction, accompanied by information loss. In specific implementation, the node to be evaluated is first... Perform a memory search on the parent node, and define the next-level parent node as... The node to be evaluated is passed to the word vector retrieval interface. and the parent node at the next higher level For each text field, its corresponding feature vector is obtained. Then, the cosine similarity algorithm is used to calculate the proximity between two vectors, yielding a basic similarity. Based on this, the semantic similarity loss is calculated by subtracting the basic similarity from the result. The feature algebraic equation for calculating the semantic similarity loss is established as follows: In the formula, This represents the calculated semantic similarity loss. This indicates the current node to be evaluated. This represents the parent node one level above the node to be evaluated. This represents the similarity function used to calculate the cosine similarity between the node to be evaluated and its parent node at the next higher level. The calculated semantic similarity loss accurately reflects the degree of semantic information loss caused by the generalization operation locally. The lower the semantic similarity loss, the smoother the transition of conceptual abstraction. This provides quantitative feedback on the value of content retention when performing cross-level generalization judgments in the subsequent process, ensuring that semantic dilution does not cause irreversible logical breaks due to blind leaps.

[0076] Combining the calculated initial semantic heat and semantic similarity loss, a generalized cross-level probability for controlling the upward climbing decision of concepts is calculated through a probabilistic exponential mapping. The generalized cross-level probability represents the likelihood that the decision algorithm allows semantics to transition to higher-level parent concepts, mimicking the probabilistic mechanism of particle state transitions in physical annealing. In specific implementation, the calculated generalized cross-level probability is defined as... Combined with initial semantic popularity With semantic similarity loss The following algebraic formula is established for calculating the generalization cross-layer probability: . In the formula, Represents the generalized cross-layer probability. This represents the semantic similarity loss. Indicates the initial semantic popularity. This represents the natural constant. It is used to calculate the generalization cross-layer probability. Next, generate a random number that is uniformly distributed in the interval between zero and one, and define the random number as... Perform numerical comparisons; when determining the random number... Lower than the generalized cross-layer probability When a decision is made to allow a concept to jump across levels, the parent node at the next higher level is selected. As the new node to be evaluated, perform pointer overwriting. If random number Greater than or equal to the generalization cross-layer probability If this occurs, it is determined that cross-layer blocking is blocked, and the current node to be evaluated is... Lock and identify as the current temporary entity This probabilistic jump control cleverly balances the strength of security defense with the value of preserving textual semantics, introducing reasonable random perturbations into the decision-making process to prevent static, regularized desensitization from being reverse-engineered.

[0077] Using a preset heat decay coefficient, the semantic heat is reduced after each iteration until it falls below a preset termination threshold, at which point the search process stops and the temporary entity obtained is established as the final target generalized entity. As the number of iterations increases, the semantic heat continuously decreases, leading to a significant reduction in the probability of cross-layer generalization. The algorithm eventually converges at the node that satisfies both safety and semantic constraints. In specific implementation, the preset heat decay coefficient is defined as... Define the current semantic popularity as And initialize the current semantic heat during the first round of calculation. After completing a single round of cross-layer judgment, the current semantic heat is updated using the following decay equation: . In the formula, Indicates the current semantic popularity. This represents a preset popularity decay coefficient between zero and one. After each decay update, the current semantic popularity is checked. Is it below the preset termination threshold? If the current semantic popularity Still higher than or equal to the termination threshold If so, return to continue iteratively performing similarity calculations and cross-layer probability determinations until the current semantic popularity is reached. Below the termination threshold At that time, forcibly exit the iteration loop and change the current temporary entity. Identified as the target generalized entity Thus, the adaptive level concept generalization calculation is smoothly completed.

[0078] The present invention also discloses a sandbox-based dynamic desensitization system for agent-oriented data interaction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the sandbox-based dynamic desensitization method for agent-oriented data interaction as described in any of the above claims.

[0079] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.

[0080] The memory can be an internal storage unit of a computer device, such as a hard disk or RAM, or an external storage device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) provided on the computer device. Furthermore, the memory can be a combination of internal storage units and external storage devices of a computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.

[0081] The present invention also discloses a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to be configured to perform the sandbox-based dynamic desensitization method for agent-oriented data interaction described in any of the above embodiments.

[0082] The computer program can be stored in a machine-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The machine-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the machine-readable medium includes, but is not limited to, the above-mentioned components.

[0083] The sandbox-based dynamic desensitization method for intelligent agent data interaction described in the above embodiments is stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the above method.

[0084] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0085] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.

Claims

1. A sandbox-based dynamic desensitization method for agent-based data interaction, characterized in that, Includes the following steps: Intercept the agent's external interaction messages in the isolation sandbox, extract the session identifier, timestamp and interaction text in the interaction messages, and generate a session logic flow graph by combining the session identifier and timestamp and based on the interaction text; Extract privacy feature attributes and identity identifiers from interactive text, reverse traverse the session logic flow graph based on privacy feature attributes, and use the reverse data flow equation to calculate the active privacy semantic attributes of the current interactive node in the session logic flow graph. Construct an association path topology with active privacy semantic attributes and identity identifiers as topology nodes, and use the association path vector metric algorithm to calculate the shortest association path distance from each active privacy semantic attribute to the identity identifier in the association path topology; When there is a target associated path distance in the shortest associated path distance that is lower than the preset safety threshold, the topology node corresponding to the target associated path distance in the associated path topology is marked as an associated blocking node; The required semantic dilution is calculated based on the deviation between the target association path distance and the security threshold. The generalization order is matched according to the semantic dilution and the cascaded concept tree constructed based on the preset category hierarchy relationship based on privacy feature attributes. The semantic smoothing dilution compensation is performed on the association blocking nodes based on the generalization order to generate desensitized candidate text. Calculate the semantic fidelity of the interactive text and the desensitized candidate text in the multidimensional semantic space, and release the desensitized candidate text when the semantic fidelity meets the preset usability threshold.

2. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 1, characterized in that, The steps of extracting privacy feature attributes and identity identifiers from interactive text, reverse traversing the session logic flow graph based on privacy feature attributes, and calculating the active privacy semantic attributes of the current interactive node in the session logic flow graph using reverse data flow equations include the following: The text entity recognition module in the isolation sandbox is invoked to perform entity tagging on the interactive text in order to extract privacy feature attributes and identity identifiers; The reverse control flow analysis algorithm is used to reverse traverse the session logic flow graph to determine the attribute lifecycle of privacy feature attributes in the current interaction node of the session logic flow graph. Configure the initial active state for privacy feature attributes within the attribute lifecycle and generate an initial active variable set; The initial set of active variables is used as the input parameter of the inverse data flow equation in the inverse control flow analysis algorithm, and the active privacy semantic attributes existing in the current interaction node are obtained by solving the inverse data flow equation.

3. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 2, characterized in that, The process of obtaining the active privacy semantic attributes existing in the current interaction node by solving the inverse data flow equation includes the following steps: The conversation logic flow graph is divided into basic interactive semantic blocks consisting of single question-and-answer interactions; Extract the privacy feature entities of the new input in each basic interactive semantic block as a local generated attribute set, and mark the privacy feature attributes of the privacy feature entities in the local generated attribute set as active states; Based on the locally generated attribute set after attribute tagging, the input active variables and output active variables of each basic interactive semantic block are calculated using the reverse survival propagation algorithm. The reverse data flow equation is established by combining the input active variables and output active variables. During the reverse traversal of the session logic flow graph, privacy feature attributes that have exceeded their lifespan are deregistered according to the attribute lifespan endpoint of the attribute lifespan, and the active state corresponding to each basic interaction semantic block is updated. Iteratively solve the inverse data flow equation until the equation converges, and output the active privacy semantic attributes after convergence.

4. The sandbox-based dynamic desensitization method for agent-oriented data interaction according to claim 3, characterized in that, The process of generating a local attribute set based on attribute tagging, calculating the input and output active variables of each basic interactive semantic block using the reverse survival propagation algorithm, and establishing the reverse data flow equation by combining the input and output active variables includes the following steps: Identify the predecessor and successor semantic nodes of each basic interaction semantic block in the session logic flow graph; Using the locally generated attribute set after attribute labeling as the local generation source, an information transmission survival matrix based on predecessor semantic nodes and successor semantic nodes is constructed. The information transmission survival matrix represents the flow relationship of the reverse survival propagation algorithm between different basic interactive semantic blocks. Based on the reverse survival propagation algorithm, a reverse data flow equation is established that uses the input active variables of the successor semantic node as the output active variables of the predecessor semantic node. Perform iterative calculations on the inverse data flow equations under a preset maximum iteration threshold constraint; When the change between two adjacent iterations of the iterative calculation is lower than the preset convergence accuracy, the iterative calculation stops, the inverse data flow equation is solved, and the active privacy semantic attributes are output.

5. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 1, characterized in that, The step of calculating the shortest path distance from each active privacy semantic attribute to the identity identifier in the association path topology using the association path vector metric algorithm includes the following steps: Retrieve the probability of direct association between active privacy semantic attributes and identity identifiers from a pre-built prior knowledge relationship graph; The direct association probability is converted into the edge association weight between topological nodes in the associated path topology; Initialize a distance update vector table pointing to the topology node corresponding to the identity identifier for each active privacy semantic attribute; The shortest associated path distance in the updated vector table is updated by using the associated path vector metric algorithm and updating the distance through information exchange between neighboring topological nodes. Based on the distance-updated vector table, and using the path convergence function, the shortest associated path distance from the topology node corresponding to each active privacy semantic attribute to the topology node corresponding to the identity identifier is calculated in the current associated path topology state.

6. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 5, characterized in that, The process of using the associated path vector metric algorithm and updating the shortest associated path distance in the vector table through information exchange between neighboring topological nodes includes the following steps: Simulate and construct topology status messages between various topology nodes in the associated path topology; Collect the latest edge association weights between topological nodes, and construct a single-hop distance matrix that reflects the local proximity relationships in the topology of associated paths based on the latest edge association weights; Control each active privacy semantic attribute corresponding to the topology node to broadcast the current distance update vector table to neighboring topology nodes through topology status messages; The loop elimination operation blocks the cyclic announcement of topology status messages in the closed path of the associated path topology. When a sudden change in the shortest associated path distance is detected in the distance update vector table, the risk reversal mechanism is triggered and the infinite distance value representing the path blockage is written as the updated shortest associated path distance into the topology status message; Based on the received topology status message and combined with the single-hop distance matrix, the distance update vector table of each topology node is updated using the path distance algorithm to complete the update calculation of the shortest associated path distance.

7. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 1, characterized in that, The process of matching the generalization order based on the semantic dilution amount and a cascaded concept tree constructed using a preset category hierarchy based on privacy feature attributes, and then performing semantic smoothing dilution compensation on associated blocking nodes based on the generalization order to generate desensitized candidate text includes the following steps: Load the preset category hierarchy of privacy feature attributes and construct a cascaded concept tree based on the category hierarchy; Establish the correspondence between semantic dilution and the depth of the cascaded concept tree hierarchy, and match to obtain the target generalization order; Based on the target generalization order, a progressively upward generalization mapping algorithm is performed on the associated blocking nodes in the cascaded concept tree to obtain the target generalized entity; Extract the syntactic skeleton of the interactive text containing the associated blocking nodes, and replace the associated blocking nodes in the interactive text with placeholder indicator tags; The target generalized entity is filled into the syntactic skeleton corresponding to the placeholder indicator label to achieve semantic smoothing dilution compensation and generate desensitized candidate text.

8. The dynamic desensitization method for agent-oriented data interaction based on a sandbox according to claim 7, characterized in that, The step of performing a progressively upward generalization mapping algorithm on the associated blocking nodes in the cascaded concept tree according to the target generalization order to obtain the target generalized entity includes the following steps: Initialize the associated blocking node in the cascaded concept tree based on the target generalization order and calculate the initial semantic heat. Calculate the semantic similarity loss between the node to be evaluated and its parent node at the next higher level; The generalized cross-layer probability is calculated by combining the initial semantic heat and semantic similarity loss. If the randomly generated value is lower than the generalized cross-layer probability, the parent node of the previous level is taken as the new node to be evaluated; otherwise, the current node to be evaluated is determined as a temporary entity. The initial semantic heat is iteratively reduced using a preset heat decay coefficient until the initial semantic heat is lower than a preset termination threshold, at which point the temporary entity is used as the target generalized entity.

9. A sandbox-based dynamic desensitization system for agent-oriented data interaction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the dynamic desensitization method for agent-oriented data interaction based on a sandbox as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When executed by a processor, this instruction causes the processor to be configured to perform a dynamic desensitization method for agent-oriented data interaction based on a sandbox, as described in any one of claims 1 to 8.