A cloud file-based intelligent file management method and system
By constructing a trusted digital archive and a four-dimensional trusted state model, the security risk control problem of cloud archive management in AI scenarios is solved, dynamic trusted governance of the entire life cycle of archives is realized, and the level of security and intelligence is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA COMSERVICE SUPPLY CHAIN MANAGEMENT CO LTD
- Filing Date
- 2026-07-01
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies are not suitable for cloud archive management in AI scenarios, and cannot effectively address new security risks such as implicit reasoning leaks, semantic cognitive pollution, and loss of trust status during the archive circulation process. They also cannot achieve dynamic and trustworthy governance of the entire life cycle of archives.
A trustworthy digital entity for archives is constructed. Through feature extraction and a four-dimensional trustworthy state model, time-series evolution calculations are performed to generate dynamic governance strategies, enabling control over the access level of archives, dynamic desensitization, limitation of inference jumps, and control over the blocking of dissemination.
It achieves comprehensive risk quantification of archive security, activity, semantics, and dissemination dimensions, accurately controls the archive access level, desensitization intensity, inference depth, and dissemination path, and improves the security and intelligence level of cloud archive management.
Smart Images

Figure CN122490590A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud archive management technology, and in particular to an intelligent archive management method and system based on cloud archives. Background Technology
[0002] With the rapid development of artificial intelligence technology, cloud-based archive management is gradually shifting from static storage to dynamic intelligent governance. Existing technologies typically manage archives using static security classification, fixed access control, and rigid isolation, relying primarily on pre-defined access control lists and simple permission verification mechanisms. However, in AI scenarios, archive data often serves as training samples or inference material in complex business processes such as multi-agent collaborative computing and cross-domain federated inference. Traditional static management models struggle to address new security risks arising during archive circulation, such as implicit inference leaks, semantic cognitive pollution, and loss of trust status, failing to achieve dynamic and trustworthy governance throughout the entire archive lifecycle. Therefore, a method is urgently needed to address the technical challenge of adapting existing technologies to security risk management in AI scenarios. Summary of the Invention
[0003] This application provides an intelligent archive management method and system based on cloud archives, which solves the technical problem that existing technologies cannot adapt to security risk control in AI scenarios.
[0004] To achieve the above objectives, this application adopts the following technical solution: Firstly, a cloud-based intelligent archive management method is provided, comprising: acquiring archive data, access logs, agent call data, and archive association data from the cloud archive system, and constructing a trusted digital archive entity based on the archive data; extracting features from the trusted digital archive entity to obtain access anomaly features, semantic offset features, propagation link features, and topological association features; constructing a four-dimensional trusted state model including a secure trusted state, a knowledge active state, a semantically stable state, and a propagation diffusion state based on the access anomaly features, agent call data, semantic offset features, and propagation link features; performing time-series evolution calculations based on the four-dimensional trusted state model to obtain the archive trusted state evolution result; calculating the archive association entropy based on the trusted state evolution result and the topological association features to obtain the archive implicit inference risk; generating a dynamic governance strategy based on the archive trusted state evolution result and the archive implicit inference risk; and implementing openness level control, dynamic desensitization, inference hop limit, and propagation blocking control on the target archive based on the dynamic governance strategy.
[0005] In conjunction with the first aspect mentioned above, one possible implementation involves constructing a trusted digital archive based on archival data, including: acquiring the original content data, structural attribute data, and semantic description data of the target archive; encoding the original content data using a hash mapping method to obtain an archive hash fingerprint; parsing the structural attribute data using a structural feature extraction method to obtain an archive structural fingerprint; vectorizing the semantic description data using a semantic embedding encoding method to obtain an archive semantic fingerprint; fusing the archive hash fingerprint, archive structural fingerprint, and archive semantic fingerprint using a nonlinear fusion mapping method to obtain an archive fused fingerprint; and constructing a trusted digital archive based on the archive fused fingerprint, archive dynamic status data, risk measurement data, and governance constraint data.
[0006] In conjunction with the first aspect mentioned above, one possible implementation involves feature extraction from the trusted digital archive to obtain access anomaly features, semantic offset features, propagation link features, and topological association features. This includes: acquiring user access logs, agent call logs, semantic interaction data, and archive association data corresponding to the target archive; performing anomaly detection on the user access logs using a time-series behavior analysis method to obtain access anomaly features; vectorizing the semantic interaction data using a semantic vector encoding method and performing offset analysis based on the differences in semantic vector distribution to obtain semantic offset features; parsing the propagation path of the agent call logs using a link tracing method to obtain propagation link features; and constructing an archive association topology graph based on the archive association data and performing node association analysis on the archive association topology graph using a graph topology analysis method to obtain topological association features.
[0007] In conjunction with the first aspect mentioned above, one possible implementation involves using a semantic vector encoding method to vectorize semantic interaction data and performing offset analysis based on the differences in semantic vector distributions to obtain semantic offset features. This includes: acquiring the original semantic data of the target file and the corresponding AI-generated semantic data; using a pre-trained semantic encoding model to vectorize and encode the original semantic data and AI-generated semantic data to obtain original semantic vectors and generated semantic vectors; constructing a corresponding semantic probability distribution based on the original semantic vectors and generated semantic vectors; using the KL divergence calculation method to calculate the distribution difference of the semantic probability distribution to obtain semantic distribution offset parameters; using the Mahalanobis distance calculation method to calculate the spatial offset of the original semantic vectors and generated semantic vectors to obtain semantic spatial offset parameters; and generating semantic offset features corresponding to the target file based on the semantic distribution offset parameters and semantic spatial offset parameters.
[0008] In conjunction with the first aspect mentioned above, in one possible implementation, a four-dimensional trustworthy state model is constructed based on access anomaly features, agent call data, semantic offset features, and propagation link features. This model includes: statistically analyzing the abnormal access frequency of access anomaly features and performing nonlinear coupling calculations based on the call frequency and sensitive association strength in the agent call data to obtain the security trustworthy state corresponding to the target file; the sensitive association strength is obtained through sensitive association analysis based on file association data; constructing a file association topology network based on propagation link features and performing topology diffusion calculations based on the activity of associated nodes, node association weights, and topology propagation distance to obtain the knowledge active state corresponding to the target file; fusing the semantic distribution offset parameter and semantic space offset parameter in the semantic offset features to obtain the semantic stable state corresponding to the target file; constructing a propagation state transition link based on agent call data and performing chain diffusion calculations based on the propagation path length, node transition probability, and link attenuation coefficient to obtain the propagation diffusion state corresponding to the target file; and performing state coupling processing on the security trustworthy state, knowledge active state, semantic stable state, and propagation diffusion state to construct the four-dimensional trustworthy state model corresponding to the target file.
[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the secure and trusted state satisfies the following formula:
[0010] in, This represents the secure and trustworthy state of the i-th target file at time t; This represents the access anomaly rate of the i-th target file at time t. This represents the weighting coefficient for abnormal access risks; This represents the sensitivity correlation of the i-th target file at time t. Indicates the sensitive correlation coupling coefficient; This represents the agent's call frequency for the i-th target file at time t; This indicates the call to the change disturbance coefficient; This represents a non-linear activation function; i represents the target file index, and t represents the current time step; This represents the intermediate time variable during the integration process. This represents the rate of change in the frequency of agent calls; This represents the secure and trusted state of the i-th target file at the previous time step. The historical state damping attenuation coefficient; In conjunction with the first aspect mentioned above, in one possible implementation, the knowledge activity state satisfies the following formula:
[0011] in, Let j represent the knowledge activity state of target file i at time t, and j represent the adjacent node number associated with the i-th target file. G represents the set of adjacent nodes corresponding to the i-th target file; G represents the topological gravity constant. This represents the association weight between the i-th target file and the j-th adjacent node; This represents the knowledge activity of the j-th adjacent node at time t; This represents the topological propagation distance between the i-th target file and the j-th adjacent node; Indicates the time decay coefficient; In conjunction with the first aspect mentioned above, in one possible implementation, the semantically stable state satisfies the following formula:
[0012] in, This represents the semantically stable state of the i-th target file at time t; Indicates the semantic distribution offset weight coefficient; Represents the original semantic probability distribution With the probability distribution of generated semantics KL divergence between them; This represents the original semantic probability distribution corresponding to the i-th target file; This represents the generated semantic probability distribution corresponding to the i-th target file; This represents the semantic space offset weight coefficient; This represents the semantic vector corresponding to the i-th target file; Represents the mean of a semantic vector; This represents the matrix transpose operation; The inverse matrix representing the semantic vector covariance matrix; In conjunction with the first aspect mentioned above, in one possible implementation, the propagation and diffusion state satisfies the following formula:
[0013] in, The propagation state of the i-th target file at time t is represented by k; the propagation link level number is represented by k; and the propagation link length is represented by L. This represents the state transition probability corresponding to the k-th propagation link; This represents the link attenuation coefficient.
[0014] In conjunction with the first aspect mentioned above, in one possible implementation, a temporal evolution calculation based on a four-dimensional trust state model is performed to obtain the archive trust state evolution result. This includes: obtaining the four-dimensional trust state corresponding to the target archive at the current moment, including: a secure trust state, a knowledge-active state, a semantically stable state, and a propagation and diffusion state; constructing a state coupling matrix corresponding to the four-dimensional trust state, and determining the state transmission coefficient and state damping coefficient between each trust state; performing coupled evolution calculations on the secure trust state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state based on the state coupling matrix, the state transmission coefficient, and the state damping coefficient to obtain the trust state change amount; performing temporal iterative calculations based on the trust state change amount and historical trust state state data to obtain the trust state evolution state vector corresponding to the target moment; and generating the archive trust state evolution result corresponding to the target archive based on the trust state evolution state vector.
[0015] In conjunction with the first aspect mentioned above, in one possible implementation, based on the credible state evolution results and topological association characteristics, the archive association entropy is calculated to obtain the archive implicit inference risk. This includes: obtaining the credible state evolution results, archive association topology graph, and node association relationship data corresponding to the target archive; performing node centrality analysis on the archive association topology graph using graph topology analysis methods to obtain node degree centrality and node betweenness centrality; constructing archive topology composite weights based on node degree centrality, node betweenness centrality, and node association weights; calculating weighted entropy values based on topology composite weights and archive association probability distribution to obtain the archive association entropy corresponding to the target archive; the archive association probability distribution is constructed based on the archive node association probability; performing coupling analysis on the inference dependency relationships between multiple associated archives using mutual information calculation methods to obtain archive cluster coupling entropy; and generating the archive implicit inference risk corresponding to the target archive based on the archive association entropy, archive cluster coupling entropy, and credible state evolution results.
[0016] In conjunction with the first aspect mentioned above, in one possible implementation, a dynamic governance strategy is generated based on the evolution of the archive's trustworthy state and the implicit inference risk of the archive. This includes: obtaining the security and trustworthiness state, knowledge-active state, semantically stable state, propagation and diffusion state, and implicit inference risk of the target archive; calculating the state norm of the security and trustworthiness state, knowledge-active state, semantically stable state, and propagation and diffusion state to obtain the comprehensive risk level of the target archive; generating an openness level control strategy for the target archive using a dynamic adjustment method based on the comprehensive risk level and the implicit inference risk of the archive; generating a dynamic desensitization strategy for the target archive using a nonlinear risk coupling calculation method based on the security and trustworthiness state and the archive's association entropy; generating a reasoning hop count restriction strategy for the target archive using a topology depth constraint method based on the propagation and diffusion state, node centrality, and inference link length; generating an output constraint strategy for the target archive using a risk index constraint method based on the comprehensive risk level, archive association entropy, and propagation and diffusion state; and fusing the openness level control strategy, dynamic desensitization strategy, reasoning hop count restriction strategy, and output constraint strategy to obtain the dynamic governance strategy for the target archive.
[0017] Secondly, a cloud-based intelligent archive management system is provided. The system includes a data acquisition device and an electronic device. The data acquisition device acquires archive data, access behavior data, agent call data, and archive association data from the cloud archive system, and constructs a trusted digital archive entity based on the archive data. The electronic device extracts features from the trusted digital archive entity to obtain access anomaly features, semantic offset features, propagation link features, and topological association features. Based on the access anomaly features, agent call data, semantic offset features, and propagation link features, a four-dimensional trusted state model is constructed, including a secure trusted state, a knowledge active state, a semantically stable state, and a propagation diffusion state. A temporal evolution calculation is performed based on the four-dimensional trusted state model to obtain the archive trusted state evolution result. Based on the trusted state evolution result and the topological association features, the archive association entropy is calculated to obtain the archive implicit inference risk. Based on the archive trusted state evolution result and the archive implicit inference risk, a dynamic governance strategy is generated. Based on the dynamic governance strategy, openness level control, dynamic desensitization, inference jump limit, and propagation blocking control are implemented on the target archive.
[0018] This application provides an intelligent archive management method and system based on cloud archives. By constructing a trusted digital entity for archives, it provides archives with a dynamically evolving data structure foundation. By extracting multi-dimensional features and constructing a four-dimensional trusted state model, it achieves comprehensive risk quantification of archive security, activity, semantics, and propagation dimensions. Through temporal evolution calculation, it captures the dynamic transmission and evolution trend of risk states. By calculating archive association entropy, it effectively identifies implicit reasoning risks in the topological structure. Finally, based on multi-dimensional risk parameters, it generates dynamic governance strategies, achieving precise control over archive openness level, desensitization strength, reasoning depth, and propagation path. This solves the technical problem that existing technologies cannot adapt to the security risk management brought about by the dynamic evolution and complex association reasoning of archives in AI scenarios, and improves the security and intelligence level of cloud archive management.
[0019] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0020] Figure 1 A system architecture diagram of a cloud-based intelligent archive management system provided in this application embodiment; Figure 2 A flowchart illustrating an intelligent archive management method based on cloud archives, provided for an embodiment of this application; Figure 3 A flowchart illustrating another intelligent archive management method based on cloud archives provided in this application embodiment; Figure 4 A flowchart illustrating another intelligent archive management method based on cloud archives provided in this application embodiment; Figure 5 A flowchart illustrating another intelligent archive management method based on cloud archives provided in this application embodiment; Figure 6 A flowchart illustrating another intelligent archive management method based on cloud archives provided in this application embodiment; Figure 7A flowchart illustrating another intelligent archive management method based on cloud archives provided in this application embodiment; Figure 8 This is a flowchart illustrating another intelligent archive management method based on cloud archives provided in an embodiment of this application. Detailed Implementation
[0021] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0022] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0023] The cloud-based intelligent archive management method provided in this application can be applied to, for example... Figure 1 In the cloud-based intelligent archive management system shown, such as Figure 1 As shown, the system includes: a data acquisition device 101 and an electronic device 102; Among them, the data acquisition device 101 is used to acquire archive data, access behavior data, intelligent agent call data and archive association data in the cloud archive system, and to construct a trusted digital archive based on the archive data; Electronic device 102 is used to extract features from trusted digital archives, obtaining access anomaly features, semantic offset features, propagation link features, and topological association features. Based on access anomaly features, agent-invoked data, semantic offset features, and propagation link features, a four-dimensional trusted state model is constructed, including a secure trusted state, a knowledge active state, a semantically stable state, and a propagation diffusion state. Temporal evolution calculations are performed based on the four-dimensional trusted state model to obtain the archive trusted state evolution results. Based on the trusted state evolution results and topological association features, the archive association entropy is calculated to obtain the archive implicit inference risk. Based on the archive trusted state evolution results and the archive implicit inference risk, a dynamic governance strategy is generated. Based on the dynamic governance strategy, openness level control, dynamic desensitization, inference jump limit, and propagation blocking control are implemented on the target archives.
[0024] To address the technical problem that existing technologies cannot be adapted to security risk management in AI scenarios, this application provides an intelligent archive management method based on cloud archives. Figure 2 A flowchart illustrating the intelligent archive management method based on cloud archives provided in this application embodiment is shown below. Figure 2 As shown, the method includes: S201. Obtain archive data, access logs, intelligent agent call data, and archive association data from the cloud archive system, and construct a trusted digital archive based on the archive data.
[0025] Among them, archive data refers to the original file content and its metadata stored in the cloud archive system; access logs record the user or system's operation behavior on the archives; intelligent agent call data refers to the call records of the archives by AI models or automated programs during training and inference; archive association data describes the topological relationships such as references and associations between archives.
[0026] In this embodiment, the system first collects the aforementioned multi-source heterogeneous data from the cloud storage layer and log service. Based on the archival data, a trusted digital entity for the archives is constructed. This trusted digital entity is not a static file copy in the traditional sense, but a data structure that encapsulates the archive's static attributes (such as content fingerprints and structural features) and dynamic states (such as access frequency and risk measurement). The system performs feature encoding and fusion on the original content, structural attributes, and semantic descriptions of the archives to generate a unique trusted identifier, and binds real-time collected dynamic data such as access behavior and risk status to this identifier, forming a one-to-one digital entity mapping. This construction method enables each archive to possess a computable and evolvable digital twin, providing a data foundation for subsequent risk quantification and dynamic governance.
[0027] It should be noted that the construction process of a trustworthy digital entity for archives is dynamic and continuous. As the archive data is updated and external access behavior changes, the status data within the digital entity will be iterated and updated in real time, thereby ensuring the timeliness of risk perception.
[0028] As an example, when a classified file is uploaded to a cloud archive system, the system automatically extracts its hash fingerprint, structural fingerprint, and semantic fingerprint, merges them to generate an archive fusion fingerprint, and combines it with its initial access permissions, security classification identifier, etc. to construct the initial trusted digital body of the file.
[0029] Based on the above steps, this step achieves unified modeling of static attributes and dynamic behaviors of archives by constructing a trusted digital archive, which solves the problems of data silos and unknown status in traditional archive management and provides an accurate data carrier for subsequent risk calculation.
[0030] S202. Extract features from the trusted digital body of the archive to obtain access anomaly features, semantic offset features, propagation link features, and topological association features.
[0031] Among them, access anomaly features are used to characterize the degree of violation or abnormality of access behavior; semantic offset features are used to characterize the degree of semantic change of file content during AI processing; propagation link features are used to characterize the flow path and hierarchy of files among agents; and topological association features are used to characterize the importance and connection relationship of files in the associated network.
[0032] In this embodiment, the system performs in-depth mining of the multidimensional data encapsulated in the trusted digital body of the archives. Using time-series behavior analysis, it performs statistical analysis and pattern recognition on user access logs to extract abnormal access features such as high-frequency access and access during abnormal periods. Using semantic vector encoding, it compares the original semantics of the archives with the semantic vectors generated or invoked by AI, calculates their distribution differences, and extracts semantic offset features. Using link tracing, it analyzes the agent call logs to reconstruct the propagation path of the archives among multiple agents and extracts propagation link features. Using graph topology analysis, it constructs an archive association topology graph, analyzes indicators such as node degree and centrality, and extracts topological association features. These four types of features characterize the risk patterns faced by the archives from different dimensions, constituting the input variables for the subsequent risk quantification model.
[0033] It should be noted that the feature extraction process employs a sliding time window mechanism, which can capture short-term mutations and long-term trends, ensuring that the features can reflect the latest risk status of the archives.
[0034] Based on the above steps, this step achieves a comprehensive understanding of the risks related to archive security, semantics, dissemination, and association through multi-dimensional feature extraction, solving the problem that single-dimensional features cannot characterize the risks in complex AI scenarios.
[0035] S203. Based on access anomaly features, agent call data, semantic offset features, and propagation link features, a four-dimensional trustworthy state model is constructed, including a secure and trustworthy state, a knowledge-active state, a semantically stable state, and a propagation and diffusion state.
[0036] Among them, the secure and trustworthy state is used to quantify the security risks faced by archives, such as unauthorized access and malicious attacks; the knowledge active state is used to quantify the popularity and value of archives in the knowledge network; the semantically stable state is used to quantify the degree to which archive content is tampered with or polluted by AI; and the propagation and diffusion state is used to quantify the scope and speed of archive diffusion in the intelligent agent link.
[0037] In this embodiment, the system maps features to a four-dimensional state space: It non-linearly couples access anomaly features with agent call frequency to calculate a secure and trustworthy state; a higher state value indicates a greater security risk. It combines propagation link features with the activity of associated nodes to calculate a knowledge-active state using a topology diffusion algorithm; a higher state value indicates more active file circulation. It fuses distribution differences and spatial distance in semantic offset features to calculate a semantically stable state; a lower state value indicates a higher risk of semantic contamination. Finally, it combines path length and transition probability in propagation link features to calculate a propagation and diffusion state using a chain diffusion model; a higher state value indicates a wider file diffusion range. These four dimensions constitute a four-dimensional trustworthy state model for the file, capable of comprehensively describing the file's current operational status and risk level.
[0038] It should be noted that the states in the four-dimensional trustworthy state model do not exist in isolation, but are coupled. For example, an increase in security risk may inhibit the active state of knowledge, and semantic instability may exacerbate the risk of propagation and diffusion. This coupling relationship will be reflected in the subsequent evolutionary calculation.
[0039] S204. Based on the four-dimensional credible state model, perform time-series evolution calculations to obtain the credible state evolution results of the archives.
[0040] Among them, temporal evolution calculation refers to predicting the trend of changes in the credible state of archives over a future period of time based on historical and current states.
[0041] In this embodiment, the system constructs a state coupling matrix to describe the interrelationships among the four dimensions of security, activity, semantics, and propagation, and sets state transmission coefficients and damping coefficients. Based on the current four-dimensional trustworthy state vector, combined with the state coupling matrix, time-series deduction is performed using differential equations or discrete iteration. During the deduction process, the system considers not only the inertia (damping effect) of each state itself, but also the transmission influence of other states on it. For example, an increase in the propagation and diffusion state may lead to an increase in the security and trustworthy state through the transmission coefficient. Through multiple rounds of iterative calculation, the system outputs the evolution result of the archive trustworthy state at future times, which includes the predicted values of the states in each dimension.
[0042] It should be noted that the temporal evolution calculation introduces a random perturbation term to simulate the impact of uncertain factors such as sudden access or abnormal calls, making the prediction results more in line with the real scenario.
[0043] Based on the above steps, this step uses time-series evolution calculations to predict the risk trends of archives, thus solving the problem that traditional static assessments lag behind the occurrence of risks.
[0044] S205. Based on the results of the credible state evolution and the topological association characteristics, calculate the archive association entropy to obtain the implicit inference risk of the archive.
[0045] Among them, the archive association entropy is used to quantify the uncertainty of archives in the association network. The higher the entropy value, the greater the possibility of obtaining information through association reasoning. The archive implicit reasoning risk refers to the risk that attackers will deduce classified information by analyzing multiple non-classified archives.
[0046] In this embodiment, the system first analyzes the degree centrality and betweenness centrality of archive nodes in the network based on topological association characteristics, and constructs a topological composite weight. The higher the weight, the stronger the pivotal role of the archive in association reasoning. Next, the system combines the archive association probability distribution and the topological composite weight to calculate the weighted entropy value, obtaining the association entropy of a single archive. Further, the system uses a mutual information calculation method to analyze the reasoning dependencies between multiple associated archives and calculate the archive cluster coupling entropy. Finally, by combining the association entropy, cluster coupling entropy, and the credible state evolution results, the system generates the implicit reasoning risk of the archives. If the evolution results show that the activity of certain key archives increases and the association entropy increases, a high implicit reasoning risk is determined to exist.
[0047] It should be noted that this step introduces topological weights and mutual information, which can accurately identify seemingly isolated file nodes that can reveal core information through correlation reasoning.
[0048] Based on the above steps, this step achieves quantitative identification of implicit reasoning risks by calculating the file association entropy, thus solving the problem that existing technologies cannot prevent the leakage of information related to multiple files.
[0049] S206. Based on the results of the evolution of the credible state of archives and the implicit reasoning risks of archives, generate dynamic governance strategies.
[0050] Among them, dynamic governance strategy refers to a set of control measures that are adaptively generated based on the current and predicted risk status of the archives.
[0051] In this embodiment, the system comprehensively analyzes the state values of each dimension and the implicit inference risk value of the archive in the archive's trust state evolution result. The state norm is calculated for the four-dimensional trust state to obtain the comprehensive risk level. Based on the risk level and implicit inference risk, specific governance strategies are generated through a preset strategy mapping model or decision tree. These strategies include: openness level control strategies (e.g., lowering the openness level), dynamic desensitization strategies (e.g., increasing desensitization intensity), inference jump limit strategies (e.g., limiting the depth of related queries), and output constraint strategies (e.g., limiting the length of output content). The system integrates these strategies to form a complete dynamic governance strategy.
[0052] It should be noted that the strategy generation process is non-linear; high-risk profiles will trigger exponentially enhanced control measures, rather than simple linear adjustments.
[0053] Based on the above steps, this step generates a dynamic governance strategy by fusing multi-dimensional risk parameters, thereby achieving precision and adaptability of control measures and solving the problem of the lack of flexibility in traditional rigid control strategies.
[0054] S207. Based on dynamic governance strategies, target files are subject to openness level control, dynamic desensitization, inference jump limit, and propagation blocking control.
[0055] Among them, openness level control refers to adjusting the scope and access permissions of archives; dynamic desensitization refers to covering or replacing sensitive fields; inference jump limit refers to limiting the depth of access to related archives through association queries; and propagation blocking control refers to cutting off the propagation path of archives between intelligent agents.
[0056] In this embodiment, the system distributes the generated dynamic governance strategy to the execution layer of the cloud archive system. Based on the strategy instructions, the execution layer modifies the metadata tags of the target archive to adjust its openness level; calls the de-identification algorithm interface to perform real-time de-identification processing on the output content; sets a threshold for the depth of related queries in the query engine to limit the number of inference hops; and blocks propagation by modifying the access control list or cutting off the call link in the propagation chain. Through the above execution operations, the system completes closed-loop governance of archive risks.
[0057] It should be noted that the execution process is real-time and dynamic. Once the risk status of the files changes, the governance strategy will automatically adjust and implement new control measures.
[0058] As an example, if the system detects that a file has a semantic pollution risk, it will immediately implement a dynamic desensitization strategy, mask the key information in the output content, and restrict the file to be accessed only by low-privilege agents, thus preventing the pollution from spreading to other files.
[0059] This application's embodiments achieve quantitative perception and trend prediction of multidimensional risks of archives by constructing a trusted digital entity and a four-dimensional trusted state model; accurately identify implicit reasoning risks by calculating the archive association entropy; and finally generate and execute dynamic governance strategies to achieve intelligent security management and control of the entire lifecycle of cloud archives.
[0060] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above-mentioned S201, which constructs a trustworthy digital entity of archives based on archive data, can be specifically implemented through the following S301 to S305, which are explained in detail below: S301. Obtain the original content data, structural attribute data, and semantic description data of the target file.
[0061] Among them, raw content data refers to the binary stream data or text content of the archives; structural attribute data refers to the archives' metadata, directory structure, file format attributes, etc.; semantic description data refers to semantic information such as keywords, summaries, and subject classifications of the archives' content.
[0062] In this embodiment, before constructing a trusted digital archive, the system first needs to extract multi-dimensional source data from the storage layer and index of the cloud archive system. For raw content data, the system reads the byte stream of the archive file; for structural attribute data, the system parses the file's header information and directory tree; for semantic description data, the system calls a natural language processing interface to extract the archive's semantic tags. This acquisition of multi-source data ensures the comprehensiveness and accuracy of subsequent fingerprint generation.
[0063] S302. The original content data is encoded using a hash mapping method to obtain the archive hash fingerprint.
[0064] Among them, hash mapping method refers to the algorithm that maps data of arbitrary length to a hash value of fixed length, such as SHA-256, MD5, etc.; file hash fingerprint refers to the hash string used to uniquely identify the content of a file.
[0065] In this embodiment, the system uses a secure hash algorithm (SHA-256) to process the original content data. The binary stream of the file is input into the hash function, and a 256-bit hash value is output as the file's hash fingerprint. This fingerprint is unique and irreversible; any minor alteration to the original content will cause a drastic change in the hash fingerprint, thereby verifying the integrity of the file content.
[0066] S303. The structural attribute data is parsed and processed using the structural feature extraction method to obtain the archive structural fingerprint; the semantic embedding coding method is used to vectorize the semantic description data to obtain the archive semantic fingerprint.
[0067] Among them, structural feature extraction methods refer to methods that parse the structure of archive metadata and generate feature vectors; semantic embedding encoding methods refer to methods that map text semantic information into high-dimensional vectors, such as Word2Vec and BERT.
[0068] In this embodiment, for the archive structure fingerprint, the system parses the hierarchical relationships and attribute fields in the structural attribute data to construct a structural feature vector. For example, attributes such as directory depth, file size, and creation time are quantified into numerical features and combined to form the structure fingerprint. For the archive semantic fingerprint, the system uses a pre-trained semantic coding model (such as BERT) to encode the semantic description data, transforming textual information into a high-dimensional dense vector to capture the deep semantic features of the archive.
[0069] It should be noted that the introduction of structural fingerprints and semantic fingerprints makes up for the deficiency that a single hash fingerprint cannot characterize the structure and semantic features of an archive, enabling digital entities to reflect more dimensional attributes of the archive.
[0070] S304. The nonlinear fusion mapping method is used to fuse the archive hash fingerprint, archive structure fingerprint and archive semantic fingerprint to obtain the archive fusion fingerprint.
[0071] Among them, the nonlinear fusion mapping method refers to the method of fusing multiple heterogeneous fingerprints into a unified feature vector through nonlinear operations, such as tensor product operation and polynomial mapping; the archive fusion fingerprint refers to the comprehensive feature identifier after fusion.
[0072] In this embodiment, the system abandons the traditional simple splicing method and uses tensor product operations for nonlinear fusion. Hash fingerprints, structural fingerprints, and semantic fingerprints are mapped to a high-dimensional feature space, and nonlinear interaction terms between them are calculated through tensor product operations to generate a high-dimensional, dense archive fusion fingerprint. This fusion method not only preserves the feature information of each dimension but also introduces nonlinear correlations between dimensions, greatly enhancing the fingerprint's resistance to cracking and its irreversibility. Even if an attacker obtains partial fingerprint information, it will be difficult to deduce the complete features of the original archive.
[0073] It should be noted that nonlinear fusion mapping can also be implemented using kernel function mapping or deep neural network feature fusion layers, and this application does not impose any restrictions on this.
[0074] As an example, the system calculates the tensor product of the hash fingerprint vector H, the structural fingerprint vector S, and the semantic fingerprint vector V to obtain the archive fusion fingerprint.
[0075] S305. Based on the fusion fingerprint of archives, dynamic status data of archives, risk measurement data and governance constraint data, construct a trustworthy digital entity of archives.
[0076] Among them, the dynamic status data of archives refers to the real-time status of archives, such as current access popularity and access frequency; the risk measurement data refers to risk indicators such as current security risk score and semantic offset; and the governance constraint data refers to constraint information such as archive classification and access control policies.
[0077] In this embodiment, the system uses the archive fusion fingerprint as a static unique identifier, and encapsulates it with real-time collected dynamic status data of the archive, calculated risk measurement data, and preset governance constraint data to construct a complete trusted digital entity for the archive. This digital entity not only contains the static feature fingerprint of the archive, but also binds to its dynamic evolution state, forming a composite data structure of static fingerprint and dynamic state.
[0078] It should be noted that the trusted digital body of archives is a dynamically updated data structure. As archives are accessed, modified, and transferred, its internal dynamic status data and risk measurement data will be updated in real time. The archive fusion fingerprint is only recalculated when the content of the archives undergoes substantial changes.
[0079] This application embodiment achieves unified modeling of the full lifecycle characteristics of archives by constructing a trusted digital entity of archives that includes static fingerprints and dynamic states, providing an accurate data carrier for subsequent risk evolution calculations.
[0080] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 4 As shown, the above S202 can be specifically implemented through the following S401 to S405, which are explained in detail below: S401. Obtain the user access log, agent call log, semantic interaction data, and file association data corresponding to the target file.
[0081] Among them, the user access log records information such as the time, frequency, and type of user access to the archive; the agent call log records the link information of the AI model or automated program calling the archive; the semantic interaction data includes the original semantic content of the archive and the semantic content generated after AI processing; and the archive association data describes the topological structure such as references and associations between archives.
[0082] In this embodiment, the system extracts the aforementioned multi-source heterogeneous data from the trusted digital archive as input sources for feature extraction. The system reads user access logs and agent call logs through a log parsing interface, extracts semantic interaction data through a semantic analysis interface, and queries archive relationship data through a topological relation database. This acquisition of multi-source data ensures the comprehensiveness and accuracy of subsequent feature extraction.
[0083] S402. Use time-series behavior analysis to detect anomalies in user access logs and obtain access anomaly characteristics.
[0084] Among them, time-series behavior analysis method refers to the method of statistical analysis and pattern recognition of user behavior based on time series; access anomaly features are used to characterize the degree of violation or anomaly of access behavior.
[0085] In this embodiment, the system performs time-series analysis on user access logs, calculating statistical indicators such as access frequency, access time distribution, and operation sequence patterns. Using a preset anomaly detection model (such as the Isolation Forest algorithm or statistical threshold determination), it identifies abnormal behaviors such as high-frequency access, access outside of working hours, and abnormal operation sequences, quantifying them into access anomaly feature vectors. For example, the system calculates the number of accesses per unit time; if it exceeds a preset threshold, it is determined to be an abnormal frequency.
[0086] It should be noted that the extraction of access anomaly features not only focuses on the anomaly of a single behavior, but also on the contextual association of the behavior sequence in order to identify covert probing attack behaviors.
[0087] S403. Use the link tracing method to analyze the propagation path of the agent's call logs and obtain the propagation link characteristics.
[0088] Link tracing refers to the technique of reconstructing the path of data flow between multiple nodes; propagation link features are used to characterize the propagation path, level and scope of files among agents.
[0089] In this embodiment, the system parses the agent call logs to construct a call relationship graph of the file among multiple agent nodes. By tracing the starting node, intermediate nodes, and ending node of the call chain, it calculates indicators such as propagation path length, propagation level depth, and node transfer probability to generate propagation link features. These features can reflect the scope and speed of file circulation within the AI system.
[0090] It should be noted that the extraction of propagation link features adopts distributed link tracing technology, which can reconstruct the complete call link across service nodes and avoid the loss of link information.
[0091] S404. Construct an archive association topology graph based on archive association data, and use graph topology analysis method to perform node association analysis on the archive association topology graph to obtain topology association features.
[0092] Among them, the archive association topology graph refers to the graph structure constructed with archives as nodes and association relationships as edges; the topology association features are used to characterize the importance of the nodes and the connection relationships of the archives in the association network.
[0093] In this embodiment, the system constructs a graph structure based on archival relationship data, where nodes represent archives and edges represent references or associations. Graph topology analysis methods are used to calculate topological indices such as degree centrality, betweenness centrality, and clustering coefficient of nodes, generating topological association features. These features reflect the hub status and association strength of archives within the knowledge network.
[0094] S405. The semantic vector encoding method is used to vectorize the semantic interaction data, and the offset analysis is performed based on the difference in semantic vector distribution to obtain semantic offset features.
[0095] Among them, semantic vector encoding refers to the method of mapping text semantic information into high-dimensional vectors; semantic offset features are used to characterize the degree of semantic change in archive content during AI processing.
[0096] In this embodiment, the system first acquires the original semantic data of the target file and the corresponding AI-generated semantic data. A pre-trained semantic encoding model (such as BERT or RoBERTa) is used to vectorize and encode the original semantic data and the AI-generated semantic data, resulting in original semantic vectors and generated semantic vectors. Based on the original semantic vectors and generated semantic vectors, a corresponding semantic probability distribution is constructed. The system uses the KL divergence calculation method to calculate the distribution difference of the semantic probability distribution, obtaining semantic distribution offset parameters; simultaneously, the Mahalanobis distance calculation method is used to calculate the spatial offset of the original semantic vectors and generated semantic vectors, obtaining semantic spatial offset parameters. Finally, based on the semantic distribution offset parameters and the semantic spatial offset parameters, semantic offset features corresponding to the target file are generated.
[0097] It should be noted that this step abandons the single cosine similarity calculation and introduces KL divergence and Mahalanobis distance as dual measures, which can accurately identify semantic probability distribution shifts and spatial location drifts, and effectively detect the cognitive pollution risk caused by AIGC content.
[0098] Optionally, semantic distribution offset parameters The calculation satisfies the following formula:
[0099] in, Represents the original semantic probability distribution. This represents the semantic probability distribution generated. This formula is used to measure the difference between two probability distributions.
[0100] Optionally, semantic space offset parameters The calculation satisfies the following formula:
[0101] in, This represents the semantic vector corresponding to the i-th target file. Let T represent the mean of the semantic vectors, and let T represent the matrix transpose operation. This represents the inverse matrix of the semantic vector covariance matrix. This formula is used to measure the distance between vectors in a high-dimensional space, eliminating dimensional interference.
[0102] Based on the above steps, this step achieves in-depth detection of semantic pollution risks through dual semantic offset analysis, solving the problem that a single similarity index cannot identify deep semantic tampering.
[0103] This application embodiment comprehensively captures the risk characteristics of archives at the levels of access behavior, semantic content, propagation path and topology through multi-dimensional feature extraction. In particular, it improves the accuracy of semantic pollution risk identification through dual semantic analysis of KL divergence and Mahalanobis distance.
[0104] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 5 As shown, the above S203 can be specifically implemented through the following S501 to S505, which are explained in detail below: S501. Statistical analysis of abnormal access frequency is performed on the abnormal access characteristics, and nonlinear coupling calculation is performed in combination with the call frequency and sensitive association strength in the agent call data to obtain the secure and trustworthy state corresponding to the target file.
[0105] Among them, the sensitivity correlation strength is obtained by sensitivity correlation analysis based on the data of the archives' correlation relationship, and is used to characterize the degree of correlation between the archives and sensitive nodes; nonlinear coupling calculation refers to the calculation method of fusing multiple risk factors through nonlinear functions.
[0106] In this embodiment, the system does not simply perform a linear weighting of the risk factors, but constructs a nonlinear dynamic model containing squared terms, product terms, and integral terms. The system first statistically analyzes the frequency of access anomalies to calculate the access anomaly rate; simultaneously, it analyzes the agent's call logs to extract the call frequency and its rate of change; and combines this with file association relationships to calculate the strength of sensitive associations. A nonlinear activation function (such as the sigmoid function) is used to couple and map these factors, generating a safe and trustworthy state. This nonlinear coupling method can capture the linkage effect between risk factors; for example, when the access anomaly rate is high and the agent's call frequency surges, the risk value will increase exponentially, rather than through simple linear superposition.
[0107] It should be noted that the calculation of the safe and reliable state introduces a historical state damping attenuation mechanism, which makes the current risk state affected by both the current factors and the inertia of the historical state, thus avoiding drastic fluctuations in the risk value.
[0108] Optionally, the secure and trusted state satisfies the following formula:
[0109] in, This represents the secure and trustworthy state of the i-th target file at time t; This represents the access anomaly rate of the i-th target file at time t. This represents the weighting coefficient for abnormal access risks; This represents the sensitivity correlation of the i-th target file at time t. Indicates the sensitive correlation coupling coefficient; This represents the agent's call frequency for the i-th target file at time t; This indicates the call to the change disturbance coefficient; This represents a non-linear activation function; i represents the target file index, and t represents the current time step; This represents the intermediate time variable during the integration process. This represents the rate of change in the frequency of agent calls; This represents the secure and trusted state of the i-th target file at the previous time step. The historical state damping attenuation coefficient is given.
[0110] It should be noted that in this formula, The term adopts a squared form, which can amplify the penalty weight of high-risk abnormal access, so that the risk increases exponentially with the increase of the abnormality rate; The term is the product of the sensitivity correlation degree and the agent's call frequency, representing the amplification effect of the sensitivity correlation risk in high-frequency call scenarios; the integral term captures the sudden changes in call frequency within a short period of time, identifying instantaneous attack behaviors; the damping term... This introduces the inertia of historical states, smoothing out risk fluctuations.
[0111] Based on the above steps, this step calculates the safe and reliable state through a nonlinear coupling model, which enables the accurate capture of the linkage effect of multiple risk factors and solves the problem that traditional linear models cannot identify high-risk linkage risks.
[0112] S502. Construct an archive association topology network based on the characteristics of the propagation link, and perform topology diffusion calculation based on the activity of associated nodes, node association weights and topology propagation distance to obtain the knowledge activity state corresponding to the target archive.
[0113] Among them, the archive association topology network refers to a network structure built with archives as nodes and association relationships as edges; topology diffusion computation refers to the computational process of simulating the attenuation of knowledge or information as it propagates in the network.
[0114] In this embodiment, the system constructs an interconnected topology network of archives based on propagation link characteristics. For any target archive node in the network, its knowledge activity state depends not only on its own access popularity but also on the activity of neighboring nodes. The system introduces a topological gravity model to calculate the contribution of neighboring nodes to the activity of the target node. The higher the activity and the greater the association weight of a neighboring node, the greater its contribution to the target node's activity; however, the greater the topological propagation distance, the greater the contribution decays with the square of the distance. By traversing all neighboring nodes and performing a weighted sum, the knowledge activity state of the target archive is obtained. This calculation method simulates the diffusion law of knowledge in the network and can identify key archives that, although their own access volume is not high, are located at the hub of knowledge propagation.
[0115] It should be noted that the calculation of knowledge activity state also introduces a time decay factor, which makes the influence of historical active behavior on the current state gradually weaken over time, thus ensuring the timeliness of the activity state.
[0116] Optionally, the knowledge activity state satisfies the following formula:
[0117] in, Let j represent the knowledge activity state of target file i at time t, and j represent the adjacent node number associated with the i-th target file. G represents the set of adjacent nodes corresponding to the i-th target file; G represents the topological gravity constant. This represents the association weight between the i-th target file and the j-th adjacent node; This represents the knowledge activity of the j-th adjacent node at time t; This represents the topological propagation distance between the i-th target file and the j-th adjacent node; Indicates the time decay coefficient; Based on the above steps, this step calculates the knowledge activity state through the topological gravity model, realizing a quantitative assessment of the value of archives in the knowledge network, and solving the problem that traditional methods only focus on single-point access volume and ignore the network propagation effect.
[0118] S503. The semantic distribution offset parameter and semantic space offset parameter in the semantic offset feature are fused and calculated to obtain the semantic stable state corresponding to the target file.
[0119] Among them, the semantic distribution offset parameter is used to characterize the difference at the semantic probability distribution level; the semantic space offset parameter is used to characterize the position offset of the semantic vector in the high-dimensional space.
[0120] In this embodiment, the system extracts semantic distribution offset parameters (such as KL divergence) and semantic space offset parameters (such as Mahalanobis distance) from semantic offset features. The semantic distribution offset parameter reflects the difference in probability distribution between the original semantics of the archive and the AI-generated semantics, enabling the identification of logical tampering or factual deletion; the semantic space offset parameter reflects the spatial drift of the semantic vector, enabling the identification of semantic drift or irrelevant illusions. The system combines these two parameters into a semantic stable state through weighted fusion. The lower the value of the semantic stable state, the higher the degree to which the archive content has been tampered with or contaminated by AI. This dual-measurement mechanism compensates for the deficiency of a single similarity index in identifying deep semantic changes.
[0121] Optionally, the semantically stable state satisfies the following formula:
[0122] in, This represents the semantically stable state of the i-th target file at time t; Indicates the semantic distribution offset weight coefficient; Represents the original semantic probability distribution With the probability distribution of generated semantics KL divergence between them; This represents the original semantic probability distribution corresponding to the i-th target file; This represents the generated semantic probability distribution corresponding to the i-th target file; This represents the semantic space offset weight coefficient; This represents the semantic vector corresponding to the i-th target file; Represents the mean of a semantic vector; This represents the matrix transpose operation; This represents the inverse matrix of the semantic vector covariance matrix.
[0123] This formula quantifies both the difference in probability distribution and the shift in spatial location, and the combination of the two enables comprehensive detection of semantic pollution.
[0124] Based on the above steps, this step achieves accurate identification of cognitive pollution of AIGC content by calculating the semantic steady state through dual-metric fusion, and solves the problem that a single similarity index cannot detect deep semantic tampering.
[0125] S504. Construct a propagation state transition link based on the data called by the intelligent agent, and perform chain diffusion calculation based on the propagation path length, node transition probability and link attenuation coefficient to obtain the propagation diffusion state corresponding to the target file.
[0126] Among them, the propagation state transition link refers to the path sequence of files flowing between multiple intelligent agent nodes; chain diffusion computation refers to the computational process of simulating the propagation of risk along the link based on the Markov chain model.
[0127] In this embodiment, the system parses agent call data and constructs a propagation state transition chain. Each chain contains multiple propagation levels, and the file has a certain transition probability at each level. The risk decreases as the propagation level increases. The system uses a chain-like diffusion model to accumulate the transition probabilities and attenuation coefficients at each level to obtain the propagation diffusion state. A higher value in the propagation diffusion state indicates a wider spread of the file within the agent network and a greater potential leakage risk. This chain-like calculation method can reconstruct the propagation trajectory of risk in complex call chains.
[0128] Optionally, the propagation and diffusion state satisfies the following formula:
[0129] in, The propagation state of the i-th target file at time t is represented by k; the propagation link level number is represented by k; and the propagation link length is represented by L. This represents the state transition probability corresponding to the k-th propagation link; This represents the link attenuation coefficient.
[0130] Based on the above steps, this step calculates the propagation and diffusion state using a chain diffusion model, thereby achieving dynamic quantification of the risk of cross-agent propagation of archives and solving the problem that traditional methods cannot track the risk of chain propagation.
[0131] S505. Perform state coupling processing on the secure and trustworthy state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state to construct a four-dimensional trustworthy state model corresponding to the target file.
[0132] Among them, state coupling processing refers to the vector encapsulation or matrix transformation of the four-dimensional credible states to construct a unified state model.
[0133] In this embodiment, the system encapsulates the security-trusted state, knowledge-active state, semantically stable state, and propagation-diffusion state calculated in the above four steps into a four-dimensional state vector, which is the four-dimensional trusted state model corresponding to the target file. This model comprehensively describes the current operational state and risk level of the file from four dimensions: security, activity, semantics, and propagation. These four dimensions do not exist in isolation but influence each other through a state coupling matrix, providing a state foundation for subsequent temporal evolution calculations.
[0134] This application's embodiments quantify the trustworthy state across four dimensions—security, activity, semantics, and propagation—through specific mathematical models. It introduces mechanisms such as nonlinear coupling, topological gravity, dual metrics, and chain diffusion to ensure the accuracy and physical interpretability of risk calculations, thereby enhancing the model's ability to characterize archival risks in complex AI scenarios.
[0135] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 6 As shown, the above S204 can be specifically implemented through the following S601 to S605, which are explained in detail below: S601. Obtain the four-dimensional credible state of the target file at the current moment.
[0136] Among them, the four-dimensional trustworthy state includes the secure trustworthy state, the knowledge active state, the semantically stable state, and the propagation and diffusion state.
[0137] In this embodiment of the application, the system reads the four-dimensional trusted state vector at the current time t from the state layer of the trusted digital archive, which constitutes the initial state input for the time-series evolution calculation.
[0138] S602. Construct the state coupling matrix corresponding to the four-dimensional credible state, and determine the state transmission coefficient and state damping coefficient between each credible state.
[0139] Among them, the state coupling matrix is used to describe the relationship of mutual influence and transmission between the four credible states; the state transmission coefficient is used to quantify the degree of influence of the state change in one dimension on the other dimension; and the state damping coefficient is used to quantify the inertial characteristics of the state itself decaying over time.
[0140] In this embodiment, the system constructs a 4x4 state coupling matrix based on historical data training or expert experience. The off-diagonal elements of the matrix are state propagation coefficients; for example, an increase in the propagation state may lead to an increase in the safe and reliable state, and this propagation relationship is characterized by the corresponding propagation coefficient. The diagonal elements of the matrix contain state damping coefficients, used to characterize the self-attenuation characteristics of the state, i.e., the impact of historical risks on the current state gradually weakens over time. This matrix construction method simulates the linkage effect of risk factors in real-world risk control scenarios.
[0141] It should be noted that the state coupling matrix is not fixed and can be dynamically adjusted based on the actual risk control effect feedback to adapt to different business scenarios.
[0142] As an example, in the state coupling matrix of the system construction, the propagation coefficient of the propagation state to the safe and reliable state is 0.3, which means that for every unit increase in propagation risk, the safety risk will increase by 0.3 units; the damping coefficient of the safe and reliable state is 0.1, which means that the historical safety risk has 10% inertial retention on the current state.
[0143] S603. Based on the state coupling matrix, state transmission coefficient, and state damping coefficient, the coupled evolution calculation of the secure and trustworthy state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state is performed to obtain the change in the trustworthy state.
[0144] Among them, coupled evolution calculation refers to calculating the change trend of a four-dimensional credible state within a small time interval based on differential dynamic equations.
[0145] In this embodiment, the system performs matrix operations on the four-dimensional reliable state vector and the state coupling matrix, and calculates the reliable state change by combining the state damping coefficient. A differential equation is used to describe the rate of change of state, which considers not only the damping attenuation of each state itself, but also the transmission influence of other states. To implement this calculation in a computer system, the continuous differential equation is discretized and transformed into a difference equation for solution. By calculating the product of the current state vector and the coupling matrix, and subtracting the damping attenuation term, the reliable state change per unit time step is obtained.
[0146] It should be noted that the coupled evolutionary computation introduces a random perturbation term to simulate the impact of uncertainties such as sudden access or abnormal calls on state evolution, making the prediction results more consistent with the real environment.
[0147] S604. Based on the change in the credible state and the historical credible state data, perform time-series iterative calculations to obtain the credible state evolution vector corresponding to the target time.
[0148] Among them, time-series iterative calculation refers to superimposing the credible state change quantity onto the current state and gradually deducing the state value at future moments.
[0149] In this embodiment, the system employs a discretized iterative formula for calculation. The current state vector is superimposed with the calculated change in the credible state to obtain the predicted state vector for the next moment. The system iteratively executes this process at preset time steps (e.g., 1 minute, 5 minutes) until the state value at the target moment is derived. This iterative calculation method can simulate the evolution trajectory of risk states over time, achieving accurate prediction of future risk trends.
[0150] It should be noted that boundary constraints are introduced during the iteration process to ensure that the predicted state value is always within a reasonable physical range (such as the [0,1] interval) to avoid numerical overflow or distortion.
[0151] S605. Based on the credible state evolution vector, generate the credible state evolution result of the target file.
[0152] Among them, the archive credibility state evolution result refers to a data structure that includes the predicted values of the state in each dimension at future time and risk trend labels.
[0153] In this embodiment, the system encapsulates the iteratively calculated trusted state evolution vector into an archive trusted state evolution result. This result not only includes the specific values of each dimension at the target time, but also generates risk trend labels (such as a sharp increase in risk, aggravated semantic pollution, etc.) based on the magnitude of the values. The system stores this evolution result in the evolution history layer of the archive trusted digital entity, providing data support from a future perspective for the generation of subsequent governance strategies.
[0154] This application embodiment simulates the linkage transmission and damping attenuation effect between multidimensional risk states by constructing a state coupling matrix and performing time-series evolution calculations. This enables accurate prediction of archival risk trends, solves the problem that traditional static risk assessment lags behind the occurrence of risks, and enhances the proactive defense capability of cloud archival management.
[0155] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 7 As shown, the above-mentioned S205, which constructs a trustworthy digital entity of archives based on archive data, can be specifically implemented through the following S701 to S706, which are explained in detail below: S701. Obtain the trusted state evolution result, file association topology diagram and node association data corresponding to the target file.
[0156] Among them, the credible state evolution result includes the predicted risk state of the archive at future moments; the archive association topology graph is a graph structure constructed with archives as nodes and association relationships as edges; the node association relationship data describes the connection strength and type between nodes.
[0157] In this embodiment, the system extracts the credible state evolution results from the evolutionary history layer of the credible digital archive, and simultaneously reads the archive association topology graph and its corresponding node association data from the topology feature library. This data forms the basis for subsequent topological entropy calculations.
[0158] It should be noted that the node association data includes not only direct associations but also indirect associations formed through multi-hop paths to ensure the comprehensiveness of the association analysis.
[0159] S702. Graph topology analysis is used to perform node centrality analysis on the archive association topology graph to obtain node degree centrality and node betweenness centrality.
[0160] Among them, the degree centrality of a node is used to characterize the number of direct associations of a file node. The larger the value, the more other files the file is directly related to. The betweenness centrality of a node is used to characterize the degree to which a file node acts as a "bridge" in the topological network. The larger the value, the more hub positions the file is in the more association paths.
[0161] In this embodiment, the system performs graph computation analysis on the archive association topology. The degree centrality of a node is calculated by counting the number of edges directly connected to the target archive node; the betweenness centrality of a node is calculated by calculating the proportion of all shortest paths in the network that pass through the target archive node. These two indicators can effectively identify archive nodes that hold an important position in the network; these nodes are often key entry points for implicit reasoning.
[0162] It should be noted that for large-scale topological graphs, the system can use parallel graph computing frameworks (such as GraphX) to accelerate the centrality computation process.
[0163] Based on the above steps, this application quantifies the structural importance of archive nodes in the topological network through node centrality analysis, solving the problem that traditional entropy calculation ignores the differences in node positions.
[0164] S703. Based on node degree centrality, node betweenness centrality, and node association weight, construct a composite weight for the archive topology.
[0165] Among them, node association weight refers to the strength of the relationship between archives, such as citation frequency and semantic similarity; archive topology composite weight is a comprehensive weight value that integrates node importance indicators and association strength.
[0166] In this embodiment, the system abandons the traditional single-weight method based solely on association probability in entropy calculation and constructs a composite weight model. The system weights and fuses node degree centrality, node betweenness centrality, and node association weights to generate a composite weight for the archive topology. A higher weight value indicates a stronger information hub role for the archive node in the association reasoning network; conversely, if this node is acquired by an attacker, the risk of leaking core information through association reasoning is also higher. This construction method allows risk quantification to focus more on key nodes, improving the accuracy of risk identification.
[0167] S704. Calculate the weighted entropy value based on the topological composite weight and the archive association probability distribution to obtain the archive association entropy corresponding to the target archive.
[0168] Among them, the probability distribution of archive association refers to the probability distribution of the target archive being associated with other archives; the archive association entropy is used to quantify the uncertainty of archives in the association network. The higher the entropy value, the greater the possibility of obtaining information through association reasoning.
[0169] In this embodiment, the system uses a weighted entropy formula for calculation. Topological composite weights are introduced as adjustment factors into the Shannon entropy formula to weight the probability distribution of file associations. The calculated file association entropy not only reflects the complexity of the association relationships but also incorporates the influence of node importance. For core nodes with high weights, even if their association probability is low, they will still contribute a high entropy value, thus highlighting their risky position.
[0170] It should be noted that weighted entropy calculation can effectively avoid the interference of high correlation between edge nodes on risk assessment, making risk indicators more focused on the real threat sources.
[0171] S705. The mutual information calculation method is used to perform coupling analysis on the inference dependency relationship between multiple related archives to obtain the coupling entropy of the archive cluster.
[0172] Mutual information is used to measure the degree of interdependence between two random variables; the archive cluster coupling entropy is used to quantify the risk of information leakage between multiple archives through joint reasoning.
[0173] In this embodiment, the system not only calculates the association entropy of individual files but also further analyzes the coupling relationship between file clusters. The system employs a mutual information calculation method to calculate the mutual information value between each pair of files in the associated file set, characterizing the strength of their inference dependencies. The association entropy and mutual information value of each file are then fused to obtain the file cluster coupling entropy. This entropy value reflects the amount of information contained in the file cluster as a whole, and can identify those individual files that appear secure but, when combined, can lead to the aggregation and leakage risk of sensitive information.
[0174] It should be noted that mutual information calculation can capture nonlinear relationships between variables, and is more accurate than simple correlation analysis.
[0175] S706. Based on the archive association entropy, archive cluster coupling entropy, and credible state evolution results, generate the implicit inference risk of the target archive.
[0176] Among them, the risk of implicit reasoning in archives refers to the possibility and degree of harm that attackers can deduce classified information by analyzing multiple non-classified archives.
[0177] In this embodiment, the system integrates the archive association entropy, archive cluster coupling entropy, and the trusted state evolution results to generate a final implicit inference risk label or score. The system weights and sums the archive association entropy and cluster coupling entropy, and combines this with risk trends (such as an upward trend in the trusted state) from the trusted state evolution results to generate an implicit inference risk value using a risk mapping function (such as a nonlinear activation function). If the evolution results indicate that an archive will be accessed frequently in the future, and its association entropy and cluster coupling entropy are both high, the system will determine that the archive has an extremely high implicit inference risk and generate a corresponding risk warning.
[0178] It should be noted that the risk generation process takes into account the time dimension, can predict implicit inference risks at future moments, and achieves early warning.
[0179] This application's embodiments address the problem that traditional entropy cannot identify core leak nodes by introducing topological composite weights to distinguish node importance; they also identify cluster coupling leak risks by quantifying implicit relationships between archives through mutual information calculation; and finally, they generate implicit inference risks by combining evolutionary trends, achieving accurate early warning of complex association inference risks.
[0180] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 8 As shown, the above-mentioned S206, which constructs a trustworthy digital entity of archives based on archive data, can be specifically implemented through the following S801 to S807, which are explained in detail below: S801. Obtain the security and trustworthiness state, knowledge-active state, semantically stable state, propagation and diffusion state, and implicit reasoning risk of the target file.
[0181] In this embodiment, the system extracts the four-dimensional credible state values at the current or predicted time from the state layer of the archive's credible digital entity, and simultaneously obtains the calculated implicit inference risk values of the archive from the risk analysis module. These multi-dimensional parameters constitute the input dataset for generating the dynamic governance strategy, ensuring that the governance strategy can comprehensively reflect the security, activity, semantics, propagation, and related inference risks of the archive.
[0182] S802. Calculate the state norm of the secure and trustworthy state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state to obtain the comprehensive risk level corresponding to the target file.
[0183] Among them, the state norm calculation refers to the calculation method that maps a multidimensional state vector into a single comprehensive value; the comprehensive risk level is used to characterize the overall risk level of the archive.
[0184] In this embodiment, the system abandons simple linear weighted summation and adopts a L2 norm calculation method to comprehensively measure the four-dimensional credible state. The system calculates the Euclidean norm of the four-dimensional state vector, which is the square root of the sum of the squares of the state values of each dimension, to obtain the comprehensive risk level. This calculation method can amplify the contribution of high-risk dimensions. For example, when the risk value of a certain dimension is extremely high, the comprehensive risk level will increase significantly, thereby avoiding the problem of low-risk dimensions masking high-risk dimensions and ensuring that high-risk files are prioritized for management.
[0185] It should be noted that the calculation of the comprehensive risk level can also introduce weighting coefficients to differentiate the risk tolerance of different dimensions according to the business scenario.
[0186] S803. Based on the comprehensive risk level and the implicit inference risk of the archives, a dynamic adjustment method for the access level is adopted to generate the access level control strategy corresponding to the target archives.
[0187] Among them, the dynamic adjustment method of openness level refers to the method of adaptively adjusting the scope and access of archives based on the risk level; the openness level control strategy is used to restrict access to archives.
[0188] In this embodiment, the system constructs an openness level control equation, using the overall risk level and implicit inference risk as input variables. The system sets a risk threshold range; when either the overall risk level or the implicit inference risk exceeds a preset threshold, the openness level of the file is automatically reduced. For example, if the overall risk level exceeds 0.8, the file's openness level is downgraded from public to internal; if it exceeds 0.9, it is downgraded to confidential. This adjustment mechanism is non-linear; the higher the risk, the greater the downgrade, thus enabling rapid response to high-risk files.
[0189] S804. Based on the secure and trustworthy state and the file association entropy, a nonlinear risk coupling calculation method is used to generate a dynamic desensitization strategy corresponding to the target file.
[0190] Among them, the nonlinear risk coupling calculation method refers to the method of nonlinearly fusing security risks and related risks to determine the desensitization intensity; the dynamic desensitization strategy is used to cover or replace sensitive fields in the archive.
[0191] In this embodiment, the system employs a hyperbolic tangent function to construct a desensitization intensity calculation model. The system weights and sums the secure and trustworthy state with the file association entropy, inputs this sum into the hyperbolic tangent function, and outputs the desensitization intensity level. The hyperbolic tangent function exhibits a gentle change in low-risk areas and a steep change in high-risk areas, meaning that low-risk files require only mild desensitization, while high-risk files require severe desensitization. This non-linear mechanism ensures both data availability and security.
[0192] It should be noted that the de-identification strategy includes various techniques such as character masking, data generalization, and data substitution. The system automatically selects the appropriate de-identification algorithm based on the de-identification strength level.
[0193] S805. Based on propagation and diffusion state, node centrality, and inference link length, a topology depth constraint method is used to generate an inference hop count limit strategy corresponding to the target file.
[0194] Among them, the topology depth constraint method refers to the method of limiting the depth of association queries based on the location of the file in the network and the risk of propagation; the inference hop count limit strategy is used to restrict attackers from obtaining sensitive information through multi-hop association inference.
[0195] In this embodiment, the system constructs a reasoning depth constraint formula, taking the propagation diffusion state, node centrality (such as betweenness centrality), and reasoning link length as input parameters. For files with high propagation diffusion state and high node centrality, the system significantly compresses the allowed number of reasoning hops. For example, if the propagation diffusion state is 0.8 and the node is the central hub, the number of reasoning hops is limited to 1 hop, meaning that querying related files through this file is prohibited. This constraint mechanism cuts off deep reasoning links at the topology level, effectively preventing the risk of implicit data leakage.
[0196] It should be noted that the reasoning jump limit strategy can be dynamically adjusted. When the risk of file dissemination decreases, the system can automatically relax the limit.
[0197] S806. Based on the comprehensive risk level, archive association entropy, and propagation diffusion state, the risk index constraint method is used to generate the output constraint strategy corresponding to the target archive.
[0198] Among them, the risk index constraint method refers to the method of limiting the amount of output content of the archive based on the comprehensive risk index; the output constraint strategy is used to limit the amount of information exposed when the archive is retrieved or output.
[0199] In this embodiment, the system constructs an index constraint model, fusing the comprehensive risk level, archive association entropy, and propagation diffusion state into a risk index. The system sets an upper limit threshold for the output length or data volume, which is negatively correlated with the risk index. When the risk index is high, the system exponentially compresses the length or precision of the output content, for example, outputting only a summary or key metadata, and prohibiting the output of the full text. This mechanism ensures that high-risk archives expose only the minimum necessary information during retrieval, reducing the risk of leakage.
[0200] S807. The open level control strategy, dynamic desensitization strategy, inference jump limit strategy and output constraint strategy are integrated to obtain the dynamic governance strategy corresponding to the target file.
[0201] Among them, fusion processing refers to integrating multiple dimensions of control strategies into a unified set of execution instructions.
[0202] In this embodiment, the system encapsulates and performs conflict detection on the sub-policies generated in the above four steps. The system checks whether there are logical conflicts between the sub-policies (such as a mismatch between the openness level and the desensitization strength). If a conflict exists, adjustments are made according to preset priority rules (such as the security priority principle). Finally, the system generates a complete dynamic governance policy that includes access control, desensitization rules, query restrictions, and output constraints, and distributes it to the execution layer. This policy can comprehensively and multidimensionally manage file risks, achieving intelligent dynamic governance.
[0203] It should be noted that the merged strategy supports visual display, and administrators can view the specific content and scope of the strategy.
[0204] This application embodiment achieves adaptive dynamic adjustment of openness level, desensitization strength, inference hop count, and output constraints by integrating risk level calculation and nonlinear strategy generation method, ensuring the accuracy and timeliness of governance strategy and improving the proactive defense capability of cloud archive management.
Claims
1. A cloud-based intelligent archive management method, characterized in that, include: Acquire archive data, access logs, intelligent agent call data, and archive association data from the cloud archive system, and construct a trusted digital archive based on the archive data; Feature extraction is performed on the trusted digital archive to obtain access anomaly features, semantic offset features, propagation link features, and topological association features; Based on the access anomaly features, the agent call data, the semantic offset features, and the propagation link features, a four-dimensional trust state model is constructed, including a secure and trustworthy state, a knowledge-active state, a semantically stable state, and a propagation and diffusion state. Based on the aforementioned four-dimensional credible state model, time-series evolution calculations are performed to obtain the archive credible state evolution results; Based on the credible state evolution results and the topological association features, the archive association entropy is calculated to obtain the archive implicit inference risk; Based on the evolution results of the archive's credible state and the implicit inference risks of the archive, a dynamic governance strategy is generated; Based on the aforementioned dynamic governance strategy, the target files are subject to openness level control, dynamic desensitization, inference jump limit, and propagation blocking control.
2. The method according to claim 1, characterized in that, The construction of a trusted digital archive based on the archive data includes: Obtain the original content data, structural attribute data, and semantic description data of the target file; The original content data is encoded using a hash mapping method to obtain the file hash fingerprint; The structural attribute data is parsed and processed using a structural feature extraction method to obtain the archive structure fingerprint; The semantic description data is vectorized using a semantic embedding encoding method to obtain the archive semantic fingerprint; A nonlinear fusion mapping method is used to fuse the archive hash fingerprint, the archive structure fingerprint, and the archive semantic fingerprint to obtain the archive fused fingerprint; Based on the aforementioned archive fusion fingerprint, archive dynamic status data, risk measurement data, and governance constraint data, a trusted digital archive entity is constructed.
3. The method according to claim 1, characterized in that, The feature extraction of the trusted digital archive yields access anomaly features, semantic offset features, propagation link features, and topological association features, including: Obtain user access logs, agent call logs, semantic interaction data, and file association data corresponding to the target file; Anomaly detection was performed on the user access logs using a time-series behavior analysis method to obtain access anomaly characteristics. The semantic interaction data is vectorized using a semantic vector encoding method, and offset analysis is performed based on the differences in semantic vector distribution to obtain semantic offset features; The propagation path of the agent's call logs is analyzed using the link tracing method to obtain propagation link characteristics; Based on the archive association data, an archive association topology graph is constructed, and a graph topology analysis method is used to perform node association analysis on the archive association topology graph to obtain topology association features.
4. The method according to claim 3, characterized in that, The semantic interaction data is vectorized using a semantic vector encoding method, and offset analysis is performed based on the differences in semantic vector distribution to obtain semantic offset features, including: Obtain the original semantic data of the target file and the corresponding AI-generated semantic data; A pre-trained semantic coding model is used to vectorize and encode the original semantic data and the artificial intelligence-generated semantic data to obtain the original semantic vector and the generated semantic vector. Based on the original semantic vector and the generated semantic vector, a corresponding semantic probability distribution is constructed; The semantic probability distribution is calculated using the KL divergence method to obtain the semantic distribution offset parameter; The Mahalanobis distance method is used to calculate the spatial offset of the original semantic vector and the generated semantic vector to obtain the semantic spatial offset parameter. Based on the semantic distribution offset parameter and the semantic space offset parameter, the semantic offset feature corresponding to the target file is generated.
5. The method according to claim 1, characterized in that, The four-dimensional trustworthy state model, constructed based on the access anomaly features, the agent call data, the semantic offset features, and the propagation link features, includes a secure and trustworthy state, a knowledge-active state, a semantically stable state, and a propagation and diffusion state. The abnormal access frequency is statistically analyzed based on the access anomaly characteristics, and nonlinear coupling calculation is performed in combination with the call frequency in the agent call data and the sensitivity association strength to obtain the security and trustworthiness state corresponding to the target file; the sensitivity association strength is obtained by sensitive association analysis based on the file association relationship data; Based on the characteristics of the propagation link, an archive association topology network is constructed, and topology diffusion is calculated according to the activity of associated nodes, node association weights and topology propagation distance to obtain the knowledge activity state corresponding to the target archive. The semantic distribution offset parameter and semantic space offset parameter in the semantic offset feature are fused and calculated to obtain the semantic stable state corresponding to the target file; Based on the data invoked by the intelligent agent, a propagation state transition link is constructed, and chain diffusion calculation is performed according to the propagation path length, node transition probability and link attenuation coefficient to obtain the propagation diffusion state corresponding to the target file. The secure and trustworthy state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state are coupled to construct a four-dimensional trustworthy state model corresponding to the target file.
6. The method according to claim 5, characterized in that, The secure and trusted state satisfies the following formula: in, This represents the secure and trustworthy state of the i-th target file at time t; This represents the access anomaly rate of the i-th target file at time t. This represents the weighting coefficient for abnormal access risks; This represents the sensitivity correlation of the i-th target file at time t. Indicates the sensitive correlation coupling coefficient; This represents the agent's call frequency for the i-th target file at time t; This indicates the call to the change disturbance coefficient; This represents a non-linear activation function; i represents the target file index, and t represents the current time step; This represents the intermediate time variable during the integration process. This represents the rate of change in the frequency of agent calls; This represents the secure and trusted state of the i-th target file at the previous time step. The historical state damping attenuation coefficient; The knowledge activity state satisfies the following formula: in, Let j represent the knowledge activity state of target file i at time t, and j represent the adjacent node number associated with the i-th target file. G represents the set of adjacent nodes corresponding to the i-th target file; G represents the topological gravity constant. This represents the association weight between the i-th target file and the j-th adjacent node; This represents the knowledge activity of the j-th adjacent node at time t; This represents the topological propagation distance between the i-th target file and the j-th adjacent node; Indicates the time decay coefficient; The semantically stable state satisfies the following formula: in, This represents the semantically stable state of the i-th target file at time t; Indicates the semantic distribution offset weight coefficient; Represents the original semantic probability distribution With the probability distribution of generated semantics KL divergence between them; This represents the original semantic probability distribution corresponding to the i-th target file; This represents the generated semantic probability distribution corresponding to the i-th target file; This represents the semantic space offset weight coefficient; This represents the semantic vector corresponding to the i-th target file; Represents the mean of a semantic vector; This represents the matrix transpose operation; The inverse matrix representing the semantic vector covariance matrix; The propagation and diffusion state satisfies the following formula: in, The propagation state of the i-th target file at time t is represented by k; the propagation link level number is represented by k; and the propagation link length is represented by L. This represents the state transition probability corresponding to the k-th propagation link; This represents the link attenuation coefficient.
7. The method according to claim 1, characterized in that, The temporal evolution calculation based on the four-dimensional trust state model yields the archive trust state evolution results, including: Obtain the four-dimensional trust state of the target file at the current moment, including: secure trust state, knowledge active state, semantically stable state, and propagation and diffusion state; Construct the state coupling matrix corresponding to the four-dimensional credible state, and determine the state transmission coefficient and state damping coefficient between each credible state; Based on the state coupling matrix, the state transmission coefficient, and the state damping coefficient, the coupled evolution calculation is performed on the secure and trustworthy state, the knowledge-active state, the semantically stable state, and the propagation and diffusion state to obtain the change in the trustworthy state. Based on the aforementioned credible state change and historical credible state data, a time-series iterative calculation is performed to obtain the credible state evolution vector corresponding to the target time. Based on the aforementioned credible state evolution vector, the credible state evolution result corresponding to the target file is generated.
8. The method according to claim 1, characterized in that, The calculation of archive association entropy based on the credible state evolution result and the topological association characteristics yields the implicit inference risk of the archives, including: Obtain the trusted state evolution results, file association topology graph, and node association data corresponding to the target file; The node centrality of the file association topology graph is analyzed using graph topology analysis method to obtain the node degree centrality and node betweenness centrality. Based on the node degree centrality, the node betweenness centrality, and the node association weight, a composite weight for the archive topology is constructed. The weighted entropy value is calculated based on the topological composite weight and the archive association probability distribution to obtain the archive association entropy corresponding to the target archive; the archive association probability distribution is constructed based on the archive node association probability. The mutual information calculation method is used to perform coupling analysis on the inference dependencies between multiple related archives to obtain the archive cluster coupling entropy; Based on the archive association entropy, the archive cluster coupling entropy, and the credible state evolution result, the implicit inference risk of the target archive is generated.
9. The method according to claim 1, characterized in that, The generation of dynamic governance strategies based on the evolution results of the archive's credible state and the implicit inference risks of the archive includes: Obtain the security and trustworthiness state, knowledge activity state, semantic stability state, propagation and diffusion state, and implicit reasoning risks of the target file; The state norm of the security and trust state, the knowledge active state, the semantic stable state, and the propagation and diffusion state are calculated to obtain the comprehensive risk level corresponding to the target file. Based on the comprehensive risk level and the implicit inference risk of the archives, an openness level control strategy corresponding to the target archives is generated using a dynamic adjustment method for openness level. Based on the secure and trustworthy state and the file association entropy, a nonlinear risk coupling calculation method is used to generate a dynamic desensitization strategy for the target file. Based on the propagation and diffusion state, node centrality, and inference link length, a topology depth constraint method is used to generate an inference hop count limit strategy corresponding to the target file. Based on the comprehensive risk level, the file association entropy, and the propagation and diffusion state, the risk index constraint method is used to generate the output constraint strategy corresponding to the target file. The openness level control strategy, the dynamic desensitization strategy, the inference jump limit strategy, and the output constraint strategy are fused together to obtain the dynamic governance strategy corresponding to the target file.
10. A cloud-based intelligent archive management system, characterized in that, The system includes: a data acquisition device and an electronic device; The data acquisition device is used to acquire archive data, access behavior data, intelligent agent call data and archive association data in the cloud archive system, and to construct a trusted digital archive based on the archive data. The electronic device is used to extract features from the trusted digital entity of the archive, obtaining access anomaly features, semantic offset features, propagation link features, and topological association features; based on the access anomaly features, the agent's invoked data, the semantic offset features, and the propagation link features, a four-dimensional trusted state model is constructed, including a secure trusted state, a knowledge active state, a semantically stable state, and a propagation diffusion state; based on the four-dimensional trusted state model, a temporal evolution calculation is performed to obtain the archive trusted state evolution result; based on the trusted state evolution result and the topological association features, the archive association entropy is calculated to obtain the archive implicit inference risk; based on the archive trusted state evolution result and the archive implicit inference risk, a dynamic governance strategy is generated; based on the dynamic governance strategy, openness level control, dynamic desensitization, inference jump limit, and propagation blocking control are implemented on the target archive.