Electronic information security storage system and method based on cloud computing

By reading user behavior logs in the cloud storage system and performing scene division and label processing, analyzing the credibility of user intentions and the similarity of behavior trajectories, and dynamically setting risk levels, the copy redundancy control problem of the existing system in multiple users and multiple scenarios is solved, and efficient security storage and resource optimization are achieved.

CN120705913AActive Publication Date: 2025-09-26BEIJING HENG YUANHUA INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510914973.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-26
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing cloud storage systems lack the ability to semantically determine whether user access operations comply with business objectives, making it difficult to effectively meet the security requirements for electronic information storage. Especially in a multi-user, multi-scenario, and dynamic access environment with sensitive resources, how to control replica redundancy based on user behavior semantics and operation trajectories becomes a key issue.

Method used

By reading user behavior logs on the cloud storage platform, combining the enterprise model to perform scenario division and label processing, analyzing the credibility of user intentions and the similarity of behavior trajectories, dynamically setting trust thresholds, generating risk levels, and reversely inferring the copy settings of data resources, dynamic perception and copy control of user operations can be achieved.

Benefits of technology

It improves the system's ability to perceive the operating environment and access motivations, enhances the recognition rate of hidden abnormal behaviors, realizes closed-loop control of access intention identification, behavior evolution evaluation and copy quantity adjustment, and improves the cloud storage system's security adaptability and resource allocation efficiency under multi-user and multi-scenario conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705913A_ABST
    Figure CN120705913A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic information security storage system and method based on cloud computing, and relates to the technical field of security storage, by introducing a behavior log and enterprise mode conjoint analysis mechanism in the step S1, the system can perform scene division and label classification on user operation behaviors, and comprehensive description of behavior semantics in a multi-dimensional context is realized. In the step S2, by constructing similarity analysis of a service semantic vector and a behavior semantic vector and combining a user behavior trajectory similarity matrix, dynamic judgment of the matching degree between the user access behavior and a service target is achieved, so that an intention judgment result with high credibility is obtained, and the recognition rate of the hidden abnormal behavior is further improved. And S3, performing comparative analysis on the intention credibility under the scene label and a dynamic credibility threshold, generating a sub-scene risk level index, and performing back-stepping on a replica redundancy control strategy in combination with behavior log data, thereby effectively suppressing the problem of redundant replica diffusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of secure storage technology, and in particular to a cloud computing-based electronic information secure storage system and method. Background Art

[0002] With the development of information technology and network communications, cloud computing, as a new generation of information infrastructure, has been widely used in a variety of fields, including big data analysis, remote collaboration, distributed operations and maintenance, and intelligent storage. Electronic information storage technology supported by cloud computing has become an important component of enterprise data management. By migrating massive amounts of structured and unstructured data to cloud platforms for unified management, high availability, high redundancy, and flexible scheduling of data resources are achieved. However, user-driven data copy control and copy redundancy optimization management have gradually become key technical directions that affect information security and storage resource efficiency in cloud environments. In particular, in environments with multiple users, multiple scenarios, and dynamic access to sensitive resources, how to control copy redundancy based on user behavior semantics and operation trajectories has become an important issue in the research of secure electronic information storage.

[0003] In existing cloud storage systems, although access control, permission configuration and data backup mechanisms have been introduced to ensure the security of electronic information, a series of core defects still exist: First, the system generally relies on static access control models and lacks the ability to semantically judge whether user access operations meet business objectives. As a result, the system lacks active perception of potential endogenous threats, making it difficult to effectively meet the security requirements of electronic information storage. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the present invention provides a cloud computing-based electronic information security storage system and method, which solves the problems in the above-mentioned background technology.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for secure storage of electronic information based on cloud computing, comprising the following steps:

[0006] S1: Read each user's behavior logs from the cloud storage platform over a historical period. Based on the enterprise model, divide each user's operation behavior into scenarios to obtain scenario sets. Based on the scenario sets, each user's behavior logs are scenario-labeled to obtain each user's sub-behavior logs under different scenario labels.

[0007] S2: Based on each user's sub-behavior logs under different scenario tags, analyze whether each user's access operations meet the business goals to obtain the intent credibility AiL. By analyzing the similarity of the operation behavior trajectories of different users in the same period of time, the intent credibility AiL is further adjusted;

[0008] S3: Based on S2, the trusted threshold PY under different scenario labels is dynamically set, and after comparison, the risk level RLV under the corresponding scenario label is generated. Combined with the user behavior logs read in the cloud storage platform, the copy settings of each data resource during the next secure storage are reversed.

[0009] Preferably, the specific steps of S1 include:

[0010] S11: Reading the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include the operation timestamp, operation target, device used, operation type, data resource level, number of accesses to each data resource, and operation path;

[0011] S12: The enterprise model includes the working hours, non-working hours, work location, non-work location, authorized equipment, unauthorized equipment, business goals and non-business goals divided for each user's position in the enterprise; by utilizing the enterprise model and combining the user behavior logs read in the cloud storage platform, the operation behavior of each user is divided into scenarios to preliminarily obtain a scene set, and then the scene tags in the scene set are combined to update the preliminarily obtained scene set.

[0012] Preferably, the specific step S1 further includes:

[0013] S13: Based on the updated scenario set in S12, the behavior logs of each user are classified into corresponding scenario tags to obtain sub-behavior logs of each user's behavior under different scenario tags.

[0014] Preferably, the specific steps of S2 include:

[0015] S21: Determine the business process from the business goal in advance, and determine the operation characteristics of the corresponding business process based on the business process. Use natural language processing technology to semantically embed the task name, business process and the operation characteristics of the corresponding business process in the expression of the business goal, so as to map the business goal corresponding to each user into a business semantic vector.

[0016] S22: Use the cloud storage platform to read the behavior logs of each user in real time again, and use the behavior logs of each user read in real time to generate a set of business semantic vectors Corresponding behavioral semantic vector Behavior Semantic Vector Used to reflect the user's current access behavior:

[0017] S23: Based on the real-time reading of each user's behavior log and the updated scene set, the current user's behavior log is classified into the corresponding scene label, and the behavior deviation rate Xbm of the current user's behavior log under the corresponding scene label is analyzed and calculated, specifically: Where n is the characteristic dimension of the current operation behavior, F i is the value of the current operation behavior under the i-th feature dimension, is the average value of previous operation behaviors under the i-th feature dimension.

[0018] Preferably, the specific step S2 further includes:

[0019] S24: Use the cosine similarity algorithm to measure business semantic vectors and behavioral semantic vector The similarity between To analyze whether each user's access operation meets the business objectives, and combined with the behavior deviation rate Xbm of each user's behavior log under the corresponding scenario label calculated in S21, analyze the intention of the current user's operation behavior to calculate the acquisition intention credibility AiL, specifically: Where α is the behavior deviation penalty coefficient.

[0020] Preferably, the specific step S2 further includes:

[0021] S25: Transform all users’ behavior semantic vectors Perform time-sharing window aggregation to construct several groups of aggregated behavior vectors for the corresponding users Aggregate behavior vectors by several groups of corresponding users Analyze the similarity of the behavior trajectories of each user in the same period of time to construct the behavior trajectory similarity matrix D between users. Each element in the behavior trajectory similarity matrix D represents the similarity of the behavior trajectories of each user in the same period of time, specifically: Where, is the similarity of the behavior trajectories between the aggregated behavior vectors of the j-th user and the k-th user in the same time window w, that is, the corresponding element in the behavior trajectory similarity matrix D; and are the aggregated behavior vectors of the jth user and the kth user in the same time window w, M is the number of consensus time windows, m = 1, 2, ..., M, , j and k are user numbers; w is the time window number;

[0022] S26: Based on the behavior trajectory similarity matrix D in S25 and its dynamics, a similarity threshold is dynamically set. If the corresponding element in the behavior trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavior trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set; wherein, the context set includes elements that exceed the corresponding dynamically set similarity threshold, and the element is the similarity of the behavior trajectories between the aggregated behavior vectors of two users. Finally, the two users corresponding to the similarity of each group of behavior trajectories are recorded as a user pair.

[0023] Preferably, the specific step S2 further includes:

[0024] S27: extracting user pairs corresponding to the current user from the context set, and counting the users of all extracted user pairs except the current user to construct a matching group;

[0025] S28: Obtain the current intention credibility AiL of each user in the matching group through the matching group, and obtain the average intention credibility AiL for the matching group through the statistical averaging algorithm avg , combined with the current user's intention credibility AiL in S24, the current user's intention credibility AiL is further adjusted, specifically: AiL'=AiL+η*AiL avg ; Where AiL' is the adjusted intent credibility of the current user, and η is the group reinforcement factor.

[0026] Preferably, the specific steps of S3 include:

[0027] S31: Based on the adjusted intent credibility AiL' of the current user, a trust threshold PY is dynamically set for the corresponding scenario label. The risk level RLV of the corresponding scenario label is calculated by comparing the adjusted intent credibility AiL' of the current user with the trust threshold PY. Specifically, Where Sigmoid(x) is the Sigmoid function,

[0028] Preferably, the specific step S3 further includes:

[0029] S32: Based on the risk level RLV obtained in S31 and in combination with the sub-behavior log of the current user's behavior under the corresponding scenario label, the replica state of the data resource corresponding to the operation behavior performed by the current user is re-corrected to obtain the replica redundancy control factor Rrc, which is specifically:

[0030]

[0031] Where Rrc ris the copy redundancy control factor of the rth data resource, Apa r is the data access pressure perception value of the rth data resource, β is the suppression coefficient, and r is the number of the data resource;

[0032] S33: Map the replica redundancy control factor Rrc to the interval (0, 1) through the Sigmoid function, and determine the number of replicas R of each data resource during the next secure storage, specifically: R r =R0*(1+I(Rrc r ≤0.5)), where R0 is the number of copies of the corresponding data resource during this secure storage, R r is the number of copies of the rth data resource at the next safe storage, I(Rrc r ≤0.5) is the indicator function.

[0033] A cloud computing-based electronic information security storage system includes a partitioning module, an adjustment module, and an update module;

[0034] The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform, and divide the operation behavior of each user into scenarios according to the enterprise model to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed with scenario labels to obtain the sub-behavior logs of each user under different scenario labels;

[0035] The adjustment module is used to analyze whether each user's access operation meets the business objectives based on the sub-behavior logs of each user under different scenario tags to obtain the intent credibility AiL. It also further adjusts the intent credibility AiL by analyzing the similarity of the operation behavior trajectories of different users in the same period of time.

[0036] The update module will dynamically set the trust threshold PY under different scenario labels, and after comparison, generate the risk level RLV under the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will reversely infer the copy settings of each data resource during the next secure storage.

[0037] The present invention provides a cloud computing-based electronic information security storage system and method, which has the following beneficial effects:

[0038] (1) By introducing the joint analysis mechanism of behavior logs and enterprise models in step S1, the system can divide user operation behaviors into scenarios and classify labels, realize the comprehensive characterization of behavior semantics in multi-dimensional contexts, and enhance the system's perception of the operating environment and access motivation. In step S2, by constructing the similarity analysis of business semantic vectors and behavior semantic vectors, combined with the user behavior trajectory similarity matrix, a dynamic judgment of the matching degree between user access behavior and business goals is achieved, thereby obtaining a highly reliable intention judgment result, further improving the recognition rate of hidden abnormal behavior. In addition, step S3 generates a scenario-specific risk level indicator RLV by comparing the intention credibility under the scenario label with the dynamic trust threshold, and combines the behavior log data to reversely infer the copy redundancy control strategy, so that the number of data copies is dynamically linked to the user behavior risk, effectively suppressing the problem of redundant copy proliferation caused by high-risk users. In short, the overall method realizes the closed-loop control of access intention identification, behavior evolution evaluation and copy number adjustment, further improving the security adaptability and resource allocation efficiency of the cloud storage system under multi-user and multi-scenario conditions.

[0039] (2) By using the difference between the user's current operation and the previous feature mean, the user's behavior change range is dynamically described, providing a mathematical basis for identifying atypical operations. Combining the cosine similarity calculation of the behavioral semantics and business semantics in S24, the intention credibility AiL is finally output, thereby achieving the trustworthy modeling of access motivation and the measurement of behavior authenticity. This method breaks through the limitation of traditional access control that can only identify "what was done" and instead determines "why it was done", effectively suppressing hidden operations under non-business goals, enhancing the controllability and auditability of data access in the cloud environment, and providing a more credible behavioral basis for subsequent risk assessment and copy control.

[0040] (3) Through S25, the behavioral semantic vectors of all users are aggregated according to the time window to construct the aggregated behavioral vectors in multiple time periods. This not only effectively smooths the short-term behavioral fluctuations, but also provides a structural basis for the comparison of behavioral trajectories across users and time. Subsequently, the behavioral trajectory similarity matrix D is constructed by cosine similarity calculation. Each element in the matrix quantitatively represents the consistency of the behavioral patterns of any two users in the same window, enabling the system to capture the behavioral collaborative characteristics of user groups in multiple time dimensions. S26 sets a similarity threshold that slides adaptively over time based on the dynamic changes of the matrix, and determines whether to include user pairs in the context set based on this. Through this mechanism, the system can automatically identify user clusters with similar behaviors and construct a high-credibility behavioral context association map, effectively enhancing the group reference support capability of individual intention credibility. Overall, this step improves the accuracy and robustness of behavioral judgment, provides a scalable and interpretable group semantic basis for the subsequent adjustment of intent credibility and the determination of risk level, and further strengthens the system's collaborative judgment capability and dynamic evolution detection capability.

[0041] (4) S33 determines the number of copies of each data resource in each secure storage, so that the copy generation behavior is no longer a static redundancy strategy, but a security decision result that is dynamically adjusted according to the behavior risk. This method further improves the system's ability to suppress copies of high-risk access behaviors, reduces the probability of redundant diffusion of sensitive data under abnormal operations, and ensures the availability and disaster recovery capabilities of data in low-risk scenarios, ultimately achieving dual optimization of data security and resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of a method for securely storing electronic information based on cloud computing according to the present invention;

[0043] Figure 2 This is an overall logic diagram of a cloud computing-based electronic information security storage method of the present invention;

[0044] Figure 3 This is a partial logic diagram of a cloud computing-based electronic information security storage method of the present invention;

[0045] Figure 4 This is a block diagram of an electronic information security storage system based on cloud computing of the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0047] Example 1

[0048] See also Figures 1 to 3 The present invention provides a method for securely storing electronic information based on cloud computing, comprising the following steps:

[0049] S1: Read each user's behavior logs from the cloud storage platform over a historical period. Based on the enterprise model, divide each user's operation behavior into scenarios to obtain scenario sets. Based on the scenario sets, each user's behavior logs are scenario-labeled to obtain each user's sub-behavior logs under different scenario labels.

[0050] S2: Based on each user's sub-behavior logs under different scenario tags, analyze whether each user's access operations meet business goals (such as approval query, batch export, routine browsing, etc.) to obtain the intent credibility AiL. By analyzing the similarity of the operation behavior trajectories of different users in the same period of time, the intent credibility AiL is further adjusted;

[0051] S3: Based on S2, dynamically set the trust threshold PY for different scenario tags. After comparison, generate the risk level RLV for the corresponding scenario tag. Combined with the user behavior logs read from the cloud storage platform, reversely infer the replica setting of each data resource for the next secure storage. The data resource replica setting is performed once each time electronic information is stored.

[0052] In this embodiment, by performing scenario-based labeling on user operations, refined behavioral semantic recognition can be achieved, ensuring that the system can identify security risks with different risk levels for the same behavior in different scenarios. At the same time, the linkage mechanism between the construction of intention credibility AIL and the similarity analysis of group behavior trajectories can dynamically perceive the evolution trend of user behavior and avoid misjudgments or missed judgments caused by a single behavioral indicator.

[0053] Ultimately, by linking the dynamically generated risk level RLV with the trusted threshold PY, the replica redundancy strategy is controlled. This suppresses the replica diffusion problem caused by potential high-risk behaviors without affecting data availability, thereby achieving a risk-driven closed-loop data replica control. For example, if an employee queries an approval document through an authorized device during working hours, the system will mark this as normal behavior and allow the generation of three replicas for high availability. However, if the employee attempts to perform a batch export operation through an external device during non-working hours, although the behavior type is the same, the system's recognition of the credibility of the intention is reduced due to changes in the scenario label and increased behavior deviation rate. Combined with group behavior comparison, a high risk level is dynamically generated to prevent data from being spread to multiple points under potential threats. This method enhances the system's adaptability to dynamic access behavior and its ability to fine-tune replica control.

[0054] Example 2

[0055] Please refer to Figure 1 , specifically: S1 specific steps include:

[0056] S11: Read the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include the operation timestamp, operation target (such as file ID, resource ID, classification label, etc.), the device used, the operation type, the data resource level, the number of accesses to each data resource, and the operation path (such as click trajectory and page jump path, which are used to reflect the complexity of the operation);

[0057] S12: The enterprise model includes the working hours, non-working hours, work location, non-work location, authorized equipment, unauthorized equipment, business goals and non-business goals divided for each user's position in the enterprise; by utilizing the enterprise model and combining the user behavior logs read in the cloud storage platform, the operation behavior of each user is divided into scenarios to preliminarily obtain a scene set, and then the scene tags in the scene set are combined to update the preliminarily obtained scene set.

[0058] The specific steps of S1 also include:

[0059] S13: Based on the updated scenario set in S12, the behavior logs of each user are classified into corresponding scenario tags to obtain sub-behavior logs of each user's behavior under different scenario tags.

[0060] It should be noted that scenario labels include but are not limited to behavioral operations performed during working hours and on authorized devices, behavioral operations performed during non-working hours and on authorized devices, behavioral operations performed during working hours and on unauthorized devices, behavioral operations performed during working hours that are subject to business objectives, and behavioral operations performed during non-working hours that are subject to business objectives, etc.

[0061] In this embodiment, the present invention achieves high-precision scenario recognition and classification of user operation behaviors by introducing multi-dimensional behavior logs and enterprise model information in step S1, further improving the contextual accuracy of subsequent behavior intention judgment and risk control.

[0062] By extracting key fields such as operation timestamp, operation path, device type, and resource access frequency through S11, and combining them with the company's institutional norms for work hours, authorized devices, and job objectives in S12, the system can differentiate the same behavior across different operating environments. Furthermore, through S13's scenario label update mechanism, the system accurately extracts sub-behavior logs from user behaviors in different labeled scenarios, forming an input dataset that meets the requirements of semantic recognition. For example, a user's behavior of "accessing an approval form from a company-authorized PC at 9:30 a.m. on a weekday" would be classified under the "work hours + authorized device + business objectives" scenario, while "downloading a batch of reports from a home device at night" would be classified under the "non-work hours + unauthorized device + high-risk export" scenario. This classification not only avoids misjudgment of similar operations but also provides a more objective, hierarchical, and interpretable basis for subsequent intent credibility identification and copy control policies. This approach effectively improves the system's understanding of access context, enhances its ability to detect and respond to anomalous behavior early, and strengthens the dynamic security adaptability of cloud storage systems.

[0063] Example 3

[0064] Please refer to Figure 1 , specifically: S2 specific steps include:

[0065] S21: Predetermine the business process from the business goal, and determine the operational characteristics of the corresponding business process based on the business process (such as querying approval forms, batch exporting reports, etc.). Use natural language processing technology to semantically embed the task name, business process and the operational characteristics of the corresponding business process in the expression of the business goal (such as word embedding methods such as Word2Vec and GloVe) to map the business goal corresponding to each user into a business semantic vector

[0066] S22: Use the cloud storage platform to read the behavior logs of each user in real time again, and use the behavior logs of each user read in real time to generate a set of business semantic vectors Corresponding behavioral semantic vector Behavior Semantic Vector Used to reflect the user's current access behavior:

[0067] S23: Based on the real-time reading of each user's behavior log and the updated scenario set, the current user's behavior log is classified into the corresponding scenario label, and the behavior deviation rate Xbm of the current user's behavior log under the corresponding scenario label is analyzed and calculated. The behavior deviation rate Xbm is used to reflect the deviation of the current user's operation behavior from the previous one, specifically:

[0068]

[0069] Where n is the characteristic dimension of the current operation behavior, that is, the behavior semantic vector The feature dimension, F i is the value of the current operation behavior under the i-th characteristic dimension (such as the number of visits, operation type, resource level, etc.), is the average value of previous operation behaviors under the i-th feature dimension;

[0070] Business semantic vector Applied to the judgment of the current operation to help identify the actual target of the access operation;

[0071] Business processes usually involve multiple operational steps, such as approval, query, export, etc. Among them, the task scheduling system may have multiple steps, each of which triggers different operations and data access.

[0072] The specific steps of S2 also include:

[0073] S24: Use the cosine similarity algorithm to measure business semantic vectors and behavioral semantic vector The similarity between To analyze whether each user's access operation meets the business objectives (such as approval query, batch export, routine browsing), and combined with the behavior deviation rate Xbm of each user's behavior log under the corresponding scenario label calculated in S21, analyze the intention of the current user's operation behavior to calculate the acquisition intention credibility AiL, specifically: Where α is the behavior deviation penalty coefficient, which is used to amplify or compress the impact of the deviation on credibility, and its range is between 0 and 1;

[0074] In this embodiment, the present invention introduces an access intention recognition mechanism based on semantic modeling in step S2, which further improves the cloud platform's ability to accurately judge the compliance and potential risks of user operations.

[0075] By semantically embedding business objectives, processes, and their operational characteristics in S21, a structured business semantic vector is constructed, enabling the system to accurately capture the target semantic intent of the task. Combined with the user behavior logs extracted in real time in S22 to generate behavioral semantic vectors, the system achieves semantic-level abstraction of the user's current operation, thereby calculating the match between the behavior and the business objective through cosine similarity in S24. Furthermore, combined with the behavior deviation rate Xbm calculated in S23, an intent credibility AiL model is constructed that comprehensively considers semantic consistency and behavioral evolution deviation, effectively avoiding the problem of static permission models being slow to respond to dynamic behavioral changes.

[0076] Example 4

[0077] Please refer to Figure 1 and Figure 3 Specifically, the steps of S2 further include:

[0078] S25: Transform all users’ behavior semantic vectors Perform time-sharing window aggregation to construct several groups of aggregated behavior vectors for the corresponding users Aggregate behavior vectors by several groups of corresponding users Analyze the similarity of the behavior trajectories of each user in the same period of time to construct the behavior trajectory similarity matrix D between users. Each element in the behavior trajectory similarity matrix D represents the similarity of the behavior trajectories of each user in the same period of time, specifically:

[0079]

[0080] Where, is the similarity of the behavior trajectories between the aggregated behavior vectors of the j-th user and the k-th user in the same time window w, that is, the corresponding element in the behavior trajectory similarity matrix D; and are the aggregated behavior vectors of the jth user and the kth user in the same time window w, M is the number of consensus time windows, m = 1, 2, ..., M, , j and k are user numbers; w is the time window number;

[0081] Specifically, time-sharing window aggregation divides user behavior data into fixed-size time periods, such as hours, days, or weeks, and then aggregates the user's behavior vectors within each time period.

[0082] The purpose of the aggregation operation is to compress multiple behavioral data into an aggregated behavior vector, which represents the overall behavioral characteristics of the user within the time window. After aggregation, the aggregated behavior vector of each time period (such as hourly or daily) will represent the average characteristics of user behavior within that period. This aggregation helps smooth short-term fluctuations into longer-term trends and avoid the impact of occasional behavioral changes on the overall analysis.

[0083] The construction of a behavioral trajectory similarity matrix D helps identify "unusually consistent" behavior patterns: if multiple different accounts repeatedly access the same resources, using similar operations, and in similar time periods; even if each account individually does not pose a threat, their overall behavior shows a coordinated trend. The similarity matrix can reveal this hidden group risk.

[0084] S26: Based on the behavior trajectory similarity matrix D in S25 and its dynamics, a similarity threshold is dynamically set. If the corresponding element in the behavior trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavior trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set; wherein, the context set includes elements that exceed the corresponding dynamically set similarity threshold, and the element is the similarity of the behavior trajectories between the aggregated behavior vectors of two users. Finally, the two users corresponding to the similarity of each group of behavior trajectories are recorded as a user pair.

[0085] In this embodiment, the present invention further introduces the time window aggregation and trajectory similarity modeling mechanism of user behavior semantics in step S2, which effectively improves the recognition ability of user behavior evolution trend and the accuracy of context association analysis.

[0086] Through S25, the behavioral semantic vectors of each user within multiple consensus time windows are aggregated in a time-sharing manner to construct a behavioral trajectory sequence. Based on this, a behavioral trajectory similarity matrix D is generated between users, accurately depicting the behavioral similarities within a specific time period. In S26, the system sets an adaptive similarity threshold based on the dynamic distribution of matrix D, automatically screening out user pairs with significantly consistent behavioral characteristics to form a context set. This approach enhances behavioral credibility based on "behavioral similarity contexts" and further optimizes the stability and robustness of individual behavioral credibility assessments.

[0087] For example, if user A and user B performed approval process operations every day from 9:00 to 10:00 a.m. over the past week, and their access frequency, resource level, and operation path similarity remained above 0.92 for a long time, and in the behavior trajectory similarity matrix D, they remained above the dynamic threshold of 0.85, the system would label them as a "user pair" and include them in the context set. When user A's identical behavior deviated slightly on a certain day, the system could use user B's behavior as a reference to correct its credibility and avoid misjudging it as an anomaly. This mechanism not only enhances the context-awareness of user behavior recognition, but also introduces group semantic consensus when user behavior drifts slightly, achieving dynamic and stable correction of credibility, significantly improving the overall system's behavior understanding accuracy and risk perception capabilities in complex access environments.

[0088] Example 5

[0089] Please refer to Figure 1 Specifically, the steps of S2 further include:

[0090] S27: extracting a user pair corresponding to the current user from the context set, and counting the users excluding the current user from all the extracted user pairs to construct a matching group, where the matching group refers to the user excluding the current user from all the extracted user pairs;

[0091] S28: Obtain the current intention credibility AiL of each user in the matching group through the matching group, and obtain the average intention credibility AiL for the matching group through the statistical averaging algorithm avg , combined with the current user's intention credibility AiL in S24, the current user's intention credibility AiL is further adjusted, specifically:

[0092] AiL'=AiL+η*AiL avg ;

[0093] Where AiL' is the adjusted intent credibility of the current user, and η is the group reinforcement factor, which is a tuning parameter used to strengthen the influence of group similarity when calculating user intent credibility, especially when multiple user behaviors similar to the target user are taken into consideration. Its range is between 0 and 1.

[0094] By updating and calculating the current user's intention credibility AiL, it can be reflected that: if the user belongs to the "common purpose of the group" (such as normal template import operation), it will not be easily misjudged even if the deviation rate is high. If the similar group has malicious purposes, its behavioral intention will be reversely suppressed.

[0095] In this embodiment, a dynamic correction mechanism for intent credibility based on group behavior similarity is introduced in S2, significantly enhancing the system's accuracy in determining user behavior credibility and the stability of identifying abnormal behavior. In S27, user pairs with highly similar behavior trajectories to the current user are extracted from the context set, and the current user is removed to construct a matching group, ensuring that the groups participating in the credibility reference have sufficient behavioral similarity and contextual consistency.

[0096] S28 uses the statistical mean of the current intention credibility of each user in the matching group, combined with the original credibility of the current user, and introduces a group reinforcement factor η for weighted correction to form a credibility output that is representative of the group. This method effectively integrates individual behavioral judgment and group consensus trends, retaining the user's own behavioral characteristics while using group convergence behavior to enhance the sensitive recognition ability of non-explicit deviations.

[0097] Especially in complex access scenarios such as ambiguous behavior or overlapping semantics, this can further reduce the risk of false positives and missed positives, improving the system's ability to discern and accurately respond to gray operations. Overall, this step establishes a credibility evolution mechanism driven by group similarity reinforcement, providing a more stable and reliable behavioral intent foundation for subsequent risk level calculation and replica control, and enhancing the cloud platform's security decision-making capabilities in complex user behavior scenarios.

[0098] Example 6

[0099] Please refer to Figure 1, specifically: S3 specific steps include:

[0100] S31: Based on the adjusted intent credibility AiL' of the current user, a trust threshold PY is dynamically set for the corresponding scenario label. The risk level RLV of the corresponding scenario label is calculated by comparing the adjusted intent credibility AiL' of the current user with the trust threshold PY. Specifically,

[0101]

[0102] Where Sigmoid(x) is the Sigmoid function, Among them, e is the Euler number, which is about 2.71828, and x here refers to ∈ is a small positive number to prevent division by zero.

[0103] Among them, the Sigmoid function has the characteristics that positive output tends to 1 and negative output tends to 0. When x in Sigmoid(x) is equal to 0, the output value of Sigmoid(x) is 0.5, which means neutral risk. When x is less than 0, the output value of Sigmoid(x) ranges from 0 to 0.5, which means low risk.

[0104] If AiL' is less than PY, then RLV is less than 0, indicating that the behavior is within the acceptable range and is low risk; if AiL' is equal to PY, then RLV is equal to 0, indicating that the behavior is critical and is neutral risk; if AiL' is greater than PY, then RLV is greater than 0, indicating that the behavior has exceeded the tolerance range and is high risk;

[0105] The specific method of dynamically setting the credibility threshold PY is as follows: according to the previously adjusted intent credibility AiL' of the user, the adjusted average intent credibility AiL' is calculated. avg and its standard deviation AiL' σ , based on the adjusted average intention credibility AiL' avg and its standard deviation AiL' σ , according to the change of the current user's adjusted intention credibility AiL', dynamically set the credibility threshold PY: PY = AiL' avg +Q*AiL' σ ; Where Q is a constant, usually ranging from 1 to 3, corresponding to different confidence levels, and the specific value is set by the user (according to actual conditions);

[0106] The specific steps of S3 also include:

[0107] S32: Based on the risk level RLV obtained in S31 and in combination with the sub-behavior log of the current user's behavior under the corresponding scenario label, the replica state of the data resource corresponding to the operation behavior performed by the current user is re-corrected to obtain the replica redundancy control factor Rrc, which is specifically:

[0108]

[0109] Where Rrc r is the copy redundancy control factor of the rth data resource, Apa r is the data access pressure perception value of the rth data resource, β is the suppression coefficient, which is an adjustment coefficient ranging from 0 to 1, and r is the number of the data resource; wherein, the data access pressure perception value Apa of the rth data resource r It refers to the result of normalizing the number of accesses to the rth data resource and mapping its value to the interval [0,1], which is used to reflect the availability of information.

[0110] S33: Map the replica redundancy control factor Rrc to the interval (0, 1) through the Sigmoid function, and determine the number of replicas R of each data resource during the next secure storage, specifically: R r =R0*(1+I(Rrc r ≤0.5)), where R0 is the number of copies of the corresponding data resource during this secure storage, R r is the number of copies of the rth data resource at the next safe storage, I(Rrc r ≤0.5) is the indicator function, when Rrc r When ≤0.5, the indicator function outputs -Rrc r , if Rrc r When >0.5, the indicator function outputs Rrc r .

[0111] In cloud computing environments, redundant copies of electronic information are primarily used to enhance data availability, support disaster recovery, and improve access performance. Traditional replication strategies are typically unified (e.g., automatically creating three copies of each piece of data, distributed across three storage zones by default). However, these strategies suffer from the following issues: storage resources are wasted on "unimportant or high-risk data"; and high-risk information is replicated multiple times, increasing exposure and the risk of leakage. Therefore, through the reverse inference method in this paper, the risk level of user behavior can be used to determine the number of redundant copies of the information allowed. Specifically, if the risk level is high, multiple copies are not suitable, and the number of copies should be limited. Conversely, if the risk level is low, additional copies can be appropriately added to improve access efficiency and availability.

[0112] In this embodiment, a risk assessment and replica control mechanism driven by intent credibility is introduced into S3, which realizes dynamic security control of the data replica generation process in the cloud environment.

[0113] Through S31, the intention credibility of the scenario corresponding to the user's current behavior is processed by Sigmoid function, and combined with the dynamically set credibility threshold to generate the risk level indicator RLV. The system can accurately measure the risk level of the user's current operation; S32 further combines the risk level with the user's historical behavior log to calculate the copy redundancy control factor Rrc of the data resource, so that the copy generation strategy can converge dynamically according to the risk.

[0114] S33 achieves precise control over the number of copies by mapping the replica redundancy control factor Rrc value to the range of 0 to 1 and matching the replica setting rules, effectively preventing sensitive data from being over-copied under high-risk behaviors. For example, if a user frequently exports sensitive documents in the "non-working hours + unauthorized device" scenario, the system identifies that the credibility of their behavioral intention is low, generates a risk level RLV, and calculates a smaller Rrc value (such as 0.2) through the suppression function. Based on this, the system limits the number of copies to 0, blocking the data replication link under this behavior. This method effectively prevents the problem of data redundancy and diffusion caused by abnormal abuse within the authority boundary, enhances the system's dynamic balance between storage security and resource optimization, and has intelligent features such as self-adaptation, scenario perception, and user behavior semantic drive.

[0115] Example 7

[0116] Please refer to Figure 4 ,Specifically: A cloud computing-based electronic information security storage system, comprising a partitioning module, an adjustment module and an update module;

[0117] The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform, and divide the operation behavior of each user into scenarios according to the enterprise model to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed with scenario labels to obtain the sub-behavior logs of each user under different scenario labels;

[0118] The adjustment module is used to analyze whether each user's access operation meets the business objectives based on the sub-behavior logs of each user under different scenario tags to obtain the intent credibility AiL. It also further adjusts the intent credibility AiL by analyzing the similarity of the operation behavior trajectories of different users in the same period of time.

[0119] The update module will dynamically set the trust threshold PY under different scenario labels, and after comparison, generate the risk level RLV under the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will reversely infer the copy settings of each data resource during the next secure storage.

[0120] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for secure storage of electronic information based on cloud computing, characterized by: The following steps are included: S1: Read each user's behavior logs from the cloud storage platform over a historical period. Based on the enterprise model, divide each user's operation behavior into scenarios to obtain scenario sets. Based on the scenario sets, each user's behavior logs are scenario-labeled to obtain each user's sub-behavior logs under different scenario labels. S2: Based on each user's sub-behavior logs under different scenario tags, analyze whether each user's access operations meet the business goals to obtain the intent credibility AiL. By analyzing the similarity of the operation behavior trajectories of different users in the same period of time, the intent credibility AiL is further adjusted; S3: Based on S2, the trusted threshold PY under different scenario labels is dynamically set, and after comparison, the risk level RLV under the corresponding scenario label is generated. Combined with the user behavior logs read in the cloud storage platform, the copy settings of each data resource during the next secure storage are reversed.

2. The method for secure electronic information storage based on cloud computing according to claim 1, characterized in that: The specific steps of S1 include: S11: Reading the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include the operation timestamp, operation target, device used, operation type, data resource level, number of accesses to each data resource, and operation path; S12: The enterprise model includes the working hours, non-working hours, work location, non-work location, authorized equipment, unauthorized equipment, business goals and non-business goals divided for each user's position in the enterprise; by utilizing the enterprise model and combining the user behavior logs read in the cloud storage platform, the operation behavior of each user is divided into scenarios to preliminarily obtain a scene set, and then the scene tags in the scene set are combined to update the preliminarily obtained scene set.

3. The method for secure electronic information storage based on cloud computing according to claim 2, characterized in that: The specific steps of S1 also include: S13: Based on the updated scenario set in S12, the behavior logs of each user are classified into corresponding scenario tags to obtain sub-behavior logs of each user's behavior under different scenario tags.

4. The method for secure electronic information storage based on cloud computing according to claim 3, characterized in that: The specific steps of S2 include: S21: Determine the business process from the business goal in advance, and determine the operation characteristics of the corresponding business process based on the business process. Use natural language processing technology to semantically embed the task name, business process and the operation characteristics of the corresponding business process in the expression of the business goal, so as to map the business goal corresponding to each user into a business semantic vector. S22: Use the cloud storage platform to read the behavior logs of each user in real time again, and use the behavior logs of each user read in real time to generate a set of business semantic vectors Corresponding behavioral semantic vector Behavior Semantic Vector Used to reflect the user's current access behavior: S23: Based on the real-time reading of each user's behavior log and the updated scene set, the current user's behavior log is classified into the corresponding scene label, and the behavior deviation rate Xbm of the current user's behavior log under the corresponding scene label is analyzed and calculated, specifically: Where n is the characteristic dimension of the current operation behavior, F i is the value of the current operation behavior under the i-th feature dimension, is the average value of previous operation behaviors under the i-th feature dimension.

5. The method for secure electronic information storage based on cloud computing according to claim 4, characterized in that: The specific steps of S2 also include: S24: Use the cosine similarity algorithm to measure business semantic vectors and behavioral semantic vector The similarity between To analyze whether each user's access operation meets the business objectives, and combined with the behavior deviation rate Xbm of each user's behavior log under the corresponding scenario label calculated in S21, analyze the intention of the current user's operation behavior to calculate the acquisition intention credibility AiL, specifically: Where α is the behavior deviation penalty coefficient.

6. The method for secure electronic information storage based on cloud computing according to claim 5, characterized in that: The specific steps of S2 also include: S25: Transform all users’ behavior semantic vectors Perform time-sharing window aggregation to construct several groups of aggregated behavior vectors for the corresponding users Aggregate behavior vectors by several groups of corresponding users Analyze the similarity of the behavior trajectories of each user in the same period of time to construct the behavior trajectory similarity matrix D between users. Each element in the behavior trajectory similarity matrix D represents the similarity of the behavior trajectories of each user in the same period of time, specifically: Where, is the similarity of the behavior trajectories between the aggregated behavior vectors of the j-th user and the k-th user in the same time window w, that is, the corresponding element in the behavior trajectory similarity matrix D; and are the aggregated behavior vectors of the jth user and the kth user in the same time window w, M is the number of consensus time windows, m = 1, 2, ..., M, , j and k are user numbers; w is the time window number; S26: Based on the behavior trajectory similarity matrix D in S25 and its dynamics, a similarity threshold is dynamically set. If the corresponding element in the behavior trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavior trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set; wherein, the context set includes elements that exceed the corresponding dynamically set similarity threshold, and the element is the similarity of the behavior trajectories between the aggregated behavior vectors of two users. Finally, the two users corresponding to the similarity of each group of behavior trajectories are recorded as a user pair.

7. The method for secure electronic information storage based on cloud computing according to claim 6, characterized in that: The specific steps of S2 also include: S27: extracting user pairs corresponding to the current user from the context set, and counting the users of all extracted user pairs except the current user to construct a matching group; S28: Obtain the current intention credibility AiL of each user in the matching group through the matching group, and obtain the average intention credibility AiL for the matching group through the statistical averaging algorithm avg , combined with the current user's intention credibility AiL in S24, the current user's intention credibility AiL is further adjusted, specifically: AiL'=AiL+η*AiL avg ; Where AiL' is the adjusted intent credibility of the current user, and η is the group reinforcement factor.

8. The method for secure electronic information storage based on cloud computing according to claim 7, characterized in that: The specific steps of S3 include: S31: Based on the adjusted intent credibility AiL' of the current user, a trust threshold PY is dynamically set for the corresponding scenario label. The risk level RLV of the corresponding scenario label is calculated by comparing the adjusted intent credibility AiL' of the current user with the trust threshold PY. Specifically, Where Sigmoid(x) is the Sigmoid function, 9. The method for secure electronic information storage based on cloud computing according to claim 8, characterized in that: The specific steps of S3 also include: S32: Based on the risk level RLV obtained in S31 and in combination with the sub-behavior log of the current user's behavior under the corresponding scenario label, the replica state of the data resource corresponding to the operation behavior performed by the current user is re-corrected to obtain the replica redundancy control factor Rrc, which is specifically: Where Rrc r is the copy redundancy control factor of the rth data resource, Apa r is the data access pressure perception value of the rth data resource, β is the suppression coefficient, and r is the number of the data resource; S33: Map the replica redundancy control factor Rrc to the interval (0, 1) through the Sigmoid function, and determine the number of replicas R of each data resource during the next secure storage, specifically: R r =R0*(1+I(Rrc r ≤0.5)), where R0 is the number of copies of the corresponding data resource during this secure storage, R r is the number of copies of the rth data resource at the next safe storage, I(Rrc r ≤0.5) is the indicator function.

10. A cloud computing-based electronic information security storage system, used to implement the cloud computing-based electronic information security storage method according to any one of claims 1 to 9, characterized in that: Including division module, adjustment module and update module; The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform, and divide the operation behavior of each user into scenarios according to the enterprise model to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed with scenario labels to obtain the sub-behavior logs of each user under different scenario labels; The adjustment module is used to analyze whether each user's access operation meets the business objectives based on the sub-behavior logs of each user under different scenario tags to obtain the intent credibility AiL. It also further adjusts the intent credibility AiL by analyzing the similarity of the operation behavior trajectories of different users in the same period of time. The update module will dynamically set the trust threshold PY under different scenario labels, and after comparison, generate the risk level RLV under the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will reversely infer the copy settings of each data resource during the next secure storage.

Citation Information

Patent Citations

  • Multi-movable-duplicate mechanism applied to distributed storage and access method thereof

    CN102752381A

  • Data sensitive behavior identification method

    CN109977222A

  • Edge collaborative replica placement method

    CN115904731A

  • Government affair data sharing system based on data security law risk control mode

    CN119989417A

  • Cloud storage dynamic optimization method and system

    CN120186155A