A cloud computing-based electronic information security storage system and method

By reading user behavior logs in the cloud storage system and performing scenario segmentation and intent credibility analysis, and dynamically setting risk levels, the problem of copy redundancy control in existing systems under multiple users and scenarios is solved, and the adaptive capability and resource optimization of secure storage are realized.

CN120705913BActive Publication Date: 2026-03-13BEIJING HENG YUANHUA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing cloud storage systems lack the ability to semantically judge whether user access operations conform to business objectives, making it difficult to effectively meet the security storage requirements of electronic information. In particular, in environments with multiple users, multiple scenarios, and dynamic access to sensitive resources, how to perform copy redundancy control based on user behavior semantics and operation trajectories has become a key issue.

Method used

By reading user behavior logs from the cloud storage platform, combining them with enterprise models to segment and tag scenarios, analyzing the credibility of user intent, dynamically setting credibility thresholds, generating risk levels, and inferring the replication settings of data resources, dynamic redundancy control is achieved.

Benefits of technology

It enhances the system's ability to perceive the operating environment and access motivations, improves the recognition rate of hidden abnormal behaviors, realizes closed-loop control of access intent recognition, behavior evolution assessment and replica quantity adjustment, and improves security adaptability and resource allocation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705913B_ABST
    Figure CN120705913B_ABST
Patent Text Reader

Abstract

This invention discloses a cloud computing-based electronic information security storage system and method, relating to the field of secure storage technology. By introducing a joint analysis mechanism of behavior logs and enterprise patterns in step S1, the system can classify and categorize user operation behaviors into scenarios, achieving a comprehensive characterization of behavioral semantics in a multi-dimensional context. In step S2, by constructing a similarity analysis between business semantic vectors and behavioral semantic vectors, combined with a user behavior trajectory similarity matrix, the system dynamically judges the matching degree between user access behavior and business objectives, thereby obtaining highly reliable intent determination results and further improving the recognition rate of hidden abnormal behaviors. Step S3 compares and analyzes the intent credibility under scenario tags with dynamic credibility thresholds to generate scenario-specific risk level indicators, and combines behavior log data to infer a replica redundancy control strategy, effectively suppressing the problem of redundant replica propagation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of secure storage technology, specifically to a cloud computing-based secure electronic information storage system and method. Background Technology

[0002] With the development of information technology and network communication, cloud computing, as a new generation of information infrastructure, has been widely applied in various fields such as big data analysis, remote collaboration, distributed operation and maintenance, and intelligent storage. Among these, electronic information storage technology supported by cloud computing has become an important component of enterprise data management. By migrating massive amounts of structured and unstructured data to the cloud platform for unified management, it achieves high availability, high redundancy, and flexible scheduling of data resources. However, user behavior-driven data replica control and replica redundancy optimization management are gradually becoming key technical directions affecting information security and storage resource efficiency in the cloud environment. Especially in the face of multi-user, multi-scenario, and dynamic access environments for sensitive resources, how to perform replica redundancy control based on user behavior semantics and operation trajectories has become an important issue in electronic information security storage research.

[0003] In existing cloud storage systems, although access control, permission configuration, and data backup mechanisms have been introduced to ensure electronic information security, a series of core defects still exist: First, the systems generally rely on static access control models and lack the ability to semantically judge whether user access operations conform to business objectives, resulting in the system lacking proactive awareness of potential endogenous threats. Therefore, it is difficult to effectively meet the security storage requirements of electronic information. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a cloud computing-based electronic information security storage system and method, which solves the problems mentioned in the background.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a cloud computing-based method for secure electronic information storage, comprising the following steps:

[0006] S1: Read the behavior logs of each user in the historical period on the cloud storage platform, combine them with the enterprise model, divide the operation behavior of each user into scenarios to obtain scenario sets, and process the behavior logs of each user into scenario tags according to the scenario sets to obtain the sub-behavior logs of each user under different scenario tags.

[0007] S2: Based on the sub-behavior logs of each user under different scenario tags, analyze whether each user's access operations meet the business objectives in order to obtain the intent credibility AiL. By analyzing the similarity of the operation behavior trajectories between users in the same period of time, further adjust the intent credibility AiL.

[0008] S3: Based on S2, dynamically set the trust threshold PY for different scenario labels, and generate the risk level RLV for the corresponding scenario label after comparison. Combined with the user behavior logs read in the cloud storage platform, the replica settings of each data resource are deduced for the next secure storage.

[0009] Preferably, step S1 specifically includes:

[0010] S11: Read the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include operation timestamp, operation target, device used, operation type, data resource level, access frequency of each data resource and operation path.

[0011] S12: The enterprise model includes the division of working hours, non-working hours, working locations, non-working locations, authorized devices, unauthorized devices, business objectives, and non-business objectives for each user's position within the enterprise. By utilizing the enterprise model and combining it with the user behavior logs read from the cloud storage platform, the user's operation behavior is divided into scenarios to initially obtain a scenario set. Then, the scenario tags within the scenario set are combined to update the initially obtained scenario set.

[0012] Preferably, step S1 further includes:

[0013] S13: Based on the updated scenario set in S12, classify each user's behavior logs into the corresponding scenario tags to obtain the sub-behavior logs of each user's behavior under different scenario tags.

[0014] Preferably, step S2 specifically includes:

[0015] S21: Pre-determine business processes from business objectives, and determine the operational features of corresponding business processes based on these processes. Using natural language processing technology, semantically embed the task names, business processes, and corresponding operational features in the description of business objectives to map the business objectives of each user into business semantic vectors.

[0016] S22: Again, use the cloud storage platform to read each user's behavior logs in real time, and use the real-time read user behavior logs to generate a set of business semantic vectors. Corresponding behavioral semantic vectors Behavioral semantic vectors Used to reflect the user's current access behavior:

[0017] S23: Based on the real-time read user behavior logs and combined with the updated scenario set, classify the current user behavior logs into the corresponding scenario tags, and analyze and calculate the behavior offset rate Xbm of each user's behavior logs under the corresponding scenario tags, specifically: In the formula, n is the feature dimension of the current operation behavior, and F i This represents the value of the current operation in the i-th feature dimension. This represents the average value of past operational behaviors in the i-th feature dimension.

[0018] Preferably, step S2 further includes:

[0019] S24: Use cosine similarity algorithm to measure business semantic vectors With behavioral semantic vectors similarity between To analyze whether each user's access operations align with business objectives, and combining this with the behavior offset rate Xbm of each user's current behavior logs under the corresponding scenario label calculated in S21, the intent of the current user's operation behavior is analyzed to calculate and obtain the intent credibility AiL, specifically: In the formula, α is the behavioral deviation penalty coefficient.

[0020] Preferably, step S2 further includes:

[0021] S25: All user behavior semantic vectors Perform time-sharing window aggregation to construct several sets of aggregated behavior vectors for the corresponding users. Through several sets of aggregated behavior vectors of the corresponding users The similarity of user behavior trajectories within the same time period is analyzed to construct a user behavior trajectory similarity matrix D. Each element in the behavior trajectory similarity matrix D represents the similarity of user behavior trajectories within the same time period, specifically: In the formula, The similarity of the behavioral trajectories between the behavior vectors of user j and user k within the same time window w is the corresponding element in the behavior trajectory similarity matrix D. and Let M be the aggregation behavior vectors of user j and user k within the same time window w, respectively, where M is the consensus time window number, m = 1, 2, ..., M. Here, j and k are both user IDs; w is the time window number;

[0022] S26: Based on the behavioral trajectory similarity matrix D in S25 and its dynamic nature, a similarity threshold is dynamically set. If the corresponding element in the behavioral trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavioral trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set. The context set includes elements that exceed the corresponding dynamically set similarity threshold. The elements are the similarity of the behavioral trajectories between the aggregated behavioral vectors of two users. Finally, the two users corresponding to the similarity of each set of behavioral trajectories are recorded as user pairs.

[0023] Preferably, step S2 further includes:

[0024] S27: Extract user pairs corresponding to the current user from the context set, and count the users in all extracted user pairs excluding the current user to construct matching groups;

[0025] S28: Obtain the current intent credibility AiL for each user within the matching group, and use a statistical averaging algorithm to obtain the average intent credibility AiL for the matching group. avg Based on the current user's intent credibility AiL in S24, the current user's intent credibility AiL is further adjusted, specifically: AiL' = AiL + η * AiL avg In the formula, AiL' represents the current user's adjusted intent credibility, and η is the group reinforcement factor.

[0026] Preferably, step S3 specifically includes:

[0027] S31: Based on the current user's adjusted intent credibility AiL', dynamically set a credibility threshold PY for the corresponding scene label. Calculate the risk level RLV for the corresponding scene label by comparing the current user's adjusted intent credibility AiL' with the credibility threshold PY. Specifically: In the formula, Sigmoid(x) is the Sigmoid function.

[0028] Preferably, step S3 further includes:

[0029] S32: Based on the risk level RLV obtained in S31, and combined with the sub-behavior logs of the current user's behavior under the corresponding scenario label, the replica status of the data resources corresponding to the current user's operation is corrected to obtain the replica redundancy control factor Rrc, specifically:

[0030]

[0031] In the formula, Rrc rAs the replica redundancy control factor for the r-th data resource, Apa r Let β be the perceived data access pressure value of the r-th data resource, β be the suppression coefficient, and r be the number of the data resource.

[0032] S33: Map the replica redundancy control factor Rrc to the interval (0, 1) using the Sigmoid function, and determine the number of replicas R for each data resource during the next secure storage, specifically: R r =R0*(1+I(Rrc) r ≤0.5), where R0 is the number of copies of the corresponding data resources during this secure storage, R r I(Rrc) is the number of replicas of the r-th data resource during the next secure storage. r ≤0.5) is an indicator function.

[0033] A cloud computing-based electronic information security storage system includes a partitioning module, an adjustment module, and an update module;

[0034] The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform. Combined with the enterprise model, it divides the operation behavior of each user into scenarios to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed by scenario tags to obtain the sub-behavior logs of each user under different scenario tags.

[0035] The adjustment module is used to analyze whether each user's access operations meet business objectives based on the sub-behavior logs of each user under different scenario tags, in order to obtain the intent credibility AiL, and further adjust the intent credibility AiL by analyzing the similarity of the operation behavior trajectories between users in the same period of time.

[0036] The update module will dynamically set the trust threshold PY for different scenario labels, and after comparison, generate the risk level RLV for the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will infer the replica settings of each data resource for the next secure storage.

[0037] This invention provides a cloud computing-based electronic information security storage system and method, which has the following beneficial effects:

[0038] (1) By introducing a joint analysis mechanism of behavior logs and enterprise patterns in step S1, the system can classify and label user operation behaviors according to scenarios, realize a comprehensive characterization of behavioral semantics in a multi-dimensional context, and enhance the system's ability to perceive the operating environment and access motivation. In step S2, by constructing a similarity analysis between business semantic vectors and behavioral semantic vectors, and combining it with a user behavior trajectory similarity matrix, the system can dynamically judge the matching degree between user access behavior and business objectives, thereby obtaining a highly credible intent judgment result and further improving the recognition rate of hidden abnormal behaviors. In addition, step S3 generates a scenario-specific risk level index (RLV) by comparing and analyzing the intent credibility under scenario labels with dynamic credibility thresholds, and combines it with behavior log data to back-infer the replica redundancy control strategy, so that the number of data replicas is dynamically linked with user behavior risk, effectively suppressing the problem of redundant replica diffusion caused by high-risk users. In summary, the overall method realizes closed-loop control of access intent recognition, behavior evolution evaluation, and replica quantity adjustment, further improving the security adaptability and resource allocation efficiency of the cloud storage system under multi-user and multi-scenario conditions.

[0039] (2) By utilizing the difference between the user's current operation and the average of previous characteristics, the magnitude of their behavioral changes is dynamically described, providing a mathematical basis for identifying atypical operations. Combining the cosine similarity calculation of behavioral semantics and business semantics in S24, the final output is the intent credibility AiL, thereby realizing the credible modeling of access motivation and the measurement of behavioral authenticity. This method breaks through the limitation of traditional access control that can only identify "what was done" and instead judges "why it was done", effectively suppressing hidden operational behaviors under non-business objectives, enhancing the controllability and auditing capabilities of data access in the cloud environment, and providing a more credible behavioral basis for subsequent risk assessment and copy control.

[0040] (3) By aggregating the behavioral semantic vectors of all users according to time windows in S25, aggregated behavioral vectors over multiple time periods are constructed. This not only effectively smooths short-term behavioral fluctuations but also provides a structured foundation for cross-user and cross-time behavioral trajectory comparison. Subsequently, a behavioral trajectory similarity matrix D is constructed through cosine similarity calculation. Each element in the matrix quantifies the consistency of behavioral patterns between any two users within the same window, enabling the system to capture the collaborative behavioral characteristics of user groups across multiple time dimensions. S26 then sets an adaptive similarity threshold that slides over time based on the dynamic changes of this matrix and determines whether to include user pairs in the context set accordingly. Through this mechanism, the system can automatically identify user clusters with converging behaviors, construct a highly reliable behavioral context association graph, and effectively enhance the group reference support capability for individual intent credibility. Overall, this step improves the accuracy and robustness of behavioral judgment, provides a scalable and interpretable group semantic foundation for subsequent adjustment of intent credibility and determination of risk levels, and further strengthens the system's collaborative judgment capability and dynamic evolution detection capability.

[0041] (4) By determining the number of copies of each data resource in each secure storage, S33 makes the copy generation behavior no longer a static redundancy strategy, but a security decision result that is dynamically adjusted according to the risk of behavior. This method further improves the system's ability to suppress copies of high-risk access behaviors, reduces the probability of redundancy spread of sensitive data under abnormal operations, and ensures the availability and disaster recovery capability of data in low-risk scenarios, ultimately achieving dual optimization of data security and resource utilization. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of a cloud computing-based electronic information security storage method according to the present invention.

[0043] Figure 2 This is an overall logic diagram of a cloud computing-based electronic information security storage method according to the present invention;

[0044] Figure 3 This is a partial logic diagram of a cloud computing-based electronic information security storage method according to the present invention;

[0045] Figure 4 This is a block diagram of an electronic information security storage system based on cloud computing according to the present invention. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Example 1

[0048] Please see Figures 1 to 3 This invention provides a cloud computing-based method for secure electronic information storage, comprising the following steps:

[0049] S1: Read the behavior logs of each user in the historical period on the cloud storage platform, combine them with the enterprise model, divide the operation behavior of each user into scenarios to obtain scenario sets, and process the behavior logs of each user into scenario tags according to the scenario sets to obtain the sub-behavior logs of each user under different scenario tags.

[0050] S2: Based on the sub-behavior logs of each user under different scenario tags, analyze whether each user's access operations meet the business objectives (such as approval query, batch export, routine browsing, etc.) to obtain the intent credibility AiL, and further adjust the intent credibility AiL by analyzing the similarity of the operation behavior trajectories between users in the same period of time.

[0051] S3: Building upon S2, dynamically sets the trust threshold PY for different scenario labels. After comparison, it generates the risk level RLV for the corresponding scenario label. Combined with user behavior logs read from the cloud storage platform, it infers the replica settings for each data resource during the next secure storage. Specifically, the data resource replica settings are executed once each time electronic information is stored.

[0052] In this embodiment, by performing scene-based labeling on user operation behavior, refined behavior semantic recognition can be achieved, ensuring that the system can identify security risks with different risk levels for the same behavior in different scenarios. At the same time, the linkage mechanism between the construction of intent credibility AiL and the analysis of similarity of group behavior trajectories can dynamically perceive the evolution trend of user behavior and avoid misjudgment or omission caused by a single behavior indicator.

[0053] Ultimately, the system controls replica redundancy strategies by dynamically generating Risk Levels (RLV) and Trust Thresholds (PY) in conjunction with the data availability threshold. This suppresses replica propagation caused by potentially high-risk behaviors without affecting data availability, thus achieving a risk-driven closed-loop data replica control system. For example, if an employee queries approval documents using an authorized device during work hours, the system marks this as normal behavior and allows the generation of three copies for high availability. However, if the same employee attempts to perform a batch export operation using an external device outside of work hours, although the behavior type is the same, the system recognizes a lower level of trust in the employee's intent due to changes in the scenario label and increased behavior offset. By combining this with group behavior comparisons, a high-risk level is dynamically generated to prevent data from being propagated across multiple points under potential threats. This method enhances the system's adaptability to dynamic access behavior and its fine-grained replica control capabilities.

[0054] Example 2

[0055] Please refer to Figure 1 Specifically, the steps in S1 include:

[0056] S11: Read the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include operation timestamps, operation targets (such as file ID, resource ID, category tags, etc.), devices used, operation types, data resource levels, access times of each data resource, and operation paths (such as click traces and page jump paths, used to reflect the complexity of the operation).

[0057] S12: The enterprise model includes the division of working hours, non-working hours, working locations, non-working locations, authorized devices, unauthorized devices, business objectives, and non-business objectives for each user's position within the enterprise. By utilizing the enterprise model and combining it with the user behavior logs read from the cloud storage platform, the user's operation behavior is divided into scenarios to initially obtain a scenario set. Then, the scenario tags within the scenario set are combined to update the initially obtained scenario set.

[0058] The specific steps in S1 also include:

[0059] S13: Based on the updated scenario set in S12, classify each user's behavior logs into the corresponding scenario tags to obtain the sub-behavior logs of each user's behavior under different scenario tags.

[0060] It should be noted that scene tags include, but are not limited to, actions performed during working hours on authorized devices, actions performed outside of working hours on authorized devices, actions performed during working hours on unauthorized devices, actions performed during working hours that are related to business objectives, and actions performed outside of working hours that are related to business objectives, etc.

[0061] In this embodiment, by introducing multi-dimensional behavior logs and enterprise mode information in step S1, the present invention achieves high-precision scene recognition and classification of user operation behavior, further improving the contextual accuracy of subsequent behavior intent judgment and risk control.

[0062] By extracting key fields such as operation timestamps, operation paths, device types, and resource access frequencies in S11, and combining them with the institutional norms set by the enterprise in S12, such as working hours, authorized devices, and job objectives, the system can differentiate and categorize the same behavior under different operating environments. Furthermore, through the scenario label update mechanism in S13, sub-behavior logs of user behavior under different label scenarios can be accurately extracted, forming an input dataset that meets semantic recognition requirements. For example, a user's behavior of "accessing approval forms via a company-authorized PC at 9:30 AM on a weekday" will be categorized under the "working hours + authorized device + business objective" scenario, while the behavior of "downloading reports in batches via a home device at night" will be categorized under the "non-working hours + unauthorized device + high-risk export" scenario. This classification not only avoids misjudging the risk of similar operations but also provides a more objective, layered, and interpretable basis for subsequent intent credibility identification and copy control strategies. This method effectively improves the system's ability to understand the access context, enhances the ability to detect and respond to abnormal behavior early, and strengthens the dynamic security adaptability of the cloud storage system.

[0063] Example 3

[0064] Please refer to Figure 1 Specifically, the steps in S2 include:

[0065] S21: Pre-determine business processes from business objectives, and based on these processes, determine the operational characteristics of the corresponding business processes (such as querying approval forms, batch exporting reports, etc.). Utilize natural language processing techniques to semantically embed the task names, business processes, and corresponding operational characteristics of the business processes in the description of the business objectives (such as word embedding methods like Word2Vec and GloVe), thereby mapping the business objectives corresponding to each user into business semantic vectors.

[0066] S22: Again, use the cloud storage platform to read each user's behavior logs in real time, and use the real-time read user behavior logs to generate a set of business semantic vectors. Corresponding behavioral semantic vectors Behavioral semantic vectors Used to reflect the user's current access behavior:

[0067] S23: Based on the real-time read user behavior logs and combined with the updated scenario set, classify the current user behavior logs into the corresponding scenario tags, and analyze and calculate the behavior offset rate Xbm of each user's behavior logs under the corresponding scenario tags. The behavior offset rate Xbm is used to reflect the deviation of the current user's operation behavior from the past, specifically:

[0068]

[0069] In the formula, n is the feature dimension of the current operation, i.e., the behavior semantic vector. Feature dimensions, F i This represents the value of the current operation behavior in the i-th feature dimension (such as access count, operation type, resource level, etc.). This is the average value of past operational behaviors in the i-th feature dimension;

[0070] Business semantic vector It is applied to the judgment of the current operation to help identify the actual target of the access operation;

[0071] Business processes typically involve multiple steps, such as approval, query, and export. Task scheduling systems may have multiple steps, each of which triggers different operations and data accesses.

[0072] The specific steps in S2 also include:

[0073] S24: Use cosine similarity algorithm to measure business semantic vectors With behavioral semantic vectors similarity between To analyze whether each user's access operations align with business objectives (such as approval queries, batch exports, and routine browsing), and combined with the behavior offset rate Xbm of each user's current behavior logs under the corresponding scenario tags calculated in S21, the intent of the current user's operation behavior is analyzed to calculate and obtain the intent credibility AiL, specifically: In the formula, α is the behavior offset penalty coefficient, which is used to amplify or compress the influence of offset on credibility, and its range is between 0 and 1;

[0074] In this embodiment, the present invention introduces an access intent recognition mechanism based on semantic modeling in step S2, which further improves the cloud platform's ability to accurately judge the compliance of user operations and potential risks.

[0075] By semantically embedding business objectives, processes, and operational features in S21, a structured business semantic vector is constructed, enabling the system to accurately capture the target semantic intent of tasks. Combined with the real-time extraction of user behavior logs in S22 to generate behavioral semantic vectors, the system achieves semantic layer abstraction of the user's current operation. This allows for the calculation of the matching degree between behavior and business objectives in S24 using cosine similarity. Furthermore, in conjunction with the behavioral offset rate Xbm calculated in S23, an intent credibility AiL model that comprehensively considers semantic consistency and behavioral evolution offset is constructed, effectively avoiding the problem of slow response from static permission models to dynamic behavioral changes.

[0076] Example 4

[0077] Please refer to Figure 1 and Figure 3 Specifically, the S2 steps also include:

[0078] S25: All user behavior semantic vectors Perform time-sharing window aggregation to construct several sets of aggregated behavior vectors for the corresponding users. Through several sets of aggregated behavior vectors of the corresponding users The similarity of user behavior trajectories within the same time period is analyzed to construct a user behavior trajectory similarity matrix D. Each element in the behavior trajectory similarity matrix D represents the similarity of user behavior trajectories within the same time period, specifically:

[0079]

[0080] In the formula, The similarity of the behavioral trajectories between the behavior vectors of user j and user k within the same time window w is the corresponding element in the behavior trajectory similarity matrix D. and Let M be the aggregation behavior vectors of user j and user k within the same time window w, respectively, where M is the consensus time window number, m = 1, 2, ..., M. Here, j and k are both user IDs; w is the time window number;

[0081] Specifically, time-segmented window aggregation divides user behavior data into fixed-size time periods, such as by hour, day, or week, and then aggregates the user behavior vectors within each time period.

[0082] The purpose of aggregation is to compress multiple behavioral data into an aggregated behavioral vector, which represents the overall behavioral characteristics of users within a time window. After aggregation, the aggregated behavioral vector for each time period (such as hourly or daily) will represent the average characteristics of user behavior within that time period. This aggregation helps to smooth short-term fluctuations into longer-term trends and avoid the impact of occasional behavioral changes on the overall analysis.

[0083] By constructing a behavioral trajectory similarity matrix D, we can help identify "abnormally consistent" behavioral patterns: if multiple different accounts repeatedly access the same resources, operate in similar ways, and within similar time periods; even if each account does not pose a threat individually, their overall behavior shows a coordinated trend; the similarity matrix can reveal this hidden group-type risk.

[0084] S26: Based on the behavioral trajectory similarity matrix D in S25 and its dynamic nature, a similarity threshold is dynamically set. If the corresponding element in the behavioral trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavioral trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set. The context set includes elements that exceed the corresponding dynamically set similarity threshold. The elements are the similarity of the behavioral trajectories between the aggregated behavioral vectors of two users. Finally, the two users corresponding to the similarity of each set of behavioral trajectories are recorded as user pairs.

[0085] In this embodiment, the present invention further introduces a time window aggregation and trajectory similarity modeling mechanism for user behavior semantics in step S2, which effectively improves the ability to identify user behavior evolution trends and the accuracy of context association analysis.

[0086] In step S25, the system aggregates the semantic vectors of each user's behavior across multiple consensus time windows in a time-sharing manner, constructing a behavior trajectory sequence. Based on this sequence, a behavior trajectory similarity matrix D is generated between users, accurately depicting the convergence of user behavior within a specific time period. In step S26, the system sets an adaptive similarity threshold based on the dynamic distribution of matrix D, automatically selecting user pairs with significantly consistent behavioral characteristics to form a context set. This approach enables enhanced behavioral credibility based on "behavioral similarity background," further optimizing the stability and robustness of individual behavioral credibility assessment.

[0087] For example, if User A and User B both performed approval process operations every day from 9:00 AM to 10:00 AM over the past week, and their access frequency, resource level, and operation path similarity consistently remained above 0.92, consistently exceeding the dynamic threshold of 0.85 in the behavior trajectory similarity matrix D, the system would mark them as a "user pair" and include them in the context set. When User A's behavior on a particular day deviates slightly, the system can use User B's behavior as a reference to correct its credibility, avoiding misjudgment as abnormal. This mechanism not only enhances the context-awareness of user behavior recognition but also introduces collective semantic consensus when user behavior slightly drifts, achieving dynamic and stable correction of credibility, significantly improving the overall system's accuracy in understanding behavior and its risk perception capabilities in complex access environments.

[0088] Example 5

[0089] Please refer to Figure 1 Specifically, the S2 steps also include:

[0090] S27: Extract user pairs corresponding to the current user from the context set, and count the users in all extracted user pairs except the current user to construct a matching group. The matching group refers to another user other than the current user from all extracted user pairs.

[0091] S28: Obtain the current intent credibility AiL for each user within the matching group, and use a statistical averaging algorithm to obtain the average intent credibility AiL for the matching group. avg Based on the current user's intent credibility AiL in S24, further adjustments are made to the current user's intent credibility AiL, specifically as follows:

[0092] AiL'=AiL+η*AiL avg ;

[0093] In the formula, AiL' represents the adjusted credibility of the current user's intent, and η is the group reinforcement factor, which is an adjustment parameter used to strengthen the influence of group similarity when calculating the credibility of user intent, especially when multiple user behaviors similar to the target user are taken into consideration, and its range is between 0 and 1.

[0094] By updating and calculating the credibility of the current user's intent (AiL), it can be seen that if the user belongs to a "group with a common purpose" (such as a normal template import operation), even if the offset rate is high, it will not be easily misjudged. If the similar group has a malicious purpose, its behavioral intent will be suppressed in the opposite direction.

[0095] In this embodiment, a dynamic intention credibility correction mechanism based on group behavior similarity is further introduced in S2, which significantly enhances the system's accuracy in judging user behavior credibility and the stability of abnormal behavior identification. In S27, user pairs that are highly similar to the current user's behavior trajectory are extracted from the context set, and the current user is removed to construct a matching group, ensuring that the group participating in the credibility reference has sufficient behavioral similarity and contextual consistency.

[0096] S28 utilizes the statistical mean of the current intent credibility of each user in the matching group, combined with the original credibility of the current user, and weights it by introducing a group reinforcement factor η to form a credibility output that is representative of the group. This method effectively integrates individual behavior judgment and group consensus trend, preserving the user's own behavioral characteristics while improving the sensitivity to non-obvious deviations by leveraging group convergence behavior.

[0097] Especially when facing complex access scenarios such as ambiguous behavior or semantic overlap, this approach can further reduce the risk of false positives and false negatives, and improve the system's ability to identify and respond to gray-area operations. Overall, this step constructs a trust evolution mechanism driven by enhanced group similarity, providing a more stable and reliable behavioral intent basis for subsequent risk level calculation and replica control, and enhancing the cloud platform's security decision-making capabilities in complex user behavior scenarios.

[0098] Example 6

[0099] Please refer to Figure 1Specifically, the S3 steps include:

[0100] S31: Based on the current user's adjusted intent credibility AiL', dynamically set a credibility threshold PY for the corresponding scene label. Calculate the risk level RLV for the corresponding scene label by comparing the current user's adjusted intent credibility AiL' with the credibility threshold PY. Specifically:

[0101]

[0102] In the formula, Sigmoid(x) is the Sigmoid function. Where e is the Euler number, with a value of approximately 2.71828, and x refers to... ∈ represents a small positive number to prevent division by zero.

[0103] The Sigmoid function has the characteristic that positive values ​​tend to be 1 and negative values ​​tend to be 0. When x in Sigmoid(x) equals 0, the output value of Sigmoid(x) is 0.5, which is neutral risk. When x is less than 0, the output value of Sigmoid(x) is in the range of 0 to 0.5, which is low risk.

[0104] If AiL' is less than PY, then RLV is less than 0, indicating that the behavior is within the acceptable range and is low risk; if AiL' is equal to PY, then RLV is equal to 0, indicating that the behavior is borderline and is neutral risk; if AiL' is greater than PY, then RLV is greater than 0, indicating that the behavior has exceeded the tolerance range and is high risk.

[0105] Specifically, the method for dynamically setting the credibility threshold PY is as follows: based on the previously adjusted intent credibility AiL' of users, the adjusted average intent credibility AiL' is calculated. avg and its standard deviation AiL' σ Based on the adjusted average intent credibility AiL' avg and its standard deviation AiL' σ Based on the changes in the user's adjusted intent credibility AiL', dynamically set the credibility threshold PY: PY = AiL' avg +Q*AiL' σ Where Q is a constant, usually taking the value 1-3, corresponding to different confidence levels, and the specific value is set by the user (according to the actual situation);

[0106] The specific steps in S3 also include:

[0107] S32: Based on the risk level RLV obtained in S31, and combined with the sub-behavior logs of the current user's behavior under the corresponding scenario label, the replica status of the data resources corresponding to the current user's operation is corrected to obtain the replica redundancy control factor Rrc, specifically:

[0108]

[0109] In the formula, Rrc r As the replica redundancy control factor for the r-th data resource, Apa r Let be the perceived data access pressure value of the r-th data resource, β be the suppression coefficient (an adjustment coefficient ranging from 0 to 1), and r be the data resource number; where Apa is the perceived data access pressure value of the r-th data resource. r This refers to the result of normalizing the number of accesses to the r-th data resource and mapping its value to the interval [0,1], which is used to reflect the availability of information.

[0110] S33: Map the replica redundancy control factor Rrc to the interval (0, 1) using the Sigmoid function, and determine the number of replicas R for each data resource during the next secure storage, specifically: R r =R0*(1+I(Rrc) r ≤0.5), where R0 is the number of copies of the corresponding data resources during this secure storage, R r I(Rrc) is the number of replicas of the r-th data resource during the next secure storage. r ≤0.5) is an indicator function, when Rrc r When ≤0.5, the indicator function outputs -Rrc. r If Rrc r When the value is greater than 0.5, the indicator function outputs Rrc. r .

[0111] In cloud computing environments, redundancy in electronic information is primarily used to enhance data availability, support disaster recovery, and improve access performance. Traditional replication strategies are typically uniform (e.g., three copies of each piece of data automatically distributed across three storage areas by default). However, traditional replication strategies suffer from several problems: storage resources are wasted on "unimportant or high-risk data"; and deploying high-risk information with multiple copies increases exposure and leakage risk. Therefore, this method uses reverse engineering to deduce the acceptable number of redundant copies based on the risk level of user behavior. High-risk information is unsuitable for multiple copy propagation and the number of copies should be controlled; conversely, low-risk information can have more copies to improve access efficiency and availability.

[0112] In this embodiment, a risk assessment and copy control mechanism based on intent credibility is introduced in S3 to achieve dynamic security control of the data copy generation process in the cloud environment.

[0113] By using the Sigmoid function to process the credibility of the user's current behavior in the corresponding scenario in S31, and combining it with a dynamically set credibility threshold, a risk level index RLV is generated. The system can accurately measure the risk level of the user's current operation. S32 further combines the risk level with the user's historical behavior logs to calculate the data resource replication redundancy control factor Rrc, so that the replication generation strategy can dynamically converge according to the risk.

[0114] S33 achieves precise control over the number of replicas by mapping the replica redundancy control factor Rrc value to the 0-1 range and matching it with replica setting rules, effectively preventing the excessive copying of sensitive data under high-risk behaviors. For example, if a user frequently exports sensitive documents in a scenario of "non-working hours + unauthorized devices," the system identifies that the credibility of their behavior is low, generates a risk level RLV, and calculates a smaller Rrc value (e.g., 0.2) using a suppression function. Based on this, the system limits the number of replicas to 0, blocking the data copying link under this behavior. This method effectively prevents the spread of data redundancy caused by abnormal abuse within the permission boundary, enhances the system's dynamic balance between storage security and resource optimization, and also possesses intelligent characteristics such as adaptability, scene awareness, and user behavior semantic drive.

[0115] Example 7

[0116] Please refer to Figure 4 Specifically: a cloud computing-based electronic information security storage system, including a partitioning module, an adjustment module and an update module;

[0117] The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform. Combined with the enterprise model, it divides the operation behavior of each user into scenarios to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed by scenario tags to obtain the sub-behavior logs of each user under different scenario tags.

[0118] The adjustment module is used to analyze whether each user's access operations meet business objectives based on the sub-behavior logs of each user under different scenario tags, in order to obtain the intent credibility AiL, and further adjust the intent credibility AiL by analyzing the similarity of the operation behavior trajectories between users in the same period of time.

[0119] The update module will dynamically set the trust threshold PY for different scenario labels, and after comparison, generate the risk level RLV for the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will infer the replica settings of each data resource for the next secure storage.

[0120] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A cloud computing-based method for secure electronic information storage, characterized in that: Includes the following steps, S1: Read the behavior logs of each user within a historical time period from the cloud storage platform. Combined with the enterprise model, classify the user's operation behavior into scenarios to obtain a scenario set. Based on the scenario set, process the user's behavior logs into scenario tags to obtain the sub-behavior logs of each user under different scenario tags. The enterprise model includes the division of each user's position within the enterprise into working hours, non-working hours, working location, non-working location, authorized device, unauthorized device, business objective, and non-business objective. Scenario tags include behavior operations performed during working hours on authorized devices, behavior operations performed during non-working hours on authorized devices, behavior operations performed during working hours on unauthorized devices, behavior operations during working hours that pertain to business objectives, and behavior operations during non-working hours that pertain to business objectives. S2: Based on the sub-behavior logs of each user under different scenario tags, analyze whether each user's access operations conform to business objectives to obtain the intent credibility (AiL). Further adjust the intent credibility (AiL) by analyzing the similarity of user operation behavior trajectories within the same time period. Specific steps in S2 include: S21: Pre-determine business processes from business objectives, and determine the operational features of corresponding business processes based on these processes. Using natural language processing technology, semantically embed the task names, business processes, and corresponding operational features in the description of business objectives to map the business objectives of each user into business semantic vectors. ; S22: Again, use the cloud storage platform to read each user's behavior logs in real time, and use the real-time read user behavior logs to generate a set of business semantic vectors. Corresponding behavioral semantic vectors Behavioral semantic vector Used to reflect the user's current access behavior: S23: Based on the real-time read user behavior logs and combined with the updated scenario set, classify the current user behavior logs into the corresponding scenario tags, and analyze and calculate the behavior offset rate Xbm of each user's behavior logs under the corresponding scenario tags, specifically: In the formula, n represents the feature dimension of the current operation. The value of the current operation behavior in the i-th feature dimension. This represents the average value of past operational behaviors in the i-th feature dimension. S24: Use cosine similarity algorithm to measure business semantic vectors With behavioral semantic vectors Similarity between This is done to analyze whether each user's access operations align with business objectives. Combined with the behavior offset rate Xbm of each user's current behavior logs under the corresponding scenario label, calculated in S21, the intent of the current user's actions is analyzed to calculate and obtain the intent credibility AiL. Specifically: In the formula, This is the penalty coefficient for behavioral deviation; S25: All user behavior semantic vectors Perform time-sharing window aggregation to construct several sets of aggregated behavior vectors for the corresponding users. Through several sets of aggregated behavior vectors of the corresponding users This study analyzes the similarity of user behavior trajectories within the same time period to construct a user behavior trajectory similarity matrix D. Each element in the behavior trajectory similarity matrix D represents the similarity of user behavior trajectories within the same time period. Specifically: In the formula, The similarity of the behavioral trajectories between the behavior vectors of user j and user k within the same time window w is the corresponding element in the behavior trajectory similarity matrix D. and Let be the aggregation behavior vectors of user j and user k within the same time window w, respectively, where M is the consensus time window number, m = 1, 2, ..., M. for The cosine similarity value between them, where j and k are user IDs; w is the time window number; S26: Based on the behavioral trajectory similarity matrix D in S25 and its dynamic nature, a similarity threshold is dynamically set. If the corresponding element in the behavioral trajectory similarity matrix D exceeds the corresponding dynamically set similarity threshold, it is included in the context set; if the corresponding element in the behavioral trajectory similarity matrix D does not exceed the corresponding dynamically set similarity threshold, it is not included in the context set. The context set includes elements that exceed the corresponding dynamically set similarity threshold. The elements are the similarity of the behavioral trajectories between the aggregated behavioral vectors of two users. Finally, the two users corresponding to the similarity of each set of behavioral trajectories are recorded as user pairs. S27: Extract user pairs corresponding to the current user from the context set, and count the users in all extracted user pairs excluding the current user to construct matching groups; S28: Obtain the current intent credibility AiL for each user within the matching group, and use a statistical averaging algorithm to obtain the average intent credibility for the matching group. Based on the current user's intent credibility AiL in S24, further adjustments are made to the current user's intent credibility AiL, specifically as follows: In the formula, The credibility of the current user's adjusted intent. As a group reinforcement factor; S3: Building upon S2, dynamically set the trust threshold PY for different scenario labels, and generate the risk level RLV for the corresponding scenario label through comparison. Combined with user behavior logs read from the cloud storage platform, the replica settings for each data resource are deduced for the next secure storage operation. Specific steps in S3 include: S31: Adjust the credibility of the intent based on the current user. Dynamically set a trust threshold PY for its corresponding scene label, and adjust the trustworthiness of the current user's intent. The risk level (RLV) under the corresponding scenario label is calculated by comparing it with the confidence threshold (PY). In the formula, For the Sigmoid function, ϵ is a small positive number; The specific steps in S3 also include: S32: Based on the risk level RLV obtained in S31, and combined with the sub-behavior logs of the current user's behavior under the corresponding scenario label, the replica status of the data resources corresponding to the current user's operation is corrected to obtain the replica redundancy control factor Rrc, specifically: ; In the formula, The replica redundancy control factor for the r-th data resource. The perceived data access pressure value for the r-th data resource. Here, r is the suppression coefficient, and r is the data resource number. S33: The replication redundancy control factor Rrc is mapped to the interval (0, 1) using the Sigmoid function, and the number of replicas R of each data resource is determined during the next secure storage, specifically as follows: ,in, This refers to the number of copies of the corresponding data resources during this secure storage. This is the number of replicas of the r-th data resource during the next secure storage. This is an indicator function.

2. The electronic information security storage method based on cloud computing according to claim 1, characterized in that: The specific steps in S1 include: S11: Read the behavior logs of each user in the historical period through the cloud storage platform. The behavior logs include operation timestamp, operation target, device used, operation type, data resource level, access frequency of each data resource and operation path. S12: By utilizing the enterprise model and combining the user behavior logs read from the cloud storage platform, the operation behavior of each user is divided into scenarios to initially obtain a scenario set. Then, the scenario tags in the scenario set are combined to update the initially obtained scenario set.

3. The electronic information security storage method based on cloud computing according to claim 2, characterized in that: The specific steps in S1 also include: S13: Based on the updated scenario set in S12, classify each user's behavior logs into the corresponding scenario tags to obtain the sub-behavior logs of each user's behavior under different scenario tags.

4. A cloud computing-based secure electronic information storage system, used to implement the cloud computing-based secure electronic information storage method according to any one of claims 1 to 3, characterized in that: This includes module division, module adjustment, and module updating; The segmentation module is used to read the behavior logs of each user in the historical period on the cloud storage platform. Combined with the enterprise model, it divides the operation behavior of each user into scenarios to obtain scenario sets. Based on the scenario sets, the behavior logs of each user are processed by scenario tags to obtain the sub-behavior logs of each user under different scenario tags. The adjustment module is used to analyze whether each user's access operations meet business objectives based on the sub-behavior logs of each user under different scenario tags, in order to obtain the intent credibility AiL, and further adjust the intent credibility AiL by analyzing the similarity of the operation behavior trajectories between users in the same period of time. The update module will dynamically set the trust threshold PY for different scenario labels, and after comparison, generate the risk level RLV for the corresponding scenario label. Combined with the user behavior logs read in the cloud storage platform, it will infer the replica settings of each data resource for the next secure storage.

Citation Information

Patent Citations

  • Cloud storage dynamic optimization method and system

    CN120186155A

  • Master Network Techniques for a Digital Duplicate

    US20220076178A1