A method and system for analyzing office vulnerabilities based on big data

Through big data analysis methods, multi-source office data is collected and processed, semantic vectors are generated, and overpriced behaviors and permission redundancy is identified, which solves the problem of the risk of overlapping permissions in the existing technology, and realizes accurate, dynamic identification and prevention of office vulnerabilities.

CN120217394BActive Publication Date: 2025-09-02HEBEI WANGXIN GOVERNMENT SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510698636.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-02
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing permission audit and process compliance analysis technology is difficult to dynamically capture the path of user behavior changes, and it is impossible to effectively identify the risk of overlapping permissions under the coordination of multiple roles, resulting in hidden office loopholes such as responsibilities redirection and abuse of permissions.

Method used

By constructing an office vulnerability analysis method based on big data, multi-source heterogeneous office data are collected, feature extraction, labeling and time synchronization are performed, and a unified office behavior semantic vector is generated. Combined with the job permission matrix and organizational hierarchical semantic understanding, risk identification and permission overlap calculation are carried out to identify risks of overprivileged access and permission redundancy.

Benefits of technology

Real-time identification of overprivileged access, responsibility reversal and permission redundancy in the office process is realized, and the ability to identify highly concealed violations is improved, and the abuse of permissions and process detours is prevented, and the rigor and penetration of process audits is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217394B_ABST
    Figure CN120217394B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of data analysis technology, and provides a method and system for analyzing office vulnerabilities based on big data, which collects multi-source heterogeneous office data, processes the multi-source heterogeneous office data, and obtains a unified semantic vector of office behavior after processing; performs risk identification based on the semantic vector, and identifies whether the current behavior is within the authorized scope by comparing it; if it is determined that the access behavior does not match the job authority, it is judged as a potential unauthorized behavior and performs secondary verification; examines whether the approver of each node matches the preset job role, and obtains the frequency of replacement of the approval node processor; detects whether the actual operator is consistent with the node set position based on the mapping relationship, triggers the responsibility jump analysis process, combines the operation behavior with the potential redundant risk in the authority structure for analysis, and calculates the authority overlap, which has the advantages of avoiding misjudgment of unconventional but reasonable behavior and improving the ability to identify highly concealed violations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data analysis technology, and in particular relates to a method and system for analyzing office vulnerabilities based on big data. Background Art

[0002] In today's digital office environments of large organizations and enterprises, with the continuous integration of business systems and the growing demand for collaborative operations, cross-access and multi-role collaboration among employees is becoming increasingly common. While this multi-source, heterogeneous, and high-frequency collaborative office model improves business efficiency, it also introduces management risks such as complex permission configuration, blurred process responsibility boundaries, and uncontrollable operational behavior. This can easily lead to hidden office vulnerabilities such as responsibility transfer, permission abuse, and process circumvention.

[0003] Existing permission auditing and process compliance analysis technologies are mostly based on static rules or periodic audits. These methods struggle to dynamically capture the changing paths of user behavior within the process chain and lack the ability to determine whether operators truly possess the authority to perform their duties. Furthermore, traditional methods generally overlook the risk of collaborative overreach caused by overlapping permissions between different positions. For example, the overlapping high permissions of multiple positions with different responsibilities over the same sensitive resources go unidentified, leading to uncontrollable opportunities for responsibility avoidance during operation substitution or task delegation.

[0004] Therefore, there is an urgent need for a multi-dimensional identification method that can integrate process behavior sequences, job responsibility mapping, authority structure similarity analysis and organizational hierarchy semantic understanding, so as to achieve real-time perception and structured identification of risks such as unauthorized access, responsibility jump and redundant authority collaboration in office processes, thereby providing organizations with more accurate, dynamic and explainable behavioral compliance guarantees. Summary of the Invention

[0005] The purpose of the embodiment of the present invention is to provide a method for analyzing office vulnerabilities based on big data, aiming to solve the problems raised in the third part of the background technology.

[0006] The embodiment of the present invention is implemented as follows: a method for analyzing office vulnerabilities based on big data, the method comprising:

[0007] Collect and process multi-source heterogeneous office data using methods including feature extraction, labeling, and time synchronization to obtain a unified semantic vector of office behavior. This semantic vector is used to provide unified data support for subsequent analysis.

[0008] Risk identification is performed based on semantic vectors. By comparing and identifying whether the current behavior is within the authorized scope, if the access behavior is determined to be inconsistent with the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed.

[0009] Compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of changes in the approval node processor;

[0010] Based on the mapping relationship, it is detected whether the actual operator is consistent with the node setting position, triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and calculating the authority overlap.

[0011] Preferably, the risk identification is performed based on the semantic vector, and by comparing and identifying whether the current behavior is within the authorized scope, if the access behavior is determined not to match the job authority, it is determined to be a potential unauthorized behavior, and the secondary verification step is performed, specifically including:

[0012] Risk identification is performed based on semantic vectors. The risk identification method is multi-level identification, which includes behavior scoring and path identification. The identification result is obtained by comparing and identifying whether the current behavior is within the authorized scope;

[0013] If the access behavior is determined to be inconsistent with the position's permissions and is not accompanied by any prior authorization operations, it is considered a potential violation of authority and a further search is conducted to determine whether there are supplementary authorization records. Such authorization records include temporary release of project-based permissions, dynamic permission inheritance generated by the collaborative approval mechanism, and proxy approval relationships.

[0014] If no valid prior authorization or permission transfer chain is found, the behavior will be marked as suspected unauthorized access, and a secondary verification will be performed to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

[0015] Preferably, the step of comparing the actual node sequence executed by each process instance with the template path one by one, examining whether the approver of each node matches the preset job role, and obtaining the frequency of changes in the approval node processor specifically includes:

[0016] Obtaining a process behavior sequence, which is used to verify that the approval path, nodes, and approvers are consistent with the original configuration, and comparing the actual node sequence executed by each process instance with the template path one by one;

[0017] Identify whether there are missing nodes, deformed paths, or skipped paths. Review whether the approver of each node matches the preset job role. Determine whether the approver has approval authority. If not, mark the node as a risky one due to identity mismatch.

[0018] Get the frequency of changes in approval node handlers. If the frequency exceeds the threshold, it is determined that the process operation may be subject to manipulation.

[0019] Preferably, the step of detecting whether the actual operator is consistent with the node set position based on the mapping relationship, triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and calculating the authority overlap specifically includes:

[0020] Establish a responsibility attribution mapping relationship for approval nodes or task nodes. The mapping relationship is generated based on the organizational structure map, job responsibility specifications, and process template configuration. The actual operator is checked based on the mapping relationship to see if the node position is consistent. If inconsistency is detected, it is marked as a responsibility deviation point.

[0021] Trigger the responsibility jump analysis process, combine operational behavior with potential redundancy risks in the permission structure, and analyze duplicate permissions for the same highly sensitive resource across multiple non-identical responsibility chains;

[0022] Calculate the degree of permission overlap, obtain the overlap threshold, and determine the permission overlap based on the overlap threshold.

[0023] Preferably, the labeling is to add multiple label dimensions to the data through a label template system before storage, and the time synchronization is to unify the time base of data timestamps in different systems.

[0024] Another object of an embodiment of the present invention is to provide a big data-based office vulnerability analysis system, the system comprising:

[0025] The data acquisition module collects and processes multi-source heterogeneous office data through feature extraction, labeling, and time synchronization. The module then obtains a unified semantic vector of office behavior, which provides unified data support for subsequent analysis.

[0026] The risk identification module identifies risks based on semantic vectors and compares the current behavior to see if it is within the authorized scope. If the access behavior is determined to be inconsistent with the job authority, it is considered a potential unauthorized behavior and undergoes secondary verification.

[0027] The process approval module compares the actual node sequence executed by each process instance with the template path one by one, checks whether the approver of each node matches the preset job role, and obtains the frequency of changes in the approval node processor;

[0028] The attribution mapping module detects whether the actual operator is consistent with the node set position based on the mapping relationship, triggers the responsibility jump analysis process, combines the operation behavior with the potential redundancy risk in the authority structure, and calculates the authority overlap.

[0029] Preferably, the risk identification module includes:

[0030] The risk identification unit performs risk identification based on semantic vectors. The risk identification method is multi-level identification, which includes behavior scoring and path identification. The identification result is obtained by comparing and identifying whether the current behavior is within the authorized scope;

[0031] If the position authority unit determines that the access behavior does not match the position authority and is not accompanied by any pre-authorization operation, it is judged as a potential unauthorized behavior and further searches for supplementary authorization records. The authorization records include temporary release of project-based permissions, dynamic permission inheritance generated by the collaborative approval mechanism, and proxy approval relationships.

[0032] The secondary verification unit, if no valid pre-authorization or permission transfer chain is found, will mark the behavior as suspected unauthorized access and perform secondary verification to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

[0033] Preferably, the process approval module includes:

[0034] The process behavior sequence unit obtains the process behavior sequence, which is used to verify that the approval path, nodes, and approvers are consistent with the original configuration one by one, and compare the actual node sequence executed by each process instance with the template path one by one;

[0035] The process approval unit identifies whether there are missing nodes, deformed paths, or skipped paths. It also examines whether the approver at each node matches the preset job role. It then determines whether the approver has the approval authority. If the approver does not have the corresponding approval authority, the node is marked as a risky node with inconsistent identity.

[0036] The change frequency unit obtains the change frequency of the approval node processor. If the change frequency exceeds the threshold, it is determined that the process operation may be subject to manipulation intervention.

[0037] Preferably, the attribution mapping module includes:

[0038] The attribution mapping unit establishes a responsibility attribution mapping relationship for the approval node or task node. The mapping relationship is generated based on the organizational structure map, job responsibility specifications, and process template configuration. The mapping relationship is used to detect whether the actual operator is consistent with the node's set position. If inconsistency is detected, it is marked as a responsibility deviation point;

[0039] The responsibility analysis unit triggers the responsibility jump analysis process, combines operational behavior with potential redundancy risks in the permission structure, and analyzes duplicate permissions for the same highly sensitive resource across multiple non-identical responsibility chains.

[0040] The authority overlap unit calculates its authority overlap, obtains the overlap threshold, and determines the authority overlap situation based on the overlap threshold.

[0041] Preferably, the labeling is to add multiple label dimensions to the data through a label template system before storage, and the time synchronization is to unify the time base of data timestamps in different systems.

[0042] The present invention provides a big data-based office vulnerability analysis method. By constructing a multidimensional behavior recognition mechanism based on a job authority matrix, an approval path structure analysis mechanism, and a permission overlap calculation model, it can effectively identify complex structural vulnerabilities hidden in office processes, such as unauthorized access, approval tampering, responsibility jumps, and collaborative abuse of power. The system not only verifies the permission boundaries of user operations at the single-point behavior level but also conducts dynamic linkage analysis based on multiple semantic relationships, such as process context, authorization transfer chains, and organizational structure maps. This avoids misjudging unconventional but reasonable behaviors and improves the ability to identify highly concealed violations.

[0043] The system introduces a joint analysis mechanism combining approval behavior sequence modeling and responsibility attribution verification, enabling accurate assessment of the completeness of approval pathways, the consistency of approver and position assignments, and the proper attribution of approval responsibilities. Combined with monitoring the frequency of approver changes and semantic analysis of organizational hierarchical relationships, the system can identify authority-circumventing behaviors such as bypassing signatures, approving by proxy, and process manipulation, further enhancing the rigor and penetration of process audits.

[0044] By constructing a resource operation permission mapping diagram and introducing a method for calculating permission overlap, we can dynamically identify redundant permissions across multiple employees for highly sensitive resources and promptly detect high-level permission overlap across departments or non-role positions. Based on the permission vector model, we calculate metrics such as Jaccard similarity and cosine similarity. Based on the precise quantification of permission similarity, we then combine actual behavioral trajectories to determine whether there are risks of collaborative collaboration, responsibility avoidance, or abuse of authority. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flowchart of a method for analyzing office vulnerabilities based on big data provided by an embodiment of the present invention;

[0046] Figure 2 A flowchart of the steps for performing risk identification based on semantic vectors, by comparing and identifying whether the current behavior is within the authorized scope, provided by an embodiment of the present invention;

[0047] Figure 3 A flowchart of the steps of checking whether the approver of each node matches the preset job role and obtaining the frequency of changes in the approval node processor, provided by an embodiment of the present invention;

[0048] Figure 4A flowchart of the steps of detecting whether the actual operator and the node set position are consistent based on the mapping relationship, triggering the responsibility jump analysis process, and calculating the degree of authority overlap provided by an embodiment of the present invention;

[0049] Figure 5 An architectural diagram of a big data-based office vulnerability analysis system provided by an embodiment of the present invention;

[0050] Figure 6 This is an architectural diagram of the risk identification module provided by an embodiment of the present invention;

[0051] Figure 7 An architectural diagram of the process approval module provided in an embodiment of the present invention;

[0052] Figure 8 This is an architectural diagram of the home mapping module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0054] It is understood that the terms "first," "second," etc., used herein may be used to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script without departing from the scope of this application.

[0055] like Figure 1 As shown, a method for analyzing office vulnerabilities based on big data is provided in an embodiment of the present invention, and the method includes:

[0056] S100, collecting multi-source heterogeneous office data, processing the multi-source heterogeneous office data, processing methods including feature extraction, labeling and time synchronization, obtaining a unified office behavior semantic vector after processing, and the semantic vector is used to provide unified data support for subsequent analysis.

[0057] In this step, multi-source, heterogeneous office data is collected. To achieve unified modeling and precise analysis of complex office behaviors, the system first collects heterogeneous data from various sources, including OA systems, ERP systems, document management platforms, permission systems, instant messaging tools, and physical access control devices. These data types vary significantly in format, structure, and time. This data is standardized through the interface adaptation and normalization modules, extracting key features such as operation type, target object, operator identity, operation location, and device information. This unified feature mapping is then achieved by combining rule templates with behavioral semantic models.

[0058] On this basis, the system labels the extracted behavioral data, assigning multiple levels of risk or operational attribute labels based on behavioral semantics, sensitivity, temporal characteristics, and contextual logic, such as "unauthorized access," "process bypass," and "sensitive document manipulation," enhancing the semantic integrity and interpretability of behavioral expressions. A time synchronization mechanism is also introduced, using NTP calibration and sliding window aggregation to ensure consistency and comparability of multi-source behavioral events on the timeline, avoiding analytical bias caused by data time drift.

[0059] Ultimately, all processed behavioral data is encoded into a unified office behavior semantic vector, integrating structured features, label attributes, and temporal context, providing a high-quality input foundation for subsequent graph modeling, path analysis, and risk identification. This mechanism not only enables the integration of cross-system and cross-terminal behavioral data, but also provides strong data support for refined behavior identification and chained risk analysis in office scenarios.

[0060] S200, risk identification is performed based on the semantic vector, and whether the current behavior is within the authorized scope is compared and identified. If it is determined that the access behavior does not match the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed.

[0061] In this step, risk identification is performed based on semantic vectors. After the office behavior semantic vectors are generated, each semantic vector is first subjected to risk identification. One of the core judgment logics is whether the access behavior is within the authorized scope. A permission baseline model is constructed based on the position permission matrix. The standard operation permissions of each position in each business system, functional module, and data object are encoded in vector form and compared one by one with the fields such as the operation object, operation type, and operator identity contained in the behavior semantic vector. If it is found that the system modules or data resources involved in the current behavior exceed the permitted scope of the position in the permission matrix, it will be preliminarily marked as "suspected unauthorized behavior."

[0062] Based on this, the behavior is not immediately deemed a violation. Instead, it enters a secondary verification mechanism, combining organizational authorization logs, process task context, and temporary permission records to further assess the legitimacy of the behavior. The system checks whether the behavior has a legitimate pre-authorization path, such as position transfer authorization, collaborative task assignment, or approval delegation. It also analyzes whether the current behavior is nested within a legitimate process node or automatically triggered by a process task. If there are no authorization records, responsibility matching paths, or process linkage evidence, the behavior will ultimately be confirmed as unauthorized access, triggering an alert, log flag, or policy response action.

[0063] This mechanism not only efficiently identifies abnormal behavior outside of the scope of permissions, but also effectively avoids false positives through contextual linkage and permission chain verification, improving the accuracy and business adaptability of unauthorized identification. This process fully demonstrates the value of semantic vectors as the core carrier of behavioral analysis in permission identification, achieving an intelligent evolution from static permission configuration to dynamic behavioral compliance.

[0064] S300 , compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of changes in the approval node processor.

[0065] In this step, the actual node sequence executed by each process instance is compared one by one with the template path. When identifying abnormal behavior in process approval, each process instance is used as an analysis unit, and the node sequence executed by the instance in actual operation is first extracted. These nodes contain detailed information such as node number, node type, processor, operation time, approval action, and approval result. The actual path is then structurally compared with the standard path preset in the process template to check for structural deviations such as skipped nodes, abnormal node sequence, process interruption, or path merging. This comparison process is implemented through a node alignment algorithm, which can identify the degree of difference between the execution path and the standard model, and preliminarily determine whether there is a risk of abnormal process direction or path deformation.

[0066] After completing the path structure comparison, each approval node is further verified to ensure that the actual processor matches the role specified in the process template. This comparison is based on the position authority mapping table and the organizational structure diagram. It analyzes whether the processor belongs to the target position, has the corresponding responsibilities, and whether the processor is an authorized agent or bypasses the level of authority. If the processor does not directly correspond to the configured position, the system will mark the node as "identity mismatch" and assign a corresponding risk weight.

[0067] S400: Detect whether the actual operator is consistent with the node set position based on the mapping relationship, trigger the responsibility jump analysis process, combine the operation behavior with the potential redundancy risk in the authority structure, and calculate the authority overlap.

[0068] In this step, the mapping relationship is used to check whether the actual operator is consistent with the position set at the node. When performing the review of process node responsibility, the actual operator and the node set position will be matched one by one based on the mapping relationship between each approval node and job responsibilities in the process template. Through the authority configuration table and the organizational structure map, it is determined whether the current processor belongs to the position range specified by the node and whether he has the corresponding responsibilities and authority. If it is found that the processor is inconsistent with the set position and there is a lack of authorization, transfer records or process role change basis, it will be judged as "suspected responsibility jump" and the responsibility jump analysis process will be triggered immediately.

[0069] In responsibility transition analysis, we not only focus on the deviation between the handler's responsibilities and the node configuration, but also further analyze the linkage between this operational behavior and the potential redundancy risks in the permission structure. We construct a permission overlap calculation model, abstracting the permission set corresponding to the current handler and the designated position into a vector form. If there is a high degree of overlap between the two, such as exceeding the set threshold of 0.8, and if there are multiple similar offside actions in the behavior path, it indicates that the node may be continuously replaced by a specific person with redundant permissions, indicating a strong risk of role avoidance or responsibility overlap.

[0070] like Figure 2 As shown, as a preferred embodiment of the present invention, the risk identification is performed based on the semantic vector, and by comparing and identifying whether the current behavior is within the authorized scope, if it is determined that the access behavior does not match the position authority, it is determined to be a potential unauthorized behavior, and the secondary verification step is performed, specifically including:

[0071] S201, performing risk identification based on semantic vectors, wherein the risk identification method is multi-level identification, which includes behavior scoring and path identification, and obtaining an identification result by comparing and identifying whether the current behavior is within the authorized scope.

[0072] In this step, risk identification is performed based on semantic vectors. After the office behavior semantic vector is generated, a multi-level risk identification process is carried out based on the vector. The identification process includes two core links: behavior scoring and path identification, which correspond to the compliance judgment of the single behavior itself and the analysis of the position and rationality of the behavior in the overall process chain. First, in the behavior scoring stage, the current behavior vector is matched with the position authority baseline vector, focusing on comparing whether the operation type, object resources and operation level are within the authorized scope. If the current behavior exceeds the standard boundary of the position in the authority matrix, it will be marked as a potential overstepping of authority, and the behavior will be assigned a basic risk score. At the same time, the dynamic weight adjustment is performed based on the contextual features such as the time period of the behavior, the location of the equipment, and the frequency of operation to form a complete behavior risk score.

[0073] After completing the behavior scoring, the path identification phase begins, analyzing the logical position of the behavior in the complete task chain or process sequence. By retrieving upstream and downstream operations under the same task instance or process number, it is determined whether the behavior conforms to the process evolution logic, whether it is at a reasonable node stage, and whether there are legal predecessor tasks as triggering basis. At the same time, by comparing the standard process template with the actual path sequence, it is identified whether there are cases where the path is bypassed, the node is replaced, or the responsibility flow is abnormal, and it is structurally determined whether the behavior is legally nested in the business process. Ultimately, the behavior scoring result and the path identification judgment jointly constitute the complete risk identification result of the behavior.

[0074] This multi-layered risk identification mechanism not only assesses behavioral compliance based on permissions but also assists in verifying the rationality and necessity of behavior within the context of process structure and semantics. This approach effectively improves the accuracy of identifying abnormal behavior and enhances the system's security identification capabilities in scenarios involving complex collaborative processes and multi-role cross-operations, providing a highly reliable data foundation for subsequent response strategy selection.

[0075] S202: If it is determined that the access behavior does not match the job authority and is not accompanied by any prior authorization operation, it is judged as a potential unauthorized behavior, and further search is performed to see whether there is a supplementary authorization record. The authorization record includes temporary release of project-based authority, dynamic authority inheritance generated in the collaborative approval mechanism, and proxy approval relationship.

[0076] In this step, if the access behavior is determined to be inconsistent with the position's permissions, and if the semantic vector doesn't detect any identifiable pre-authorization actions, such as approval process invocation, position adjustment, or task assignment, the behavior will be preliminarily identified as a "potential overreach." At this point, the behavior is not immediately classified as a violation, but instead enters the supplementary authorization verification phase to further verify the legitimacy and compliance of the behavior.

[0077] All supplementary authorization records related to the current behavior will be retrieved to check whether the operator has obtained dynamic authorization within the specified time window. Supplementary authorization mainly includes three situations: the first is the temporary release of project-based permissions, such as temporarily opening specific system module permissions within a limited period due to participation in a cross-departmental project; the second is the dynamic permission inheritance generated by the collaborative approval mechanism, that is, in the process collaboration scenario, the operator obtains phased access authorization through collaborative tasks; the third is the agency approval relationship, that is, in the business delegation or temporary agency scenario, the target operation is performed by other authorized employees, and the legality, timeliness and authorization content boundaries of the agency relationship will be verified.

[0078] By searching and matching these supplementary authorization records, we can determine whether the unauthorized behavior is supported by a reasonable authorization chain. If a match is found, the behavior is removed from the risk list or its risk level is adjusted. If no valid authorization record exists, the unauthorized access is confirmed, and the risk level, behavior context, and system module are recorded, and the subsequent policy response process is initiated. This mechanism ensures that unauthorized access identification not only relies on static permission matching but also has dynamic authorization adaptability, improving the system's ability to balance actual business flexibility with anomaly identification accuracy.

[0079] S203: If no valid pre-authorization or permission transfer chain is found, the behavior is marked as suspected unauthorized access, and a secondary verification is performed to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

[0080] In this step, if no valid pre-authorization or permission chain is found, and if the system fails to identify the access behavior as a legitimate operation within the scope of the position's authority during the initial comparison, and fails to identify any clear pre-authorization path or permission chain, such as project authorization, proxy authorization, or process node task trigger, the behavior will be marked as suspected unauthorized access. At this point, it will not be immediately judged as a violation, but will enter the secondary verification process, where the behavior will be contextualized and semantically matched within the context of the task process to further verify its legitimacy.

[0081] During the secondary verification process, first check whether the current operation is bound to an existing task process instance. By analyzing the process number, task ID and node information in the semantic vector, identify whether the behavior is a valid node operation in a business process. If the task ID attached to the behavior successfully matches the process record, it will further compare the current node status of the process, the execution conditions of the operation node and the process direction to confirm whether the current process really requires the user to perform the operation at this stage. If it is found that the behavior is in the normal evolution path of the process, and although the operator has not authorized it in the permission matrix, his behavior is allowed or assigned in the process rules, the system regards it as process-driven authorization, and the risk label will be downgraded or lifted.

[0082] If a clear correlation between the behavior and any process node cannot be identified, or if the process is in a state where the action should not be performed, such as ended, withdrawn, or redirected, the action will be further confirmed as an unfounded unauthorized violation, retaining the full context and triggering a risk control response. By introducing a secondary verification mechanism for process context, it is possible to accurately distinguish between abnormal business behavior and actual illegal operations, preventing the static permission model from misjudging the dynamic process authorization mechanism, while also enhancing the ability to identify process-driven unauthorized violations.

[0083] like Figure 3 As shown, as a preferred embodiment of the present invention, the steps of comparing the actual node sequence executed by each process instance with the template path one by one, examining whether the approver of each node matches the preset job role, and obtaining the frequency of changes in the approval node processors specifically include:

[0084] S301, obtaining a process behavior sequence, which is used to verify that the approval path, nodes, and approvers are consistent with the original configuration one by one, and to compare the actual node sequence executed by each process instance with the template path one by one.

[0085] In this step, we obtain the process behavior sequence. During the process audit, we first extract the execution trajectory of each process instance from the business system to generate a complete process behavior sequence. This behavior sequence uses nodes as the basic unit and includes fields such as node number, node type, operation time, processor identity, approval action, and node status change. It is chronologically sorted according to the actual execution order. By extracting this sequence, the system can fully reproduce the process's execution path, providing basic data support for subsequent structural comparison and role review.

[0086] The extracted sequence of actual process actions is then compared node by node with the standard template path bound to that process. This comparison not only examines the consistency of node order and number with the process structure, but also the matching of node function types, such as the correct distinction between review, approval, and countersignature steps. Furthermore, the system examines whether the person handling each node aligns with the role pre-defined in the template configuration. Any discrepancies between the approver's identity and pre-defined responsibilities, or if a node is redirected by an unauthorized person, will be flagged as a process deviation or role mismatch risk.

[0087] By comparing the process behavior sequence with the template path node by node, we can identify various abnormal process phenomena such as path deformation, node missing, and approver replacement, and then determine whether there is a risk of tampering, audit evasion, or responsibility evasion in the process execution.

[0088] S302, identify whether there are any missing nodes, deformed or skipped paths, and examine whether the approver of each node matches the preset job role. Determine whether the approver has the approval authority through the review. If there is no corresponding approval authority, it will be marked as a risk node with inconsistent identity.

[0089] In this step, we identify whether there are any missing nodes, deformed paths, or skipped instances. When performing a process compliance review, we conduct a structural-level comparative analysis of each actually executed node in the process instance to identify any abnormalities such as missing nodes, deformed paths, or skipped instances. Missing nodes refer to the omission of approval steps that should be in the template during process execution; path deformation refers to artificial adjustments to the process path, such as non-linear jumps, skipping of co-signing links, or disordered node order; skipping refers to a certain approval node that must be executed being directly jumped to the next node without any substantive processing. Based on the process template comparison algorithm, the node execution sequence and logical structure are verified one by one, and the accuracy of the identification is ensured in combination with the process version record.

[0090] Based on structural verification, the system further examines whether the person handling each approval node aligns with the position role associated with that node in the process template. This step combines the organizational structure chart, the position and authority mapping table, and the current permissions snapshot to verify whether the approver has the responsibilities and permissions for the corresponding position. If the approver is in an unauthorized position, bypasses the level of authority, or their permission system does not cover the required operational permissions for the current node, the system will mark them as a risk node with "identity mismatch" and assign a score based on the degree of deviation, which will be included in the process audit results.

[0091] Through the above-mentioned process structure consistency comparison and processor authority compliance review, it is not only possible to identify whether there are abnormal operations such as human detours and path avoidance during the process operation, but also to accurately judge the rationality of the approver in the organizational responsibilities and authority system, and effectively prevent and control structural approval risks such as process modification, authority abuse and role substitution.

[0092] S303: Obtain the frequency of changes in the approval node processor. If the frequency exceeds a threshold, it is determined that the process operation may be subject to manipulation intervention.

[0093] In this step, we obtain the frequency of changes in approval node handlers. Process behavior monitoring continuously tracks changes in handlers for each approval node, calculating the frequency of handler changes within the same node or process instance. This frequency is calculated by chronologically sorting and clustering the handlers associated with each approval action in the process instance log. By combining node numbers with operation times, we can accurately identify frequent changes in handlers. For nodes with multiple approver changes, we further analyze the change period, each approver's position, organizational hierarchy, and their relationship to the original approver to construct a complete approver change path map.

[0094] If the frequency of changes in handlers at a particular node exceeds a system-defined dynamic threshold, which is dynamically adjusted based on process importance, node sensitivity, and historical baseline behavior, a risk warning mechanism will be triggered, initially identifying suspected manipulation, interference, or circumvention of the process. Typical risk scenarios include bypassing the responsible person for approval, circumventing audits through job substitution, or exploiting redundant permissions to frequently switch executors within the same node to reduce audit sensitivity.

[0095] By continuously monitoring and dynamically modeling the frequency of changes in approval node handlers, we can not only detect abnormal operational control behaviors in a timely manner, but also further explore collaborative violation patterns by combining job structure and collaborative trajectories, thereby effectively identifying hidden operational risks where the process appears normal on the surface but is actually out of control, providing key judgment basis for process auditing and security management.

[0096] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of detecting whether the actual operator is consistent with the node set position based on the mapping relationship, triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and calculating the authority overlap specifically include:

[0097] S401: Establish a responsibility attribution mapping relationship for the approval node or task node. The mapping relationship is generated based on the organizational structure chart, job responsibility specifications, and process template configuration. The actual operator is detected based on the mapping relationship to see if the node setting position is consistent. If inconsistency is detected, it is marked as a responsibility deviation point.

[0098] In this step, a responsibility mapping relationship is established for approval or task nodes. During the process audit, a responsibility mapping relationship is established for each approval or task node. This mapping relationship is constructed from three pieces of data: first, the preset requirements for node positions in the process template, which clearly define which position or person within the scope of responsibility should perform the node; second, the organizational structure map, which reflects the operator's affiliation and management level in the current organizational structure; and third, the job responsibility specifications, which define the authority boundaries and operational responsibilities of each position in different business scenarios. Through this multi-source data-fusion responsibility mapping mechanism, the system can establish an accurate correlation between each actual operation and its intended responsible person.

[0099] When actual operation records are generated during the process, the system matches the operator's identity with the pre-defined position at that node. If the current handler is found not to be within the designated position or scope of responsibilities, and there is no delegation or authorization record to support the role substitution, the node is marked as a "responsibility deviation point." This mark not only serves as a local risk signal for process risk scoring but also serves as a core basis for further path analysis and role accountability audits, identifying any systemic trends of responsibility avoidance, unauthorized approvals, or process manipulation.

[0100] S402 triggers the responsibility jump analysis process, combines the operation behavior with the potential redundancy risk in the permission structure, and analyzes the repeated permissions of the same highly sensitive resource in multiple non-identical responsibility chains.

[0101] In this step, the responsibility jump analysis process is triggered. When it is detected that the actual operator of a certain approval node is inconsistent with his preset job responsibilities and has been marked as a responsibility deviation point, the responsibility jump analysis process will be triggered immediately. This process not only checks whether there is any offside behavior in the operation of the node, but also further integrates and analyzes the behavior with the redundancy risks that may exist in the authority structure. Specifically, the system extracts the permission set of the current operator and the original position of the node from the permission configuration table, and vector models the two in dimensions such as system modules, functional permissions, and data objects, and calculates the degree of permission overlap through algorithms such as cosine similarity or Jaccard coefficient. If the current operator has a high degree of overlap with the set position permissions, it will be further determined whether it is a responsibility substitution phenomenon caused by permission redundancy.

[0102] On this basis, the analysis scope is expanded to construct a permissions map for the highly sensitive resource, scanning for other employees not in the same chain of responsibility who possess similar access or operational permissions, with particular attention paid to duplicate permissions across departments or within the same level of responsibility. If multiple individuals from different lines of responsibility are found to have similar permissions for a highly sensitive resource, such as a financial system, core data tables, or approval interfaces, and there is no necessary collaborative logic to support actual business operations, the resource will be marked as a "high-risk object for duplicate authorization" and the user combination involved will be placed under key monitoring to prevent responsibility jumping, abuse of authority, or audit evasion caused by redundant permissions.

[0103] S403: Calculate the authority overlap, obtain the overlap threshold, and determine the authority overlap according to the overlap threshold.

[0104] In this step, the degree of permission overlap is calculated. When reviewing the permission structure, the system abstracts the set of permissions for each employee or position into a vector model, constructing a unified permission feature space based on resource dimensions (such as system modules, functional interfaces, data tables, and approval actions). Each employee's permission is encoded as a vector in this space, with each dimension representing the access level or operational capability for a specific resource. For permission-based systems (such as read-only, edit, approve, and delete), the system assigns different numerical weights to the corresponding dimensions, creating a more expressive permission vector.

[0105] When determining whether there is a risk of overlapping permissions between two employees or positions, the system calculates the similarity between their permission vectors using algorithms such as cosine similarity and the Jaccard coefficient, resulting in a quantitative "permission overlap" metric. This degree of permission overlap reflects the actual degree of similarity between the two in terms of functional overlap and resource sharing, with higher values ​​indicating closer permission structures. The system pre-sets a dynamic overlap threshold (such as 0.75 or 0.8) based on historical data, job responsibility models, and security policies. This threshold automatically adjusts based on business sensitivity, system importance, or position level.

[0106] Once the calculated result exceeds this overlap threshold, the system determines "high degree of overlap" and, based on organizational relationships, further assesses whether there is a potential risk of overlapping responsibilities, collaborative overreach, or responsibility transfer. If the overlapping individuals are from different responsibility chains or non-collaborative positions, and the overlapping resources are highly sensitive, the system will flag this as a permission configuration anomaly, providing a foundation for clearing permissions, constraining collaboration, and responding to risks.

[0107] like Figure 5 As shown, an embodiment of the present invention provides a big data-based office vulnerability analysis system, the system comprising:

[0108] The data collection module 100 is used to collect multi-source heterogeneous office data and process the multi-source heterogeneous office data. The processing methods include feature extraction, labeling and time synchronization, and obtain the processed unified office behavior semantic vector. The semantic vector is used to provide unified data support for subsequent analysis.

[0109] In this system, the data acquisition module 100 collects multi-source, heterogeneous office data. To achieve unified modeling and accurate analysis of complex office behaviors, the system first collects heterogeneous data from various sources, including OA systems, ERP systems, document management platforms, permission systems, instant messaging tools, and physical access control devices. These data types vary significantly in format, structure, and time. This data is standardized through the interface adaptation and normalization module, extracting key features such as operation type, target object, operator identity, operation location, and device information. This unified feature mapping is achieved by combining rule templates with behavioral semantic models.

[0110] On this basis, the system labels the extracted behavioral data, assigning multiple levels of risk or operational attribute labels based on behavioral semantics, sensitivity, temporal characteristics, and contextual logic, such as "unauthorized access," "process bypass," and "sensitive document manipulation," enhancing the semantic integrity and interpretability of behavioral expressions. A time synchronization mechanism is also introduced, using NTP calibration and sliding window aggregation to ensure consistency and comparability of multi-source behavioral events on the timeline, avoiding analytical bias caused by data time drift.

[0111] Ultimately, all processed behavioral data is encoded into a unified office behavior semantic vector, integrating structured features, label attributes, and temporal context, providing a high-quality input foundation for subsequent graph modeling, path analysis, and risk identification. This mechanism not only enables the integration of cross-system and cross-terminal behavioral data, but also provides strong data support for refined behavior identification and chained risk analysis in office scenarios.

[0112] The risk identification module 200 is used to identify risks based on semantic vectors. It compares and identifies whether the current behavior is within the authorized scope. If it is determined that the access behavior does not match the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed.

[0113] In this system, the risk identification module 200 performs risk identification based on semantic vectors. After completing the generation of office behavior semantic vectors, it first performs risk identification on each semantic vector. One of the most core judgment logics is whether the access behavior is within the authorized scope. A permission baseline model is constructed based on the position authority matrix. The standard operation permissions of each position in each business system, functional module and data object are encoded in vector form and compared one by one with the fields such as operation object, operation type, operator identity, etc. contained in the behavior semantic vector. If it is found that the system modules or data resources involved in the current behavior exceed the permitted scope of the position in the authority matrix, it will be preliminarily marked as "suspected unauthorized behavior."

[0114] Based on this, the behavior is not immediately deemed a violation. Instead, it enters a secondary verification mechanism, combining organizational authorization logs, process task context, and temporary permission records to further assess the legitimacy of the behavior. The system checks whether the behavior has a legitimate pre-authorization path, such as position transfer authorization, collaborative task assignment, or approval delegation. It also analyzes whether the current behavior is nested within a legitimate process node or automatically triggered by a process task. If there are no authorization records, responsibility matching paths, or process linkage evidence, the behavior will ultimately be confirmed as unauthorized access, triggering an alert, log flag, or policy response action.

[0115] This mechanism not only efficiently identifies abnormal behavior outside of the scope of permissions, but also effectively avoids false positives through contextual linkage and permission chain verification, improving the accuracy and business adaptability of unauthorized identification. This process fully demonstrates the value of semantic vectors as the core carrier of behavioral analysis in permission identification, achieving an intelligent evolution from static permission configuration to dynamic behavioral compliance.

[0116] The process approval module 300 is used to compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of replacement of the approval node processor.

[0117] In this system, the process approval module 300 compares the actual node sequence executed by each process instance with the template path one by one. When identifying abnormal behavior in process approval, each process instance is used as an analysis unit, and the node sequence executed by the instance in actual operation is first extracted. These nodes contain detailed information such as node number, node type, processor, operation time, approval action, and approval result. The actual path is then structurally compared with the standard path preset in the process template to check whether there are structural deviations such as node skipping, abnormal node sequence, process interruption, or path merging. This comparison process is implemented through a node alignment algorithm, which can identify the degree of difference between the execution path and the standard model, and preliminarily determine whether there is a risk of abnormal process direction or path deformation.

[0118] After completing the path structure comparison, each approval node is further verified to ensure that the actual processor matches the role specified in the process template. This comparison is based on the position authority mapping table and the organizational structure diagram. It analyzes whether the processor belongs to the target position, has the corresponding responsibilities, and whether the processor is an authorized agent or bypasses the level of authority. If the processor does not directly correspond to the configured position, the system will mark the node as "identity mismatch" and assign a corresponding risk weight.

[0119] The attribution mapping module 400 is used to detect whether the actual operator is consistent with the node set position based on the mapping relationship, trigger the responsibility jump analysis process, combine the operation behavior with the potential redundancy risk in the authority structure, and calculate the authority overlap.

[0120] In this system, the attribution mapping module 400 detects whether the actual operator is consistent with the node setting position based on the mapping relationship. When performing the process node responsibility attribution review, the actual operator and the node setting position will be matched one by one based on the mapping relationship between each approval node and job responsibilities in the process template. Through the authority configuration table and the organizational structure map, it is determined whether the current handler belongs to the position range specified by the node and whether he has the corresponding responsibility authority. If it is found that the handler is inconsistent with the set position and there is a lack of authorization, transfer records or process role change basis, it will be judged as "responsibility jump suspicion" and the responsibility jump analysis process will be triggered immediately.

[0121] In responsibility transition analysis, we not only focus on the deviation between the handler's responsibilities and the node configuration, but also further analyze the linkage between this operational behavior and the potential redundancy risks in the permission structure. We construct a permission overlap calculation model, abstracting the permission set corresponding to the current handler and the designated position into a vector form. If there is a high degree of overlap between the two, such as exceeding the set threshold of 0.8, and if there are multiple similar offside actions in the behavior path, it indicates that the node may be continuously replaced by a specific person with redundant permissions, indicating a strong risk of role avoidance or responsibility overlap.

[0122] like Figure 6 As shown, as a preferred embodiment of the present invention, the risk identification module 200 includes:

[0123] The risk identification unit 201 is used to perform risk identification based on semantic vectors. The risk identification method is multi-level identification, which includes behavior scoring and path identification. The identification result is obtained by comparing and identifying whether the current behavior is within the authorized scope.

[0124] In this module, the risk identification unit 201 performs risk identification based on the semantic vector. After generating the office behavior semantic vector, a multi-level risk identification process is carried out based on the vector. The identification process includes two core links: behavior scoring and path identification, which correspond to the compliance judgment of the single behavior itself and the analysis of the position and rationality of the behavior in the overall process chain. First, in the behavior scoring stage, the current behavior vector is matched with the position authority baseline vector, focusing on comparing whether the operation type, object resources and operation level are within the authorized scope. If the current behavior exceeds the standard boundary of the position in the authority matrix, it will be marked as a potential overstepping of authority, and the behavior will be assigned a basic risk score. At the same time, the dynamic weight correction is performed based on the contextual features such as the time period of the behavior, the location of the equipment, and the frequency of operation to form a complete behavior risk score.

[0125] After completing the behavior scoring, the path identification phase begins, analyzing the logical position of the behavior in the complete task chain or process sequence. By retrieving upstream and downstream operations under the same task instance or process number, it is determined whether the behavior conforms to the process evolution logic, whether it is at a reasonable node stage, and whether there are legal predecessor tasks as triggering basis. At the same time, by comparing the standard process template with the actual path sequence, it is identified whether there are cases where the path is bypassed, the node is replaced, or the responsibility flow is abnormal, and it is structurally determined whether the behavior is legally nested in the business process. Ultimately, the behavior scoring result and the path identification judgment jointly constitute the complete risk identification result of the behavior.

[0126] This multi-layered risk identification mechanism not only assesses behavioral compliance based on permissions but also assists in verifying the rationality and necessity of behavior within the context of process structure and semantics. This approach effectively improves the accuracy of identifying abnormal behavior and enhances the system's security identification capabilities in scenarios involving complex collaborative processes and multi-role cross-operations, providing a highly reliable data foundation for subsequent response strategy selection.

[0127] The position authority unit 202 is used to judge the access behavior as a potential unauthorized behavior if it is determined that the access behavior does not match the position authority and is not accompanied by any prior authorization operation, and further search whether there is a supplementary authorization record. The authorization record includes the temporary release of project-based authority, dynamic authority inheritance generated in the collaborative approval mechanism, and the proxy approval relationship.

[0128] In this module, if the position authority unit 202 determines that an access behavior does not match the position authority, when an access behavior is determined to be outside the standard authority range of the operator's position and no identifiable pre-authorization operations, such as approval process invocation, position adjustment, or task assignment, are detected in the semantic vector, the behavior will be preliminarily identified as a "potential unauthorized behavior." At this point, the behavior is not immediately classified as a violation, but instead enters the supplementary authorization verification phase to further verify the rationality and compliance of the behavior.

[0129] All supplementary authorization records related to the current behavior will be retrieved to check whether the operator has obtained dynamic authorization within the specified time window. Supplementary authorization mainly includes three situations: the first is the temporary release of project-based permissions, such as temporarily opening specific system module permissions within a limited period due to participation in a cross-departmental project; the second is the dynamic permission inheritance generated by the collaborative approval mechanism, that is, in the process collaboration scenario, the operator obtains phased access authorization through collaborative tasks; the third is the agency approval relationship, that is, in the business delegation or temporary agency scenario, the target operation is performed by other authorized employees, and the legality, timeliness and authorization content boundaries of the agency relationship will be verified.

[0130] By searching and matching these supplementary authorization records, we can determine whether the unauthorized behavior is supported by a reasonable authorization chain. If a match is found, the behavior is removed from the risk list or its risk level is adjusted. If no valid authorization record exists, the unauthorized access is confirmed, and the risk level, behavior context, and system module are recorded, and the subsequent policy response process is initiated. This mechanism ensures that unauthorized access identification not only relies on static permission matching but also has dynamic authorization adaptability, improving the system's ability to balance actual business flexibility with anomaly identification accuracy.

[0131] The secondary verification unit 203 is used to mark the behavior as suspected unauthorized access if no valid pre-authorization or permission transfer chain is found, and then perform secondary verification to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

[0132] In this module, if the secondary verification unit 203 does not find a valid pre-authorization or permission transfer chain, and the system does not find the access behavior to be a legitimate operation within the scope of the job authority in the initial comparison, and fails to identify any clear pre-authorization path or permission transfer chain, such as project authorization, proxy authorization, or process node task trigger, the behavior will be marked as suspected unauthorized access. At this time, it will not be immediately judged as an illegal operation, but will enter the secondary verification process, and the behavior will be contextualized and semantically matched within the context of the task process to further verify its rationality.

[0133] During the secondary verification process, first check whether the current operation is bound to an existing task process instance. By analyzing the process number, task ID and node information in the semantic vector, identify whether the behavior is a valid node operation in a business process. If the task ID attached to the behavior successfully matches the process record, it will further compare the current node status of the process, the execution conditions of the operation node and the process direction to confirm whether the current process really requires the user to perform the operation at this stage. If it is found that the behavior is in the normal evolution path of the process, and although the operator has not authorized it in the permission matrix, his behavior is allowed or assigned in the process rules, the system regards it as process-driven authorization, and the risk label will be downgraded or lifted.

[0134] If a clear correlation between the behavior and any process node cannot be identified, or if the process is in a state where the action should not be performed, such as ended, withdrawn, or redirected, the action will be further confirmed as an unfounded unauthorized violation, retaining the full context and triggering a risk control response. By introducing a secondary verification mechanism for process context, it is possible to accurately distinguish between abnormal business behavior and actual illegal operations, preventing the static permission model from misjudging the dynamic process authorization mechanism, while also enhancing the ability to identify process-driven unauthorized violations.

[0135] like Figure 7 As shown, as a preferred embodiment of the present invention, the process approval module 300 includes:

[0136] The process behavior sequence unit 301 is used to obtain the process behavior sequence, which is used to verify the consistency of the approval path, nodes, and approvers with the original configuration one by one, and compare the actual node sequence executed by each process instance with the template path one by one.

[0137] In this module, the process behavior sequence unit 301 obtains the process behavior sequence. During the process audit, the execution trajectory of each process instance is first extracted from the business system to generate a complete process behavior sequence. This behavior sequence is based on nodes and includes fields such as node number, node type, operation time, processor identity, approval action, and node status change, sorted by time according to the actual execution order. By extracting this sequence, the system can fully reproduce the process's execution path, providing basic data support for subsequent structural comparison and role review.

[0138] The extracted sequence of actual process actions is then compared node by node with the standard template path bound to that process. This comparison not only examines the consistency of node order and number with the process structure, but also the matching of node function types, such as the correct distinction between review, approval, and countersignature steps. Furthermore, the system examines whether the person handling each node aligns with the role pre-defined in the template configuration. Any discrepancies between the approver's identity and pre-defined responsibilities, or if a node is redirected by an unauthorized person, will be flagged as a process deviation or role mismatch risk.

[0139] By comparing the process behavior sequence with the template path node by node, we can identify various abnormal process phenomena such as path deformation, node missing, and approver replacement, and then determine whether there is a risk of tampering, audit evasion, or responsibility evasion in the process execution.

[0140] The process approval unit 302 is used to identify whether there are missing nodes, deformed paths or skipped paths, and to examine whether the approver of each node matches the preset job role. It is then determined whether the approver has the approval authority. If there is no corresponding approval authority, the node is marked as a risk node with inconsistent identity.

[0141] In this module, the process approval unit 302 identifies whether there are any missing nodes, deformed paths, or skipped nodes. When performing a process compliance review, a structural-level comparative analysis will be performed on each actually executed node in the process instance to identify any abnormalities such as missing nodes, deformed paths, or skipped nodes. Node missing refers to the omission of the approval steps that should be in the template during process execution; path deformation is manifested as the process path being artificially adjusted, such as non-linear jumps, skipping of the co-signing link, or disordered node order; skipping refers to a certain approval node that must be executed being directly jumped to the next node without any substantive processing. Based on the process template comparison algorithm, the node execution sequence and logical structure are verified one by one, and the accuracy of the identification is ensured in combination with the process version record.

[0142] Based on structural verification, the system further examines whether the person handling each approval node aligns with the position role associated with that node in the process template. This step combines the organizational structure chart, the position and authority mapping table, and the current permissions snapshot to verify whether the approver has the responsibilities and permissions for the corresponding position. If the approver is in an unauthorized position, bypasses the level of authority, or their permission system does not cover the required operational permissions for the current node, the system will mark them as a risk node with "identity mismatch" and assign a score based on the degree of deviation, which will be included in the process audit results.

[0143] Through the above-mentioned process structure consistency comparison and processor authority compliance review, it is not only possible to identify whether there are abnormal operations such as human detours and path avoidance during the process operation, but also to accurately judge the rationality of the approver in the organizational responsibilities and authority system, and effectively prevent and control structural approval risks such as process modification, authority abuse and role substitution.

[0144] The change frequency unit 303 is used to obtain the change frequency of the approval node processor. If the change frequency exceeds a threshold, it is determined that the process operation may be subject to manipulation intervention.

[0145] In this module, the change frequency unit 303 obtains the frequency of changes in the approval node's handlers. During process behavior monitoring, changes in handlers for each approval node are continuously tracked, and the frequency of handler changes within the same node or process instance is calculated. This frequency is calculated by chronologically sorting and clustering the handlers for each approval action in the process instance log. By combining node numbers and operation times, it accurately identifies whether there are frequent changes in handler behavior. For nodes with multiple approver switches, the change period, each approver's position, organizational hierarchy, and their relationship with the original approver are further extracted to construct a complete approver change path map.

[0146] If the frequency of changes in handlers at a particular node exceeds a system-defined dynamic threshold, which is dynamically adjusted based on process importance, node sensitivity, and historical baseline behavior, a risk warning mechanism will be triggered, initially identifying suspected manipulation, interference, or circumvention of the process. Typical risk scenarios include bypassing the responsible person for approval, circumventing audits through job substitution, or exploiting redundant permissions to frequently switch executors within the same node to reduce audit sensitivity.

[0147] By continuously monitoring and dynamically modeling the frequency of changes in approval node handlers, we can not only detect abnormal operational control behaviors in a timely manner, but also further explore collaborative violation patterns by combining job structure and collaborative trajectories, thereby effectively identifying hidden operational risks where the process appears normal on the surface but is actually out of control, providing key judgment basis for process auditing and security management.

[0148] like Figure 8 As shown, as a preferred embodiment of the present invention, the attribution mapping module 400 includes:

[0149] The attribution mapping unit 401 is used to establish a responsibility attribution mapping relationship for the approval node or task node. The mapping relationship is generated based on the organizational structure chart, job responsibility specifications and process template configuration. The mapping relationship is used to detect whether the actual operator is consistent with the node set position. If inconsistency is detected, it is marked as a responsibility deviation point.

[0150] In this module, the attribution mapping unit 401 establishes a responsibility attribution mapping relationship for approval nodes or task nodes. During the process audit, a responsibility attribution mapping relationship is established for each approval node or task node. This mapping relationship is constructed by combining three pieces of data: first, the preset requirements for node positions in the process template, which clearly define which position or person within the scope of responsibility should execute the node; second, the organizational structure map, which reflects the operator's affiliation and management level in the current organizational structure; and third, the position responsibility specifications, which define the authority boundaries and operational responsibilities of each position in different business scenarios. Through this multi-source data-fusion responsibility mapping mechanism, the system can establish an accurate correlation between each actual operation and its intended responsible person.

[0151] When actual operation records are generated during the process, the system matches the operator's identity with the pre-defined position at that node. If the current handler is found not to be within the designated position or scope of responsibilities, and there is no delegation or authorization record to support the role substitution, the node is marked as a "responsibility deviation point." This mark not only serves as a local risk signal for process risk scoring but also serves as a core basis for further path analysis and role accountability audits, identifying any systemic trends of responsibility avoidance, unauthorized approvals, or process manipulation.

[0152] The responsibility analysis unit 402 is used to trigger the responsibility jump analysis process, combine the operation behavior with the potential redundancy risk in the authority structure, and analyze the repeated permissions of the same highly sensitive resource on multiple non-identical responsibility chains.

[0153] In this module, the responsibility analysis unit 402 triggers the responsibility jump analysis process. When it detects that the actual operator of a certain approval node is inconsistent with his preset job responsibilities and has been marked as a responsibility deviation point, the responsibility jump analysis process will be triggered immediately. This process not only checks whether there is any offside behavior in the operation of the node, but also further integrates and analyzes the behavior with the redundancy risks that may exist in the authority structure. Specifically, the system extracts the authority set of the current operator and the original position of the node from the authority configuration table, and vector models the two in dimensions such as system modules, functional permissions, and data objects, and calculates the authority overlap through algorithms such as cosine similarity or Jaccard coefficient. If the current operator has a high degree of overlap with the set position authority, it will be further determined whether it is a responsibility substitution phenomenon caused by authority redundancy.

[0154] On this basis, the analysis scope is expanded to construct a permissions map for the highly sensitive resource, scanning for other employees not in the same chain of responsibility who possess similar access or operational permissions, with particular attention paid to duplicate permissions across departments or within the same level of responsibility. If multiple individuals from different lines of responsibility are found to have similar permissions for a highly sensitive resource, such as a financial system, core data tables, or approval interfaces, and there is no necessary collaborative logic to support actual business operations, the resource will be marked as a "high-risk object for duplicate authorization" and the user combination involved will be placed under key monitoring to prevent responsibility jumping, abuse of authority, or audit evasion caused by redundant permissions.

[0155] The authority overlap unit 403 is used to calculate the authority overlap, obtain the overlap threshold, and determine the authority overlap situation according to the overlap threshold.

[0156] In this module, the permission overlap unit 403 calculates the degree of permission overlap. When reviewing the permission structure, the system abstracts the permission set for each employee or position into a vector model, constructing a unified permission feature space based on resource dimensions (such as system modules, functional interfaces, data tables, and approval actions). Each employee's permission is encoded as a vector in this space, with each dimension representing the access level or operational capability for a specific resource. For a privileged permission system (such as read-only, edit, approve, and delete), the system assigns different numerical weights to the corresponding dimensions, thereby forming a more expressive permission vector.

[0157] When determining whether there is a risk of overlapping permissions between two employees or positions, the system calculates the similarity between their permission vectors using algorithms such as cosine similarity and the Jaccard coefficient, resulting in a quantitative "permission overlap" metric. This degree of permission overlap reflects the actual degree of similarity between the two in terms of functional overlap and resource sharing, with higher values ​​indicating closer permission structures. The system pre-sets a dynamic overlap threshold (such as 0.75 or 0.8) based on historical data, job responsibility models, and security policies. This threshold automatically adjusts based on business sensitivity, system importance, or position level.

[0158] Once the calculated result exceeds this overlap threshold, the system determines "high degree of overlap" and, based on organizational relationships, further assesses whether there is a potential risk of overlapping responsibilities, collaborative overreach, or responsibility transfer. If the overlapping individuals are from different responsibility chains or non-collaborative positions, and the overlapping resources are highly sensitive, the system will flag this as a permission configuration anomaly, providing a foundation for clearing permissions, constraining collaboration, and responding to risks.

[0159] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are performed:

[0160] Collect and process multi-source heterogeneous office data using methods including feature extraction, labeling, and time synchronization to obtain a unified semantic vector of office behavior. This semantic vector is used to provide unified data support for subsequent analysis.

[0161] Risk identification is performed based on semantic vectors. By comparing and identifying whether the current behavior is within the authorized scope, if the access behavior is determined to be inconsistent with the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed.

[0162] Compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of changes in the approval node processor;

[0163] Based on the mapping relationship, it is detected whether the actual operator is consistent with the node setting position, triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and calculating the authority overlap.

[0164] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps:

[0165] Collect and process multi-source heterogeneous office data using methods including feature extraction, labeling, and time synchronization to obtain a unified semantic vector of office behavior. This semantic vector is used to provide unified data support for subsequent analysis.

[0166] Risk identification is performed based on semantic vectors. By comparing and identifying whether the current behavior is within the authorized scope, if the access behavior is determined to be inconsistent with the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed.

[0167] Compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of changes in the approval node processor;

[0168] Based on the mapping relationship, it is detected whether the actual operator is consistent with the node setting position, triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and calculating the authority overlap.

[0169] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0170] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0171] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0172] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

[0173] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for analyzing office vulnerabilities based on big data, characterized in that: The method comprises: Collect and process multi-source heterogeneous office data using methods including feature extraction, labeling, and time synchronization to obtain a unified semantic vector of office behavior. This semantic vector is used to provide unified data support for subsequent analysis. Risk identification is performed based on semantic vectors. By comparing and identifying whether the current behavior is within the authorized scope, if the access behavior is determined to be inconsistent with the job authority, it is judged as a potential unauthorized behavior and a secondary verification is performed. Compare the actual node sequence executed by each process instance with the template path one by one, check whether the approver of each node matches the preset job role, and obtain the frequency of changes in the approval node processor; Specifically, it includes: obtaining the process behavior sequence, which is used to verify the consistency of the approval path, nodes, and approvers with the original configuration one by one, and comparing the actual node sequence executed by each process instance with the template path one by one; identifying whether there are missing nodes, deformed paths, or skipped paths, and reviewing whether the approver of each node matches the preset job role. Through review, it is determined whether the approver has the approval authority. If the approver does not have the corresponding approval authority, the node is marked as a risk node with inconsistent identity; obtaining the frequency of changes in the approval node processor. If the change frequency exceeds the threshold, it is determined that the process operation may be subject to manipulation and intervention; Based on the mapping relationship, the actual operator is detected to see if the node's set position is consistent. This triggers the responsibility jump analysis process, combines the operation behavior with the potential redundancy risk in the authority structure, and calculates the authority overlap. Specifically, it includes: establishing a responsibility attribution mapping relationship for approval nodes or task nodes, the mapping relationship is generated based on the organizational structure map, job responsibility specifications and process template configuration, and detecting whether the actual operator is consistent with the node set position based on the mapping relationship. If inconsistency is detected, it is marked as a responsibility deviation point; triggering the responsibility jump analysis process, combining the operation behavior with the potential redundancy risk in the authority structure for analysis, and analyzing the repeated permissions of the same highly sensitive resources on multiple non-identical responsibility chains; calculating the degree of permission overlap, obtaining the overlap threshold, and determining the permission overlap based on the overlap threshold.

2. The method for analyzing office vulnerabilities based on big data according to claim 1, characterized in that: The risk identification based on semantic vectors is performed by comparing and identifying whether the current behavior is within the authorized scope. If the access behavior is determined to be inconsistent with the job authority, it is judged as a potential unauthorized behavior and a secondary verification step is performed, specifically including: Risk identification is performed based on semantic vectors. The risk identification method is multi-level identification, which includes behavior scoring and path identification. The identification result is obtained by comparing and identifying whether the current behavior is within the authorized scope; If the access behavior is determined to be inconsistent with the position's permissions and is not accompanied by any prior authorization operations, it is considered a potential violation of authority and a further search is conducted to determine whether there are supplementary authorization records. Such authorization records include temporary release of project-based permissions, dynamic permission inheritance generated by the collaborative approval mechanism, and proxy approval relationships. If no valid prior authorization or permission transfer chain is found, the behavior will be marked as suspected unauthorized access, and a secondary verification will be performed to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

3. The method for analyzing office vulnerabilities based on big data according to claim 1, characterized in that: The labeling is to add multiple label dimensions to the data through the label template system before it is stored in the warehouse, and the time synchronization is to unify the time reference of data timestamps in different systems.

4. A big data-based office vulnerability analysis system, characterized in that: The system comprises: The data acquisition module collects and processes multi-source heterogeneous office data through feature extraction, labeling, and time synchronization. The module then obtains a unified semantic vector of office behavior, which provides unified data support for subsequent analysis. The risk identification module identifies risks based on semantic vectors and compares the current behavior to see if it is within the authorized scope. If the access behavior is determined to be inconsistent with the job authority, it is considered a potential unauthorized behavior and undergoes secondary verification. The process approval module compares the actual node sequence executed by each process instance with the template path one by one, checks whether the approver of each node matches the preset job role, and obtains the frequency of changes in the approval node processor; Specifically, it includes: a process behavior sequence unit, which obtains the process behavior sequence. The process behavior sequence is used to verify the consistency of the approval path, nodes, and approvers with the original configuration one by one, and compare the actual node sequence executed by each process instance with the template path one by one; a process approval unit, which identifies whether there are missing nodes, deformed paths, or skipped paths, and examines whether the approver of each node matches the preset job role. Through the review, it determines whether the approver has the approval authority. If the approver does not have the corresponding approval authority, the node is marked as a risk node with inconsistent identity; a change frequency unit, which obtains the change frequency of the approval node processor. If the change frequency exceeds the threshold, it is determined that the process operation may be subject to manipulation and intervention; The attribution mapping module detects whether the actual operator is consistent with the node's set position based on the mapping relationship, triggers the responsibility jump analysis process, combines the operation behavior with the potential redundancy risk in the authority structure, and calculates the authority overlap; Specifically, it includes: an attribution mapping unit, which establishes a responsibility attribution mapping relationship for the approval node or task node. The mapping relationship is generated based on the organizational structure map, job responsibility specifications and process template configuration, and detects whether the actual operator is consistent with the node set position based on the mapping relationship. If inconsistency is detected, it is marked as a responsibility deviation point; a responsibility analysis unit, which triggers the responsibility jump analysis process, combines the operation behavior with the potential redundant risks in the authority structure for analysis, and analyzes the repeated permissions of the same highly sensitive resources on multiple non-identical responsibility chains; a permission overlap unit, which calculates its permission overlap, obtains the overlap threshold, and determines the permission overlap based on the overlap threshold.

5. The big data-based office vulnerability analysis system according to claim 4 is characterized in that: The risk identification module includes: The risk identification unit performs risk identification based on semantic vectors. The risk identification method is multi-level identification, which includes behavior scoring and path identification. The identification result is obtained by comparing and identifying whether the current behavior is within the authorized scope; If the position authority unit determines that the access behavior does not match the position authority and is not accompanied by any pre-authorization operation, it is judged as a potential unauthorized behavior and further searches for supplementary authorization records. The authorization records include temporary release of project-based permissions, dynamic permission inheritance generated by the collaborative approval mechanism, and proxy approval relationships. The secondary verification unit, if no valid pre-authorization or permission transfer chain is found, will mark the behavior as suspected unauthorized access and perform secondary verification to determine whether the operation is part of a task process and whether the current process status requires the user to perform this operation.

6. The big data-based office vulnerability analysis system according to claim 5, characterized in that: The labeling is to add multiple label dimensions to the data through the label template system before it is stored in the warehouse, and the time synchronization is to unify the time reference of data timestamps in different systems.

Citation Information

Patent Citations

  • Data security capability detection method and system

    CN118427843A

  • Government affair data sharing system based on data security law risk control mode

    CN119989417A

  • Financial data security regulation and control management method and system

    CN119991046A