Business process multi-view deviation detection method and system based on privacy protection

By establishing a conditional random field CRF model and a privacy access control model, integrating data flow and control flow to identify deviations in business processes, the problems of low data utilization and privacy leakage in the existing technology are solved, and efficient deviation detection and privacy protection are achieved.

CN120296792AInactive Publication Date: 2025-07-11ANHUI WATER CONSERVANCY TECHN COLLEGE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510454853.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing deviation detection methods are difficult to maximize data utilization while ensuring data privacy and security, and there is a risk of data leakage of private information.

Method used

A multi-view deviation detection method for business processes based on privacy protection is adopted, and a conditional random field CRF model is established to identify named entities, combining the privacy access control model and data decision logic, data flow and control flow are integrated to identify deviations in business processes.

Benefits of technology

It realizes that under the conditions of protecting data privacy, maximizes data utilization, improves the comprehensiveness and accuracy of deviation detection, detects deviations in data flow and control flow, and reduces false positives and data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296792A_ABST
    Figure CN120296792A_ABST
Patent Text Reader

Abstract

The invention discloses a business process multi-view deviation detection method and system based on privacy protection, and the method comprises the steps: building a conditional random field CRF model for named entity recognition, and extracting important data for business process research; establishing a privacy access control model based on identity and destination to obtain a data mode with privacy protection; a business process is sensed by adopting a data decision based on fusion of a data flow and a control flow, and deviation in the business process is identified by monitoring activity, resources, data and composite movement classification of dimensions corresponding to data operation and event logs in the business process and combining business decision logic analysis. By means of the method, under the condition that data privacy security is guaranteed, the data utilization rate is maximized, and deviation detection is conducted on the business process through multiple perspectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to process mining technology, and more particularly to a multi - perspective deviation detection method and system for business processes based on privacy protection. Background Art

[0002] Process mining technology can extract valuable information from event logs commonly generated by modern information systems. This technology provides new means for process discovery, monitoring, and improvement in various application fields. In process mining, the Directly Follows Mining method is an important technique that is used to mine the directly - follows relationships between activities from event logs. The core of this method lies in identifying which activities often occur consecutively in the process, thereby revealing the structure and behavior patterns of the process. However, due to insufficient information system design or system upgrades, the event logs generated by these systems may be inconsistent with the existing models. We call this kind of inconsistency deviation. Detecting this deviation can verify and extend business process models and accordingly improve business processes. In addition, deviation detection is also a popular topic in enterprises because its application fields are diverse, such as fraud detection, intrusion detection, and abnormal transaction detection in e - commerce, etc.

[0003] Existing anomaly detection methods mainly focus on control flow, point anomalies, and strive to avoid false alarms when unexpected events occur. In the past few years, there have been more and more transformations in the design, engineering, and mining of processes, from a purely control - flow perspective to a more integrated model, in which data and decisions are also explicitly considered. In fact, it is difficult to distinguish between legal and illegal behaviors without knowing the context of data access. In the digital economy era, data opening has become an inevitable trend. In recent years, under the framework of process mining, a large number of methods have been developed to analyze event logs. However, the impact of privacy rules on the design and organizational application of process mining technology has been largely ignored, resulting in the irresponsible use of personal data. For example, healthcare information systems contain highly sensitive information, and healthcare regulations usually require protecting data privacy. Complying with strict privacy requirements may lead to a decrease in the utility of data used for analysis. Introducing data information in the business process deviation detection process can improve the accuracy of deviation detection. However, these data information often contains personal sensitive information. Data security and privacy protection issues are inevitable during the data opening process. It is of great significance to explore how to achieve a balance between maximizing data utilization and privacy protection during the deviation detection process. However, there is less research in this direction currently. Existing methods either focus on the data view and check data access according to security policies, or focus on checking the deviation of activities according to the activities required to execute the business process. However, analyzing user behavior from these perspectives alone will not only lead to some deviations not being detected or false positives, but also disclose the privacy data of data providers.

[0004] Deviation detection has become a key research project in many business processes because it can help enterprises prevent fraud and monitor anomalies. However, existing deviation detection methods either only focus on the deviation between control flow activities and cannot detect the deviation caused by data, or introduce too much data for deviation detection and disclose personal privacy information. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to maximize data utilization rate under the condition of ensuring data privacy and security, and conduct deviation detection on business processes from multiple perspectives.

[0006] The present invention realizes the solution of the above technical problems through the following technical means:

[0007] The present invention provides a multi-perspective deviation detection method for business processes based on privacy protection, including the following steps:

[0008] S1. Establish a conditional random field (CRF) model for named entity recognition, select a feature set from complex and diverse unstructured data information by using compound entity construction rules, and extract data important for business process research;

[0009] S2. Establish a privacy access control model based on identity and purpose, authenticate the data visitors and set access permissions to obtain a data pattern with privacy protection;

[0010] S3. Adopt data decision-making based on the fusion of data flow and control flow to perceive business processes, identify deviations in business processes through monitoring the compound movement classification of activities, resources, data, and data operations and event logs in business processes corresponding dimensions, and combining with business decision logic analysis.

[0011] Further, the establishment of the conditional random field (CRF) model in S1 is specifically as follows:

[0012] Assume that X represents a data sequence and Y represents entity categories, where x = {x1, x2, x3... x n} is an observation sequence, and y = {y1, y2, y3... y n} is a state sequence. P(y|x) represents the conditional probability distribution of outputting y under the condition of given x; CRF uses a first-order linear chain random field model to represent as follows

[0013]

[0014] Among them, t k , s l represents a feature function, and λ k , μ l represents the corresponding weight. Z(x) represents a normalization factor used to constrain the conditional probability;

[0015] The CRF model is used for named entity recognition, regarded as a sequence labeling problem; each sentence to be recognized is used as an observation sequence, and each word in the sentence is used as a symbol, and a category label is assigned to each symbol.

[0016] The described feature set includes: language symbol features, suffix features, keyword features, dictionary features, and rule optimization.

[0017] Further, the S2 includes the following steps:

[0018] S21. Establish a data pattern, including complete information IP * and conditional access information IP + and over-generalized information IP × ;

[0019] S22. Construct a purpose tree table, and the specific method is as follows:

[0020] (1) According to the designed tree table structure, collect the data to be filled;

[0021] (2) Establish a database that meets the tree structure to store or update the tree table data;

[0022] (3) Fill the collected data into the tree table and ensure that the hierarchical relationship of each node is correct.

[0023] S23. According to the data pattern and purpose tree table constructed in steps S21 and S22, establish a matching algorithm based on the access purpose.

[0024] Further, the S23 includes the following steps:

[0025] S231. The data visitor submits the attributes that can prove his own identity to the system, and the system analyzes whether he has the right to access; if there is no right to access, the access is rejected, otherwise, continue to step S232;

[0026] S232. Judge the access purpose, and judge whether the access purpose AP meets the expected purpose IP. If it does not meet, the visitor is rejected; otherwise, continue to step S233;

[0027] S233. According to the result of the purpose matching, judge the data pattern obtained by the user for access, and return the corresponding data pattern to the visitor.

[0028] Further, the S3 includes the following steps:

[0029] S31. Use the privacy-protected data information as the input, and establish an activity view according to the resources, access data, and data operations executed for accessing the data pattern in the business process.

[0030] S32. Define the deviation set of the business process and the event log based on multiple perspectives; for each trace in the event log, classify the composite movement of the log activity a i in terms of activity, resource, and data, obtain the corresponding deviation set and record its deviation cost Cost;

[0031] S33. Aggregate all the deviation sets obtained in step S32 to get the total deviation set Deviations Set;

[0032] S34. Extract the business process rule constraint RC, formulate a decision table according to the business rule constraint, and further analyze whether the total deviation set is a deviation that satisfies the business process decision logic from perspectives other than activity, resource, and data. After correction, output the business process deviation cost Cost and the total deviation set DeviationsSet under data decision perception with privacy protection.

[0033] Furthermore, the S31 includes the following steps:

[0034] S311. Authenticate the data visitor according to the algorithm in step S2 and set the access permission to obtain the data mode with privacy protection Data = {data patter1, data patter2, data patter3};

[0035] S312. Construct the business process activity view AV according to the business process M ,

[0036] AV M = {R i , S i , A i , ATP = (c, r, u, d)}, where R i is the resource set for executing the process activity A i ; S i is the data set participating in the data mode involved in the process activity A i ; ATP is the data operation on the process activity A i on the data S i ATP = {(create (c), read (r), update (u), delete (d)}.

[0037] S313. Construct the business log activity view AV according to the event log L ,

[0038] AV L = {r i , s i , a i , atp = (c, r, u, d)}, where ri Resource set s for executing log activity a i ; s i Dataset involved in the data pattern for log activity a i ; atp is the data operation on the log activity a i on the data s i The data operation atp = {(create (c), read (r), update (u), delete (d)}.

[0039] Furthermore, the S32 specifically includes:

[0040] S321. Define the deviation resource set Count res = {r i}, the deviation activity set Count act = {a i}, the deviation dataset Count dat = {s i}, the deviation data operation set Count do = {atp i};

[0041] S322. According to the synchronization of activities, resources, data, and data operations in the log activity view AV L and the process activity view AV M , record the corresponding deviation sets and deviation cost Cost.

[0042] Furthermore, the total deviation set of S33 is specifically:

[0043] Deviations Set = Count res ∪ Count act ∪ Count dat ∪ Count do .

[0044] Furthermore, the S34 includes the following steps:

[0045] S341. According to the business process constraint rules, refine the relevant elements of the decision table DMN, including the decision name Name, the input attribute set I of the decision table, the output attribute set O, the input range function infacet, the output range function orange, and the default assignment function odef, and initialize the decision table;

[0046] S342. Use the input range function infacet to verify whether each input attribute data (Input dataattributes) meets the rule constraints S-FEEL conditions specified by the business process, record any data that does not meet the conditions, and generate a warning or error log;

[0047] S343. Make a decision judgment according to the logic defined in the decision table DMN, in combination with the input attributes and rule constraints; and correct the deviation cost Cost and the deviation set according to the judgment.

[0048] S344. Output the accumulated deviation cost Cost and the total deviation set DeviationsSet of the business process with data decision awareness under privacy protection.

[0049] The present invention also provides a multi-perspective deviation detection system for business processes based on privacy protection. When the system runs, the above method is applied, including the following modules:

[0050] A data extraction module, which is used to establish a conditional random field (CRF) model for named entity recognition, select a feature set from complex and diverse unstructured data information by using composite entity construction rules, and extract data important for business process research.

[0051] A business process control flow and data flow fusion module, which is used to establish a privacy access control model based on identity and purpose, authenticate the data visitors and set access permissions, and obtain a data mode with privacy protection.

[0052] A deviation detection module, which is used to sense the business process by using data decisions based on the fusion of data flow and control flow, identify the deviations in the business process by monitoring the composite movement of the activities, resources, data, data operations and corresponding dimensions of the event log in the business process, and combining with business decision logic analysis.

[0053] The advantages of the present invention are as follows:

[0054] (1) A purpose-based privacy access control algorithm is proposed to maximize the availability of data under the condition of protecting privacy. A decision table and an activity view are introduced into the traditional business process for data logic decision to guide the business process. The proposed multi-perspective deviation detection algorithm can detect deviations in the data flow, control flow and privacy layer, improving the comprehensiveness and detection accuracy of process deviation detection.

[0055] (2) The random field model is used to extract data attributes important for business processes from unstructured data, improving the efficiency and accuracy of entity recognition. Description of the Drawings

[0056] Figure 1 It is a flow schematic diagram of the multi-perspective deviation detection method for business processes based on privacy protection of the present invention;

[0057] Figure 2 It is a schematic diagram of the CRF linear chain structure of an embodiment of the present invention;

[0058] Figure 3Schematic diagram for identifying important medical attribute entities in the embodiments of the present invention;

[0059] Figure 4 Schematic diagram of the medical purpose tree in the embodiments of the present invention;

[0060] Figure 5 Schematic diagram of the matching algorithm process based on access purposes in the embodiments of the present invention; Detailed implementation manners

[0061] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Embodiment 1

[0063] Organizations often use process models and security policies to describe the normative behavior of business systems and the legitimate use of data. However, in practice, organizations may allow users to deviate from the specified behavior to effectively handle unexpected situations. However, this function may be abused, increasing the risk of harmful data leakage. In addition, insiders may use their permissions to obtain sensitive information for personal or economic gain. In actual operations, the real process behavior often deviates from the expected process, which often opens the way for fraud or performance problems. In healthcare information systems, highly sensitive information is involved (such as patient personal information, medical history, diagnosis results). Detecting business process deviations, preventing data leakage, and protecting privacy are meaningful things.

[0064] This embodiment provides a multi-perspective deviation detection method for business processes based on privacy protection, as Figure 1 shown, including the following steps:

[0065] S1. Establish a conditional random field (CRF) model for named entity recognition, select a feature set from complex and diverse unstructured data information using composite entity construction rules, and extract data attributes important for business process research;

[0066] The specific establishment of the conditional random field (CRF) model is as follows:

[0067] Assume that X represents the data sequence and Y represents the entity category, where x = {x1, x2, x3... x n} is the observation sequence, and y = {y1, y2, y3... y n} is the state sequence, and P(y|x) represents the conditional probability distribution of output y given x; CRF uses a first-order linear chain random field model to represent as follows

[0068]

[0069] where t k , s l represents the feature function, and λ k , μ l represents the corresponding weight. Z(x) represents the normalization factor, which is used to constrain the conditional probability; the feature function and the corresponding weights determine the final linear chain random field.

[0070] The CRF model is used for named entity recognition, regarded as a sequence labeling problem; each sentence to be recognized is used as an observation sequence, and each word in the sentence is used as a symbol, and a class label is assigned to each symbol. The simplest model of CFR is a chain structure, such as Figure 2 shown. Given a training dataset

[0071] D = {(X1, Y1), (X2, Y2), … (X n , Y n )}, by training on the training set, the conditional probability u l , s l that meets the requirements can be obtained. CRF can also use Viterbi decoding to obtain the best state sequence y * = argmaxP(y|x).

[0072] In medical data, electronic medical records are unstructured texts and lack a unified expression standard. In this embodiment, the CRF model is used to perform named entity recognition on electronic medical records, and the medical record data is converted into a structured form that can be recognized by a computer. Named entity recognition in Chinese electronic medical records mainly refers to identifying entities such as disease names, treatment methods, and drugs in the medical records. The main problems in medical record entity recognition are: there are many types of medical record entities, the entity length has no limit, the expressions are not unified, there are a large number of aliases and abbreviations; in different situations, the lengths of medical record entity characters are different, such as cold and influenza. For the characteristics of electronic medical record texts, a method combining the CRF model and rules is used to extract entity names. The following features are selected for recognition in this embodiment:

[0073] (1) Linguistic symbol features: Since Chinese does not have obvious space segmentation like English, before entity recognition, the corpus is segmented by the ICTCLAS word segmentation system of the Chinese Academy of Sciences, and the segmentation results are used as linguistic symbol features.

[0074] (2) Suffix features: Disease names often end with words like "disease", such as diabetes. Drugs often end with "pill", "disintegrating tablet", "element", etc., such as penicillin and enteric-coated aspirin tablets. Treatment methods often end with "surgery", such as tumor resection. These special suffixes are regarded as one of the features.

[0075] (3) Keyword features: By analyzing medical records, certain keywords are followed by disease names or symptoms, such as "suffering from", "appearing", "accompanied by", "found", etc., such as accompanied by chest tightness and found elevated blood sugar. These keywords are used as the boundaries for entity recognition.

[0076] (4) Dictionary features.

[0077] (5) Rule optimization, as shown in Table 1 below

[0078] Table 1 Construction rules for some composite entities

[0079]

[0080] In this embodiment, 2000 medical record data are obtained from a certain tertiary hospital in Anhui Province, and a Chinese electronic medical record corpus of 567,843 characters is constructed. As Figure 3 shown, the method combining CRF and rules is used to identify important medical attribute entities respectively. The important medical domain attribute entities identified are patient basic information, and the important medical business process attributes are obtained through conditional random field CRF plus feature extraction of medical records, including patient basic information I (Patient Information), inspection item P (Inspection item), symptoms S (Symptoms), triage type T (Triage type), treatment protocol R (Treatment protocols), and medical history M (Medical History).

[0081] S2. Establish a privacy access control model based on identity and purpose, authenticate the data visitors and set access permissions to obtain a data pattern with privacy protection. In this embodiment, a purpose-based privacy access control PBAC (Purpose Based Access Control) model is proposed. The core of the model is that the privacy data provided by the data owner matches the access purpose AP (Access Purpose) proposed by the data user according to the intended purpose IP (Intended Purpose) set by the data owner, and then the data user is authorized to access. In order to ensure both the high quality and privacy of the data and extract more valuable information under the premise of privacy protection, a conditional purpose-based access control model CPBAC (Conditional Purpose-based Access Control) is proposed. It includes the following steps:

[0082] S21. Establish a data pattern, including complete information IP * , conditional access information IP + , over-generalized information IP × . As Figure 4 shown, for the medical information provided for a patient, the set of intended purposes for allowing data use is

[0083] IP = <AIP, PIP> = <{Medical treatment}{Scientific research, adjuvant treatment}>, where the set of intended purposes for allowing access is AIP ↓ = {Medical treatment, Clinical treatment, Internal medicine treatment, Surgery treatment, Adjuvant treatment}, and the set of data purpose that cannot be accessed is

[0084]

[0085] Conditional purpose IP + = P - IP * - IP × = {Self-access, Patient access, Relative access}, and the set of actual accessible data purposes When and only when the user's access purpose The user can obtain the data.

[0086] According to the access purpose of the visitor in different purpose sets (IP * , IP × , IP + ), the data pattern that the data visitor can access is different, which is used to process and balance data availability (Data Utility) and privacy confidentiality (Privacy Protection), and improve data security in the big data era. When the data visitor's access purpose Visitors can obtain the complete information pattern 1; when the access purpose Visitors can conditionally obtain the information pattern 2; when the access purpose Visitors can obtain the over-generalized information pattern 3.

[0087] S22. Construct a purpose tree table, and the specific steps are as follows:

[0088] (1) According to the designed tree table structure, collect the data to be filled;

[0089] (2) Establish a database that meets the tree structure to store or update the tree table data;

[0090] (3) Fill the collected data into the tree table and ensure that the hierarchical relationship of each node is correct.

[0091] In this embodiment, the medical purpose tree table is shown in the following table

[0092]

[0093]

[0094] pub1 = PID + '00' + AIP_code + PIP_code

[0095] = '1110001' + '00' + '0010010011' + '1001000100'

[0096] = '11100010000100100111001000100'

[0097] pub2 = PID + '01' + AIP_code + PIP_code

[0098] = '1110001' + '01' + '0010010011' + '1001000100'

[0099] = '11100010100100100111001000100'

[0100] pub3 = PID + '11' + AIP_code + PIP_code

[0101] = '1110001' + '11' + '0010010011' + '1001000100'

[0102] = '11100011100100100111001000100'

[0103] Through the above analysis, 3 public keys can be obtained. The first byte of the public key is the 7-digit PID of the patient and two data mode identifiers (CondBit), where 00 represents data mode 1; 01 represents data mode 2; 11 represents data mode 3 and the AIP_code and PIP_code of equal length. Similarly, the data visitor applies to the system for a private key, and this process is similar to the purpose matching phase. It is necessary to judge the access purpose AP according to the visitor role RIP, and determine which version of the private key to return according to the matching result of AP and IP, where (CondBit = 00, 01, 11)

[0104] S23. Based on the data mode and purpose tree table constructed in steps S21 and S22, establish a matching algorithm based on the access purpose. As Figure 5 shown, the specific steps are as follows:

[0105] S231. The data visitor submits the attributes (PID) that can prove his own identity to the system, and the system analyzes whether he has the right to access; if there is no right to access, the access is rejected, otherwise, proceed to step S232;

[0106] S232. Judge the access purpose to determine whether the access purpose AP meets the expected purpose IP. If not, the visitor is rejected; otherwise, proceed to step S233;

[0107] S233. According to the result of purpose matching, judge the data mode obtained by the user for access, and return the corresponding data mode to the visitor.

[0108] When the visitor accesses the patient's medical information, the system generates an access purpose according to his role, and performs purpose matching to match the access purpose with the expected purpose. Given that AIP_code = 0x093 = (0010010011)2 and PIP_code = 0x244 = (1001000100)2, according to the purpose matching algorithm, we can get:

[0109]

[0110] If the user's access purpose is ap = attending physician of internal medicine and ap_code = 0x2, then

[0111]

[0112] Therefore, the matching result is Permit, and the system returns the original medical data mode 1 and the private key corresponding to pub1 to the visitor. Identity ID public key: 11100010100100100111001000100'

[0113] User private key: 84862017576510131310889759908

[0114] Decrypted data: Data mode 1

[0115] This embodiment proposes a matching algorithm based on roles and access purposes, which can not only avoid affecting the key access, but also effectively prevent the occurrence of malicious access data leakage. Selecting to verify the identity of the visitor with the role ID to determine whether there is access right, and then encrypting the data based on the conditional access purpose and the expected purpose as the identity key can effectively ensure that the system user can access the patient information only if the authentication is passed and the access purpose matches the expected purpose.

[0116] The specific algorithm is shown in the following table. Lines 1-2 set the expected access purpose data set and the prohibited access purpose set according to the data provider's wishes. Lines 3-5 provide the data accessor with the ID number and access role information RIP. If RIP meets the role attribute Attribute allowed by the data provider, the access purpose compliance verification is performed, otherwise the user does not have access to the data. Lines 6-10 calculate the binary and hexadecimal encoding of the allowed access purpose set, prohibited access purpose set, and conditional purpose set based on the purpose tree table. There is no intersection between the purpose sets and they fully cover the purpose tree. Lines 11-14 determine whether the access purpose matches the allowed access purpose and output the key and complete data information data mode 1. Lines 15-18 determine whether the access purpose matches the conditional purpose set. If so, the corresponding key and data mode 2 are returned. Otherwise, lines 19-21 analyze the degree of matching between the access data and the prohibited purpose set and return the corresponding key and data mode 3.

[0117]

[0118] In the field of business process management (BPM), the integration of business processes and related data is a key issue, because the execution of business processes is often constrained by data. Especially in data- and decision-intensive contexts, it is very useful to connect activities and data operations to detect anomalies in business processes from multiple angles. In this embodiment, the concepts of activity view and decision table (DMN) are extended to bridge the gap between process and data conceptual design, and the specified process activities and related data are linked.

[0119] S3. The present invention relates to the field of business process management, and in particular to complex and changeable business process environments, a data decision-aware business process based on the fusion of data flow and control flow is proposed. This process monitors the composite mobile classification of activities, resources, data and data operations in the business process and the corresponding dimensions of the event log, and after combining the business decision logic analysis, it can accurately identify deviations in the business process, thereby improving the efficiency and accuracy of the business process. The implementation steps are as follows:

[0120] S31. Establish an activity view based on the resources, accessed data, and data operations performed in the activity access data pattern of the business process. The specific steps are as follows:

[0121] S311. Authenticate the data visitor according to the algorithm in step S2 and set the access permissions to obtain the data pattern with privacy protection Data = {data patter1, data patter2, data patter3}.

[0122] S312. Construct the business process activity view AV M ,

[0123] AV M = {R i , S i , A i , ATP = (c, r, u, d)}, where R i is the resource set for executing the process activity A i ; S i is the data set participating in the data pattern involved in the process activity A i ; ATP is the data operation on the process activity A i on the data S i ATP = {(create (c), read (r), update (u), delete (d)}.

[0124] S313. Construct the business log activity view AV based on the event log L ,

[0125] AV L = {r i , s i , a i , atp = (c, r, u, d)}, where r i is the resource set for executing the log activity a i ; s i is the data attribute set participating in the data pattern involved in the log activity a i ; atp is the data operation on the log activity a i on the data s i atp = {(create (c), read (r), update (u), delete (d)}.

[0126] S32. Define the composite movement classification and deviation set of the business process and event log from multiple perspectives; perform the composite movement classification of activity, resource, and data for each log activity a i in each trace of the event log, obtain the corresponding deviation set and record its deviation cost Cost. The specific steps are as follows:

[0127] S321. Define the deviation resource set Count res ={r i}, the deviation activity set Count act ={a i}, the deviation data set Count dat ={s i}, the deviation data operation set Count do ={atp i}.

[0128] S322. According to the synchronization status of activities, resources, data, and data operations in the log activity view AV L and the process activity view AV M , record the corresponding deviation sets and deviation costs Cost. Classify according to specific different situations as follows:

[0129] S3221. If in the log activity view AV L and the process activity view AV M , the activities, resources, and data are completely synchronized, then the composite move classification record is CompositeMove = ((r i , R i ), (s i , S i ), (a i , A i ))

[0130] The deviation set is Deviations Set = φ;

[0131] S3222. If in the log activity view AV L and the process activity view AV M , the log activity a i , the log resource r i are synchronized and aligned with the process activity A i and the process resource R i , then

[0132] (1) When the data that should be executed in the system is missing in the log activity view, s i = <<, and the expected activity is executed by a legal role, and the activity data in the log activity view is missing, then the composite move classification record is CompositeMove = ((r i , R i ), (<<, S i ), (a i , A i )) i ), and the data operation type in the log activity view is empty, denoted as atp(s i) = φ, the missing operation may correspond to data update or security check. Skipping data update indicates that the data may be unreliable, and the deviation data set is Count dat = {(<<, S i )}. The deviation cost is Cost = Cost + 2

[0133] (2) When the data that should be executed in the system is missing in the process activity view, S i = <<, then the composite movement classification record is CompositeMove = ((r i , R i ), (s i , <<), (a i , A i )), and the deviation data set is

[0134] Count dat = Count dat ∪{(s i , <<)}; then record its deviation cost as:

[0135] (a) The data operation in the log activity view conforms to the data operation in the process activity view atp(s i ) ∈ ATP(S i ), and the deviation cost is Cost = Cost + 1;

[0136] (b) The data operation in the log activity view does not conform to the data operation in the process activity view The deviation cost is Cost = Cost + 2, and the deviation data operation set Count do = {s i};

[0137] S3223. If in the log activity view AV L and the process activity view AV M , the log resource r i , the log data s i and the process resource R i , the process data S i are synchronized and aligned, then the composite movement classification record is

[0138] CompositeMove = ((r i , R i ), (s i , S i ), (<<, A i ))), then record its deviation cost as:

[0139] (1) The data operation in the log activity view conforms to the data operation in the process activity view atp(s i) ∈ ATP(S i ), the deviation cost is Cost = Cost + 2;

[0140] (2) The data operations in the log activity view do not conform to the data operations in the process activity view The deviation cost is Cost = Cost + 3, and the deviation data operation set Count do = Count do ∪ {s i};

[0141] S3224. If in the log activity view AV L and the process activity view AV M , the log resource r i is synchronized and aligned with the process resource R i , but the activities and data in the log do not meet the business process specifications, then the composite move classification record is CompositeMove = ((r i , R i ), (<<, S i ), (<<, A i ))), the deviation activity set is Count act = {(<<, A i )}, the deviation data set Count dat = Count dat ∪ {(<<, S i )}, and the deviation cost is Cost = Cost + 2;

[0142] S3225. If in the log activity view AV L and the process activity view AV M , the log activities are synchronized and aligned with the process activities, and the activities and data operations in the log activity view AV L are completed by the wrong resources that do not conform to the process activity view AV M , then,

[0143] (1) When the wrong resources perform synchronous composite moves on the activities and data in the process activity view AV M , it is recorded as CompositeMove = (Wrong(r i ), R i ), (s i , S i ), (a i , A i ))). Since different roles obtain different data patterns, although the executor is wrong, but completing the reasonable data operations and activity executions specified by the system causes less harm to the business process, the deviation resource set Count res = {(Wrong(r i),R i )}, the deviation cost is recorded as Cost = Cost + 1.

[0144] (2) When the data operations that should be executed in the system are missing from the data log, but the expected activities are executed by unauthorized roles, the error resource executes the activities a i missing data s i The composite movement classification is recorded as

[0145] CompositeMove = (Wrong(r i ),R i ),(<<,S i ),(a i ,A i ))). Since this embodiment adopts identity authentication access permission restrictions, such deviations may correspond to unauthorized roles having no access rights to system data operations, and such deviations pose a relatively small threat. The deviation data set Count dat = Count dat ∪{(<<,S i ), and the deviation cost is Cost = Cost + 4;

[0146] (3) When an error resource accesses data s i prohibited by the process for the process activity view A i , the composite movement classification is recorded as CompositeMove = (Wrong(r i ),R i ),(s i ,<<),(a i ,A i ))), and the deviation data set is

[0147] Count dat = Count dat ∪{(s i ,<<)}, the deviation resource set Count res = Count res ∪{(Wrong(r i ),R i )}, then its deviation cost is recorded as:

[0148] (a) The data operation of the log activity view conforms to the data operation of the process activity view atp(s i ) ∈ ATP(S i ), and the deviation cost Cost = Cost + 4;

[0149] (b) The data operation of the log activity view does not conform to the data operation of the process activity view The deviation data operation set Countdo = Count do ∪ {s i}, then the deviation cost Cost = Cost + 5.

[0150] S3226. If in the log activity view AV L and the process activity view AV M the missing resources in the log and process for the wrong resources (Wrong(r i ), R i ) for the activities (A i = <<, a i = a) ∪ (A i = A, a i = <<),

[0151] (1) For the activity a inserted in the log i execute the data s allowed by the access process i = S i , the composite movement classification is recorded as CompositeMove = (Wrong(r i ), R i ), (s i , S i ), (a i , <<)), the deviation resource set

[0152] Count res = Count res ∪ {(Wrong(r i ), R i )}, the deviation activity set is Count act = Count act ∪ {(a i , <<)}, the deviation cost Cost = Cost + 4.

[0153] (2) For the activity a inserted in the log i lacking the data s allowed by the execution process i = φ, the composite movement classification is recorded as CompositeMove = (Wrong(r i ), R i ), (<<, S i ), (a i , <<)), the deviation resource set

[0154] Count res = Count res ∪ {(Wrong(r i ), R i )}, the deviation activity set is Count act = Countact ∪{(a i , <<)}, Count dat = Count dat ∪{(<<, S i ))} Then the deviation cost is Cost = Cost + 4.

[0155] (3) Insert an activity a that is not allowed in the process into the log i And access data s that does not meet the business process i , The composite movement classification is recorded as

[0156] CompositeMove = (Wrong(r i ), R i ), (s i , <<), (a i , <<)), The deviation resource set

[0157] Count res = Count res ∪{(Wrong(r i ), R i )}, The deviation activity set is Count act = Count act ∪{(a i , <<)}, Count dat = Count dat ∪{(s i , <<)), Then the recorded cost deviation is:

[0158] (a) The data operation in the log activity view conforms to the data operation in the process activity view atp(s i ) ∈ ATP(S i ), Then the deviation cost is Cost = Cost + 4.

[0159] (b) The data operation in the log activity view does not conform to the data operation in the process activity view Then the deviation cost is Cost = Cost + 5.

[0160] (4) For process activity A i Access the data s that is not allowed in the process i , The composite movement classification is recorded as CompositeMove = (Wrong(r i ), R i ), (s i , <<), (<<, A i ))), The deviation resource set

[0161] Count res = Countres ∪{(Wrong(r i ),R i )}, the deviation activity set is Count act = Count act ∪{(<<, A i )}, Count dat = Count dat ∪{(s i , <<)}, then the recorded cost deviation is:

[0162] (a) The data operation of the log activity view conforms to the data operation of the process activity view. atp(s i ) ∈ ATP(S i ), then the deviation cost is Cost = Cost + 4.

[0163] (b) The data operation of the log activity view does not conform to the data operation of the process activity view Then the deviation cost is Cost = Cost + 5.

[0164] (5) Insert an activity a that is not allowed in the process into the log i and access data s that is not allowed in the process i Then the composite move classification is recorded as CompositeMove = (Wrong(r i ), R i ), (s i , <<), (a i , <<)). The deviation resource set Count res = Count res ∪{(Wrong(r i ), R i )}, the deviation activity set is Count act = Count act ∪{(a i , <<)}, the deviation data set Count dat = Count dat ∪{(s i , <<)}. Then record its deviation cost as:

[0165] (a) The data operation of the log activity view conforms to the data operation of the process activity view. atp(s i ) ∈ ATP(S i ), the deviation cost is Cost = Cost + 4;

[0166] (b) The data operation of the log activity view does not conform to the data operation of the process activity view The deviation data operation set Count do= Count do ∪ {s i}, the deviation cost is Cost = Cost + 5.

[0167] (6) For process activity A i Access data s not allowed by the process i , the composite move classification is recorded as CompositeMove = (Wrong(r i ), R i ), (s i , <<), (<<, A i ))), the deviation resource set Count res = Count res ∪{(Wrong(r i ), R i )}, the deviation activity set is Count act = Count act ∪{(<<, A i )}, the deviation data set Count dat = Count dat ∪{(s i , <<)}; The recorded deviation cost is:

[0168] (a) If the data operation in the log activity view conforms to the data operation in the process activity view atp(s i ) ∈ ATP(S i ), the deviation cost is Cost = Cost + 5;

[0169] (b) If the data operation in the log activity view does not conform to the data operation in the process activity view The deviation data operation set Count do = Count do ∪{s i}, the deviation cost is Cost = Cost + 6.

[0170] S3227. If the resource absence in the log activity view AV L in the business process affects multi-perspective deviation detection as follows:

[0171] (1) When the activity and data meet the business process specifications, the composite move classification is recorded as

[0172] CompositeMove = (<<, R i ), (s i , S i ), (a i , A i )), the deviation resource set Count res = Countres ∪{(<<,R i )}, the deviation cost is Cost = Cost + 1;

[0173] (2) When data is missing and the activity meets the process specifications, the composite movement classification is recorded as

[0174] CompositeMove = (<<,R i ),(<<,S i ),(a i ,A i ))), the deviation resource set Count res = Count res ∪{(<<,R i )}, the deviation data set is Count dat = Count dat ∪{(<<,S i )}, the deviation cost is Cost = Cost + 2;

[0175] (3) When the activity meets the process specifications and accesses data not allowed by the process for the activity, the composite movement type is CompositeMove = (<<,R i ),(s i ,<<),(a i ,A i ))), the deviation resource set Count res = Count res ∪{(<<,R i )}, the deviation data set is Count dat = Count dat ∪{(s i ,<<)}, then record the deviation cost as:

[0176] (a) The data operation of the log activity view conforms to the data operation of the process activity view atp(s i ) ∈ ATP(S i ), the deviation cost is Cost = Cost + 3;

[0177] (b) The data operation of the log activity view does not conform to the data operation of the process activity view The deviation data operation set Count do = Count do ∪{s i} The deviation cost is Cost = Cost + 4.

[0178] S33. Aggregate all the deviation sets in the log activity view into

[0179] Deviations Set = Countres ∪Count act ∪Count dat ∪Count do , further detect whether these deviations meet the process rule constraints.

[0180] When there are deviations in resources, data, activities, and data operations, further determine through decision logic whether it is an expected behavior that meets the rule constraints or a deviation. When business rules or constraints need to be introduced for decision-making in the process, introduce a decision table, input relevant data into the decision table, and output the decision result through decision logic and return it to the business process decision-making activity. Elevate the business process to this richer, data-aware level. It is more conducive to detecting decision logic deviations in the business process. The specific steps are as follows:

[0181] S34. Extract the business process rule constraints RC (Rules and regulations constraints), formulate a decision table according to the business rule constraints, and further analyze whether there are deviations that meet the business process decision logic from the perspectives of activities, resources, and data, so as to further improve the accuracy of deviation detection. The specific steps are as follows:

[0182] S341. According to the business process constraint rules, refine the relevant elements of the decision table DMN, including the decision name Name, the input attribute set I of the decision table, the output attribute set O, the input range function infacet, the output range function orange, and the default assignment function odef, and initialize the decision table;

[0183] S342. Use the input range function infacet to verify whether each input attribute data (Input dataattributes) meets the rule constraints S-FEEL conditions specified by the business process, record any data that does not meet the conditions, and generate a warning or error log;

[0184] S343. According to the logic defined in the decision table DMN, combine the input attributes and rule constraints to make a decision judgment; and correct the deviation cost Cost and the deviation set according to the judgment;

[0185] The specific decision judgment is as follows:

[0186] (1) Through decision logic analysis, there are deviations in the composite movement classification activities, resources, and data operations in the log and process model, but they may meet the business process rule constraints and are favorable operations taken to cope with environmental changes. These behaviors are reasonably deleted from the deviations. The specific operations are as follows:

[0187] (a) When the resource deviation meets the business process rule constraint conditions, that is the deviation cost Cost = Cost - 1;

[0188] (b) The activity deviation satisfies the constraints of the business process rules, that is The deviation cost Cost = Cost - 1;

[0189] (c) The data deviation and the data operation deviation satisfy the constraints of the business process rules

[0190]

[0191] The deviation cost Cost = Cost - 2;

[0192] (2) Through the decision logic analysis of the log and the process model, the resources, data, and activities satisfy the process view. When there is no deviation between the log activity view and the process activity view, but the behavior that does not satisfy the constraints of the business process rules is recorded as a decision logic error The deviation cost Cost = Cost + 3;

[0193] S354. Output the cumulative deviation cost Cost of the business process with data decision perception under privacy protection and the total deviation set DeviationssSet.

[0194] Embodiment 2

[0195] It should be further noted that, based on the same inventive concept, the embodiment of the present invention also provides a multi - perspective deviation detection system for business processes based on privacy protection. When the system runs, it executes the method described in Embodiment 1, including the following modules:

[0196] The data extraction module is used to establish a conditional random field (CRF) model for named entity recognition, select a feature set from complex and diverse unstructured data information by using composite entity construction rules, and extract data attributes important for business process research;

[0197] The business process control flow and data flow fusion module is used to establish a privacy access control model based on identity and purpose, authenticate the data visitors and set access permissions to obtain a data mode with privacy protection;

[0198] The deviation detection module is used to detect the deviation between the business process and the log from multiple perspectives based on the fusion of the data flow and the control flow.

[0199] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi - perspective deviation detection method for business processes based on privacy protection, characterized in that, It includes the following steps: S1. Establish a conditional random field (CRF) model for named entity recognition, select a feature set using composite entity construction rules from complex and diverse unstructured data information, and extract data important for business process research; S2. Establish a privacy access control model based on identity and purpose, authenticate data visitors and set access permissions to obtain a data mode with privacy protection; S3. Adopt data decision-making based on the fusion of data flow and control flow to perceive the business process, identify deviations in the business process by monitoring the composite movement classification of corresponding dimensions of activities, resources, data, and data operations and event logs in the business process, and combining with business decision logic analysis; 2. The multi-perspective deviation detection method for business processes based on privacy protection according to claim 1, wherein The specific method of establishing the conditional random field (CRF) model described in S1 is as follows: Suppose X represents a data sequence and Y is an entity category, where x = {x1, x2, x3…x n} is the observation sequence, y = {y1, y2, y3…y n} is the state sequence, and P(y|x) represents the conditional probability distribution of outputting y given x; CRF uses a first-order linear chain random field model to represent as follows where t k , s l denotes the characteristic function, λ k , μ l denote the corresponding weights; Z(x) denotes the normalization factor for constraining the conditional probability; The CRF model performs named entity recognition, regarded as a sequence labeling problem; each sentence to be recognized is regarded as an observation sequence, and each word in the sentence is regarded as a symbol, and a category label is assigned to each symbol; The feature set includes: language symbol features, suffix features, keyword features, dictionary features, and rule optimization; 3. The multi-perspective deviation detection method for business processes based on privacy protection according to claim 1, wherein S2 includes the following steps: S21. Establish a data model, including complete information IP * . Conditional access information IP + . Overly generalized information IP × ; S22. Construct a purpose tree table, and the specific method is as follows: (1) According to the designed tree table structure, collect the data to be filled; (2) Establish a database that meets the tree structure to store or update the tree table data; (3) Fill the collected data into the tree table and ensure that the hierarchical relationship of each node is correct; S23. Based on the data mode and purpose tree table constructed in steps S21 and S22, establish a matching algorithm based on the access purpose; 4. The multi-perspective deviation detection method for business processes based on privacy protection according to claim 3, wherein S23 includes the following steps: S231. The data visitor submits the attributes that can prove their identity to the system, and the system analyzes whether they have access rights; if there are no access rights, the access is rejected, otherwise, proceed to step S232; S232. Judge the access purpose to determine whether the access purpose AP meets the expected purpose IP. If not, the visitor is rejected; otherwise, proceed to step S233; S233. According to the result of purpose matching, judge the data mode obtained by the user for access and return the corresponding data mode to the visitor; 5. The multi - perspective deviation detection method for business processes based on privacy protection according to claim 1, wherein, S3 includes the following steps: S31. Take the data information with privacy protection as the input, and establish an activity view based on the resources, accessed data, and data operations executed for accessing the data mode in the business process; S32. Define the deviation set of the business process and the event log based on multiple perspectives; for each trace in the event log, perform composite movement classification of the log activity a i on activities, resources, and data to obtain the corresponding deviation set and record its deviation cost Cost; S33. Aggregate all the deviation sets obtained in step S32 to obtain the total deviation set Deviations Set; S34. Extract the business process rule constraint RC, formulate a decision table according to the business rule constraint, and further analyze whether the total deviation set is a deviation that meets the business process decision logic from perspectives other than activities, resources, and data. After correction, output the business process deviation cost Cost and the total deviation set DeviationsSet under data decision perception with privacy protection; 6. The method for multi-perspective deviation detection of business processes based on privacy protection according to claim 5, characterized in that S31 includes the following steps: S311. Authenticate the data visitor according to the algorithm in step S2 and set the access permissions to obtain a data pattern with privacy protection Data = {data patter1, data patter2, data patter3}; S312. Construct a business process activity view AV according to the business process M , AV M ={R i , S i , A i , ATP = (c, r, u, d)}, where R i is the set of resources for executing process activity A i ; S i is the data set participating in the data schema involved in process activity A i ; ATP is the data operation on process activity A i on data S i ATP = {(create (c), read (r), update (u), delete (d)}; S313. Construct a business log activity view AV based on the event log L , AV L ={r i , s i , a i , atp = (c, r, u, d)}, where r i is the resource set for executing the log activity a i ; s i is the data set participating in the data schema involved in the log activity a i ; atp is the data operation on the log activity a i on the data s i atp = {(create (c), read (r), update (u), delete (d)}.

7. The multi - perspective deviation detection method for business processes based on privacy protection according to claim 6, wherein The specific content of S32 includes: S321. Define the deviation resource set Count res = {r i}, the deviation activity set Count act = {a i}, the deviation data set Count dat = {s i}, the deviation data operation set Count do = {atp i}; S322. According to the log activity view AV L and the process activity view AV M In, record the corresponding deviation set and deviation cost Cost based on the synchronization status of activities, resources, data, and data operations.

8. The multi-perspective deviation detection method for business processes based on privacy protection according to claim 7, characterized in that The total deviation set in S33 is specifically: Deviations Set=Count res ∪Count act ∪Count dat ∪Count do 。 9. The multi-perspective deviation detection method for business processes based on privacy protection according to claim 7, characterized in that, S34 includes the following steps: S341. Refine the relevant elements of the decision table DMN according to the business process constraint rules, including the decision name Name, the input attribute set I of the decision table, the output attribute set O, the input range function infacet, the output range function orange, and the default assignment function odef, and initialize the decision table; S342. Use the input range function infacet to verify whether each input attribute data (Input data attributes) meets the rule constraints S-FEEL conditions specified by the business process, record any data that does not meet the conditions, and generate warning or error logs; S343. Make a decision judgment according to the logic defined in the decision table DMN, combine the input attributes and rule constraints, and correct the deviation cost Cost and the deviation set; S344. Output the cumulative deviation cost Cost and the total deviation set DeviationsSet of the business process with data decision perception under privacy protection.

10. A multi - perspective deviation detection system for business processes based on privacy protection, characterized in that, When the system runs, it executes the method described in any one of claims 1-9, including the following modules: The data extraction module is used to establish a conditional random field CRF model for named entity recognition, select a feature set from complex and diverse unstructured data information by using the composite entity construction rule, and extract data important for business process research; The business process control flow and data flow fusion module is used to establish a privacy access control model based on identity and purpose, authenticate the data visitor and set the access permissions to obtain a data pattern with privacy protection; The deviation detection module is used to sense the business process by using data decisions based on the fusion of data flow and control flow, identify the deviations in the business process by monitoring the composite movement of activities, resources, data, and data operations and event logs in the business process in corresponding dimensions, and combining business decision logic analysis.

Citation Information

Patent Citations

  • Object-centered business process violation inspection method and system based on Tokine replay

    CN119477231A