System and method for calculating path sequence similarity of a business process model unfolding tree
By constructing a path network diagram and calculating the edit distance of the path sequence, the problem of inefficient business process model similarity measurement methods in the prior art is solved, and the accuracy and efficiency of data migration are achieved.
Patent Information
- Application Number
- CN202211266905.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-10-17
AI Technical Summary
The existing business process model similarity measurement methods have insufficient computing efficiency and accuracy, and it is difficult to assist in data migration between old and new systems.
The path sequence similarity calculation method is adopted for the business process model expansion tree, and the path network diagram is built, and the path search is performed using Petri network or BPMN/UML modeling tools, and the edit distance and similarity of the path sequence are calculated to realize data migration.
Improves the accuracy and computing efficiency of data migration, reduces the time and errors of manual statistical path sequences, and helps maintenance and developers reduce duplicate labor.
Smart Images

Figure CN115629787B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data migration, and particularly relates to a path sequence similarity calculation system and method for a business process model unfolding tree. Background Art
[0002] In actual data migration projects, in order to solve the problems of data mapping and partial reuse of program modules, it is necessary to determine the similarity between two business process models. However, some mainstream business process model similarity measurement methods have various defects, either unable to completely process various structures, or having low calculation efficiency, or the similarity calculation results being difficult to intuitively understand. Therefore, a new behavior similarity measurement method is needed to assist data migration between old systems and new systems, and improve the accuracy and operation efficiency of data migration. Summary of the Invention
[0003] The purpose of the present invention is to provide a path sequence similarity calculation system and method for a business process model unfolding tree, to assist data migration between old systems and new systems, and improve the accuracy and operation efficiency of data migration.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:[[]]END]]
[0005] The first aspect of the present invention provides a path sequence similarity calculation method for a business process model unfolding tree, including:
[0006] Extract keywords from the requirement document of the system software of the data to be migrated out, and construct a first business process model based on the keywords in the requirement document;
[0007] Extract keywords from the requirement document of the system software of the data to be migrated in, and construct a second business process model based on the keywords in the requirement document;
[0008] Unfold the first business process model and the second business process model respectively to form path network graphs G1 and G2; use the business process model unfolding into path sequence algorithm to perform path search on the path network graphs G1 and G2, and respectively obtain path sequences S i (i = 1, 2,..., n) and path sequences S j (j = 1, 2,..., m);
[0009] Compare path sequences S i and path sequences S j to obtain the maximum path length l between the two, calculate the edit distance D between path graph G1 and path graph G2; calculate the similarity between path sequences S i and path sequences S j based on the maximum path length l and the edit distance D;
[0010] Complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in according to the similarity between the path sequences S i and the path sequence S j
[0011] Preferably, the method for constructing the first business process model and the first business process model based on the keywords in the requirements document includes:
[0012] Construct a complete workflow in the order of the keywords;
[0013] Convert the constructed workflow through symbolic expressions and use a modeling tool to construct the corresponding business process model.
[0014] Preferably, the modeling tool includes, such as, Petri net, BPMN process modeling tool or UML modeling tool.
[0015] Preferably, the method for expanding the first business process model and the second business process model into path network diagrams G1 and G2 respectively includes:
[0016] Use Petri net to construct the corresponding business process model; use the Python language to read the business process model, use an xml parsing tool to parse the business process model, and create the parsed content into a graph G.
[0017] Preferably, use the path search algorithm for expanding the business process model into path sequences to perform path search on the path network diagrams G1 and G2, and obtain the path sequences S i (i = 1, 2,..., n) and the path sequence S j (j = 1, 2,..., m); the method includes:
[0018] Construct the start transition T s and the end transition T e on the path network diagram G, and obtain the positions of the start transition T s and the end transition T e ; the path network diagram G includes the path network diagrams G1 and G2;
[0019] Perform a forward node breadth-first search on the path network diagram G from the start transition T s to the end transition T e Te to obtain the forward shortest path s′ 主 ;
[0020] Perform a reverse node breadth-first search on the path network diagram G from the end transition T e to the start transition T s to obtain the reverse shortest path s″主 ;
[0021] According to the forward shortest path s′ 主 and the reverse shortest path s″ 主 The composite backbone path s is fitted 主 ; Remove the composite backbone path s from the path network graph G 主 to obtain the residual network graph G′; Repeat to extract the composite backbone path s in the residual network graph G′ 主 ;
[0022] Select node E in the residual network graph G′ to perform a breadth - first search on the path network graph G, from the starting edge node μ connected to the selected node E to the ending transition T e to obtain the shortest path s′ 残 and from the ending edge node v connected to the selected node E to the starting transition T s to obtain the shortest path s″ 残 ; According to the shortest path s′ 残 and the shortest path s″ 残 The shortest path s is fitted 残 ;
[0023] The shortest path s obtained from the path network graph G1 残 and the composite backbone path s 主 are denoted as the path sequence S i ;
[0024] The shortest path s obtained from the path network graph G2 残 and the composite backbone path s 主 are denoted as the path sequence S j .
[0025] Preferably, the method for performing a breadth - first search on the path network graph G includes:
[0026] Visit the initial node a, mark the initial node a as visited, and add the initial node a to the queue; Terminate the algorithm when the nodes added to the queue are empty;
[0027] When the nodes added to the queue are non - empty, take the head node b of the queue; Search for the adjacent nodes c of the head node b;
[0028] If there are no adjacent nodes c of the head node b, terminate the algorithm; If there are adjacent nodes c of the head node b, then visit the node c and mark it as visited, add the node c to the queue, and repeat the iteration to search for the adjacent nodes of the head node.
[0029] Preferably, the method for calculating the edit distance D between the path graph G1 and the path graph G2 includes:
[0030] Select each path sequence S in turn iOne path in it and each path sequence S j Perform calculation to obtain the edit distance D ij ;
[0031] Based on the edit distance D ij The average edit distance D i , calculate and obtain the edit distance D between the path graph G1 and the path graph G2, and the expression formula is:
[0032]
[0033]
[0034] In the formula, m represents the number of path sequences S j The number of, n represents the number of path sequences S i The number of.
[0035] Preferably, calculate and obtain the similarity between the path sequence S i And the path sequence S j Based on the maximum path length l and the edit distance D, and the expression formula is:
[0036]
[0037] The formula is: Sim represents the similarity between the path sequence S i And the path sequence S j Between.
[0038] In the second aspect, the present invention provides a path sequence similarity calculation system for a business process model expansion tree, which is characterized by including:
[0039] A model construction module, which is used to extract keywords from the requirement document of the system software of the data to be migrated out, and construct the first business process model based on the keywords in the requirement document; extract keywords from the requirement document of the system software of the data to be migrated in, and construct the second business process model based on the keywords in the requirement document;
[0040] A path network graph construction module, which is used to expand the first business process model and the second business process model respectively to form a path network graph G1 and a path network graph G2;
[0041] A search module, which is used to perform path search on the path network graph G1 and the path network graph G2 by using the business process model expansion into path sequence algorithm, and respectively obtain the path sequence S i (i = 1, 2,..., n) and the path sequence S j (j = 1, 2,..., m);
[0042] A similarity calculation module, which is used to compare the path sequence S i And the path sequence Sj Obtain the maximum path length l between the two, and calculate the edit distance D between path graph G1 and path graph G2; based on the maximum path length l and the edit distance D, obtain the path sequence S i and the path sequence S j the similarity between them;
[0043] The data migration module is used to complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in according to the similarity between the path sequence S i and the path sequence S j the similarity between them.
[0044] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the path sequence similarity calculation method are implemented.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] The present invention uses the algorithm of expanding the business process model into a path sequence to search for paths in path network graph G1 and path network graph G2, and respectively obtains path sequences S i (i = 1, 2,..., n) and path sequence S j (j = 1, 2,..., m); the relevant algorithm can be used to automatically obtain the path sequence, reducing the time and errors of manual statistics of the path sequence, and realizing the efficient acquisition of the path sequence.
[0047] The present invention compares path sequence S i and path sequence S j Obtain the maximum path length l between the two, and calculate the edit distance D between path graph G1 and path graph G2; based on the maximum path length l and the edit distance D, obtain the path sequence S i and path sequence S j the similarity between them; according to the similarity between path sequence S i and path sequence S j the similarity between them to complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in; perform more accurate data migration or program module reuse, which can help maintenance and development personnel reduce repetitive labor. Description of the Drawings
[0048] Figure 1 is a flowchart of the path sequence similarity calculation method for the business process model expansion tree provided in Embodiment 1 of the present invention. Detailed Embodiments
[0049] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0050] Embodiment 1
[0051] As Figure 1 shown, a method for calculating the path sequence similarity of a tree-expanded business process model includes:
[0052] Extract keywords from the requirement document of the system software of the data to be migrated out, and construct the first business process model based on the keywords in the requirement document; extract keywords from the requirement document of the system software of the data to be migrated in, and construct the second business process model based on the keywords in the requirement document.
[0053] The method for constructing the first business process model and the first business process model based on the keywords in the requirement document includes:
[0054] Construct a complete workflow in the order of the keywords; convert the constructed workflow through symbolic expressions, and use a modeling tool to construct the corresponding business process model; the modeling tool includes, for example, Petri net, BPMN process modeling tool or UML modeling tool.
[0055] The method for respectively expanding the first business process model and the second business process model to form path network diagrams G1 and G2 includes:
[0056] Use Petri net to construct the corresponding business process model; use the Python language to read the business process model, use an xml parsing tool to parse the business process model, and create the parsed content into a graph G.
[0057] Adopt the path sequence algorithm for expanding the business process model to perform path search on the path network diagrams G1 and G2, and respectively obtain the path sequences S i (i = 1, 2,..., n) and the path sequence S j (j = 1, 2,..., m); the method includes:
[0058] Construct the start transition T s and the end transition T e on the path network diagram G, and obtain the positions of the start transition T s and the end transition T e ; the path network diagram G includes the path network diagrams G1 and G2;
[0059] Perform a forward node breadth-first search on the path network diagram G from the start transition T s to the end transition T e Te to obtain the forward shortest path s′主 ;
[0060] From the end transition T e to the start transition T s Perform a reverse node breadth-first search on the path network graph G to obtain the reverse shortest path s″ 主 ;
[0061] According to the forward shortest path s′ 主 and the reverse shortest path s″ 主 Fit to obtain the composite backbone path s 主 ; Remove the composite backbone path s 主 from the path network graph G to obtain the residual network graph G′; Repeat extracting the composite backbone path s 主 ;
[0062] Select a node E in the residual network graph G′ to perform a breadth-first search on the path network graph G. From the starting edge node μ connected to the selected node E to the end transition T e Obtain the shortest path s′ 残 ; From the ending edge node v connected to the selected node E to the start transition T s Obtain the shortest path s″ 残 ; According to the shortest path s′ 残 and the shortest path s″ 残 Fit to obtain the shortest path s 残 ;
[0063] Record the shortest path s 残 obtained from the path network graph G1 and the composite backbone path s 主 as the path sequence S i ;
[0064] Record the shortest path s 残 obtained from the path network graph G2 and the composite backbone path s 主 as the path sequence S j .
[0065] The method for performing a breadth-first search on the path network graph G includes:
[0066] Visit the initial node a, mark the initial node a as visited, and add the initial node a to the queue; Terminate the algorithm when the nodes added to the queue are empty;
[0067] When the nodes added to the queue are non-empty, take the head node b of the queue; Search for the adjacent node c of the head node b;
[0068] If there is no adjacent node c for the head node b, terminate the algorithm; If there is an adjacent node c for the head node b, then visit the node c and mark it as visited, add the node c to the queue, and repeat the iteration to search for the adjacent nodes of the head node.
[0069] A method for calculating the edit distance D between path graph G1 and path graph G2 includes:
[0070] Select each path in path sequence S in turn i and each path sequence S j to calculate the edit distance Di j ;
[0071] Based on the average edit distance D ij of the edit distance D i calculate the edit distance D between path graph G1 and path graph G2, and the expression formula is:
[0072]
[0073]
[0074] In the formula, m represents the number of path sequences S j and n represents the number of path sequences S i .
[0075] Compare path sequence S i and path sequence S j to obtain the maximum path length l between the two. Calculate the similarity between path sequence S i and path sequence S j based on the maximum path length l and the edit distance D, and the expression formula is:
[0076]
[0077] The formula is: Sim represents the similarity between path sequence S i and path sequence S j .
[0078] Complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in according to the similarity between path sequence S i and path sequence S j .
[0079] Embodiment 2
[0080] A path sequence similarity calculation system for a business process model expansion tree. The system provided in this embodiment can be applied to the method described in Embodiment 1 and includes:
[0081] A model construction module, configured to extract keywords from the requirement documents of the system software for the data to be migrated out, and construct a first business process model based on the keywords in the requirement documents; extract keywords from the requirement documents of the system software for the data to be migrated in, and construct a second business process model based on the keywords in the requirement documents.
[0082] A path network graph construction module, configured to expand the first business process model and the second business process model respectively to form a path network graph G1 and a path network graph G2.
[0083] A search module, configured to perform path search on the path network graph G1 and the path network graph G2 by using the business process model expansion into path sequence algorithm, and respectively obtain path sequences S i (i = 1, 2,..., n) and path sequences S j (j = 1, 2,..., m).
[0084] A similarity calculation module, configured to compare path sequence S i and path sequence S j to obtain the maximum path length l between the two, calculate the edit distance D between the path graph G1 and the path graph G2; calculate the similarity between the path sequence S i and path sequence S j based on the maximum path length l and the edit distance D.
[0085] A data migration module, configured to complete the data migration between the system software for the data to be migrated out and the system software for the data to be migrated in according to the similarity between the path sequence S j and path sequence S j .
[0086] Embodiment 3
[0087] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the path sequence similarity calculation method described in Embodiment 1 are implemented.
[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0089] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that realize the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0092] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for calculating the similarity of path sequences for a business process model unfolding tree, characterized in that, Including: Extract keywords from the requirements document of the system software from which data is to be migrated out, and construct a first business process model based on the keywords in the requirements document; Extract keywords from the requirements document of the system software into which data is to be migrated in, and construct a second business process model based on the keywords in the requirements document; Unfold the first business process model and the second business process model respectively to form path network diagrams and path network diagrams ; Use the algorithm of unfolding the business process model into a path sequence to perform path search on the path network diagrams and path network diagrams to obtain path sequences and path sequences ; Comparison path sequence and path sequence Obtain the maximum path length of the two , calculate the path graph and path graph The edit distance between ; Based on the maximum path length and the edit distance calculate to obtain the path sequence and the path sequence between the similarities; Complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in according to the similarity between the path sequence and the path sequence .
2. The path sequence similarity calculation method for the business process model unfolding tree according to claim 1, characterized in that, The method for constructing the first business process model and the first business process model based on the keywords in the requirements document includes: Construct a complete workflow in the order of the keywords; Convert the constructed workflow through symbolic expressions, and use a modeling tool to construct a corresponding business process model.
3. The method for calculating the path sequence similarity of the business process model expansion tree according to claim 1, wherein, The modeling tool includes, such as, Petri net, BPMN process modeling tool or UML modeling tool.
4. The method for calculating the path sequence similarity of the business process model expansion tree according to claim 1, characterized in that Unfolding the first business process model and the second business process model respectively to form path network diagrams and path network diagrams The method includes: Construct the corresponding business process model using Petri nets; read the business process model using the Python language, parse the business process model using an xml parsing tool, and create a graph from the parsed content 。 5. The method for calculating the path sequence similarity of the business process model expansion tree according to claim 1, characterized in that Using the algorithm of expanding the business process model into a path sequence to perform path search on the path network graph and the path network graph to obtain the path sequences and the path sequence ; The method includes: Construct start transitions and end transitions on the path network graph and obtain the positions of the start transitions and end transitions ; The path network graph includes path network graph and path network graph ; From the start transition to the end transition perform a forward node breadth-first search on the path network graph to obtain the forward shortest path ; From the end transition to the start transition perform a reverse node breadth-first search on the path network graph to obtain the reverse shortest path ; According to the forward shortest path and the reverse shortest path fitted to the composite backbone path ; from the path network diagram remove the composite backbone path to obtain the residual network diagram ; repeat extracting the composite backbone path in the residual network diagram ; In the residual network diagram Select nodes For the path network diagram Perform a breadth-first search starting from the selected nodes To the starting edge node connected To the end transition Obtain the shortest path From the end edge node connected to the selected nodes To the start transition To the start transition Obtain the shortest path ; According to the shortest path And the shortest path The shortest path fitted ; The shortest path obtained from the path network diagram and the composite backbone path are recorded as a path sequence ; ; The shortest path obtained from the path network graph and the composite backbone path are denoted as a path sequence . .
6. The method for calculating the similarity of path sequences of a business process model expansion tree according to claim 1, characterized in that For the path network graph The method for performing breadth-first search includes: Access the initial node and mark the initial node as visited, add the initial node to the queue; terminate the algorithm when the node added to the queue is empty; When the enqueued node is not empty, take the head node of the queue ; Search for the head node of the queue of the adjacent nodes ; If the head node of the queue has no adjacent nodes then terminate the algorithm; if the head node of the queue has adjacent nodes then visit the node and mark it as visited, add the node to the queue, and repeat the iteration to find the adjacent nodes of the head node of the queue.
7. The method for calculating the path sequence similarity of the tree expanded from the business process model according to claim 1, characterized in that Computation path graph and path graph The edit distance between The method includes: Select each path sequence in turn One path in each path sequence and calculate the edit distance ; Based on the edit distance the average edit distance , calculate to obtain the path graph and the path graph the edit distance between , and the expression formula is: ; ; In the formula, represents the number of expressed as a path sequence represents the number of expressed as a path sequence 8. The method for calculating the path sequence similarity of the business process model expansion tree according to claim 1, characterized in that Based on the maximum path length and the edit distance calculate to obtain the path sequence and the path sequence The similarity between them is expressed by the formula: ; The formula is: represented as a path sequence and the path sequence the similarity between them.
9. A path sequence similarity calculation system for a business process model unfolding tree, characterized in that, Including: A model construction module, configured to extract keywords from the requirements document of the system software from which data is to be migrated out, and construct a first business process model based on the keywords in the requirements document; Extract keywords from the requirements document of the system software into which data is to be migrated in, and construct a second business process model based on the keywords in the requirements document; A path network diagram construction module for expanding the first business process model and the second business process model respectively to form path network diagrams and path network diagrams ; A search module, which is used to perform path search on the path network graph and the path network graph by using the algorithm of expanding the business process model into a path sequence, and respectively obtain the path sequence and the path sequence ; A similarity calculation module for comparing path sequences and path sequences to obtain the maximum path length between the two , and calculate the edit distance between path graph and path graph ; Based on the maximum path length and the edit distance calculate to obtain the path sequence and the path sequence between the similarities; A data migration module, which is used to complete the data migration between the system software of the data to be migrated out and the system software of the data to be migrated in according to the similarity between and the path sequence .
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the path sequence similarity calculation method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Service path design method and device, electronic equipment and storage medium
CN111221508A
Complex business process-oriented reusable component mining method
CN113254013A