A General Workflow Modeling Method for Oceanographic Research Vessels
Through general workflow modeling methods and knowledge graph feature learning, the complex workflow and difficulty in information exchange between laboratories on marine scientific research ships were solved, and unified management and efficient collaborative operations were achieved between laboratories.
Patent Information
- Application Number
- CN202311452792.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-03
AI Technical Summary
Due to the independent experimental information management system, the onboard laboratories on the marine scientific research ship have difficulty in exchanging work information, chaotic material resource management, complex work flow and strong flexibility, making it difficult to achieve unified management and collaborative operations.
The general workflow modeling method is adopted to construct the workflow modeling and expression of laboratory detection tasks through unified resource description and workflow modeling, and a knowledge graph feature learning method is introduced to realize automatic matching of process nodes and intelligent recommendation.
It realizes unified management of information and materials inside and outside the laboratory, improves the ability to cooperate and collaborate between laboratories, and improves the efficiency of laboratory work.
Smart Images

Figure CN117455323B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of ship information systems, and in particular to a general workflow modeling method for oceanographic research vessels. Background Art
[0002] At present, the on-board laboratories of oceanographic research vessels mainly include geophysical laboratories, basic laboratories, inorganic geochemistry laboratories, organic geochemistry laboratories, marine science laboratories, paleomagnetism laboratories, microbiology laboratories, drilling technology laboratories, etc. Their functions cover many disciplines such as geological science, marine science, biological science, meteorological science, and environmental science. The workload of data collection, analysis, and management is large, the personnel operation is complex, and data storage and sharing are also greatly restricted. There are usually many intersections in the work processes among on-board laboratories. There are complex material and data exchanges in the work of multiple on-board laboratories. At the same time, the detection process is characterized by a large number of experiments and strong process flexibility. Each laboratory uses its own independent experimental information management system, resulting in difficult work information exchange and chaotic material resource management. By unifying the workflow modeling among laboratories, it provides consistent and convenient work planning support and work progress monitoring functions for large-scale multi-laboratory workflows. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a general workflow modeling method for oceanographic research vessels in view of the defects in the prior art.
[0004] The technical solution adopted by the present invention to solve its technical problems is: a general workflow modeling method for oceanographic research vessels, including the following steps:
[0005] S1. Uniformly describe the experimental resources including data, materials, and experimental equipment of each laboratory on the oceanographic research vessel, and generate references or representations of relevant resources;
[0006] Specifically as follows:
[0007] 1.1) Decompose all experimental resources into several distributable individual resources, the distribution number of each individual resource is greater than or equal to 1, and the access of each individual resource by external nodes is exclusive;
[0008] 1.2) Define the available number for each individual resource, and the available number represents the distributable quantity, that is, the quantity that can be accessed simultaneously;
[0009] 1.3) Determine the relationship between the available number of individual resources and the distribution mode and node dependency according to the available number situation of individual resources, as follows:
[0010] For individual resources that can be infinitely distributed, obtain their distribution methods by setting the access timing relationship between external nodes;
[0011] For the individually distributed resources, their distribution methods are obtained by setting the parallel execution relationships among external nodes;
[0012] For the resources with limited distribution and non - replenishment, by setting the branch selection among external nodes, once a branch is selected, it will cause the invalidation of other branches;
[0013] For the resources with limited distribution and replenishment, by setting the node relationships of sequential execution among external nodes where the order is irrelevant;
[0014] Based on the four types of resource - dependency relationships in step 1.3), five basic operations of resources can be constructed: 1. Produce resources, generating new resources and registering them in the system; 2. Modify resources, including modifying content, status, or even deleting; 3. Consume resources; 4. Access resources; 5. Replenish resources, such as returning samples after lending.
[0015] Related resources include sample resources, document resources, and other objective existences that can satisfy the above five basic operations;
[0016] All objective existences with the above - mentioned five basic operations can be abstracted into the resource concept.
[0017] The abstracted document resources in the system represent various semi - structured or unstructured data (as well as structured data carried in text form), and are organized and managed in file form. In principle, document resources are all infinitely distributed.
[0018] S2. Conduct unified workflow modeling for the intra - laboratory and inter - laboratory work processes, construct the workflow modeling expression for laboratory testing tasks, and associate the relevant resource references described in S1;
[0019] The unified workflow modeling defines the process as a directed graph G, and its formal expression is
[0020] G = {V, E, R, I, O, f}
[0021] Where:
[0022] R is the set of individual resources; all input interfaces are defined as I, all output interfaces are defined as O, and each interface represents an individual resource, indicating that this interface inputs / outputs this type of resource;
[0023] All input interfaces are defined as I, all output interfaces are defined as O, and there exists a mapping f: I ∪ O → R;
[0024] The node set is defined as V, and each node v contains two groups of interface sequences for input and output, that is
[0025]
[0026] The edge set is defined as E, and each edge is defined as a quadruple (starting node u, starting output interface o, ending node, ending input interface), with the constraint that the starting output interface and the ending input interface have the same resource type, that is
[0027]
[0028] Take u = {I u , O u}, v = {I v , O v}, there is o ∈ O u , i ∈ I v and f(o) = f(i).
[0029] Establish constraint relationships on the graph relying on the resource description method described in S1 to prevent deadlocks;
[0030] 1. Experiment to detect the timing constraint relationship between tasks;
[0031] 2. Experiment to detect the resource constraint relationship between tasks;
[0032] 3. The dynamic allocation relationship of experimental resources; for example, the activity of sample cutting will produce several samples for each laboratory to perform subsequent detections respectively. In this case, the proposed description scheme needs to be able to express both the static concept of "there is a sample allocation behavior" and the relationship between the "sample cutting" activity and each sub-activity; and it also needs to be able to carry the dynamic concept of "actively applying for samples" during the execution process, support the experimenters of sub-activities to actively apply for the samples produced by the parent activity to perform experiments, and record the whereabouts of the samples;
[0033] 4. Selective jump between experimental detection activities; on this basis, higher-order control flows such as loops and branches can be realized.
[0034] S3. Orchestrate the business workflow modeled in S2 to generate a workflow model that can be recognized and traced by the system;
[0035] Introduce the knowledge graph feature learning method (KGE, Knowledge Graph Embedding) to achieve automatic matching and intelligent recommendation of process nodes. Represent the workflow directed graph using knowledge graph triples:
[0036] Represent the numerical values in the workflow directed graph in the form of: (entity, attribute, attribute value);
[0037] Represent the facts in the workflow directed graph as triples <h, r, t>;
[0038] The loss function fr(h,t), where h,t is the vectorized representation of the two entities h and t of the triple.<h,r,t> When it is established, the expected fr(h,t) is minimum
[0039] Objective function: minΣ<h,r,t> ∈O fr(h,t), where O represents the set of all facts.
[0040] Using an embedding-based method, the nodes and edges in the knowledge graph are embedded in a low-dimensional vector space, and the knowledge graph is used to enrich the representation of item / user. Through distance-based translational models, when two entities belong to the same triple<h,r,t> , their vector representations should be close to each other in the projected space, and the loss function is:
[0041] fr(h,t)=||Wr,1h-Wr,2t|| uses the 1-norm.
[0042] The automatic matching of process nodes predicts possible workflow arrangements based on the input and output of the process nodes that have been included in the workflow, and pushes them to the user according to the degree of relevance.
[0043] The intelligent process recommendation will associate the subsequent process nodes that the user may need based on the entered work node template, and push them to the user according to the degree of relevance.
[0044] The correlation is predicted based on the workflow created historically, and is characterized in that the correlation is calculated based on factors such as the frequency of historical co-occurrence, the matching degree of input and output resources, and an artificially designed expert knowledge base, and the intelligence level continues to increase with user use.
[0045] S4. The laboratory personnel perform the inspection work, and the work status is synchronized to the display system and presented to the management personnel and customers.
[0046] The beneficial effects produced by the present invention are:
[0047] 1. The present invention can describe and manage all objects in the laboratory (including data, materials, laboratory equipment, etc.) in a unified manner, introduce available numbers and lock mechanisms, abstract the laboratory workflow into a directed graph, and construct a unified workflow model; realize the unified management of the exchange of information and materials between laboratories on marine scientific research ships, which is conducive to improving the coordination and collaborative work capabilities between laboratories;
[0048] 2. The present invention uses a knowledge graph intelligent recommendation method to automatically match the input and output of the process nodes included in the workflow to predict possible workflow arrangements, and push them to users based on relevance, thereby improving the efficiency of laboratory work. Brief Description of the Drawings
[0049] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0050] Figure 1 is the flowchart of the method according to the embodiment of the present invention;
[0051] Figure 2 is an example diagram of the laboratory workflow according to the embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the workflow node prediction according to the embodiment of the present invention. Detailed Embodiments
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] The on-board laboratories of oceanographic research vessels generally include the following laboratories: basic geology laboratory, paleomagnetism laboratory, geophysics laboratory, inorganic geochemistry laboratory, organic geochemistry laboratory, drilling technology laboratory, hydrate laboratory, microbiology laboratory and marine science laboratory. These laboratories jointly complete tasks such as collection, data sorting and analysis, and report generation.
[0055] The workflow of the on-board laboratories of oceanographic research vessels is divided into three parts: First, the marine science laboratory and the geophysics laboratory collect routine data. The information collected includes ocean depth, current profile, meteorology, water temperature information, etc. Then comes the drilling process and sample collection. During the drilling process, the geophysics laboratory and the drilling technology laboratory obtain well logging data and drilling fluid in real time, and then analyze them and provide timely feedback for service decision-making during the drilling process. The cores collected during the drilling process are cut and coded by the basic geology laboratory to make core thin sections. In addition, the samples collected include microbial samples, sediment samples, and hydrate samples collected by ROV. The samples will be sent to laboratories including paleomagnetism laboratory, hydrate laboratory, basic geology laboratory, inorganic chemistry laboratory, microbial laboratory, and organic geochemistry laboratory for experimental analysis. Among them, the paleomagnetism laboratory analyzes the remanent magnetization state of magnetic rocks from the aspect of the influence of the geomagnetic field to obtain information such as the characteristics of the geomagnetic field that caused its magnetization and the paleogeographic location of the plate. The microbial laboratory cultivates and observes the microbial components in the samples to analyze their microbial characteristics. Then the organic geochemistry laboratory and the inorganic geochemistry laboratory will analyze the material composition and its content of the samples from different chemical composition perspectives respectively. The relevant geological information obtained by the basic geology laboratory will also participate in the data interpretation process of these laboratories to improve the accuracy of data interpretation. A part of the analyzed data can be analyzed by the geophysics laboratory for wellbore data.
[0056] As Figure 1 shown, the present invention provides a general workflow modeling method for oceanographic research vessels, including the following steps:
[0057] S1. Uniformly describe the experimental key factors such as data, materials, experimental equipment, etc., and generate references or representations of relevant resources in the system;
[0058] Specifically as follows:
[0059] 1.1) Decompose all experimental resources into several distributable individual resources, the distribution number of each individual resource is greater than or equal to 1, and the access of each individual resource by external nodes is exclusive;
[0060] 1.2) Define the available number for each individual resource, and the available number represents the distributable quantity, that is, the quantity that can be accessed simultaneously;
[0061] 1.3) Determine the relationship between the available number of individual resources, the distribution mode, and node dependencies according to the available number situation of individual resources, as follows:
[0062] For individual resources that can be infinitely distributed, obtain their distribution methods by setting the access timing relationship between external nodes;
[0063] For the individually distributed resources, their distribution methods are obtained by setting the parallel execution relationships among external nodes;
[0064] For the resources with limited distribution and non - replenishment, by setting the branch selection among external nodes, once a branch is selected, it will cause the invalidation of other branches;
[0065] For the resources with limited distribution and replenishment, by setting the node relationships of sequential execution among external nodes where the order is irrelevant;
[0066] Based on the four types of resource - dependency relationships in step 1.3), five basic operations of resources can be constructed: 1. Produce resources, generate new resources and register them into the system; 2. Modify resources, including modifying content, status or even deleting; 3. Consume resources; 4. Access resources; 5. Replenish resources, such as returning the samples after lending.
[0067] The related resources include sample resources, document resources and other objective existences that can satisfy the above five basic operations;
[0068] All objective existences with the above - mentioned five basic operations can be abstracted into the resource concept.
[0069] The abstracted document resources in the system represent various semi - structured or unstructured data (as well as structured data carried in text form), and are organized and managed in the form of files. In principle, document resources are all infinitely distributed.
[0070] S2. Conduct unified workflow modeling for the intra - laboratory and inter - laboratory work processes, construct the modeling expression of laboratory testing workflows, and associate the relevant resource references described in S1;
[0071] Such as Figure 2 , which is an example diagram of the laboratory workflow. A completed task is decomposed into several small tasks, the sequence is deconstructed, and finally the workflow is obtained.
[0072] A directed graph is introduced, and based on the resource description method described in S1, constraint relationships on the graph are established to prevent deadlocks, enabling the following expressions: 1. Detecting the temporal relationship between activities. For example, detecting that detection A must occur after detection B; 2. Detecting the resource constraint relationship between activities. For example, there is a resource constraint relationship between multiple detections performed on a sample in sequence; 3. The output of the experiment may have dynamic allocation. For example, the activity of sample cutting will produce several samples for each laboratory to perform subsequent detections respectively. In this case, the proposed description scheme needs to be able to express both the static concept of "there is a sample allocation behavior" and the relationship between the "sample cutting" activity and each sub-activity; and it also needs to be able to carry the dynamic concept of "actively applying for samples" during the execution process, support the experimenters of sub-activities to actively apply for the samples produced by the parent activity to perform experiments, and record the whereabouts of the samples; 4. Supporting selective jumps between detection activities.
[0073] Based on the unified node modeling method, a directed graph feature graph is introduced, and the resource description method is established on the directed graph, which can prevent deadlock problems, detect the temporal relationship between activities, detect the existing resource constraint relationship between activities, and implement more advanced control flows such as loops and branches.
[0074] The modeling method defines the process as a directed graph, and its formal expression is
[0075] G = {V, E, R, I, O, f} #(1)
[0076] Where:
[0077] The set of resource types is defined as R;
[0078] All input interfaces are defined as I, all output interfaces are defined as O, and there exists a mapping f: I ∪ O → R;
[0079] The point set is defined as V, each node contains two groups of interface sequences of input and output, and each interface is defined as a resource type, representing that this interface inputs / outputs resources of this type, that is
[0080]
[0081] The edge set is defined as E, and each edge is defined as a quadruple (starting node, starting output interface, ending node, ending input interface), and it is restricted that the starting output interface and the ending input interface have the same resource type, that is
[0082]
[0083] Take u = {I u , O u}, v = {I v , O v}, and there is o ∈ Ou , where \(i\in I\) v and \(f(o)=f(i)\).
[0084] The process is characterized by the following basic operations: 1. Activate the process; 2. Consume the candidate input resources of the consumption node and activate the node or ignore the input and forcefully activate the node; 3. Enter the resources and bind them to the output interface of the node; 4. Terminate the node execution.
[0085] The node is characterized by the following basic states: 1. Newly created, representing that the process is newly created and not submitted for review; 2. Available, representing that the process application for review has passed; 3. In execution, representing that the process is in execution; 4. Execution completed, entering this state when all node states in the process are execution completed or unavailable; 5. Terminated, entering this state when all node states in the process are execution completed, unavailable or terminated, and at least one node state is terminated.
[0086] S3. Use the workflow orchestration system to orchestrate the business workflow modeled in S2 to generate a workflow model that can be recognized and traced by the system;
[0087] The knowledge graph feature learning method (KGE, Knowledge Graph Embedding) is introduced to realize the automatic matching and intelligent recommendation of process nodes. Represent the workflow directed graph using knowledge graph triples:
[0088] Knowledge is represented as SPO (Subject - Predicate - Object) triples, where P corresponds to the predicate in the subject - verb - object structure and is divided into two forms, one is property and the other is relation, that is, it can be subdivided into two forms: (entity 1, relation, entity 2) and (entity 1, property, property value).
[0089] Represent the numerical values in the workflow directed graph in the form of: (entity, property, property value);
[0090] Represent the facts in the workflow directed graph as triples \(\langle h,r,t\rangle\);
[0091] The loss function \(f_r(h,t)\), where \(h,t\) are the vectorized representations of the two entities \(h\) and \(t\) of the triple. When the fact \(\langle h,r,t\rangle\) holds, it is expected that \(f_r(h,t)\) is minimized
[0092] The objective function: \(\min\sum_{\langle h,r,t\rangle\in O}f_r(h,t)\), where \(O\) represents the set of all facts.
[0093] Using an embedding-based method, the nodes and edges in the knowledge graph are embedded in a low-dimensional vector space, and the knowledge graph is used to enrich the representation of item / user. Through distance-based translational models, when two entities belong to the same triple<h,r,t> , their vector representations should be close to each other in the projected space, and the loss function is:
[0094] fr(h,t)=||Wr,1h-Wr,2t||
[0095] like Figure 3 , the automatic matching of process nodes will guess the possible workflow arrangement based on the input and output of the process nodes included in the workflow, and push it to the user according to the degree of relevance.
[0096] Intelligent process recommendations will associate the subsequent process nodes that the user may need based on the entered work node template, and push them to the user according to the degree of relevance.
[0097] The association degree is inferred based on the historically created workflows. Its characteristic is that it calculates the association degree based on factors such as the frequency of historical co-occurrence and the matching degree of input and output resources, combined with the manually designed expert knowledge base, and its intelligence level continues to improve with user use.
[0098] S4. The laboratory personnel perform the testing work, and the work status is synchronized to the system and presented to the management personnel and customers.
[0099] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
Claims
1. A general workflow modeling method for oceanographic research vessels, characterized in that, it includes the following steps: S1. Uniformly describe the experimental resources including data, materials, and experimental equipment in each laboratory of the oceanographic research vessel to generate references or representations of relevant resources; Specifically as follows: 1.1) Decompose all experimental resources into several distributable individual resources, where the distribution number of each individual resource is greater than or equal to 1, and the access of external nodes to each individual resource is exclusive; 1.2) Define the available number for each individual resource, and the available number represents the distributable quantity, that is, the quantity that can be accessed simultaneously; 1.3) Determine the relationship between the available number of individual resources and the distribution mode and node dependencies according to the available number situation of individual resources, as follows: For individual resources that can be infinitely distributed, obtain their distribution methods by setting the access timing relationship between external nodes; For multi-distributed individual resources, obtain their distribution methods by setting the parallel execution relationship between external nodes; For resources with limited distribution and non-supplementary, obtain their distribution methods by setting branch selection between external nodes. After selecting one branch, it will cause the invalidation of other branches; For resources with limited distribution and supplementary, obtain their distribution methods by setting the node relationship of sequential execution but independent of order between external nodes; Based on the four resource dependency relationships in step 1.3), five basic operations of resources can be constructed: produce resources, generate new resources and register them in the system; Modify resources, including modifying content, status, and even deleting; Consume resources; Access resources; Supplement resources, including external supplementary resources and the return after sample lending; Related resources include sample resources, document resources, and other objective existences that can satisfy the above five basic operations; S2. Uniformly model the workflow inside and between laboratories, construct the workflow modeling expression of laboratory detection tasks, and associate the relevant resource references described in S1; The unified workflow modeling defines the process as a directed graph G, and its formal expression is G = {V, E, R, I, O, f} Where: R is the set of individual resources; all input interfaces are defined as I, all output interfaces are defined as O, and each interface represents an individual resource, indicating that this interface inputs / outputs this type of resource; All input interfaces are defined as I, all output interfaces are defined as O, and there exists a mapping f: I ∪ O → R; the node set is defined as V, and each node v contains two groups of interface sequences of input and output, that is The edge set is defined as E, and each edge is defined as a quadruple: the starting node u, the starting output interface o, the ending node v, and the ending input interface i; It is required that the starting output interface and the ending input interface have the same resource type, that is Take \(u = \{I u , O u \}\), \(v=\{I v , O v \}\), there exists \(o\in O u , i\in I v and \(f(o)=f(i)\); S3. Orchestrate the business workflow modeled in S2 to generate a workflow model that can be identified and traced by the system; S4. Experimental personnel perform detection business work, and the working status is synchronized to the display system and presented to managers and customers.
2. The general workflow modeling method for oceanographic research vessels according to claim 1, characterized in that, in step S2, based on the S1 resource description method, establish constraint relationships on the directed graph G to prevent deadlocks; The constraint relationships include: Temporal constraint relationships between experimental detection tasks; Resource constraint relationships between experimental detection tasks; Dynamic allocation relationships of experimental resources; Selective jumps between experimental detection activities.
3. The general workflow modeling method for an oceanographic research vessel according to claim 1, characterized in that, in step S3, the workflow directed graph is represented using knowledge graph triples: The numerical values in the workflow directed graph are represented in the form of: (entity, attribute, attribute value); The facts in the workflow directed graph are represented as triples <h, r, t>; The loss function fr(h, t), where h and t are the vectorized representations of the two entities h and t of the triple. When the fact <h, r, t> holds, it is expected that fr(h, t) is minimized.
4. The general workflow modeling method for an oceanographic research vessel according to claim 1, characterized in that, in step S3, a knowledge graph feature learning method is adopted to achieve automatic matching and intelligent recommendation of process nodes; Using an embedding-based method, the nodes and edges in the knowledge graph directed graph are embedded in a low-dimensional vector space, and then through a distance-based matching model, when two entities h and t belong to the same triple <h, r, t>, their vector representations Wr h and Wr t should be close to each other in the projected space, and the loss function is: fr(h,t) = ||Wr h -Wr t || Based on the inputs and outputs of the process nodes already included in the workflow, predict possible workflow arrangements for automatic matching of process nodes, and push them to the user in the order of relevance; Based on the entered work node templates, associate the subsequent process nodes that the user may need for intelligent recommendation of process nodes, and push them to the user in the order of relevance; The relevance is predicted based on the historically created workflow data, that is, the relevance is calculated based on two influencing factors: the frequency of historical co-occurrence and the matching degree of input and output resources.
Citation Information
Patent Citations
Knowledge service innovation method driven by polymorphic knowledge graph
CN112199515A