Marine transportation business process automation implementation method based on large language model
Through methods based on large language model and reinforcement learning model, automatic scripts of maritime business processes are generated, which solves the problems of long process mining cycle and high cost in the existing technology, and realizes efficient and readable automated script generation, which is highly adaptable.
Patent Information
- Application Number
- CN202510544762.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the automation of maritime business process, the method of obtaining log information through transforming software systems for process mining has problems of long cycles and high costs.
Using a method based on a large language model, a task order list is generated by collecting multimodal data, a reinforcement learning model is used to allocate activities, and an RPA automation script is generated to avoid transformation of the original system.
It realizes rapid mining of processes without modifying the existing system, reduces costs, improves work efficiency, ensures efficient and compatibility of automated scripts, strong adaptability, and good readability of generated scripts.
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of process automation, and particularly relates to a method for realizing the automation of the sea freight business process. Background Art
[0002] The sea freight business involves a large number of repetitive operations, such as booking order processing, customs declaration document generation, freight settlement, etc. These processes usually need to be operated across multiple systems (such as ERP, email, Excel, customs declaration platform). In the past, these processes relied on manual execution, which was not only inefficient but also prone to delays or losses due to operational errors.
[0003] With the development of sea freight related technologies, process automation has become a key means to improve efficiency and reduce human errors. Robot Process Automation (RPA) technology can simulate manual operations through software robots and automatically complete repetitive tasks, which is a research hotspot in this field. How to perform accurate process mining is the key to realizing process automation. The existing operation method is to track the status and events of specific data through pre-set buried point logs in the business software system. For example, the Chinese patent application with the publication number CN118585575A discloses a process mining method, device, electronic device and medium based on a graph database. It extracts case identifiers and information from event logs, stores activities in chronological order, and mines process edge information, and finally stores all information in the graph database to achieve process mining. However, this method requires secondary transformation of the software system, which has a long cycle and high transformation cost. Summary of the Invention
[0004] The present invention proposes a method for realizing the automation of the sea freight business process based on a large language model, and its purpose is to solve the problems of long cycle and high cost existing in obtaining log information through transforming the existing software system for process mining.
[0005] The technical solution of the present invention is as follows: A method for realizing the automation of the sea freight business process based on a large language model, the steps include: Step S1, collect multimodal data while the operator performs the target process operation, and obtain a task sequence list of the target process based on the multimodal data; Step S2, group the task items in the task sequence list; Step S3, use the large language model to assign activity names to each group; Step S4, use the reinforcement learning model to assign a callable tool to each activity name respectively; Step S5: Input the task sequence list, the activity names corresponding to each task item, and the tools corresponding to each activity name into the large language model to obtain process orchestration data, and then convert the process orchestration data into RPA automation scripts.
[0006] As a further improvement of the method for automating the maritime business process based on the large language model, in step S1, the method for collecting multimodal data is as follows: Real-time capture the multimodal data during the operation of the operator. The multimodal data includes audio data, screen image data, and UI interaction data. When the operator triggers a recording operation, record the current multimodal data and add it to the temporary list.
[0007] As a further improvement of the method for automating the maritime business process based on the large language model: In step S1, obtaining the task sequence list of the target process based on multimodal data means: Read each piece of multimodal data in the temporary list. One piece of multimodal data corresponds to one task item, and send each piece of multimodal data to the large language model that supports multimodality to obtain the task description information corresponding to each task item.
[0008] As a further improvement of the method for automating the maritime business process based on the large language model, the specific process of step S2 is as follows: Step S2-1: Convert the task description information of each task item into an embedding vector; Step S2-2: Calculate the semantic similarity between the embedding vectors; Step S2-3: Group the task items in a clustering manner based on the semantic similarity between the embedding vectors and the preset clustering parameters.
[0009] As a further improvement of the method for automating the maritime business process based on the large language model: In step S3, for the current group, input the task description information of all task items in the group and the corresponding time into the large language model, so that the large language model, under the guidance of the prompt words, combines the pre-set set of activity names to assign the best activity name to the group.
[0010] As a further improvement of the method for automating the maritime business process based on the large language model: Use the RPA process operation historical data to train the reinforcement learning model: Define each task item, the context information of the task item, and the activity to which the task item belongs in the historical data as a structured state. The context information includes one or more previous and subsequent task items for the targeted task item. Define the tools used by each task item in the historical data as actions, and define the task completion rate and execution efficiency corresponding to the implementation of the actions of each task item as rewards, thereby obtaining a training set; then use the training set to train the reinforcement learning model to learn the association relationship between the tools and activities. When using the trained reinforcement learning model to allocate tools, the input of the reinforcement learning model is an activity name and the context corresponding to the activity name. The context includes the previous and subsequent task items and the activity names corresponding to the task items, and the output is the tool recommended for the activity.
[0011] As a further improvement of the method for implementing the automation of the maritime business process based on the large language model: in step S5, add the process choreography data template and examples to the prompt words input to the large language model, and also add the usage documents of all the selected tools, so as to drive the large language model to output the process choreography data in the format of structured text.
[0012] As a further improvement of the method for implementing the automation of the maritime business process based on the large language model: the process choreography data includes a process input interface and several nodes arranged in sequence and corresponding one-to-one to the task items; The process input interface receives external incoming data when the RPA automation script runs; The nodes include the tool call interface, node input interface, node processing function, and node output interface corresponding to the task items; the node input interface is used to dock with the process input interface or the output interface of the previous node to obtain data, and the obtained data is passed into the node processing function for processing; the tool call interface is used to pass the result data processed by the node processing function into the called tool to complete the automation operation of the current task item; the node output interface is used to receive the result data output by the node processing function within the current node and the data returned by the tool call interface.
[0013] As a further improvement of the method for implementing the automation of the maritime business process based on the large language model, it also includes: Step S6: Test the RPA automation script obtained in step S5, and modify the RPA automation script for the abnormal situations that occur.
[0014] As a further improvement of the method for implementing the automation of the maritime business process based on the large language model, it also includes: Step S7: Apply the modified RPA automation script to the production environment, continuously monitor the running effect, and optimize the generation process of the RPA automation script based on the monitoring situation; Collect the following data during monitoring: the success rate and error rate of the automated tasks, the reduction ratio of the process automation execution time relative to the original execution time, and the satisfaction degree feedback by users; update the clustering parameters used in the grouping in step S2, the set of activity names used by the large language model in step S3, and the historical data used for training the reinforcement learning model based on the collected data, retrain the reinforcement learning model, and optimize the generation process of the RPA automation script.
[0015] Compared with the prior art, the present invention has the following positive effects: 1. The present invention uses a multimodal large model to parse the task sequence list obtained from the actual operation process, so that the process can be mined without modifying the original system, shortening the mining time, reducing the cost, and eliminating the potential hazards brought by system modification.
[0016] 2. The process from the task sequence list to the generation of RPA automation scripts in the present invention is automatically realized by a program and a large language model. RPA developers do not need to understand the current specific shipping business and the operation process of the software system in detail. They only need to maintain the prompt words of the large model and the corresponding various data to complete the generation of the scripts, thus greatly improving the work efficiency. At the same time, relying on the excellent semantic processing ability of the large language model, it can process complex processes more accurately and efficiently, with stronger adaptability.
[0017] 3. The present invention introduces a reinforcement learning model to recommend tools for each activity. The advantage of the reinforcement learning model is that it pays more attention to the relationship between structured states, actions, and rewards. Since the structured state includes not only the information of the current task item but also the context information in historical data, the tools recommended by the reinforcement learning model are not only based on a specific activity name itself, but also based on the previous and subsequent tasks and activities, and more importantly, based on the completion rate in historical data to fully evaluate the compatibility of the connection combinations of different tools before and after, so as to recommend the optimal tool, solve the failure problem caused by the compatibility between tools during the automated connection of adjacent tasks, and ensure that the obtained automation script is optimal.
[0018] 4. The process orchestration data of the present invention is in a structured text format. The node includes the tool call interface, node input interface, node processing function, and node output interface of the current task item. It can not only quickly complete the conversion from process orchestration data to RPA automation scripts, but also has good readability, which is beneficial for developers to discover problems in time and modify and optimize.
[0019] 5. After the automation script is put into use, the present invention further optimizes grouping, activity naming, and the historical data of the reinforcement learning model based on the collected data, and continuously iterates to improve the automated execution effect. Detailed implementation manners
[0020] The technical solution of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0021] A method for automating the shipping business process based on a large language model, the steps include: Step S1: While the operator is performing the target process operation, collect multimodal data, and obtain a task sequence list of the target process based on the multimodal data.
[0022] The target process is usually a highly repetitive process common in the shipping business, such as: generating a shipping report, invoice processing, approval flow, extracting customer information, data entry, or data cleaning, etc. It can also be a process involving cross-system operations, such as the process of extracting information from an email and updating it to the booking system.
[0023] The method of collecting multimodal data is as follows: Real-time capture the multimodal data during the operator's operation. The multimodal data includes audio data, screen image data, and UI interaction data (such as web interaction). When the operator triggers a recording operation (which can be button-triggered or voice-triggered), record the current multimodal data and add it to the temporary list.
[0024] Obtaining the task sequence list of the target process based on the multimodal data means: Read each piece of multimodal data in the temporary list. One piece of multimodal data corresponds to one task item. Send each piece of multimodal data to a large language model that supports multimodality to obtain the task description information corresponding to each task item. The task description information includes the task name, operation instructions, software window, etc.
[0025] Table 1: Task sequence list of a certain process.
[0026] Task Item Number Task Name Software Window Operation Instructions 1 Open Approval Form Approval System Open the form page of the "Approval System" 2 Fill in the approval form Approval System Enter approval information in the form 3 Submit Approval Form Approval System Click the submit button to complete form submission 4 Open Email Email System Open Outlook to view emails awaiting approval 5 Reply to Email Email System Reply to the customer's email to confirm the approval status 6 View Approval Records Approval System View historical approval records in the approval system In this step, the open-source model "visualglm-6b" (a multimodal large language model) from Tsinghua University is used, combined with an audio processing library (such as speech_recognition) and video frame extraction technology (cv2.VideoCapture of OpenCV) to process the captured data.
[0027] Step S2: Group the task items in the task sequence list.
[0028] The specific process is as follows: Step S2-1: Convert the task description information of each task item into an embedding vector.
[0029] An embedding vector is a way to represent the semantic information of text (or other data) using a high-dimensional real-valued vector. Through a pre-trained language model (such as BERT, GPT, word2vec, or text-embedding-ada-002), the semantics of the text can be encoded into a fixed-length vector.
[0030] For example, for two task descriptions: Task 1: "Open the approval form"; Task 2: "Fill in the form content"; After being processed by an embedding model (such as the embedding API of GPT), they may be converted into the following vectors: Vector for Task 1: [0.85, 0.27, 0.63…]; Vector for Task 2: [0.83, 0.25, 0.65…].
[0031] Step S2-2: Calculate the semantic similarity between each embedding vector.
[0032] In this embodiment, the cosine similarity between the embedding vectors is used as the semantic similarity.
[0033] Each dimension of the vector is a real number, representing the projection of the text on certain semantic features.
[0034] The calculation method of the cosine similarity between any two embedding vectors A and B is: .
[0035] Table 2: Semantic similarity between the 6 task items in Table 1.
[0036] Task Item Number 1 2 3 4 5 6 1 1 0.99 0.98 0.15 0.14 0.97 2 0.99 1 0.98 0.16 0.15 0.96 3 0.98 0.98 1 0.14 0.13 0.95 4 0.15 0.16 0.14 1 0.99 0.16 5 0.14 0.15 0.13 0.99 1 0.15 6 0.97 0.96 0.95 0.16 0.15 1 Step S2-3: Group the task items based on the semantic similarity between the embedding vectors and the preset clustering parameters.
[0037] The clustering parameters include: the minimum number of groups min_samples = 2, and the similarity threshold eps = 0.05. During the clustering process, texts with similar semantics will be mapped to a similar vector space.
[0038] Clustering can use clustering methods such as k-means.
[0039] In this embodiment, task items 1, 2, 3, 5 are grouped into Group 1, and task items 4, 5 are grouped into Group 2.
[0040] Table 3: Grouping results.
[0041] Group Number Task Item Number Remarks 1 1, 2, 3, 6 Tasks related to approval 2 4, 5 Tasks related to email interaction Step S2 has the following advantages: 1. High flexibility: Based on semantic similarity, it adapts to different task descriptions and language expressions.
[0042] 2. Strong precision: Capturing deep semantic information through embedding vectors, avoiding grouping errors caused by surface text differences.
[0043] 3. Wide range of applicable scenarios: It can be used in complex business processes, especially in scenarios with diverse task descriptions.
[0044] In this step, through a grouping strategy based on task semantic similarity, combined with text embedding and clustering algorithms, efficient and accurate task grouping can be achieved. This method is particularly suitable for business processes with complex and diverse task descriptions, such as approval processes, customer service processes, etc. Combining the actual business needs and the capabilities of pre-trained language models, the grouping effect can be further optimized.
[0045] Step S3: Use a large language model to assign activity names to each group.
[0046] For the current group, input the task description information of all task items within the group and the corresponding time into the large language model, so that the large language model, under the guidance of the prompt words and in combination with the pre-set activity name set, assigns the best activity name to the group.
[0047] The prompt words in this step should clearly define the input and expected output, and reference the activity name set as a knowledge base to provide reference for the large language model.
[0048] In this embodiment, the activity name of Group 1 is "Approval Operation", and the activity name of Group 2 is "Email Communication".
[0049] Step S4: Use a reinforcement learning model to assign a callable tool to each activity name respectively.
[0050] The reinforcement learning model is Deep Q-Network. Using the RPA process operation historical data, the reinforcement learning model is trained through experience replay and target network: Define each task item, the context information of the task item, and the activity to which the task item belongs in the historical data as a structured state. The context information includes more than one previous and next task item for the targeted task item. Define the tools used by each task item in the historical data as actions, and define the task completion rate and execution efficiency corresponding to the implementation of the actions of each task item as rewards, thus obtaining a training set. Then use the training set to train the reinforcement learning model to learn the association relationship between tools and activities.
[0051] The tool refers to the connectors available in the RPA system, which need to be sorted out in advance. For example: - Email processing: Outlook, Sendmail.
[0052] - Data entry / management: SharePoint, Google Forms.
[0053] - Approval process: Microsoft Forms, Planner.
[0054] - Notification services: RSS, Teams.
[0055] When using the trained reinforcement learning model to assign tools, the input of the reinforcement learning model is an activity name and the context corresponding to the activity name. The context includes the previous and subsequent task items and the activity names corresponding to the task items, and the output is the tool recommended for the activity.
[0056] In this embodiment, the tool assigned for "approval operation" is Microsoft Forms, and the tool assigned for "email communication" is Outlook.
[0057] Step S5: Input the task sequence list, the activity names corresponding to each task item, and the tools corresponding to each activity name into the large language model to obtain process orchestration data, and then convert the process orchestration data into an RPA automation script.
[0058] In this step, add the process orchestration data template and example (in xml format) to the prompt words input to the large language model, and also add the usage documents of all the selected tools, so as to drive the large language model to output the process orchestration data in the format of structured text.
[0059] The process orchestration data includes a process input interface and several nodes arranged in sequence and corresponding to the task items one by one. The process input interface receives external input data when the RPA automation script runs. Each node contains a tool call interface, a node input interface, a node processing function, and a node output interface for the corresponding task item. The node input interface is used to dock with the process input interface or the output interface of the previous node to obtain data, and pass the obtained data to the node processing function for processing. The tool call interface is used to pass the result data processed by the node processing function to the called tool to complete the automated operation of the current task item. The node output interface is used to receive the result data output by the node processing function within the current node and the data returned by the tool call interface.
[0060] If there is no need to process data within the node, the node processing function can directly pass the received data to the tool call interface and the node output interface.
[0061] When converting the process orchestration data into an RPA automation script, it is preferred to use a program for conversion. Since the format of the process orchestration data is structured text, the conversion to an RPA automation script can be completed through string concatenation and template application.
[0062] Step S6: Test the RPA automation script obtained in step S5, and modify the RPA automation script for the abnormal situations that occur.
[0063] It should be ensured during testing that: - Each task is triggered and executed correctly.
[0064] - The connector is seamlessly integrated with the business system.
[0065] - The output result meets the expectations.
[0066] Step S7: Apply the modified RPA automation script to the production environment, continuously monitor the running effect, and optimize the generation process of the RPA automation script based on the monitoring situation.
[0067] Collect the following data during monitoring: the success rate and error rate of the automation tasks, the reduction ratio of the process automation execution time compared to the original execution time, and the satisfaction degree of the user feedback, etc. Update the clustering parameters in step S2, the set of activity names in step S3, and the historical data in step S4 based on the collected data, retrain the reinforcement learning model, and optimize the generation process of the RPA automation script.
[0068] It should be noted that for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. The scope of the present invention is defined by the claims rather than the above description.
Claims
1. A method for implementing the automation of the maritime transportation business process based on a large language model, characterized in that: The steps include: Step S1: Collect multimodal data while the operator performs the target process operation, and obtain a task sequence list of the target process based on the multimodal data; Step S2: Group the task items in the task sequence list; Step S3: Use the large language model to assign activity names to each group; Step S4: Use the reinforcement learning model to assign a callable tool to each activity name respectively; Step S5: Input the task sequence list, the activity names corresponding to each task item, and the tools corresponding to each activity name into the large language model to obtain process choreography data, and then convert the process choreography data into an RPA automation script.
2. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 1, wherein, In step S1, the method for collecting multimodal data is: real-time capture of multimodal data during the operator's operation. The multimodal data includes audio data, screen image data, and UI interaction data. When the operator triggers the recording operation, the current multimodal data is recorded and added to the temporary list.
3. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 1, characterized in that: In step S1, obtaining the task sequence list of the target process based on the multimodal data means: reading each piece of multimodal data in the temporary list. One piece of multimodal data corresponds to one task item, and each piece of multimodal data is respectively sent to the large language model that supports multimodality to obtain the task description information corresponding to each task item.
4. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 3, wherein, The specific process of step S2 is: Step S2-1: Convert the task description information of each task item into an embedding vector; Step S2-2: Calculate the semantic similarity between the embedding vectors; Step S2-3: Group the task items in a clustering manner based on the semantic similarity between the embedding vectors and the preset clustering parameters.
5. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 3, characterized in that: In step S3, for the current group, input the task description information of all task items in the group and the corresponding time into the large language model, so that the large language model, under the guidance of the prompt words, combines the pre-set activity name set to assign the best activity name to the group.
6. The method for implementing the automation of the maritime transportation business process based on the large language model according to claim 1, wherein: Use the RPA process operation historical data to train the reinforcement learning model: Define each task item, the context information of the task item, and the activity to which the task item belongs in the historical data as a structured state. The context information includes more than one previous and subsequent task item for the targeted task item. Define the tools used by each task item in the historical data as actions, and define the task completion rate and execution efficiency corresponding to the implementation of the actions of each task item as rewards, so as to obtain a training set; then use the training set to train the reinforcement learning model to learn the association relationship between the tool and the activity; When using the trained reinforcement learning model to assign tools, the input of the reinforcement learning model is an activity name and the context corresponding to the activity name. The context includes the previous and subsequent task items and the activity names corresponding to the task items, and the output is the tool recommended for the activity.
7. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 1, wherein: In step S5, add the process choreography data template and examples to the prompt words input into the large language model, and also add the usage documents of all the selected tools, driving the large language model to output the process choreography data in the format of structured text.
8. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 1, wherein: The process choreography data includes a process input interface and several nodes arranged in sequence and corresponding to task items one by one; The process input interface receives externally incoming data during the operation of the RPA automation script; The nodes include a tool call interface, a node input interface, a node processing function, and a node output interface corresponding to the corresponding task items; the node input interface is used to dock with the process input interface or the output interface of the previous node to obtain data, and the obtained data is passed into the node processing function for processing; The tool call interface is used to pass the result data processed by the node processing function into the called tool to complete the automated operation of the current task item; The node output interface is used to receive the result data output by the node processing function within the current node and the data returned by the tool call interface.
9. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 1, characterized in that, It also includes: Step S6: Test the RPA automation script obtained in step S5, and modify the RPA automation script for the abnormal situations that occur.
10. The method for realizing the automation of the maritime transportation business process based on the large language model according to claim 9, wherein, It also includes: Step S7: Apply the modified RPA automation script to the production environment, continuously monitor the operation effect, and optimize the generation process of the RPA automation script based on the monitoring situation; The following data is collected during monitoring: the success rate and error rate of the automated tasks, the reduction ratio of the process automation execution time relative to the original execution time, and the satisfaction feedback by users; based on the collected data, update the clustering parameters used in the grouping in step S2, the set of activity names used by the large language model in step S3, and the historical data used for training the reinforcement learning model, retrain the reinforcement learning model, and optimize the generation process of the RPA automation script.
Citation Information
Patent Citations
Method for constructing pre-training subject corpus of digital human teacher multi-mode large language model
CN117076693A
RPA process automatic construction method and system combining large language model and reinforcement learning
CN117634867A
Training data set construction method and device, equipment, storage medium and program product
CN117669774A
Multi-modal personalized content generation method
CN118260483A
Session type recommendation method and system based on semantic similarity and clustering model
CN118312606A