Robot process automation method based on deep learning
By building multimodal prediction models, identifying classification models and anomaly detection models, and integrating them into robot process automation programs, the problem of difficulty in handling multimodal data and evaluating the performance of automation processes in the existing technology is solved, and a more efficient and intelligent automated process is achieved.
Patent Information
- Application Number
- CN202510223660.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art is difficult to construct multimodal prediction models, identification classification models and anomaly detection models, it is difficult to integrate these models into robot process automation programs, and it is difficult to evaluate the performance of robot process automation processes.
By obtaining structured and unstructured data, preprocessing and feature extraction, generating multimodal feature vectors, using deep learning algorithms to build multimodal prediction models, identify classification models and anomaly detection models, integrate these models into robot process automation programs, and conduct business process testing and performance evaluation.
The robot process automation program can handle more complex data types, improve the intelligence level of automation processes, significantly improve the execution efficiency of business processes, reduce the possibility of human error, reduce operational costs, and adapt to changes in different business needs.
Smart Images

Figure CN120145005A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a robot process automation method based on deep learning. Background Art
[0002] Robot Process Automation (RPA) is a technology that uses software robots or robot simulations to integrate human interactions in digital systems to execute business processes. These robots can learn and repeat a series of tasks, thereby improving efficiency, reducing errors, and lowering operating costs. Traditional RPA usually relies on explicit programming rules to process structured data and application interfaces. Deep learning is a subset of machine learning that uses multi-layer neural networks to simulate the way the human brain processes information. With the development of deep learning technology, it provides new solutions for robot process automation.
[0003] The following problems exist in the prior art, including: it is difficult to construct multi-modal prediction models, recognition and classification models, and anomaly detection models to analyze business data; it is difficult to integrate various constructed models into robot process automation programs respectively; and finally, it is difficult to evaluate the performance of robot process automation processes. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a robot process automation method based on deep learning to solve the above problems. For this reason, the first aspect of the present invention provides a robot process automation method based on deep learning, including the following steps:
[0005] S1: Obtain a data acquisition request sent by a client, and collect business data. The client has a robot process automation program, and the business data includes: structured data and unstructured data;
[0006] S2: Preprocess the collected structured data, including: data cleaning, denoising, and standardization; preprocess the collected unstructured data, including: text tokenization, stop word removal, image enhancement, grayscale conversion, and denoising;
[0007] S3: Perform feature extraction processing on the collected and preprocessed structured data and unstructured data; generate a multi-modal feature vector according to the results of feature extraction;
[0008] S4: Construct a multi-modal prediction model using a deep learning algorithm based on the multi-modal feature vector; construct a recognition and classification model by using a deep learning algorithm; analyze abnormal situations in business data by constructing an anomaly detection model; integrate the multi-modal prediction model, the recognition and classification model, and the anomaly detection model into the robot process automation program respectively and conduct business process tests;
[0009] S5: Comprehensively evaluate the performance of the robotic process automation process, and iteratively update the robotic process automation process according to the evaluation results.
[0010] Preferably, the step S1 includes the following steps:
[0011] Receive the data acquisition request sent by the client using the API interface or the web server, wherein the request content includes: the type and source of the business data;
[0012] Automatically identify the type and source of the business data in the request content by using the deep learning model;
[0013] By obtaining the business data sample set, the business data sample set includes various business data types and sources; wherein, the business data types include: structured business data and unstructured business data; the structured business data includes: fields and data values in the database or table; the unstructured business data includes: text, video images, and voices; the business data sources include: databases and web pages;
[0014] Label each business data sample, label the data type and source, and input the labeled business data sample set into the deep learning model for training; according to the trained deep learning model and input the request content into the deep learning model, automatically identify the type and source of the business data in the request content.
[0015] According to the type and source of the business data automatically identified by the deep learning model, simulate the manual operation process by using the robotic process automation (RPA) program; connect the RPA program to the database or web page; by executing the RPA script, automatically complete the data acquisition operation and output the result, and the result is the business data collected from the database or web page.
[0016] Preferably, the step of performing feature extraction processing on the collected and preprocessed structured data and unstructured data in the step S3 includes the following steps:
[0017] Perform feature extraction according to the collected and preprocessed structured data, including: numerical features and categorical variable features; perform normalization or standardization processing on the numerical features; use the one-hot encoding method to convert the categorical variables into a set of binary features, where each categorical variable feature represents a category;
[0018] According to the preprocessed text, extract word embedding or sentence embedding features by using natural language processing technology;
[0019] Based on the preprocessed video images, use the computer vision technology OpenCV to process the video image data and extract video image features including: color features, texture features, and shape features;
[0020] Based on the preprocessed language, segment the audio signal into frames with a length of 25 milliseconds, and extract language features such as Mel Frequency Cepstral Coefficients, spectral centroid, and spectral bandwidth from each frame of audio.
[0021] Preferably, in step S3, according to the results of feature extraction, perform feature fusion to generate a multi-modal feature vector, including the following steps:
[0022] Perform feature alignment on the extracted structured data features, text features, video image features, and speech features using timestamps, and use the principal component analysis method for feature dimensionality reduction;
[0023] Use the attention mechanism to calculate the weights of each feature for the structured data features, text features, video image features, and speech features after feature alignment and feature dimensionality reduction, and perform feature fusion on the weighted features to emphasize important features; perform joint representation learning through a shared network layer or a cross-modal loss function, and by constructing a multi-layer perceptron network, input each feature after joint representation learning into the multi-layer perceptron network to generate a multi-modal feature vector.
[0024] Preferably, in step S4, construct a multi-modal prediction model according to the multi-modal feature vector using a deep learning algorithm, including the following steps:
[0025] According to the generated multi-modal feature vector, construct a multi-modal prediction model using a deep learning model to analyze and predict business data;
[0026] Input the multi-modal feature vector into the multi-modal prediction model for training;
[0027] Through the real-time collected and preprocessed business data, and extract the new multi-modal vector of the real-time collected business data; input the new multi-modal vector into the trained multi-modal prediction model for analysis and prediction and output the prediction result, and the prediction result is the change trend of the predicted business data.
[0028] Preferably, in step S4, construct an identification and classification model by using a deep learning algorithm, including the following steps:
[0029] Construct an identification and classification model by using a deep learning algorithm, input the multi-modal feature vector into the identification and classification model for training, and during the training process, perform supervised learning using label data and optimize the identification and classification model through the backpropagation algorithm and the gradient descent method;
[0030] Input the real-time collected and preprocessed business data into the trained recognition and classification model; the recognition and classification model performs behavior pattern recognition and classification on the input business data.
[0031] Preferably, in step S4, analyzing the abnormal conditions in the business data by constructing an anomaly detection model includes the following steps:
[0032] Construct an anomaly detection model according to the deep learning algorithm; input the multi-modal feature vectors into the anomaly detection model for training; during the training process, through iterative optimization, enable the anomaly detection model to learn the distribution characteristics of normal business data;
[0033] Input the real-time collected and preprocessed business data into the trained anomaly detection model. The anomaly detection model evaluates the input business data according to the learned normal data distribution characteristics to detect whether the real-time business data is abnormal.
[0034] Preferably, in step S4, integrating the multi-modal prediction model, the recognition and classification model, and the anomaly detection model into the robotic process automation program respectively and conducting business process testing includes the following steps:
[0035] Use the designer of the robotic process automation tool to create the framework of the automation process, including: the starting point, ending point, intermediate steps, and decision points of the business process;
[0036] Deploy the trained multi-modal prediction model, recognition and classification model, and anomaly detection model to the server respectively;
[0037] Design the automated business process in the designer of the robotic process automation tool;
[0038] Through the API interface, embed the calls to the multi-modal prediction model, recognition and classification model, and anomaly detection model in the robotic process automation script, and design an exception handling mechanism in the automated business process; when a call fails or an unforeseen result is returned in one of the models, give an exception warning, otherwise continue to call various models for analysis;
[0039] Conduct individual tests on each component in the business process and conduct integration tests on the business process; judge whether all components work together by testing the entire robotic process automation business process; optimize the robotic process automation business process according to the test results; deploy the tested and optimized robotic process automation business process to the production environment and start automatically executing the business process;
[0040] According to the business process of the deployed robotic process automation, by monitoring the running status of the business process of the robotic process automation in real time, key performance indicators are collected, including: processing time, error rate, and response time.
[0041] Preferably, step S5 includes the following steps:
[0042] Evaluate according to the trained multi-modal prediction model, recognition and classification model, and anomaly detection model respectively, and calculate the F1 value of the key performance indicator respectively; by calculating the model stability coefficient formula:
[0043] Obtain the model stability coefficient M; where F1 D represents the F1 value of the multi-modal prediction model; F1 S represents the F1 value of the recognition and classification model; F1 J respectively represent the F1 values of the anomaly detection model;
[0044] Obtain the business process efficiency η by calculating the ratio of the processing time of the robotic process automation business process to the processing time of the manual business process;
[0045] By calculating the business process stability formula:
[0046] Obtain the business process stability C; where Wr represents the error rate of the robotic process automation business process, and T represents the response time of the robotic process automation business process;
[0047] By comprehensively evaluating and calculating the performance formula of the robotic process automation process: N = α * M + β * η + γ * C
[0048] Obtain the robotic process automation process performance value N; where α, β, and γ respectively represent the weight coefficients of the model stability coefficient, business process efficiency, and business process stability; the sum of α, β, and γ is 1;
[0049] According to the calculation result of the robotic process automation process performance, determine whether to iteratively update the business process of the robotic process automation; when the real-time calculated robotic process automation process performance value is lower than the average value of the historical performance value, iteratively update the business process of the robotic process automation, otherwise do not perform iterative update.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] By extracting and integrating features from structured and unstructured data, the present invention generates multi-modal feature vectors, enabling robotic process automation programs to handle more complex data types; breaking through the limitation that traditional RPA can only process regular and structured data; the multi-modal prediction model, identification and classification model, and anomaly detection model constructed by the present invention using deep learning algorithms can analyze business data more accurately, improving the intelligent level of automated processes;
[0052] By integrating various models constructed into the robotic process automation program, the present invention automates tasks that originally required manual intervention, significantly improving the execution efficiency of business processes; the automated process reduces the possibility of human error, lowers operating costs, and improves the execution efficiency of business processes;
[0053] By comprehensively evaluating the performance of the robotic process automation process and performing iterative updates, the present invention adapts to changes in different business requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0055] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0057] Please refer to Figure 1 As shown, the first aspect embodiment of the present invention provides a deep learning-based robotic process automation method, including the following steps:
[0058] S1: Obtain a data acquisition request sent by the client, collect business data, the client has a robotic process automation program, and the business data includes: structured data and unstructured data;
[0059] S2: Preprocess the collected structured data, including data cleaning, denoising and standardization; preprocess the collected unstructured data, including text segmentation, stop word removal, image enhancement, grayscale and denoising;
[0060] S3: Perform feature extraction on the collected and preprocessed structured data and unstructured data; perform feature fusion based on the feature extraction results to generate a multi-modal feature vector;
[0061] S4: Construct a multimodal prediction model using a deep learning algorithm based on the multimodal feature vector; construct a recognition and classification model by using a deep learning algorithm; analyze anomalies in business data by constructing anomaly detection models; integrate the multimodal prediction model, recognition and classification model, and anomaly detection model into the robotic process automation program and conduct business process testing;
[0062] S5: Comprehensively evaluate the performance of the RPA process and iteratively update the RPA process based on the evaluation results.
[0063] Specifically, ensure that the client has installed and configured the Robotic Process Automation (RPA) program. The RPA program should be able to identify and trigger data acquisition requests and automatically start when business data needs to be collected. Automatically capture business data from databases, APIs, and other sources through the RPA program. Preprocessing of structured data in business data includes: removing duplicate, missing, or invalid data records, identifying and correcting outliers or noise in the data, and converting the data into a unified format and unit. Preprocessing of text data includes: splitting text data into words or phrases, removing common words that do not contribute much to the meaning of the text, such as "的" and "了"; preprocessing of video image data includes: enhancing image contrast, converting color images to grayscale images, and using filters to remove noise in the image; preprocessing of voice data for denoising. Extract features from structured data and unstructured data and fuse the extracted structured features with unstructured features to generate a multimodal feature vector. Construct a multimodal prediction model, a recognition and classification model, and anomaly detection model and integrate them into the RPA program respectively. By conducting business process testing, we ensure that the RPA program can correctly perform tasks such as data acquisition, preprocessing, feature extraction, model prediction, and anomaly detection. By comprehensively evaluating the performance of the robotic process automation process, we can iterate and update the RPA process.
[0064] In this embodiment, step S1 includes the following steps:
[0065] Using an API interface or a Web server to receive a data acquisition request sent by a client, wherein the request content includes: the type and source of the business data;
[0066] Automatically identify the types and sources of business data in the request content by using a deep learning model;
[0067] By obtaining a business data sample set, the business data sample set includes various business data types and sources; among them, the business data types include: structured business data and unstructured business data; the structured business data includes: fields and data values in a database or table; the unstructured business data includes: text, video images, and voice; the business data sources include: databases and web pages;
[0068] Label each business data sample, label the data type and source, and input the labeled business data sample set into the deep learning model for training; according to the trained deep learning model and input the request content into the deep learning model, automatically identify the types and sources of business data in the request content;
[0069] According to the types and sources of business data automatically identified by the deep learning model, simulate the manual operation process by using a robotic process automation (RPA) program; connect the RPA program to a database or a web page; by executing the RPA script, automatically complete the data collection operation and output the result, and the result is the business data collected from the database or the web page.
[0070] Specifically, use an API interface or a web server as the front-end receiving point to receive a data acquisition request from a client, and the request content should contain information clearly indicating the types and sources of the required business data. Parse the received request to extract the types and source information of the business data. Prepare and train a deep learning model in advance, and this model can identify and classify the types and sources of business data. Input the parsed request content, that is, the types and sources of business data, into the deep learning model, and the deep learning model identifies the input content and outputs the identified business data types and sources. By collecting business data samples of various types and sources, form a business data sample set, and label each business data sample to clearly label the data type and source. Input the labeled business data sample set into the deep learning model for training to enable the model to accurately identify the types and sources of business data. According to the types and sources of business data identified by the deep learning model, configure the corresponding RPA program. The RPA program should be able to simulate the manual operation process and connect to the corresponding database or web page. Execute the RPA script to automatically complete the data collection operation. During the collection process, the RPA program will extract the required business data from the database or web page according to predefined rules and logic. The RPA program provides the collected business data as the output result to the client or subsequent processing processes.
[0071] In this embodiment, the step S3 performs feature extraction processing on the collected and preprocessed structured data and unstructured data, including the following steps:
[0072] Performing feature extraction according to the collected and preprocessed structured data, including: numerical features and categorical variable features; normalizing or standardizing the numerical features; using the one-hot encoding method to convert the categorical variables into a set of binary features, where each categorical variable feature represents a category;
[0073] Extracting word embedding or sentence embedding features according to the preprocessed text by using natural language processing techniques;
[0074] Processing the preprocessed video images by using the computer vision technology OpenCV to extract video image features, including: color features, texture features and shape features;
[0075] Segmenting the audio signal into frames with a length of 25 milliseconds according to the preprocessed language, and extracting language features such as Mel-frequency cepstral coefficients, spectral centroid, and spectral bandwidth from each frame of audio.
[0076] Specifically, for the feature extraction of structured data, numerical features usually represent some quantifiable metrics or attributes. According to normalization, using the maximum and minimum values of numerical features, the values of numerical features are scaled to the interval [0, 1]. For each column of features, the min-max function is used for scaling. According to the standardization method, through the mean and standard deviation of numerical features, the features are scaled into a standard normal distribution, with a mean of 0 and a variance of 1 after scaling. Even if the data does not follow a normal distribution, the standardization method can still be used. Categorical variables represent data with a finite number of categories. One-hot encoding is to convert categorical variables into binary features. Each categorical variable feature represents a category. For a categorical variable with N categories, one-hot encoding converts it into N binary features. Each feature corresponds to a category, and only one feature is 1, and the rest are 0. For the feature extraction of unstructured data, the feature extraction of text data includes: word embedding or sentence embedding features; using pre-trained models such as Word2Vec, GloVe, or BERT to convert words or sentences in the text into low-dimensional vector representations. The feature extraction of video image data includes: extracting color features in the image, including: RGB values, HSV values, etc.; extracting texture features, including: LBP local binary pattern features, Gabor filter features, etc.; extracting shape features in the image, including edge features, contour features, etc. The feature extraction of speech data includes: segmenting the audio signal into frames of a fixed length, which is 25 milliseconds in this embodiment and can be dynamically adjusted according to the actual application situation; the extracted speech features include: Mel frequency cepstral coefficients, spectral centroid, and spectral bandwidth, etc. According to the above, structured data features include: the one-hot encoding of the numerical feature values after normalization or standardization and categorical variable features; text features include: word embedding feature vectors or sentence embedding feature vectors; video image features include: RGB values, HSV values, LBP local binary pattern features, Gabor filter features, edge intensity, and contour edge points; speech features include: Mel frequency cepstral coefficients, spectral centroid, and spectral bandwidth.
[0077] In this embodiment, in step S3, according to the results of feature extraction, feature fusion is performed to generate a multi-modal feature vector, including the following steps:
[0078] Align the extracted structured data features, text features, video image features, and speech features using timestamps, and use the principal component analysis method for feature dimensionality reduction;
[0079] Using the attention mechanism, calculate the weights of the structured data features, text features, video image features, and speech features after feature alignment and dimensionality reduction, and perform feature fusion on the weighted features to emphasize important features; perform joint representation learning through a shared network layer or a cross-modal loss function, and by constructing a multi-layer perceptron network, input each feature after joint representation learning into the multi-layer perceptron network to generate a multi-modal feature vector.
[0080] Specifically, when processing multi-modal data, since different sensors or data sources may generate data at different time points, it is necessary to use timestamps to align this data to the same time axis. Extract the timestamp information of each modal data, and according to the timestamp information, use appropriate alignment methods such as interpolation and resampling to align each modal data to the same time axis. Perform standardization processing on the aligned modal data to eliminate the dimension difference. Calculate the covariance matrix of each modal data, and perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; select the first k principal components according to the size of the eigenvalues to construct a dimensionality-reduced feature space; project each modal data into the dimensionality-reduced feature space to obtain the dimensionality-reduced feature representation. Perform the calculation of the attention mechanism on the dimensionality-reduced modal features to obtain the weight of each feature. Perform weighted fusion on the modal features according to the weights to obtain the fused feature representation. Design and construct a shared network layer or a cross-modal loss function to achieve joint representation learning. Use the modal features after joint representation learning as the input of the MLP network. Construct an MLP network, including an input layer, multiple hidden layers, and an output layer. The hidden layer can use non-linear activation functions such as ReLU for non-linear transformation. Train and optimize the MLP network to obtain the final multi-modal feature vector.
[0081] In this embodiment, the step S4 of constructing a multi-modal prediction model based on the multi-modal feature vector using a deep learning algorithm includes the following steps:
[0082] According to the generated multi-modal feature vector, construct a multi-modal prediction model by using a deep learning model to analyze and predict business data;
[0083] Input the multi-modal feature vector into the multi-modal prediction model for training;
[0084] Through the real-time collected and preprocessed business data, and extract the new multi-modal vector of the real-time collected business data; input the new multi-modal vector into the trained multi-modal prediction model for analysis and prediction and output the prediction result, where the prediction result is the change trend of the predicted business data.
[0085] Specifically, according to the characteristics of business data and prediction requirements, select an appropriate deep learning model to construct a multi-modal prediction model. For time series data, recurrent neural networks or their variants such as long short-term memory network (LSTM), gated recurrent unit (GRU), etc. can be selected; for image data, convolutional neural networks can be selected; for text data, word embedding models or Transformer, etc. can be selected. By designing the architecture of the multi-modal prediction model, ensure that it can receive and process multi-modal feature vectors; the model architecture can include an input layer, multiple hidden layers such as convolutional layers, recurrent layers, fully connected layers, etc. and an output layer. According to the requirements of the prediction task, select an appropriate loss function; for regression tasks, mean squared error or mean absolute error, etc. can be selected; for classification tasks, cross-entropy loss, etc. can be selected. Pair the generated multi-modal feature vectors as input data with corresponding business data labels such as the change trend of historical business data. Divide the data set into a training set, a validation set, and a test set for model training, validation, and testing. Input the training set into the multi-modal prediction model for model training. During the training process, minimize the loss function by adjusting the model parameters to improve the prediction accuracy of the model. Use the validation set to evaluate the performance of the model and select the best model parameters and architecture. Collect real-time business data and perform preprocessing operations such as data cleaning, feature extraction, etc.; convert the preprocessed business data into new multi-modal vectors and input the new multi-modal vectors into the trained multi-modal prediction model for prediction analysis. The multi-modal prediction model outputs a prediction result, that is, the change trend of real-time business data, according to the input new multi-modal vectors.
[0086] In this embodiment, in step S4, constructing an identification and classification model by using a deep learning algorithm includes the following steps:
[0087] Construct an identification and classification model by using a deep learning algorithm, input the multi-modal feature vectors into the identification and classification model for training, and during the training process, perform supervised learning by using label data and optimize the identification and classification model by using the backpropagation algorithm and the gradient descent method;
[0088] Input the real-time collected and preprocessed business data into the trained identification and classification model; the identification and classification model performs behavior pattern recognition and classification on the input business data.
[0089] Specifically, an identification and classification model is constructed by using deep learning algorithms. Appropriate deep learning frameworks are selected, such as TensorFlow, PyTorch, etc. These frameworks provide all the tools and functions required for constructing and training neural networks. According to the characteristics of the multimodal feature vectors and the requirements of the classification task, the architecture of the identification and classification model is designed. The model may include an input layer, a feature extraction layer, a fully connected layer, and an output layer. The output layer usually uses the softmax function for multi-classification or the sigmoid function for binary classification. A loss function and an optimizer are selected to update the model parameters through the backpropagation algorithm to minimize the loss function. A dataset of multimodal feature vectors with labels is collected and organized, where the labels represent the behavior pattern categories corresponding to each feature vector. The dataset is preprocessed, including data cleaning, normalization, standardization, etc., to ensure data quality and improve model training efficiency. The preprocessed multimodal feature vectors and their corresponding labels are input into the identification and classification model. The backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters. The model parameters are updated by the optimizer to minimize the loss function. The above steps are repeated until the performance of the model on the validation set reaches stability or the preset number of training epochs ends. Real-time business data is collected and preprocessing operations, such as data cleaning and feature extraction, are performed to generate new multimodal feature vectors. The preprocessed real-time business data, i.e., the new multimodal feature vectors, are input into the trained identification and classification model. The model performs forward propagation on the input data, calculates the prediction probabilities for each behavior pattern category, and selects the behavior pattern category with the highest probability as the classification result.
[0090] In this embodiment, in step S4, the abnormal conditions in the business data are analyzed by constructing an anomaly detection model, including the following steps:
[0091] An anomaly detection model is constructed according to deep learning algorithms; the multimodal feature vectors are input into the anomaly detection model for training; during the training process, through iterative optimization, the anomaly detection model learns the distribution characteristics of normal business data;
[0092] The preprocessed business data obtained by real-time collection is input into the trained anomaly detection model. The anomaly detection model evaluates the input business data according to the learned normal data distribution characteristics to detect whether the real-time business data is abnormal.
[0093] Specifically, select a deep learning framework, including: TensorFlow, PyTorch, etc., for building and training the anomaly detection model. According to the characteristics of the multimodal feature vectors and the requirements of the anomaly detection task, design the architecture of the anomaly detection model. Common architectures include: autoencoders, variants of generative adversarial networks, density-based networks, etc. These models can learn the distribution characteristics of normal data. For autoencoders, the reconstruction error is usually used as the loss function, that is, the difference between the input data and the data reconstructed by the model. For variants of generative adversarial networks, it may be necessary to define the loss functions of the generator and discriminator, as well as possible additional loss functions. Select a suitable optimizer for updating the model parameters through iterative optimization. Collect and organize the multimodal feature vectors of normal business data for training the anomaly detection model. Preprocess the normal business data, including data cleaning, normalization, standardization, etc., to ensure data quality. Input the preprocessed normal business data into the anomaly detection model. Through iterative optimization, the model learns the distribution characteristics of normal data. This usually involves minimizing the loss function so that the model can accurately reconstruct the input data. Use a part of the normal business data that has not participated in training to verify the performance of the model and ensure that the model can accurately reconstruct or generate normal data. Collect business data in real time and perform preprocessing operations to generate new multimodal feature vectors. Input the preprocessed real-time business data into the trained anomaly detection model. The anomaly detection model evaluates the input data according to the learned normal data distribution characteristics. For autoencoders, calculate the reconstruction error between the input data and the data reconstructed by the model; if the reconstruction error exceeds the preset threshold, the input data is considered abnormal; for variants of generative adversarial networks, it may be necessary to judge whether the input data is abnormal according to the difference between the data generated by the generator and the input data, the output of the discriminator, or other metrics.
[0094] In this embodiment, in step S4, integrating the multimodal prediction model, the recognition and classification model, and the anomaly detection model into the robotic process automation program respectively and performing business process testing includes the following steps:
[0095] Use the designer of the robotic process automation tool to create a framework for the automated process, including: the starting point, ending point, intermediate steps, and decision points of the business process;
[0096] Deploy the trained multimodal prediction model, recognition and classification model, and anomaly detection model to the server respectively;
[0097] Design an automated business process in the designer of the robotic process automation tool;
[0098] Embed the calls to the multi-modal prediction model, recognition and classification model, and anomaly detection model in the robotic process automation script through the API interface, and design an exception handling mechanism in the automated business process; when a call fails or returns unforeseen results in one of the models, issue an exception warning, otherwise continue to call various models for analysis;
[0099] Individually test each component in the business process and conduct an integration test on the business process; determine whether all components work together by testing the entire robotic process automation business process; optimize the robotic process automation business process based on the test results; deploy the tested and optimized robotic process automation business process to the production environment and start automatically executing the business process;
[0100] Based on the deployed robotic process automation business process, collect key performance indicators, including: processing time, error rate, and response time, by monitoring the running status of the robotic process automation business process in real time.
[0101] Specifically, using the designer of the robotic process automation (RPA) tool, first construct the basic framework of the automation process, including: clarifying the starting point, ending point, intermediate steps, and decision points of the business process. Deploy the trained multi-modal prediction model, recognition and classification model, and anomaly detection model to the server. These models are the core of the automation process, responsible for processing and analyzing data and providing decision support. In the designer of the RPA tool, design the specific automation business process according to business requirements, including: defining the operations of each step, data flow direction, and the order of model calls, etc. Through the API interface, embed the calls to the multi-modal prediction model, recognition and classification model, and anomaly detection model in the RPA script; this means that in the automation process, when these models are needed for analysis or prediction, the RPA system can automatically call the corresponding models and obtain the results. At the same time, design an exception handling mechanism; when the call of one of the models fails or returns unforeseen results, an exception warning can be issued, and corresponding measures can be taken, such as retrying, skipping this step, or triggering other alternative processes. Conduct separate tests on each component in the business process to ensure that each component can work properly and meet the expected requirements; on the basis of component testing, conduct integration testing on the business process, including: testing the interfaces and interactions between each component to ensure that the entire business process can work together and achieve the expected effect. According to the test results, optimize the business process of robotic process automation, including: adjusting the order of model calls, optimizing the data processing process, improving the exception handling mechanism, etc. Deploy the tested and optimized business process of robotic process automation to the production environment and start automatically executing the business process. By real-time monitoring the running status of the business process of robotic process automation, collect key performance indicators, including: processing time, error rate, and response time.
[0102] In this embodiment, step S5 includes the following steps:
[0103] Evaluate according to the trained multi-modal prediction model, recognition and classification model, and anomaly detection model respectively, and calculate the F1 value of the key performance indicator respectively; by calculating the model stability coefficient formula:
[0104] Obtain the model stability coefficient M; where F1 D represents the F1 value of the multi-modal prediction model; F1 S represents the F1 value of the recognition and classification model; F1 J respectively represent the F1 values of the anomaly detection model;
[0105] Obtain the business process efficiency η by calculating the ratio of the processing time of the robotic process automation business process to the processing time of the manual business process;
[0106] By calculating the business process stability formula:
[0107] Obtain the business process stability C; where Wr represents the error rate of the robotic process automation business process, and T represents the response time of the robotic process automation business process;
[0108] By comprehensively evaluating the performance formula of the robotic process automation process: N = α * M + β * η + γ * C
[0109] Obtain the robotic process automation process performance value N; where α, β, and γ respectively represent the weight coefficients of the model stability coefficient, business process efficiency, and business process stability; the sum of α, β, and γ is 1;
[0110] According to the calculation result of the robotic process automation process performance, determine whether to iteratively update the business process of the robotic process automation; when the real-time calculated robotic process automation process performance value is lower than the mean value of the historical performance values, iteratively update the business process of the robotic process automation, otherwise do not perform iterative update.
[0111] Specifically, calculate the model stability coefficient, business process efficiency, and business process stability respectively through formulas, and perform comprehensive calculation to obtain the robotic process automation process performance value according to the weight coefficients of the model stability coefficient, business process efficiency, and business process stability, which are 0.4, 0.3, and 0.3 respectively; the values of the weight coefficients are dynamically adjusted according to the actual situation and historical data. Calculate the mean value of the historical robotic process automation RPA process performance values and compare the real-time calculated RPA process performance value with the mean value of the historical performance values. If the real-time performance value is lower than the mean value of the historical performance values, it indicates that there is room for improvement in the RPA process and iterative update is required; otherwise, continue with the current process unchanged.
[0112] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A deep learning-based robotic process automation method, characterized in that: The following steps are involved: S1: Obtaining a data acquisition request sent by a client, and collecting business data, wherein the client has a robotic process automation program, and the business data includes: structured data and unstructured data; S2: Preprocess the collected structured data, including data cleaning, denoising and standardization; preprocess the collected unstructured data, including text segmentation, stop word removal, image enhancement, grayscale and denoising; S3: Perform feature extraction on the collected and preprocessed structured data and unstructured data; perform feature fusion based on the feature extraction results to generate a multi-modal feature vector; S4: Construct a multimodal prediction model using a deep learning algorithm based on the multimodal feature vector; construct a recognition and classification model by using a deep learning algorithm; analyze anomalies in business data by constructing anomaly detection models; integrate the multimodal prediction model, recognition and classification model, and anomaly detection model into the robotic process automation program and conduct business process testing; S5: Comprehensively evaluate the performance of the RPA process and iteratively update the RPA process based on the evaluation results.
2. The deep learning-based robotic process automation method according to claim 1, characterized in that: The step S1 comprises the following steps: Using an API interface or a Web server to receive a data acquisition request sent by a client, wherein the request content includes: the type and source of the business data; Automatically identify the type and source of business data in request content by leveraging deep learning models; By acquiring a business data sample set, the business data sample set includes various business data types and sources; wherein the business data types include: structured business data and unstructured business data; the structured business data includes: fields and data values in a database or table; the unstructured business data includes: text, video images and voice; the business data sources include: databases and web pages; Label each business data sample, mark the data type and source, and input the labeled business data sample set into the deep learning model for training; according to the trained deep learning model, input the request content into the deep learning model to automatically identify the type and source of the business data in the request content; Based on the type and source of business data automatically identified by the deep learning model, the manual operation process is simulated by using the Robotic Process Automation (RPA) program; the RPA program is connected to the database or web page; and the data collection operation is automatically completed and the results are output by executing the RPA script. The results are business data collected from the database or web page.
3. The deep learning-based robotic process automation method according to claim 1, characterized in that: The step S3 performs feature extraction on the collected and pre-processed structured data and unstructured data, including the following steps: Extract features based on the collected and preprocessed structured data, including: numerical features and categorical variable features; normalize or standardize the numerical features; use the one-hot encoding method to convert the categorical variables into a set of binary features, where each categorical variable feature represents a category; Based on the preprocessed text, natural language processing technology is used to extract word embedding or sentence embedding features; According to the preprocessed video image, the computer vision technology OpenCV is used to process the video image data and extract the video image features including color features, texture features and shape features; According to the preprocessed language, the audio signal is divided into frames of 25 milliseconds in length, and the language features of Mel-frequency cepstral coefficients, spectrum centroid, and spectrum bandwidth are extracted from each frame of audio.
4. The deep learning-based robotic process automation method according to claim 3, characterized in that: In step S3, feature fusion is performed according to the result of feature extraction to generate a multimodal feature vector, which includes the following steps: Based on the extracted structured data features, text features, video image features, and voice features, feature alignment is performed using timestamps, and feature dimension reduction is performed using principal component analysis. The attention mechanism is used to calculate the weight of each feature of the structured data features, text features, video image features and speech features after feature alignment and feature dimensionality reduction, and the weighted features are fused to emphasize important features; joint representation learning is performed through shared network layers or cross-modal loss functions, and by building a multi-layer perceptron network, the various features after joint representation learning are input into the multi-layer perceptron network to generate a multimodal feature vector.
5. The deep learning-based robotic process automation method according to claim 1, characterized in that: In step S4, a multimodal prediction model is constructed using a deep learning algorithm according to the multimodal feature vector, including the following steps: Based on the generated multimodal feature vectors, a multimodal prediction model is constructed by using a deep learning model to analyze and predict business data; Input the multimodal feature vector into the multimodal prediction model for training; By collecting and preprocessing business data in real time, and extracting new multimodal vectors of the business data collected in real time; inputting the new multimodal vectors into the trained multimodal prediction model for analysis and prediction and outputting the prediction results, the prediction results are the changing trends of the predicted business data.
6. The deep learning-based robotic process automation method according to claim 1, characterized in that: In step S4, a recognition and classification model is constructed by using a deep learning algorithm, including the following steps: By using deep learning algorithms to build recognition and classification models, multimodal feature vectors are input into the recognition and classification models for training. During the training process, supervised learning is performed using labeled data and the recognition and classification models are optimized using back propagation algorithms and gradient descent methods. The business data collected and pre-processed in real time is input into the trained recognition and classification model; the recognition and classification model recognizes and classifies the behavior patterns of the input business data.
7. The deep learning-based robotic process automation method according to claim 1, characterized in that: In step S4, the abnormal situation in the business data is analyzed by building an abnormality detection model, including the following steps: Build an anomaly detection model based on the deep learning algorithm; input the multimodal feature vector into the anomaly detection model for training; during the training process, through iterative optimization, the anomaly detection model learns the distribution characteristics of normal business data; The business data collected and preprocessed in real time is input into the trained anomaly detection model. The anomaly detection model evaluates the input business data based on the learned normal data distribution characteristics to detect whether the real-time business data is abnormal.
8. The deep learning-based robotic process automation method according to claim 1, characterized in that: In step S4, the multimodal prediction model, the recognition classification model and the anomaly detection model are respectively integrated into the robotic process automation program and the business process test is performed, which includes the following steps: Use the RPA tool’s designer to create the framework of the automated process, including: the starting point, end point, intermediate steps, and decision points of the business process; Deploy the trained multimodal prediction model, recognition and classification model, and anomaly detection model to the server respectively; Design automated business processes in the RPA tool’s designer; Through the API interface, calls to multimodal prediction models, recognition and classification models, and anomaly detection models are embedded in the robotic process automation script, and an exception handling mechanism is designed in the automated business process; when a call fails or returns an unpredictable result in one of the models, an exception warning is issued, otherwise various models continue to be called for analysis; Test each component in the business process individually and conduct integration testing on the business process; determine whether all components work together by testing the entire RPA business process; optimize the RPA business process based on the test results; deploy the tested and optimized RPA business process to the production environment and start automating the business process; According to the business process of the deployed RPA, the running status of the business process of the RPA is monitored in real time to collect key performance indicators, including processing time, error rate and response time.
9. The deep learning-based robotic process automation method according to claim 1, characterized in that: The step S5 comprises the following steps: The trained multimodal prediction model, recognition classification model and anomaly detection model are evaluated respectively, and the key performance indicator F1 value is calculated respectively; the model stability coefficient formula is calculated as follows: Get the model stability coefficient M; where F1 D Represents the F1 value of the multimodal prediction model; F1 S Indicates the F1 value of the recognition classification model; F1 J They represent the F1 value of the anomaly detection model respectively; The business process efficiency η is obtained by calculating the ratio of the processing time of the RPA business process to the processing time of the manual business process; By calculating the business process stability formula: The business process stability C is obtained; wherein Wr represents the error rate of the robotic process automation business process, and T represents the response time of the robotic process automation business process; The performance formula for calculating the RPA process through comprehensive evaluation is: N = α*M+β*η+γ*C The robot process automation process performance value N is obtained; wherein α, β and γ represent the weight coefficients of the model stability coefficient, the business process efficiency and the business process stability respectively; the sum of α, β and γ is 1; According to the calculation results of the robotic process automation process performance, determine whether to iteratively update the business process of the robotic process automation; when the real-time calculated robotic process automation process performance value is lower than the average of the historical performance values, the business process of the robotic process automation is iteratively updated, otherwise iterative update is not performed.
Citation Information
Patent Citations
Data fusion method and device, electronic equipment and storage medium
CN114386509A
Automatic data management system based on intelligent identification
CN115860697A
UI automatic test system and method based on RPA robot
CN118227506A
Automatic log decomposition process mining method based on RPA
CN118643471A
RPA service processing method and system based on artificial intelligence
CN118887044A
Cited By
Data cataloguing method, system and equipment based on artificial intelligence and storage medium
CN120951074A
An artificial intelligence-based data cataloging method, system, device and storage medium
CN120951074B