Big data preprocessing method and system based on machine learning
By cleaning and feature extraction of the operation log data of the RPA system, and combining with machine learning models for abnormal detection, the shortcomings of the RPA system in log data processing and abnormal detection are solved, the stability and efficiency of the system are improved, and the smooth progress of business processes are ensured.
Patent Information
- Application Number
- CN202510750257.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing RPA system lacks effective analysis methods when processing operation log data and cannot accurately detect abnormal situations, resulting in insecure system stability and efficiency, affecting the normal operation of business processes.
By obtaining the operation log data of the RPA system and performing data cleaning, the feature extraction algorithm is used to extract text semantic features and operation sequence features, and the machine learning model is called for joint modeling, generating exception detection results, and adjusting task execution strategies.
It improves the stability and efficiency of the RPA system, ensures smooth business processes, dynamically adapts to the system operation situation, and optimizes task execution processes.
Smart Images

Figure CN120256837A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of RPA data processing, and in particular, to a big data preprocessing method and system based on machine learning. Background Art
[0002] In the current era of rapid development of digital business, the RPA system has been widely used because it can automate the execution of repetitive tasks and greatly improve work efficiency. It realizes the automation of various business processes by simulating human operations, reduces manual intervention, and reduces the error probability.
[0003] However, under the background of the existing RPA system technology, there are lack of effective means for processing and analyzing the log data generated during the system operation. A large amount of operation log data is often in the original and messy state, valuable information cannot be extracted from it, and it is difficult to comprehensively understand the operation status of the system. In addition, the existing technology cannot accurately detect abnormal situations in the system, and it is even more difficult to reasonably adjust the task execution strategy according to the abnormal situations, resulting in problems such as chaotic task execution and reduced efficiency when the system faces abnormalities, affecting the normal operation of the business process.
[0004] In summary, the existing RPA system technology has deficiencies in processing operation log data, detecting abnormalities, and adjusting task execution strategies. There is an urgent need for a more effective technical solution to solve these problems in order to improve the stability and efficiency of the RPA system. Summary of the Invention
[0005] The embodiments of the present application provide a big data preprocessing method and system based on machine learning, which are used to improve the stability and efficiency of the RPA system and ensure the smooth progress of the business process.
[0006] In the first aspect, the embodiments of the present application provide a big data preprocessing method based on machine learning, which is applied to a big data preprocessing system. The method includes: obtaining a set of operation log data of the RPA system, where the set of operation log data includes multiple session interaction records; performing data cleaning processing on the set of operation log data to obtain a preprocessed set of operation log data; performing feature extraction processing on the preprocessed set of operation log data based on a preset feature extraction algorithm to obtain the text semantic features and operation sequence features of each session interaction record; calling a machine learning model to perform joint modeling processing on the text semantic features and the operation sequence features to generate an abnormal detection result of the session interaction record, and adjusting the task execution strategy of the RPA system based on the abnormal detection result.
[0007] In the second aspect, the embodiments of the present application provide a big data preprocessing system, including: A processor; A storage device on which a computer program is stored, When the computer program is executed by the processor, the processor implements any one of the above-mentioned machine learning-based big data preprocessing methods.
[0008] An embodiment of the present application provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the machine learning-based big data preprocessing method are implemented.
[0009] In the embodiment of the present application, first, the operation log data set of the RPA system is obtained and merged for data cleaning processing, effectively removing noise and incomplete data, improving the data quality, and making the preprocessed operation log data set more accurate and reliable. Secondly, based on a preset feature extraction algorithm, the text semantic features and operation sequence features of each session interaction record are extracted, which can accurately describe the characteristics of the interaction record from different dimensions and provide a comprehensive and detailed basis for anomaly detection. Then, a machine learning model is called to perform joint modeling on the text semantic features and operation sequence features, which can fully explore the potential relationships between the features and generate accurate anomaly detection results. Finally, according to the anomaly detection results, the task execution strategy of the RPA system is adjusted, which can dynamically adapt to the system operation conditions, optimize the task execution process, improve the stability and efficiency of the RPA system, and ensure the smooth progress of the business process. Description of the Drawings
[0010] Figure 1 It is a flowchart of a machine learning-based big data preprocessing method provided by an embodiment of the present application.
[0011] Figure 2 It is a schematic diagram of the basic structure of a big data preprocessing system provided by an embodiment of the present application. Detailed Embodiments
[0012] To make the above objects, features, and advantages of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0013] Refer to Figure 1 As shown in the figure, which is a flowchart of a machine learning-based big data preprocessing method provided by an embodiment of the present application, and this method can be applied to a big data preprocessing system. As Figure 1 shown, this method includes step 110-step 140.
[0014] It should be noted that in the embodiments of the present application, the acquisition of the operation log data of the RPA system strictly adheres to the principles of privacy protection and user authorization to ensure the legality and transparency of data processing. When a user first uses the system, the scope, purpose, and usage method of data collection are informed through a clear privacy policy and user agreement, including recording session interaction logs for system optimization and anomaly detection, and the relevant functions can only be enabled after the user actively checks and authorizes consent. The system follows the principle of minimum necessity, only collecting operation events and response data directly related to the business process, and avoiding personal sensitive information. For data fields that may contain user identifiers (such as user IDs), de-identification technology is used for anonymization to ensure that the log records cannot be associated with specific personal identities. A strict data access permission control mechanism is established within the enterprise, and only specific operation and risk control personnel are authorized to access the original logs, and all data processing behaviors are traced through audit logs. Users can view the types of data collected through system settings at any time and have the right to request the export or deletion of their personal operation records. In addition, the system regularly pushes data usage reports to users, explaining the purpose and results of log analysis, and continuously maintaining the users' right to know and right to choose.
[0015] Step 110: Obtain a set of operation log data of the RPA system, where the set of operation log data includes multiple session interaction records.
[0016] In the embodiments of the present application, taking an RPA system used within an enterprise as an example, this RPA system is responsible for processing various business processes, such as order processing, data entry, etc. Over a period of time, the system will record numerous interaction operations between users and the system, and these records constitute the set of operation log data. For example, in the order processing process, user A submits an order, and the system records user A's operation event and the corresponding system response event; user B modifies an order, which is also recorded. These various operation events and system response events of different users at different times together form multiple session interaction records in the set of operation log data. The session interaction records store all the key information during the operation of the RPA system and provide the original material for subsequent data processing and analysis. Then, the system will integrate the data scattered in each session interaction record to form a complete set of operation log data that can be processed subsequently.
[0017] Step 120: Perform data cleaning processing on the set of operation log data to obtain a preprocessed set of operation log data.
[0018] In this example, each session interaction record consists of at least one user operation event and the corresponding system response event. The data cleaning processing includes deleting redundant operation events, correcting inconsistent system response events, and filling in missing session interaction records. Based on this, step 120 includes: Step 121: Traverse each session interaction record in the set of running log data, identify the logical conflict events between user operation events and system response events in the session interaction record, where the logical conflict events include events where user operation events are not correctly reflected by system response events; perform context correlation analysis on the logical conflict events, determine the target system response event corresponding to the logical conflict event, and perform logical correction processing on the target system response event based on a preset rule library to obtain a corrected system response event.
[0019] Taking the order processing process as an example, in the set of running log data, there is a session interaction record showing that the user submitted an order to purchase a commodity, but the system response event shows that the order processing failed and the error message has nothing to do with the commodity. This belongs to a logical conflict event where the user operation event is not correctly reflected by the system response event. At this time, the system will perform context correlation analysis on this logical conflict event to view other relevant operations and system responses before and after the order is submitted. For example, it is analyzed that the inventory of the commodity was sufficient before the order was submitted, but suddenly insufficient after the order was submitted. According to the preset rule library, the rule library clearly stipulates that when the commodity inventory is insufficient, the system should correctly prompt that the order failed due to insufficient inventory. Based on this rule library, the target system response event is corrected to clearly prompt the information of insufficient commodity inventory, thereby obtaining a corrected system response event. In this way, it can ensure the correct logical relationship between the system response and the user operation, providing an accurate data basis for subsequent analysis and processing.
[0020] In a preferred embodiment, the performing logical correction processing on the target system response event based on a preset rule library to obtain a corrected system response event includes: Step 1211: Obtain the semantic description text of the user operation event corresponding to the target system response event, perform intention parsing processing on the semantic description text to obtain the target operation intention of the user operation event.
[0021] For the above example of order processing, obtain the semantic description text of the user operation event corresponding to the target system response event, such as "the user submitted an order to purchase commodity X". Perform intention parsing processing on this semantic description text, and analyze the keywords and grammatical structures in the text through natural language processing technology. For example, identify the core verb "purchase" and the operation object "commodity X", thereby determining that the target operation intention of the user operation event is to purchase commodity X. This target operation intention is an important basis for subsequent operations. It clarifies the user's operation purpose, enabling the system to better understand the user's needs and perform corresponding processing.
[0022] Step 1212: Match the corresponding expected response event template in the rule base based on the target operation intention. The expected response event template is pre-configured with the standard response event format corresponding to the target operation intention.
[0023] Optionally, based on the determined target operation intention of purchasing product X, perform a match in the rule base. The rule base is pre-configured with expected response event templates corresponding to various target operation intentions. For the purchase operation intention, the expected response event template may include the response format for successful purchase, such as "The order has been successfully submitted, and product X will be shipped within XX time", and the response format for failed purchase, such as "Due to XX reason, the order for purchasing product X has failed". By searching in the rule base, find the expected response event template corresponding to the target operation intention of purchasing product X, which provides a standard format for subsequent judgment and correction of the system response event.
[0024] Step 1213: Compare the format of the target system response event with the expected response event template. If the response format of the target system response event is inconsistent with the expected response event template, perform format conversion processing on the target system response event according to the format of the expected response event template.
[0025] Optionally, compare the target system response event "Order processing failed and the error message is not related to this product" with the expected response event template. It is found that the format of the target system response event is inconsistent with the format of failed purchase in the expected response event template. The expected response event template requires clearly indicating that the failure reason is related to the product, while the current target system response event does not meet this format requirement. Therefore, perform format conversion processing on the target system response event according to the format of the expected response event template, and convert it to "Due to insufficient stock of product X, the order for purchasing product X has failed", so that it meets the standard format requirement for subsequent logical consistency verification with the user operation event.
[0026] Step 1214: Perform logical consistency verification on the format-converted target system response event and the user operation event. If the verification passes, use the format-converted target system response event as the corrected system response event.
[0027] Optionally, perform a logical consistency check on the target system response event after format conversion, "The order to purchase product X failed due to insufficient stock of product X", and the user operation event, "The user submitted an order to purchase product X". By analyzing the logical relationship between the two, determine whether the reason for the order failure (insufficient product stock) is reasonably associated with the user's purchase operation. After verification, it is found that the logical relationship is reasonable and the verification passes. At this time, use the target system response event after format conversion as the corrected system response event to complete the correction of the system response event and ensure the accuracy and logical consistency of the data.
[0028] Step 122: Traverse each session interaction record in the set of running log data, and detect whether there are consecutive and repeated redundant operation events in the session interaction record. If so, perform duplicate removal processing on the redundant operation events, retain the first user operation event, and delete subsequent repeated redundant operation events.
[0029] In the set of running log data for order processing, traverse each session interaction record. For example, it is found that in one session interaction record, the user submitted the same order to purchase product Y multiple times in a short period. These consecutive and repeated order submission operations are redundant operation events. The system will perform duplicate removal processing on these redundant operation events, retain the operation event of the first user submitting an order to purchase product Y, and delete subsequent repeated order submission operation events. This can reduce redundant information in the data, improve data quality, and avoid interference caused by duplicate data during subsequent analysis and processing.
[0030] Step 123: Traverse each session interaction record in the set of running log data, and detect whether there is a missing timestamp between the user operation event and the system response event in the session interaction record. If so, perform interpolation filling processing on the missing timestamp based on the timestamp sequence of adjacent session interaction records.
[0031] Continuing with the order processing example, when traversing the set of running log data, it is found that the timestamp of the user operation event (submitting an order) in a certain session interaction record is missing. At this time, the system will view the timestamp sequence of adjacent session interaction records. For example, the timestamp of the user submitting another order in the previous adjacent session interaction record is 10:00, and the timestamp of the system's response to this order in the subsequent adjacent session interaction record is 10:05. By analyzing this time sequence and considering the general time interval and logical order of order processing, using a suitable interpolation method, such as linear interpolation, it is estimated that the missing timestamp of the user submitting the order can be 10:02. Based on this method, perform interpolation filling processing on the missing timestamp to make the time information in the session interaction record complete, ensure the coherence and accuracy of the data, and provide reliable data for subsequent time series-based analysis and processing.
[0032] Step 124: Based on the session interaction records that have completed system response event correction, duplicate removal of redundant operation events, and interpolation filling processing, obtain the preprocessed operation log data set.
[0033] Optionally, after performing system response event correction, duplicate removal of redundant operation events, and interpolation filling processing on each session interaction record in the operation log data set, integrate these processed session interaction records to obtain the preprocessed operation log data set. This preprocessed operation log data set removes error information and redundant information in the original data, supplements missing information, and significantly improves the data quality, providing a high-quality data basis for subsequent feature extraction and model processing. For example, after the above processing, the session interaction records related to order processing are more accurate and complete, and can better reflect the operations and responses in the actual business process.
[0034] Step 130: Based on a preset feature extraction algorithm, perform feature extraction processing on the preprocessed operation log data set to obtain the text semantic features and operation sequence features of each session interaction record.
[0035] As an implementation, the performing feature extraction processing on the preprocessed operation log data set based on a preset feature extraction algorithm to obtain the text semantic features and operation sequence features of each session interaction record includes: Step 131: Perform text word segmentation processing on the user operation events in the session interaction record to obtain multiple operation word units, and perform part-of-speech tagging processing on the multiple operation word units to identify the core operation verbs and auxiliary operation objects in the user operation events.
[0036] Taking a session interaction record of order processing as an example, the user operation event is "submit an order to purchase a computer". Perform word segmentation processing on this text to obtain multiple operation word units, such as "submit", "an", "purchase", "computer", "order". Then perform part-of-speech tagging processing on these operation word units. "Submit" and "purchase" are tagged as verbs, "an" is a quantifier, and "computer" and "order" are nouns. By analyzing these parts of speech, identify "submit" and "purchase" as the core operation verbs, and "computer" and "order" as the auxiliary operation objects. The above core operation verbs and auxiliary operation objects can accurately reflect the key information of the user operation event and provide a basis for generating text semantic features later.
[0037] Step 132: Invoke a pre-trained language encoding model to perform semantic encoding processing on the core operation verbs and the auxiliary operation objects to generate a text vector representation of the user operation event, and use the text vector representation as the text semantic feature.
[0038] Among them, a language encoding model that has been pre-trained on a large amount of text data is called, such as the BERT model. The identified core operation verbs "submit" and "purchase", and the auxiliary operation objects "computer" and "order" are input into the model. The model will perform semantic encoding processing on these words, and generate corresponding vector representations for each word according to the context information and semantic relationships of the words in a large amount of text. For example, "submit" may be encoded as a vector containing multiple numerical values, and "purchase", "computer", and "order" also have their respective corresponding vectors. Then these vectors are combined or further processed to generate a text vector representation of the user operation event, which can comprehensively reflect the semantic information of the user operation event, and use it as a text semantic feature to provide semantic-level input for subsequent machine learning models.
[0039] Step 133: Perform operation sequence parsing processing on the system response events in the session interaction record, identify the set of atomic operation instructions included in the system response events, and perform sequence pattern encoding processing on the set of atomic operation instructions to generate the operation sequence encoding of the system response events.
[0040] In order processing, the system response event may be "verify order information, check inventory, generate shipping note, arrange logistics". Perform operation sequence parsing processing on this system response event to identify the set of atomic operation instructions included therein, such as "verify order information", "check inventory", "generate shipping note", "arrange logistics". Each atomic operation instruction represents a basic operation unit. As an optional embodiment, the performing sequence pattern encoding processing on the set of atomic operation instructions to generate the operation sequence encoding of the system response event includes: Step 1331: Traverse each atomic operation instruction in the set of atomic operation instructions to determine the operation type identifier and the set of operation parameters of the atomic operation instruction.
[0041] Further, traverse the above set of atomic operation instructions. For the atomic operation instruction "verify order information", determine that its operation type identifier is "order verification", and the set of operation parameters may include order number, customer information, etc. For the atomic operation instruction "check inventory", the operation type identifier is "inventory check", and the set of operation parameters includes product name, inventory quantity, etc. In the above manner, the operation type identifier and the set of operation parameters of each atomic operation instruction are determined for subsequent encoding processing.
[0042] Step 1332: Map the atomic operation instruction to a preset operation type space based on the operation type identifier to generate a type encoding vector of the atomic operation instruction.
[0043] In an embodiment of the present application, an operation type space is preset. For example, "order verification" is mapped to a region in the operation type space, and according to the coding rules of this region, a type coding vector of the "order verification" atomic operation instruction is generated. This type coding vector is a multi-dimensional vector, and its dimensions and values are determined according to the definition and coding rules of the operation type space. Similarly, other atomic operation instructions such as "inventory check" are also mapped to the operation type space to generate their respective type coding vectors, and in this way, the operation type information is digitally represented.
[0044] Step 1333: Perform parameter type parsing processing on each operation parameter in the operation parameter set, determine the data type identifier and parameter value range of the operation parameter, and generate a parameter coding vector of the operation parameter based on the data type identifier and the parameter value range.
[0045] For the order number in the operation parameter set of "verifying order information", its data type identifier is parsed as "string type", and the parameter value range can be a string of the corresponding length composed of numbers and letters. According to the definition of the data type identifier and the parameter value range, a parameter coding vector of the operation parameter of the order number is generated. For other operation parameters such as customer information, the same processing is also performed to generate their respective parameter coding vectors, and these parameter coding vectors can accurately represent the type and value range information of the operation parameters.
[0046] Step 1334: Concatenate the type coding vector and the parameter coding vector to obtain an instruction coding vector of the atomic operation instruction, and perform serialization processing on the instruction coding vector based on the execution order of the atomic operation instruction in the system response event to generate the operation sequence coding.
[0047] Furthermore, the type coding vectors and parameter coding vectors of each atomic operation instruction are concatenated. For example, the type coding vector of "verifying order information" and the parameter coding vectors of order number, customer information, etc. are concatenated together to obtain an instruction coding vector of the "verifying order information" atomic operation instruction. Then, according to the execution order of the atomic operation instructions in the system response event, such as the order of "verifying order information", "checking inventory", "generating a shipping order", "arranging logistics", serialization processing is performed on these instruction coding vectors. These vectors are arranged and combined in order to generate an operation sequence coding of the system response event, and this operation sequence coding can comprehensively reflect the operation order and specific operation content information of the system response event.
[0048] Step 134: Perform time dimension alignment processing on the operation sequence coding and the text vector representation to obtain the operation sequence feature of the session interaction record.
[0049] In the session interaction record of order processing, the generated operation sequence encoding and text vector representation are aligned in the time dimension. Since there is a chronological order between user operation events and system response events, through time dimension alignment, the operation sequence encoding and text vector representation can correspond to each other in time. For example, the text vector representation corresponding to the user's order submission operation is aligned in time with the operation sequence encoding of the earliest "verify order information" operation in the system response, and subsequent operation sequence encodings and text vector representations are also aligned in sequence according to the time order. Through this alignment process, the operation sequence features of the session interaction record are obtained, and these features can comprehensively reflect the association and operation order between user operations and system responses in the time dimension.
[0050] Step 140: Invoke a machine learning model to perform joint modeling on the text semantic features and the operation sequence features, generate the anomaly detection result of the session interaction record, and adjust the task execution strategy of the RPA system based on the anomaly detection result.
[0051] In a preferred embodiment, the invoking a machine learning model to perform joint modeling on the text semantic features and the operation sequence features, and generating the anomaly detection result of the session interaction record includes: Step 141: Input the text semantic features into the text encoding branch of the machine learning model, and generate a text high-order feature vector through multi-layer non-linear transformation.
[0052] The machine learning model in the embodiment of the present application is a deep neural network model with multiple branches and layers. The generated text semantic features are input into the text encoding branch of the model, and this text encoding branch includes multiple non-linear transformation layers, such as fully connected layers and activation function layers. After the text semantic feature vector enters these layers, it first passes through the fully connected layer. The fully connected layer performs a linear transformation on the input vector according to a preset weight matrix, and then performs a non-linear transformation through the activation function. For example, using the ReLU activation function, the result of the linear transformation is non-linearly converted to enhance the model's ability to express semantic information. After multiple layers of the above non-linear transformation, the text semantic features are gradually converted into text high-order feature vectors, and this high-order feature vector can represent the complex information in the text semantics more deeply.
[0053] Step 142: Input the operation sequence features into the sequence encoding branch of the machine learning model, and generate a sequence high-order feature vector through temporal convolutional operation.
[0054] Optionally, the operation sequence features are input into the sequence encoding branch of the machine learning model. The sequence encoding branch employs temporal convolution operations, where the temporal convolution operations perform sliding convolutions on the operation sequence feature vectors along the temporal dimension. For example, the size of the convolution kernel can be 3, and it takes 3 consecutive vector values each time along the temporal dimension of the operation sequence encoding vector for convolution calculation. Through the convolution calculation, local features and patterns of the operation sequence along the temporal dimension are extracted. After multiple temporal convolution operations, the operation sequence features are transformed into sequence high-order feature vectors, which can better reflect the variation laws and features of the operation sequence over time.
[0055] Step 143: Perform cross-modal attention interaction processing on the text high-order feature vectors and the sequence high-order feature vectors to determine the association weight matrix between the text high-order feature vectors and the sequence high-order feature vectors.
[0056] Further, step 143 includes: Step 1431: Use the text high-order feature vectors as the first query vectors, the sequence high-order feature vectors as the first key vectors and the first value vectors, and calculate the first attention scores between the first query vectors and the first key vectors; perform normalization processing on the first attention scores to obtain the first attention weight distribution between the text high-order feature vectors and the sequence high-order feature vectors; perform weighted summation processing on the first value vectors based on the first attention weight distribution to obtain the first cross-modal context vectors corresponding to the text high-order feature vectors; perform residual connection processing on the first cross-modal context vectors and the text high-order feature vectors to obtain the updated text high-order feature vectors.
[0057] In this step, the text high-order feature vectors are used as the first query vectors, and the sequence high-order feature vectors are used as the first key vectors and the first value vectors. By performing operations such as calculating the dot product between the query vectors and the key vectors, the first attention scores are obtained. For example, using the dot product formula to calculate the sum of the products of each element in the text high-order feature vectors and the corresponding elements in the sequence high-order feature vectors to obtain the first attention scores. Then, normalization processing is performed on this score, such as using the softmax function for normalization, to obtain the first attention weight distribution between the text high-order feature vectors and the sequence high-order feature vectors, which represents the degree of attention of the text high-order feature vectors to different parts of the sequence high-order feature vectors.
[0058] Based on the first attention weight distribution, perform a weighted sum operation on the first value vector (i.e., the sequence high-order feature vector), that is, multiply the weight distribution by the corresponding elements of the sequence high-order feature vector and then sum them up to obtain the first cross-modal context vector corresponding to the text high-order feature vector. This cross-modal context vector contains information related to the text high-order feature vector obtained from the sequence high-order feature vector. Next, perform a residual connection operation on the first cross-modal context vector and the text high-order feature vector, that is, add their corresponding elements to obtain the updated text high-order feature vector. This residual connection helps the model better learn and retain the information in the original text high-order feature vector while integrating the relevant information from the sequence high-order feature vector.
[0059] Step 1432: Use the sequence high-order feature vector as the second query vector, and the text high-order feature vector as the second key vector and the second value vector, and calculate the second attention score between the second query vector and the second key vector; perform a normalization process on the second attention score to obtain the second attention weight distribution between the sequence high-order feature vector and the text high-order feature vector; based on the second attention weight distribution, perform a weighted sum operation on the second value vector to obtain the second cross-modal context vector corresponding to the sequence high-order feature vector; perform a residual connection operation on the second cross-modal context vector and the sequence high-order feature vector to obtain the updated sequence high-order feature vector.
[0060] Similarly, use the sequence high-order feature vector as the second query vector, and the text high-order feature vector as the second key vector and the second value vector. Through a similar calculation method, first calculate the dot product and other operations between the second query vector and the second key vector to obtain the second attention score. Then perform a normalization process on the second attention score, for example, use the softmax function again, to obtain the second attention weight distribution between the sequence high-order feature vector and the text high-order feature vector. This weight distribution reflects the degree of attention of the sequence high-order feature vector to different parts of the text high-order feature vector.
[0061] According to the second attention weight distribution, perform a weighted sum on the second value vector (i.e., the text high-order feature vector) to obtain the second cross-modal context vector corresponding to the sequence high-order feature vector. Finally, perform a residual connection on the second cross-modal context vector and the sequence high-order feature vector, that is, add their corresponding elements, to obtain the updated sequence high-order feature vector. In this way, the sequence high-order feature vector also integrates the relevant information from the text high-order feature vector, further enriching its own feature representation.
[0062] Step 1433: Generate a third query vector based on the updated text high-order feature vector, and generate a third key vector based on the updated sequence high-order feature vector; calculate the bidirectional cross-attention score matrix between the third query vector and the third key vector; perform row-column bidirectional normalization processing on the bidirectional cross-attention score matrix to generate the correlation weight matrix between the text high-order feature vector and the sequence high-order feature vector.
[0063] Further, use the updated text high-order feature vector to generate a third query vector, and the updated sequence high-order feature vector to generate a third key vector. Calculate the bidirectional cross-attention score matrix between the third query vector and the third key vector, which means calculating not only the attention scores in the text-to-sequence direction but also the attention scores in the sequence-to-text direction, and combining these scores into a matrix. For example, each element in the matrix represents the attention score between a certain dimension of the text high-order feature vector and a certain dimension of the sequence high-order feature vector.
[0064] Next, perform row-column bidirectional normalization processing on this bidirectional cross-attention score matrix. In the row direction, by normalizing the elements in each row, the sum of the elements in each row is made equal to 1; in the column direction, the elements in each column are also normalized so that the sum of the elements in each column is equal to 1. After the above row-column bidirectional normalization processing, a correlation weight matrix between the text high-order feature vector and the sequence high-order feature vector is generated. This correlation weight matrix can accurately reflect the degree of mutual correlation between the two modal feature vectors, providing an important basis for subsequent weighted fusion.
[0065] Step 144: Perform weighted fusion processing on the text high-order feature vector and the sequence high-order feature vector based on the correlation weight matrix to generate a joint feature representation vector; In an exemplary embodiment, the performing weighted fusion processing on the text high-order feature vector and the sequence high-order feature vector based on the correlation weight matrix to generate a joint feature representation vector includes: Step 1440: Decompose the associated weight matrix into a text attention weight vector and a sequence attention weight vector; perform weighted scaling processing on the text high-order feature vector based on the text attention weight vector to obtain a scaled text feature vector; perform weighted scaling processing on the sequence high-order feature vector based on the sequence attention weight vector to obtain a scaled sequence feature vector; perform element-wise addition on the scaled text feature vector and the scaled sequence feature vector to obtain an initial fusion feature vector; perform layer normalization processing on the initial fusion feature vector to eliminate the feature scale difference and obtain a normalized fusion feature vector; input the normalized fusion feature vector into a fully connected layer for feature compression processing to generate the joint feature representation vector.
[0066] Specifically, the associated weight matrix is decomposed into a text attention weight vector and a sequence attention weight vector according to certain rules. For example, it is divided according to the row or column information of the matrix, and the weight information related to the text high-order feature vector is extracted to form the text attention weight vector, and the weight information related to the sequence high-order feature vector is extracted to form the sequence attention weight vector.
[0067] Then, use the text attention weight vector to perform weighted scaling processing on the text high-order feature vector, that is, multiply each element of the text attention weight vector by the corresponding element of the text high-order feature vector to obtain a scaled text feature vector. Similarly, use the sequence attention weight vector to perform weighted scaling on the sequence high-order feature vector to obtain a scaled sequence feature vector.
[0068] Next, perform element-wise addition on the scaled text feature vector and the scaled sequence feature vector, that is, add the elements at the corresponding positions of the two vectors to obtain an initial fusion feature vector. Since the scales of different feature vectors may vary, this may affect the training and performance of the model, so layer normalization processing is performed on the initial fusion feature vector. Layer normalization is to perform normalization operations on each dimension of the vector. By subtracting the mean and dividing by the standard deviation and other calculations, the feature scale difference is eliminated to obtain a normalized fusion feature vector.
[0069] Finally, input the normalized fusion feature vector into a fully connected layer for feature compression processing. The fully connected layer performs matrix multiplication operations with the input vector through a weight matrix, compresses the high-dimensional fusion feature vector into a lower-dimensional space, and generates a joint feature representation vector. This joint feature representation vector fuses the feature information of text semantics and operation sequences, providing a more comprehensive and effective feature representation for subsequent anomaly detection.
[0070] Step 145: Input the combined feature representation vector into the classifier layer of the machine learning model, and output the anomaly detection result of the session interaction record, where the anomaly detection result is used to indicate whether there is an operation logic anomaly or a system response anomaly in the session interaction record.
[0071] Furthermore, input the generated combined feature representation vector into the classifier layer of the machine learning model. The classifier layer can be a simple fully connected layer or a more complex classification model, such as a softmax classifier. The classifier layer performs calculations and judgments on the combined feature representation vector according to the pre-trained weight parameters. For example, the weight matrix in the classifier layer performs a matrix multiplication operation with the combined feature representation vector, and then the result is converted into a probability distribution through an activation function (such as the softmax function). This probability distribution represents the likelihood that the session interaction record belongs to different categories (normal or abnormal). Based on the probability distribution, the model can output a decision result, that is, the anomaly detection result of the session interaction record. If the output result indicates that the probability of belonging to the abnormal category exceeds a certain threshold (for example, 0.5), it is determined that there is an operation logic anomaly or a system response anomaly in the session interaction record; otherwise, if the probability of belonging to the normal category is higher, it is determined that the session interaction record is normal. This anomaly detection result can help the RPA system timely discover problems during the operation process and provide a basis for adjusting the task execution strategy.
[0072] As an implementation, adjusting the task execution strategy of the RPA system based on the anomaly detection result includes: Step 146: Statistically analyze the target operation event distribution and anomaly type distribution of the abnormal session interaction records in the anomaly detection result, and generate an abnormal event statistical report; determine the set of frequently abnormal operation events and the set of risk anomaly types in the RPA system based on the abnormal event statistical report.
[0073] In the order processing scenario, statistical analysis is performed on the anomaly detection results. First, all session interaction records determined to be anomalies are traversed, and the occurrence frequency and distribution of target operation events among them are counted. For example, it is found that the "submit order" operation appears abnormally frequently in the anomaly session interaction records, and at the same time, different types of anomalies are recorded, such as the number of occurrences and distribution of anomaly types like "the order fails due to insufficient inventory but no correct prompt is given" and "the order information verification fails but the reason is unknown". Based on these statistical information, an anomaly event statistical report is generated. The report details the abnormal occurrence frequencies of various target operation events and the distribution of different anomaly types. Based on this anomaly event statistical report, the set of frequently abnormal operation events and the set of risk anomaly types in the RPA system are further determined. The set of frequently abnormal operation events may include operations such as "submit order" and "modify order", which appear frequently in the anomaly session interaction records. The set of risk anomaly types may include those anomaly types that have a greater impact on the business and a relatively high occurrence frequency, such as the anomaly type of "the order fails due to insufficient inventory but no correct prompt is given", which may lead to customer loss.
[0074] Step 147: For each frequently abnormal operation event in the set of frequently abnormal operation events, insert a pre-check rule into the task execution process of the RPA system. The pre-check rule is used to activate a secondary confirmation process when the frequently abnormal operation event is detected.
[0075] It can be understood that for the "submit order" operation in the set of frequently abnormal operation events, a pre-check rule is inserted into the order processing task execution process of the RPA system. When the system detects that the user initiates the "submit order" operation, the pre-check rule will be triggered first. For example, the pre-check rule will check the integrity of the product inventory information, user order information, etc. If a situation that may lead to an anomaly is found, such as the product inventory approaching the critical value or there being partial missing in the order information, the secondary confirmation process will be activated. In the secondary confirmation process, the system will pop up a prompt box to the user, informing the user of the possible problems and asking the user to confirm again whether to submit the order. In this way, frequent abnormal situations caused by user misoperations or incomplete information can be avoided, improving the accuracy and success rate of order processing.
[0076] Step 148: For each risk anomaly type in the set of risk anomaly types, configure an anomaly fuse mechanism in the task execution process of the RPA system. The anomaly fuse mechanism is used to pause the current task process and activate a manual intervention process when a preset number of the risk anomaly types are continuously detected.
[0077] It can be understood that for the risk anomaly type of "order failure due to insufficient inventory but without correct prompt" in the risk anomaly type set, an exception fuse mechanism is configured in the order processing task execution process of the RPA system. A preset quantity is set. For example, when this risk anomaly type is continuously detected 3 times, the exception fuse mechanism will be triggered. When the exception fuse mechanism is triggered, the current order task process being processed will be paused to avoid continuing to process orders that may cause more errors or adverse effects. At the same time, the system will activate the manual intervention process to notify relevant staff to intervene and handle. The staff can check the inventory management system, confirm the actual inventory situation, correct the system prompt information, etc. to solve the problems causing risk anomalies and ensure the normal progress of subsequent order processing.
[0078] In a non-limiting embodiment, after adjusting the task execution strategy of the RPA system based on the anomaly detection result, it further includes: obtaining the target session interaction record set corresponding to the anomaly detection result, extracting the target text semantic features and target operation sequence features of the target session interaction record, and generating a derived training sample set; optimizing and training the machine learning model based on the derived training sample set, calculating the gradient difference vector between the current model parameters and the model parameters after derived training; adjusting the weight distribution of the machine learning model according to the gradient difference vector to generate an optimized machine learning model; deploying the optimized machine learning model to the real-time detection module of the RPA system to perform anomaly detection on the newly added session interaction records, and updating the derived training sample set based on the detection result.
[0079] In the order processing scenario, obtain the corresponding target session interaction record set according to the anomaly detection result. For example, those session interaction records determined to be abnormal and some representative normal session interaction records. For these target session interaction records, extract their target text semantic features and target operation sequence features again. Use the same feature extraction method as before, such as performing text tokenization and part-of-speech tagging on user operation events to generate text semantic features, and performing operation sequence parsing and encoding on system response events to generate operation sequence features. Combine these newly extracted features to generate a derived training sample set. Use the derived training sample set to optimize and train the machine learning model.
[0080] During the training process, the model calculates the gradient difference vector between the current model parameters and the model parameters after training with the derived training sample set. This gradient difference vector reflects the direction and degree of change in the model parameters after training with the new sample set. Based on the gradient difference vector, the weight distribution of the machine learning model is adjusted through an optimization algorithm (such as the stochastic gradient descent algorithm). The adjusted weight distribution enables the model to better adapt to the new sample data, thereby generating an optimized machine learning model. The optimized machine learning model is deployed to the real-time detection module of the RPA system. When new session interaction records enter the system, the real-time detection module uses the optimized model for anomaly detection. If new anomalies or new representative normal situations are detected, these new session interaction records are added to the derived training sample set to further update the sample set, so as to continuously optimize the model and improve the accuracy and adaptability of the model.
[0081] In another non-limiting embodiment, after adjusting the task execution strategy of the RPA system based on the anomaly detection result, it further includes: real-time collecting the operation log data of the RPA system after executing the adjusted task execution strategy, and extracting the session interaction record set after the strategy adjustment; performing anomaly recheck processing on the session interaction record set after the strategy adjustment to generate a recheck anomaly rate and an anomaly type distribution change matrix; calculating the difference coefficient between the recheck anomaly rate and the historical anomaly rate, and analyzing the frequency change gradient of various anomalies in the anomaly type distribution change matrix; if the difference coefficient exceeds the preset threshold or the frequency change gradient indicates the spread of anomaly types, adjusting the activation condition of the pre-check rule and the pause threshold of the anomaly fuse mechanism; generating a strategy optimization instruction based on the adjusted activation condition and pause threshold, driving the RPA system to execute strategy iterative update and output a strategy adjustment effectiveness report.
[0082] In the order processing system, the operation log data of the RPA system after executing the adjusted task execution strategy is collected in real time. For example, after inserting the pre-check rule and configuring the anomaly fuse mechanism, the system continuously records the relevant information of user operations and system responses. The session interaction record set after the strategy adjustment is extracted from these operation log data. Anomaly recheck processing is performed on this set, that is, the machine learning model is used again to detect anomalies in these session interaction records, and the recheck anomaly rate is statistically calculated, that is, the proportion of abnormal session interaction records in the total session interaction records.
[0083] Meanwhile, generate an exception type distribution change matrix, which records the change in the occurrence frequency of different exception types before and after the policy adjustment. Calculate the difference coefficient between the re-inspection exception rate and the historical exception rate. For example, the difference coefficient is obtained by dividing the difference between the two by the historical exception rate. Analyze the frequency change gradient of each type of exception in the exception type distribution change matrix to observe which exception types have an increasing or decreasing occurrence frequency. If the difference coefficient exceeds a preset threshold, such as exceeding 0.1, or the frequency change gradient indicates that the occurrence frequency of certain exception types is increasing rapidly, indicating the spread of exception types, it means that the current policy adjustment effect may not be ideal and further adjustment is required. At this time, adjust the activation conditions of the pre-check rules, such as relaxing or tightening the threshold for inventory inspection, and adjust the pause threshold of the exception fusing mechanism, such as changing the number of times of continuously detecting risk exception types from 3 times to 4 times. Generate a policy optimization instruction based on the adjusted activation conditions and pause thresholds and send it to the RPA system. The RPA system performs policy iterative update according to this instruction and updates the relevant rules and mechanisms in the task execution process.
[0084] Finally, the system outputs a policy adjustment effectiveness report, which details the exception detection situation, changes in task execution efficiency, etc. before and after the policy adjustment, so as to evaluate the effect of the policy adjustment and further optimize the system.
[0085] It can be understood that those skilled in the art can perform optimization and supplementary processing based on natural language processing, sequence pattern encoding, and cross-modal attention mechanism in the prior art when implementing the above technical solutions of the embodiments of the present application. Specifically, for the construction of the rule base in data cleaning, an expected response template matching mechanism can be established by combining the domain knowledge graph and the historical exception event library, and a standard response format can be dynamically generated through entity relationship extraction and logical reasoning to avoid the correction failure caused by the staticization of the rule base.
[0086] When encoding the operation sequence, the operation type space can be defined with reference to the instruction classification system of the industrial standard OPCUA. A lightweight encoder is used to perform normalized mapping on the parameter types, and a fixed-length parameter encoding vector is generated by combining a hash function to process string-type parameters, ensuring the dimensional consistency of different data types.
[0087] For cross-modal attention interaction, a multi-head attention mechanism can be introduced and a dynamic dimension projection layer can be set. The dimensional difference between the text vector and the sequence encoding is eliminated through adaptive weight allocation, and at the same time, a gated residual network is combined to control the feature fusion ratio.
[0088] For the implementation of the classifier layer, a hierarchical softmax structure can be adopted and the Focal Loss loss function can be introduced to optimize the problem of sample imbalance, and the exception determination threshold can be dynamically adjusted in combination with the sliding window mechanism.
[0089] In the policy adjustment phase, a policy optimizer can be constructed based on the reinforcement learning framework. The triggering conditions of the pre-check rules and the fusing mechanism are quantified through the Q-learning algorithm, and the policy parameters are updated in real time in combination with online learning to ensure the dynamic adaptation of the exception handling logic to the business scenarios.
[0090] By introducing the above-mentioned existing technologies for optimizing and improving the implementation of the technical solution, the processing from log cleaning, feature fusion to policy iteration can be fully realized, enabling those skilled in the art to clearly and completely implement the embodiments of the present application.
[0091] In the embodiments of the present application, first, the operation log data set of the RPA system is obtained and merged for data cleaning processing, effectively removing noise and incomplete data, improving the data quality, and making the preprocessed operation log data set more accurate and reliable. Secondly, based on the preset feature extraction algorithm, the text semantic features and operation sequence features of each session interaction record are extracted, which can accurately describe the characteristics of the interaction record from different dimensions and provide a comprehensive and detailed basis for anomaly detection. Then, a machine learning model is called to perform joint modeling on the text semantic features and operation sequence features, which can fully explore the potential relationships between the features and generate accurate anomaly detection results. Finally, according to the anomaly detection results, the task execution policy of the RPA system is adjusted, which can dynamically adapt to the system operation conditions, optimize the task execution process, improve the stability and efficiency of the RPA system, and ensure the smooth progress of the business process.
[0092] See Figure 2 As shown in the figure, the figure is a schematic diagram of the basic structure of a big data preprocessing system 200 provided by the embodiments of the present application. The big data preprocessing system 200 includes: A processor 201; A storage device 202, on which a computer program 2020 is stored; When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the big data preprocessing methods based on machine learning.
[0093] On this basis, a readable storage medium is provided. Programs or instructions are stored on the readable storage medium, and when the programs or instructions are executed by a processor, the steps of the above method are implemented.
[0094] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
Claims
1. A big data preprocessing method based on machine learning, characterized in that, Including: Obtain the operation log data set of the RPA system, where the operation log data set includes multiple session interaction records; Perform data cleaning processing on the operation log data set to obtain a preprocessed operation log data set; Based on a preset feature extraction algorithm, perform feature extraction processing on the preprocessed operation log data set to obtain the text semantic features and operation sequence features of each session interaction record; Call a machine learning model to perform joint modeling processing on the text semantic features and the operation sequence features, generate the anomaly detection result of the session interaction record, and adjust the task execution strategy of the RPA system based on the anomaly detection result.
2. The method according to claim 1, wherein Each session interaction record consists of at least one user operation event and a corresponding system response event. The data cleaning processing includes deleting redundant operation events, correcting inconsistent system response events, and filling in missing session interaction records. The performing data cleaning processing on the operation log data set to obtain a preprocessed operation log data set includes: Traverse each session interaction record in the operation log data set, identify the logical conflict events between the user operation event and the system response event in the session interaction record, where the logical conflict events include events where the user operation event is not correctly reflected by the system response event; perform context correlation analysis on the logical conflict events, determine the target system response event corresponding to the logical conflict events, and perform logical correction processing on the target system response event based on a preset rule library to obtain a corrected system response event; Traverse each session interaction record in the operation log data set, detect whether there are consecutive repeated redundant operation events in the session interaction record. If so, perform duplicate removal processing on the redundant operation events, retain the first user operation event, and delete subsequent repeated redundant operation events; Traverse each session interaction record in the operation log data set, detect whether there is a timestamp missing between the user operation event and the system response event in the session interaction record. If so, perform interpolation filling processing on the missing timestamp based on the timestamp sequence of adjacent session interaction records; Based on the session interaction records that have completed system response event correction, completed redundant operation event duplicate removal, and completed interpolation filling processing, obtain the preprocessed operation log data set.
3. The method according to claim 2, wherein The performing logical correction processing on the target system response event based on a preset rule library to obtain a corrected system response event includes: Obtain the semantic description text of the user operation event corresponding to the target system response event, perform intent parsing processing on the semantic description text to obtain the target operation intent of the user operation event; Match the corresponding expected response event template in the rule library based on the target operation intent, and the expected response event template is pre-configured with the standard response event format corresponding to the target operation intent; Compare the format of the target system response event with the expected response event template. If the response format of the target system response event is inconsistent with the expected response event template, perform format conversion processing on the target system response event according to the format of the expected response event template; Perform logical consistency verification on the format-converted target system response event and the user operation event. If the verification passes, use the format-converted target system response event as the corrected system response event.
4. The method according to claim 1, wherein Perform feature extraction processing on the preprocessed operation log data set based on a preset feature extraction algorithm to obtain the text semantic features and operation sequence features of each session interaction record, including: Perform text tokenization on the user operation events in the session interaction record to obtain multiple operation word units, and perform part-of-speech tagging on the multiple operation word units to identify the core operation verbs and auxiliary operation objects in the user operation events; Call a pre-trained language encoding model to perform semantic encoding on the core operation verb and the auxiliary operation object to generate a text vector representation of the user operation event, and use the text vector representation as the text semantic feature; Perform operation sequence parsing on the system response events in the session interaction record to identify the set of atomic operation instructions included in the system response events, and perform sequence pattern encoding on the set of atomic operation instructions to generate the operation sequence encoding of the system response events; Perform time dimension alignment processing on the operation sequence encoding and the text vector representation to obtain the operation sequence features of the session interaction record.
5. The method according to claim 4, characterized in that, The performing sequence pattern encoding processing on the set of atomic operation instructions to generate the operation sequence encoding of the system response event includes: Traverse each atomic operation instruction in the set of atomic operation instructions to determine the operation type identifier and the set of operation parameters of the atomic operation instruction; Map the atomic operation instruction to a preset operation type space based on the operation type identifier to generate a type encoding vector of the atomic operation instruction; Perform parameter type parsing on each operation parameter in the set of operation parameters to determine the data type identifier and the parameter value range of the operation parameter, and generate a parameter encoding vector of the operation parameter based on the data type identifier and the parameter value range; Perform splicing processing on the type encoding vector and the parameter encoding vector to obtain an instruction encoding vector of the atomic operation instruction, and perform serialization processing on the instruction encoding vector based on the execution order of the atomic operation instruction in the system response event to generate the operation sequence encoding.
6. The method according to claim 1, wherein The calling a machine learning model to perform joint modeling on the text semantic features and the operation sequence features to generate the anomaly detection result of the session interaction record includes: Input the text semantic features into the text encoding branch of the machine learning model to generate text high-order feature vectors through multi-layer non-linear transformation; Input the operation sequence features into the sequence encoding branch of the machine learning model, and generate sequence high-order feature vectors through temporal convolution operations; Perform cross-modal attention interaction processing on the text high-order feature vectors and the sequence high-order feature vectors to determine the correlation weight matrix between the text high-order feature vectors and the sequence high-order feature vectors; Perform weighted fusion processing on the text high-order feature vectors and the sequence high-order feature vectors based on the correlation weight matrix to generate a joint feature representation vector; Input the joint feature representation vector into the classifier layer of the machine learning model, and output the anomaly detection result of the session interaction record, where the anomaly detection result is used to indicate whether there is an operation logic anomaly or a system response anomaly in the session interaction record.
7. The method according to claim 6, characterized in that The performing cross-modal attention interaction processing on the text high-order feature vectors and the sequence high-order feature vectors to determine the correlation weight matrix between the text high-order feature vectors and the sequence high-order feature vectors includes: Take the text high-order feature vectors as the first query vectors, take the sequence high-order feature vectors as the first key vectors and the first value vectors, and calculate the first attention scores between the first query vectors and the first key vectors; perform normalization processing on the first attention scores to obtain the first attention weight distribution between the text high-order feature vectors and the sequence high-order feature vectors; perform weighted summation processing on the first value vectors based on the first attention weight distribution to obtain the first cross-modal context vectors corresponding to the text high-order feature vectors; perform residual connection processing on the first cross-modal context vectors and the text high-order feature vectors to obtain updated text high-order feature vectors; Take the sequence high-order feature vectors as the second query vectors, take the text high-order feature vectors as the second key vectors and the second value vectors, and calculate the second attention scores between the second query vectors and the second key vectors; perform normalization processing on the second attention scores to obtain the second attention weight distribution between the sequence high-order feature vectors and the text high-order feature vectors; perform weighted summation processing on the second value vectors based on the second attention weight distribution to obtain the second cross-modal context vectors corresponding to the sequence high-order feature vectors; perform residual connection processing on the second cross-modal context vectors and the sequence high-order feature vectors to obtain updated sequence high-order feature vectors; Generate a third query vector based on the updated text high-order feature vectors, and generate a third key vector based on the updated sequence high-order feature vectors; calculate the bidirectional cross-attention score matrix between the third query vector and the third key vector; perform row-column bidirectional normalization processing on the bidirectional cross-attention score matrix to generate the correlation weight matrix between the text high-order feature vectors and the sequence high-order feature vectors.
8. The method according to claim 6, wherein The performing weighted fusion processing on the text high-order feature vectors and the sequence high-order feature vectors based on the correlation weight matrix to generate a joint feature representation vector includes: Decompose the associated weight matrix into a text attention weight vector and a sequence attention weight vector; Perform weighted scaling processing on the text high-order feature vector based on the text attention weight vector to obtain a scaled text feature vector; Perform weighted scaling processing on the sequence high-order feature vector based on the sequence attention weight vector to obtain a scaled sequence feature vector; Perform element-wise addition processing on the scaled text feature vector and the scaled sequence feature vector to obtain an initial fusion feature vector; Perform layer normalization processing on the initial fusion feature vector to eliminate the feature scale difference and obtain a normalized fusion feature vector; Input the normalized fusion feature vector into a fully connected layer for feature compression processing to generate the joint feature representation vector.
9. The method according to claim 1, wherein The adjusting the task execution strategy of the RPA system based on the anomaly detection result includes: Statistically analyze the target operation event distribution and anomaly type distribution of abnormal session interaction records in the anomaly detection result to generate an anomaly event statistical report; Determine a set of frequently abnormal operation events and a set of risk anomaly types in the RPA system based on the anomaly event statistical report; For each frequently abnormal operation event in the set of frequently abnormal operation events, insert a pre-check rule into the task execution process of the RPA system, and the pre-check rule is used to activate a secondary confirmation process when the frequently abnormal operation event is detected; For each risk anomaly type in the set of risk anomaly types, configure an anomaly fusing mechanism in the task execution process of the RPA system, and the anomaly fusing mechanism is used to pause the current task process and activate a manual intervention process when a preset number of the risk anomaly types are continuously detected.
10. A big data preprocessing system, characterized in that, Including: A processor; A storage device storing a computer program, and when the computer program is executed by the processor, the processor implements the machine learning-based big data preprocessing method according to any one of claims 1-9.
Citation Information
Patent Citations
Credit and credential application compatible adaptation method and system based on cloud computing
CN118444979A
RPA service data anomaly detection method and detection system fused with AI model
CN118897751A
Virtual machine anomaly detection method and device, equipment, storage medium and program product
CN119622715A
Database abnormal behavior detection method and device, electronic equipment and storage medium
CN119670072A
Automated system and method for detection and remediation of anomalies in robotic process automation environment
US20230039566A1
Cited By
Full-process information tracing method and system for asset management data
CN121961615A
An asset management data whole-process information tracing method and system
CN121961615B