RPA full-process automatic declaration data processing method and system
By constructing a business logic mapping matrix and a node processing time prediction model, and combining RPA tools for bidirectional adaptation and resource scheduling of data and process nodes, the problem of insufficient logical understanding in existing RPA application data processing methods is solved, achieving full-process automation and efficient and accurate application data processing.
Patent Information
- Application Number
- CN202511383360.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing RPA data processing methods lack a deep understanding of the overall logical relationship of the data, making it impossible to achieve fully automated processing. In particular, manual intervention is required when data is missing or format is incorrect, resulting in low efficiency and insufficient accuracy.
By receiving the set of application data, a business logic mapping matrix is constructed to achieve bidirectional adaptation between data and process nodes. Resources and data transmission channels are allocated using a node processing time prediction model. RPA tools are called to perform data association deduction, realize automatic data supplementation and format correction, and finally generate an application data submission package and push it to the target system.
It has achieved full automation of the data processing process, improving processing efficiency and accuracy, reducing manual intervention, and lowering error rates and labor costs.
Smart Images

Figure CN120874803A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to an RPA (Robotic Process Automation) method and system for fully automated data processing of application submissions. Background Technology
[0002] In the context of the rapid development of digital government and business services, various application procedures are becoming increasingly numerous, covering many areas such as tax declaration, project application, and qualification application. Traditional data processing methods rely heavily on manual operation, requiring staff to manually collect, organize, and verify various application forms and supporting documents submitted by applicants. This method is not only inefficient but also prone to human error, such as data entry errors, omissions, and non-standard formats, which can disrupt the application process and affect the timeliness and accuracy of applications.
[0003] With the emergence of Robotic Process Automation (RPA) technology, some declaration processes have begun to explore the use of RPA for automation. However, existing RPA methods for processing declaration data are mostly limited to simple data entry and form filling, lacking a deep understanding and processing capability of the overall logical relationships within the declaration data. Existing methods struggle to effectively handle complex business logic within declaration datasets, such as the relationships between different data points and the correspondence between data and declaration process nodes. Furthermore, manual intervention is often required when dealing with issues like missing data or format discrepancies, preventing full-process automation and limiting the effectiveness and scope of RPA technology in declaration data processing. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a fully automated RPA (Robotic Process Automation) method for processing application data, the method comprising: The system receives a set of application data, which includes various application form data and supporting document data submitted by the applicant. The application form data and supporting document data are associated with each other, and the association identifier is used to uniquely associate the corresponding application data of the same applicant. Retrieve a preset RPA application process template, construct a business logic mapping matrix based on the content attributes of various types of data in the application data set, and perform bidirectional adaptation between various types of data and nodes in the RPA application process template through the business logic mapping matrix to obtain the bidirectional adaptation result of process nodes. Based on the bidirectional adaptation results of the process nodes, and combined with the historical processing data of each RPA application process node, a node processing time prediction model is constructed. Based on the node processing time prediction model, RPA computing resources and data transmission channels are allocated, and a node resource prediction scheduling scheme is generated. For RPA application process nodes with missing data or format deviations in the bidirectional adaptation results of the process nodes, the RPA tool is called based on the node resource prediction and scheduling scheme to perform cross-node data association derivation, supplement the missing data or correct the format deviation data, and obtain the node data after association derivation. Based on the business logic sequence of the RPA application process template, the node data without data problems in the bidirectional adaptation results of the process nodes are integrated with the node data after association deduction to generate an application data submission package. The application data submission package is then pushed to the target application system to obtain application submission feedback data.
[0005] In another aspect, embodiments of the present invention also provide an RPA full-process automated reporting data processing system, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the above-described method.
[0006] Based on the above, this embodiment of the invention receives a set of application data containing application form data and supporting document data, each with a unique association identifier. It then retrieves a pre-defined RPA application process template and constructs a business logic mapping matrix. This achieves efficient bidirectional adaptation between various data types and application process nodes, accurately grasps the intrinsic connection between data and processes, and effectively solves the problem of insufficient understanding of business logic in traditional methods. Based on the adaptation results and historical processing data, a node processing time prediction model is constructed, thereby generating a node resource prediction and scheduling scheme. This rationally allocates RPA computing resources and data transmission channels, improving resource utilization efficiency and ensuring the smooth operation of the application process. For nodes with data problems, the RPA tool is invoked to perform cross-node data association deduction, achieving automatic data supplementation and format correction, reducing manual intervention, and improving the accuracy and completeness of data processing. Finally, data is integrated according to the business logic sequence to generate an application data submission package and pushed to the target application system to obtain application submission feedback data. This completes the full automation of the application data processing process, significantly improving the processing efficiency and quality of application business, and reducing labor costs and error rates. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of the execution flow of the RPA full-process automated data processing method provided in this embodiment of the invention.
[0008] Figure 2 This is a schematic diagram of exemplary hardware and software components of the RPA full-process automated declaration data processing system provided in this embodiment of the invention. Detailed Implementation
[0009] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an embodiment of the RPA fully automated application data processing method provided by the present invention. The following is a detailed description of the RPA fully automated application data processing method.
[0010] Step S110: Receive the application data set, which includes various application form data and supporting document data submitted by the applicant. The application form data and supporting document data are associated with each other. The association identifier is used to uniquely associate the corresponding application data of the same applicant.
[0011] This example uses a monthly VAT declaration scenario for enterprises, with the declarant being an enterprise engaged in production and operation activities. The declaration data set covers various types of data submitted by the enterprise through the online declaration platform. The declaration form data includes the main VAT tax return form, supplementary documents one to five, and the VAT tax reduction and exemption declaration details form, etc.; the supporting documents data includes scanned copies of VAT special invoices for deduction, proof of input tax transfer, and electronic versions of contracts and agreements corresponding to tax-exempt items, etc.
[0012] The association identifier uses the enterprise's Unified Social Credit Code, which is recorded in the header of the application form data and the metadata information of the supporting documents. This association identifier allows for the precise linking of all application form data and supporting document data submitted by the same enterprise, ensuring the accuracy of data attribution during subsequent data processing.
[0013] Step S111: Deploy a distributed data receiving cluster for the RPA tool. The distributed data receiving cluster contains multiple parallel receiving nodes, each of which is configured with an independent network port for simultaneously receiving data submitted by multiple claiming entities.
[0014] In the process of receiving monthly VAT declaration data from enterprises, the deployed distributed data receiving cluster consists of multiple parallel receiving nodes. These parallel receiving nodes are distributed across different physical servers and work collaboratively through cluster management software. Each parallel receiving node is configured with an independent network port, which is assigned a different port identifier in the network configuration and is always active, capable of receiving external data transmission requests at any time.
[0015] The setup of multiple parallel receiving nodes allows the receiving cluster to process data submission requests from different companies simultaneously, avoiding congestion on a single node due to excessive data traffic during peak submission periods, thus ensuring the efficiency and stability of data reception.
[0016] Step S112: Configure dedicated data receiving protocols for the application form data and the supporting document data respectively. The application form data adopts a dedicated form data transmission protocol, which supports field-level data transmission; the supporting document data adopts a file fragmentation transmission protocol, which supports uploading large file fragments.
[0017] Different dedicated data receiving protocols are configured for different data types in enterprise VAT declarations. For declaration form data, a dedicated form data transmission protocol is used. During data transmission, this protocol treats each data field in the form as an independent transmission unit, with each unit containing information such as field name, data type, length, and value. Upon receiving the data, the receiving end can accurately parse the content of each field based on this information, ensuring the integrity and accuracy of the form data.
[0018] For supporting documentation, due to the large file size of some documents, such as large collections of scanned invoices or long-term contracts, a file chunking transmission protocol is used. This protocol splits large files into multiple chunks according to a preset size, each chunk containing information such as a chunk number, the total number of chunks, and a file identifier. After receiving all chunks, the receiving end can reassemble them into a complete file based on this information, ensuring a high success rate for large file transmission.
[0019] Step S113: When the reporting entity initiates a data submission request, the receiving and scheduling module of the RPA tool allocates the data submission request to the parallel receiving node of the corresponding data receiving protocol according to the data type in the data submission request.
[0020] When a company initiates a VAT declaration data submission request through the declaration terminal, the data submission request message first enters the receive and schedule module of the RPA tool. The receive and schedule module performs preliminary parsing of the message and extracts the data type identifier field, which clearly indicates whether the submitted data belongs to the declaration form or supporting documentation.
[0021] For example, in step S1131: the receiving and scheduling module of the RPA tool listens to the data submission request port of the applicant in real time. When a data submission request is received, the data type identifier is extracted from the header field of the data submission request message. The data type identifier is used to distinguish between application form data and supporting material data.
[0022] The receiving and scheduling module continuously monitors the configured data submission request port through a pre-set network listening program. When a submission request from the applicant arrives at the port, the listening program immediately captures the request message. The receiving and scheduling module parses the message header to locate the data type identifier field. This field uses a pre-set encoding rule, such as a specific string or numeric code, to correspond to the two types of data: application form data and supporting document data, thereby achieving accurate differentiation of data types.
[0023] Step S1132: Retrieve the receiving node status monitoring ledger, which includes the current number of receiving tasks, remaining receiving capacity, and supported data receiving protocol types for each parallel receiving node.
[0024] After determining the data type, the receiving scheduling module calls the receiving node status monitoring ledger via an interface. This ledger is maintained in real time by the cluster management system and updated with node status information at fixed intervals. The ledger records the unique identifier of each parallel receiving node, the corresponding current number of receiving tasks (i.e., the number of data submission requests being processed by the node), the remaining receiving capacity (i.e., the maximum number of tasks the node can still handle), and the supported data receiving protocol types, such as a dedicated transmission protocol for form data or a file fragmentation transmission protocol.
[0025] Step S1133: Select parallel receiving nodes that support the receiving protocol corresponding to the data type in the data submission request. From these parallel receiving nodes, prioritize the selection of parallel receiving nodes whose current number of receiving tasks does not exceed the node's normal receiving range and whose remaining receiving capacity can fully accommodate the data volume corresponding to the current data submission request as the target receiving node.
[0026] Based on the data type identifier and corresponding receiving protocol, the receiving scheduling module filters all parallel receiving nodes that support that protocol from the ledger, forming a candidate node list. The nodes in the candidate node list are further filtered, checking whether the current number of receiving tasks for each node is within its normal receiving range, which is preset based on the node's hardware performance; simultaneously, it checks whether the node's remaining receiving capacity is greater than or equal to the amount of data corresponding to the current data submission request, ensuring that the node has sufficient capacity to process the request, and includes eligible nodes in the target receiving node candidate pool.
[0027] Step S1134: If there are multiple target receiving nodes that meet the conditions, calculate the network delay between each target receiving node and the terminal submitted by the applicant, and select the target receiving node with the smallest network delay.
[0028] When two or more nodes in the target receiving node candidate pool meet the criteria, the receiving scheduling module initiates a network latency detection procedure. This procedure sends probe packets to each candidate node, records the round-trip time of the packet from the submitting terminal to the node, and calculates the network latency accordingly. The network latency of all candidate nodes is compared, and the node with the lowest latency value is selected as the final target receiving node to reduce latency during data transmission.
[0029] Step S1135: Send a task allocation instruction to the target receiving node. The task allocation instruction includes the data submission request identifier of the reporting entity, the data type, and the receiving parameter requirements. At the same time, update the current number of receiving tasks for the target receiving node in the receiving node status monitoring log.
[0030] The receiving scheduling module generates a task allocation instruction. This instruction includes a unique identifier for the data submission request from the submitting entity, used to associate it with subsequent data reception; the data type submitted to ensure the node uses the correct protocol for reception; and reception parameter requirements, such as data transmission timeout and verification method. After the instruction is sent to the target receiving node via the internal communication channel, the receiving scheduling module immediately updates the current number of receiving tasks for that node in the receiving node's status monitoring log, incrementing its value by one.
[0031] Step S1136: Return the connection information of the target receiving node to the submitting terminal of the applicant. The connection information includes the node IP address, port number and protocol version, guiding the applicant to establish a data transmission connection with the target receiving node.
[0032] The receiving and scheduling module extracts the target receiving node's IP address, port number for receiving this type of data, and supported protocol versions from the node's configuration information. This information is then encapsulated into a connection response message and sent to the submitting terminal of the submitting entity. Based on the returned connection information, the submitting terminal starts the corresponding client program, establishes a TCP / IP connection with the target receiving node, and prepares for data transmission.
[0033] Step S1137: The receiving scheduling module tracks the data receiving progress of the target receiving node in real time. If the target receiving node experiences a receiving delay, the unfinished receiving task will be transferred to other idle parallel receiving nodes of the same protocol.
[0034] The receiving scheduling module obtains data reception progress information, such as the percentage of received data to the total data volume and the reception rate, through periodic communication with the target receiving node. The receiving progress is compared with a preset progress threshold. If the progress does not reach the threshold within a specified time, it is determined to be a reception delay. At this time, the receiving scheduling module selects an idle node from the parallel receiving nodes of the same protocol, sends a task transfer instruction to the original target node, transfers the unfinished receiving task to the new node, updates the receiving node status monitoring log, and simultaneously notifies the submitting terminal of the reporting entity to switch to the new node.
[0035] Step S1138: Record the process information of each data submission request allocation, including the data submission request identifier, data type, target receiving node identifier, and allocation time, to form a receiving scheduling log.
[0036] After completing task allocation, the receiving scheduling module writes information such as the unique identifier of this data submission request, data type description, unique identifier of the target receiving node, and the specific time of task allocation (accurate to the second) into the receiving scheduling log. The log is stored in a structured format, with each record containing fixed fields, facilitating subsequent log analysis, troubleshooting, and data statistics, and ensuring the traceability of the receiving scheduling process.
[0037] Step S114: The parallel receiving node receives the application form data uploaded by the applicant, parses the field structure of the application form data, extracts the preset association identifier field in the application form data, and records the content of the association identifier field and the corresponding application form data reception time.
[0038] After receiving the application form data uploaded by the enterprise, the parallel receiving node initiates the form data parsing program. This program parses the field structure of the transmitted form data according to the specifications of the dedicated form data transmission protocol. During the parsing process, the program identifies each field in the form, determining its position, data format, and other information.
[0039] After parsing the field structure, the program will locate the preset associated identifier field, namely the enterprise's Unified Social Credit Code field, and extract its specific content. Simultaneously, the system will record the time the application form data is received, accurate to the millisecond level, for subsequent traceability and management of the data reception process.
[0040] Step S115: The parallel receiving node receives the supporting document data uploaded by the applicant, reassembles the fragmented file according to the file fragmentation transmission protocol, and extracts the association identifier from the reassembled file metadata, which includes file naming information and file attribute information.
[0041] For the supporting documentation data uploaded by enterprises, the parallel receiving nodes receive each file fragment according to the requirements of the file fragmentation transmission protocol. During the reception process, the parallel receiving nodes verify the integrity of each fragment to ensure that the received fragment data is not corrupted. Once all fragments have been received, the nodes initiate a file reassembly process, piecing together the fragments into a complete supporting documentation file based on information such as the fragment number and the total number of fragments.
[0042] After the reorganization is complete, the program reads the file's metadata, which includes file naming information such as the company's abbreviation in the file name, and file attribute information such as creation time and modification time. The unified social credit code of the enterprise is extracted from the above metadata and used as the basis for linking the supporting document data with the corresponding enterprise.
[0043] Step S116: Group the application form data and supporting document data corresponding to the same association identifier into a group of application data units, with each application data unit corresponding to a single application data of an applicant.
[0044] After obtaining the respective association identifiers for the application form data and supporting document data, the system categorizes all received data. By comparing the content of the association identifiers, application form data and supporting document data with the same unified social credit code are grouped into the same application data unit.
[0045] Each declaration data unit contains all the form data and supporting documents submitted by the enterprise for this VAT declaration, forming a complete set of data for a single declaration by the enterprise.
[0046] Step S117: Summarize all declared data units to form a declared data set, add a receiving summary identifier to the declared data set, the receiving summary identifier includes the parallel receiving node number, the receiving completion time and the total number of declared data units, and record the name of the declaring entity and the submitting terminal information corresponding to each declared data unit, and complete the receiving of the declared data set.
[0047] The system aggregates all categorized application data units, integrating them into a final application data set. During the aggregation process, a receiving aggregation identifier is added to this application data set. The parallel receiving node number in the receiving aggregation identifier records the identification information of each node participating in this data reception; the reception completion time records the time when all data units were received; and the total number of application data units reflects the number of data units received in this instance.
[0048] Simultaneously, the system records the corresponding applicant name for each data unit, extracted from the header information of the application form data. The submitting terminal information includes the device identifier and network address of the submitting terminal, obtained from the header of the data submission request. After these operations are completed, the process of receiving the application data set is finished.
[0049] Step S120: Retrieve the preset RPA application process template, construct a business logic mapping matrix based on the content attributes of various types of data in the application data set, and perform bidirectional adaptation between various types of data and nodes in the RPA application process template through the business logic mapping matrix to obtain the bidirectional adaptation result of process nodes.
[0050] After receiving the declaration data set, the process moves to data and workflow node adaptation. First, a pre-set enterprise VAT declaration RPA workflow template is retrieved, which is pre-defined according to the standardized procedures for VAT declaration. Next, the content attributes of various data types in the declaration data set are analyzed, such as the declaration items involved, data format, and data source.
[0051] Based on these content attributes, a business logic mapping matrix is constructed, which contains information on the relationship between data attributes and process nodes. This matrix is used to perform bidirectional adaptation between various types of data and process nodes, determining which process node each type of data should be matched to for processing, and what data each process node needs to process, ultimately yielding the bidirectional adaptation results for process nodes.
[0052] Step S121: Retrieve a preset RPA application process template. The RPA application process template contains multiple RPA application process nodes connected in series according to the application business logic. Each RPA application process node is marked with a business logic description. The business logic description includes the application business link corresponding to the node, the functional positioning of the data to be processed, and the data flow relationship with upstream and downstream nodes.
[0053] Retrieve a pre-set enterprise VAT declaration RPA process template from the system's process template library. This enterprise VAT declaration RPA process template connects multiple RPA declaration process nodes according to the business logic sequence of VAT declaration, including data verification node, input tax calculation node, output tax calculation node, tax payable calculation node, and declaration form generation node.
[0054] Each RPA reporting process node has a detailed business logic description. For example, the business logic description of the data verification node is as follows: It corresponds to the initial verification stage of the reporting business. The data to be processed is the basic field data of all reporting forms. Its function is to check the standardization and completeness of the data format. Its upstream node is the data receiving node, and its downstream nodes are the input tax calculation node and the output tax calculation node. The data flow relationship is that the data that has passed the verification is transmitted to the two downstream nodes respectively.
[0055] Step S122: Perform content attribute parsing on the various application form data and supporting material data in the application data set, and extract the business function tags for each type of data. The business function tags include the application business actions supported by the data, the business decision information they carry, and the associated upstream and downstream data flows.
[0056] Content attribute analysis is performed on various types of data in the declaration data set. For declaration form data, such as Supplementary Information 1 to the VAT tax return, the content of fields such as taxable sales amount and output tax amount is analyzed to determine that the declaration business action supported by this data is the statistics and declaration of output tax amount, the business decision information it carries is the basis for calculating output tax amount, the upstream data flow is the enterprise sales ledger data, and the downstream data flow is the tax payable calculation node.
[0057] For supporting documentation data, such as scanned copies of VAT invoices for deduction, key information such as invoice code, invoice number, and deductible tax amount is analyzed to determine that the supported declaration business action is input tax deduction declaration, the business decision information it carries serves as the legal basis for input tax deduction, the upstream data flow is the invoice issuer's data, and the downstream data flow is the input tax calculation node. Through the above analysis, corresponding business function tags are extracted for each type of data.
[0058] Step S123: Based on the business logic descriptions of all RPA application process nodes and the business function tags of various types of data, construct a business logic mapping matrix. The row dimension of the business logic mapping matrix is the RPA application process node identifier, the column dimension is the data business function tag, and the matrix elements are the business logic matching degree between the node and the data.
[0059] Collect business logic descriptions and business function tags for all RPA reporting process nodes, and construct a business logic mapping matrix based on these. The rows of the matrix sequentially arrange the identifiers of each RPA reporting process node, such as data verification node identifiers and input tax calculation node identifiers; the columns sequentially arrange the business function tags for various types of data, such as output tax declaration support tags and input tax deduction basis tags.
[0060] Each element in the matrix represents the business logic matching degree between the corresponding RPA application process node and the data business function tag. The matching degree is determined by comparing the degree of overlap between the data function positioning to be processed in the business logic description of the node and the supporting business actions, decision information, etc. in the data business function tag. The higher the degree of overlap, the higher the matching degree.
[0061] Step S124: Calculate the business logic matching degree between each type of data and each RPA application process node through the business logic mapping matrix, and determine the RPA application process node with the highest business logic matching degree as the initial data adaptation node.
[0062] The matching degree is calculated using a business logic mapping matrix. For each type of data, the matrix element value at the intersection of the column corresponding to its business function label and the row corresponding to each RPA application process node identifier is examined to obtain the business logic matching degree between that type of data and each node.
[0063] Among all nodes, the node with the highest matching degree is selected and identified as the initial matching node for that type of data. For example, the data in the output tax-related forms has the highest matching degree with the output tax calculation node, so the output tax calculation node is identified as the initial matching node for that type of data.
[0064] Step S125: Analyze the data flow relationship between the initial adaptation node and the upstream and downstream nodes, and determine whether this type of data can be transferred to the downstream node through the initial adaptation node and meet the business logic requirements of the downstream node. If it can, the initial adaptation node is determined as the target adaptation node; if it cannot, the matching dimension of the business logic mapping matrix is readjusted, the potential business function tags of the data are added, and the matching degree is recalculated until the target adaptation node is determined or the data to be adapted is marked.
[0065] For data for which a preliminary adaptation node has been identified, further analysis is needed to determine the data flow relationship between this preliminary adaptation node and its upstream and downstream nodes. For example, if the preliminary adaptation node for a certain type of data is the input tax calculation node, it is necessary to examine the business logic requirements of its downstream nodes, such as the tax payable calculation node, to determine whether this type of data, after being processed by the input tax calculation node, can meet the requirements of the tax payable calculation node regarding data format, data content, etc.
[0066] If the analysis results show that the requirements are met, the input tax calculation node is identified as the target matching node for this type of data. If the requirements are not met, such as the data lacking certain derived information required by the tax payable calculation node, the matching dimensions of the business logic mapping matrix are readjusted, and potential business function tags for the data are added, such as potential tags generated from derived information. Then, the matching degree is recalculated, and the initial matching node is re-determined and analyzed until a target matching node that meets the needs of downstream nodes is found. If the requirements are still not met after multiple adjustments, the data is marked as data to be processed.
[0067] Step S126: Associate the data of all identified target adaptation nodes with the corresponding RPA application process nodes, record the business function tags of the data to be adapted and the reasons for non-adaptation, and form a two-way adaptation result of the process nodes including node identifiers, corresponding data lists and data records to be processed.
[0068] For each type of data with identified target adaptation nodes, an association is established with the corresponding RPA application process nodes. The correspondence between data identifiers and node identifiers is recorded in the form of an association table. For data to be adapted and processed, its business function tags, such as data type tags and business domain tags, are recorded in detail, as well as the reasons for non-adaptation, such as data format incompatibility with the requirements of all nodes or lack of key related information.
[0069] By integrating this related information, data lists, and pending data records, a bidirectional adaptation result for process nodes is formed. This bidirectional adaptation result displays the adapted data and the data that failed to adapt for each RPA application process node.
[0070] Step S130: Based on the bidirectional adaptation results of the process nodes, construct a node processing time prediction model by combining the historical processing data of each RPA application process node, allocate RPA computing resources and data transmission channels based on the node processing time prediction model, and generate a node resource prediction scheduling scheme.
[0071] Based on the bidirectional adaptation results of the process nodes, the data that each RPA application process node needs to process is clarified. Simultaneously, historical processing data of similar data from previous processing by each node is retrieved, and this data is used to construct a node processing time prediction model. This model predicts the time required for each node to process the current data. Then, based on the predicted time and system resource status, computing resources and data transmission channels are allocated to each node, ultimately generating a node resource prediction and scheduling scheme.
[0072] Step S131: Extract the amount of adaptation data corresponding to each RPA application process node from the bidirectional adaptation results of the process nodes, and at the same time retrieve the historical processing data of each RPA application process node. The historical processing data includes the historical adaptation data amount, corresponding processing time and resource consumption records.
[0073] From the bidirectional adaptation results of the process nodes, for each RPA application process node, the total amount of adaptation data corresponding to it is calculated, i.e., the amount of adaptation data. This amount of adaptation data can be measured by the storage size of the data or the number of records. At the same time, historical processing data of each RPA application process node is retrieved from the system's historical database.
[0074] In the historical processing data, the historical adaptation data volume records the amount of data processed by the node in the past; the corresponding processing time records the time spent processing the corresponding historical adaptation data volume; and the resource consumption record includes information such as processor resources, memory resources, and network bandwidth used during the processing.
[0075] Step S132: Extract features from historical processing data, use the extracted features as training samples to input into the regression model, and build a node processing time prediction model. The extracted features include the amount of historical adaptive data, the number of data fields, the data association dimension, and the corresponding processing time.
[0076] Feature extraction is performed on the retrieved historical processing data. First, the amount of historical adaptation data is extracted, which reflects the scale of the historical processing data. Then, the total number of independent fields contained in each batch of historical processing data is counted as the data field quantity feature, which reflects the complexity of the data structure.
[0077] Next, the number of correlation links between each batch of historical processing data and other RPA application process node data is counted, serving as a data correlation dimension feature. This data correlation dimension feature reflects the tightness of the correlation between data. The extracted features, together with the corresponding historical processing times, form a training sample, which is then input into a regression model to construct a node processing time prediction model.
[0078] Step S1321: Clean the historical processing data of each RPA application process node, extract the historical adaptation data volume from the cleaned historical processing data, and use the number of data bytes as the quantitative indicator of the historical adaptation data volume; extract the number of data fields, and count the total number of independent fields contained in each batch of historical processing data; extract the data association dimension, and count the number of association links between each batch of historical processing data and other RPA application process node data.
[0079] The historical processing data of each RPA application process node is cleaned to remove duplicate data, abnormal data, and data with excessive missing values. After cleaning, the amount of historical adaptation data is extracted from the data and quantified using the number of bytes, that is, the total number of storage bytes for each batch of historical data is counted.
[0080] When extracting the number of data fields, the structure of each batch of historical processed data is analyzed to count the total number of independent fields contained therein, with each independent field name corresponding to a counting unit. When extracting the data correlation dimension, the number of correlation links is counted by analyzing the reference relationships and transmission relationships between each batch of historical processed data and other node data, with each independent correlation path counted as one correlation link.
[0081] Step S1322: Use the historical adaptation data volume, number of data fields, and data association dimension as model input features, and the corresponding historical processing time as model output labels to construct a training sample set.
[0082] The volume of historical data extracted after cleaning, the number of data fields, and the dimension of data association were determined as the input features of the model. The historical processing time corresponding to each batch of historical data was used as the output label of the model, i.e., the target value predicted by the model.
[0083] Input features and output labels are matched one-to-one in batches to form multiple training samples. These training samples are stored in a fixed format, and each sample contains a corresponding combination of input features and an output label value. For example, a batch of historically processed data might have a specific number of bytes of historically adapted data, a certain number of data fields, a certain number of data association dimensions, and a specific historical processing time period, thus constituting a complete training sample. All of these samples are then combined to form a training sample set.
[0084] Step S1323: Divide the training sample set into a training subset and a validation subset, wherein the training subset is used for model training and the validation subset is used for model performance validation.
[0085] The constructed training sample set is then partitioned. The partitioning process uses a preset ratio, for example, dividing the sample set into a training subset and a validation subset according to a certain ratio. Samples are selected through random sampling during partitioning to ensure that both the training and validation subsets can well represent the feature distribution of the overall sample set.
[0086] The training subset contains most of the sample data and is mainly used for model parameter learning and iterative training; the validation subset contains a small portion of the sample data and is used to evaluate the model's performance during training and to promptly identify problems such as overfitting or underfitting.
[0087] Step S1324: Select the gradient boosting regression model as the base model, input the training subset into the base model for iterative training, and adjust the hyperparameters of the base model according to the prediction error of the validation subset in each iteration. The hyperparameters include the learning rate, the number of decision trees, and the tree depth.
[0088] A gradient boosting regression model was chosen as the base model for constructing the node processing time prediction model. This gradient boosting regression model improves prediction performance by constructing multiple decision trees and performing ensemble learning. A training subset was input into the base model to initiate the iterative training process.
[0089] In each training iteration, the model constructs decision trees and updates parameters based on the input training data. Simultaneously, a validation subset is input into the currently trained model, and the prediction error between the model's predictions for the validation subset and the actual output labels is calculated. Based on this prediction error, the model's hyperparameters are adjusted, such as reducing the learning rate to slow down training, increasing the number of decision trees to increase model complexity, or adjusting tree depth to control the growth scale of the decision trees, until the model performance reaches an optimal state.
[0090] Step S1325: When the prediction error of the validation subset is less than the preset error threshold, stop the model training and obtain the final node processing time prediction model.
[0091] During model iterative training, the prediction error of the validation subset is continuously monitored. A preset error threshold is established, which is an acceptable error range determined based on historical model training experience and actual business needs. When, after a certain iteration, the calculated prediction error of the validation subset is less than the preset error threshold, it indicates that the model's current performance meets the requirements.
[0092] At this point, the iterative training process of the model is stopped, and the current model parameters are saved to obtain the final node processing time prediction model. This node processing time prediction model can predict the processing time of RPA application process nodes relatively accurately based on the input feature data.
[0093] Step S1326: Perform a generalization test on the final node processing time prediction model. Select historical processing data that was not used in training as test samples. Input the test samples into the node processing time prediction model to obtain the predicted processing time. Calculate the deviation rate between the predicted processing time and the actual processing time. If the deviation rate is lower than the preset deviation threshold, the node processing time prediction model is confirmed to be usable. If the deviation rate is higher than the preset deviation threshold, readjust the training samples or model structure until the generalization of the node processing time prediction model meets the requirements.
[0094] To ensure the node processing time prediction model has good generalization ability and can be applied to different data scenarios, the final model is subjected to generalization tests. A portion of the historical processing data that was not used in model training and validation is selected as test samples. These sample data also include input features such as the amount of historical adaptation data, the number of data fields, and the data association dimensions, as well as the corresponding actual processing time.
[0095] The input features of the test samples are fed into the node processing time prediction model to obtain the predicted processing time output by the model. The deviation rate between the predicted processing time and the actual processing time is calculated. The deviation rate is calculated as the ratio of the difference between the two to the actual processing time. If the deviation rate is lower than a preset deviation threshold, it indicates that the model can maintain good prediction performance even on unseen data, confirming that the node processing time prediction model is usable. If the deviation rate is higher than the preset deviation threshold, it is necessary to readjust the selection range of training samples or modify the model structure, such as increasing the number of training samples or adjusting hyperparameters. Then, the model should be retrained and tested until the model's generalization meets the requirements.
[0096] Step S133: Input the current amount of adapted data, number of data fields, and data association dimension of each RPA application process node into the node processing time prediction model to obtain the predicted processing time of each RPA application process node.
[0097] From the bidirectional adaptation results of process nodes, extract the amount of adaptation data that each RPA application process node needs to process, count the number of independent data fields contained in this adaptation data, and analyze and determine the number of association links between the data and the data of other nodes, i.e., the data association dimension.
[0098] The three feature data points are organized according to the input format required by the node processing time prediction model, and then input into the model respectively. The model processes and calculates the input feature data, and outputs the predicted processing time for each RPA application process node. This predicted processing time reflects the estimated time required for the node to process the currently adapted data.
[0099] Step S134: Retrieve the available computing resource ledger of the RPA system. The available computing resource ledger includes the computing speed of idle processors, the read and write speed of memory, and the capacity of storage resources. At the same time, retrieve the real-time transmission rate and channel load record of available data transmission channels.
[0100] The resource management module of the RPA system retrieves the available computing resource ledger. This ledger updates the usage status and performance parameters of various computing resources in the system in real time. Among them, the processing speed of idle processors reflects the number of instructions that the processor can process per unit time; the memory read and write speed indicates how fast the memory can perform data reading and writing operations; and the storage resource capacity shows the amount of space currently available for data storage.
[0101] At the same time, relevant information about available data transmission channels is retrieved from the system's network management module, including the real-time transmission rate of each channel, i.e., the amount of data that can be transmitted per unit time, and the channel load record, which reflects the current busyness of the channel, such as the proportion of bandwidth used to the total bandwidth.
[0102] Step S135: Based on the predicted processing time of each RPA application process node and the available computing resource ledger, allocate corresponding computing resource specifications and quantities to each RPA application process node.
[0103] Analyzing the predicted processing time of each RPA application process node, nodes with longer predicted processing times typically require processing larger datasets or more complex data structures, thus demanding higher computing resources. Based on information such as idle processor processing speed, memory read / write speed, and storage capacity recorded in the available computing resource ledger, appropriate computing resources are allocated to each node.
[0104] For example, for input tax calculation nodes with long prediction processing times and a large number of data fields, allocate processors with high computing speeds, memory with fast read and write speeds, and sufficient storage resources; for data verification nodes with shorter prediction processing times, allocate relatively basic computing resource specifications. At the same time, based on the number of nodes and resource requirements, rationally determine the allocation quantity of each specification of computing resources to ensure the rationality and efficiency of resource allocation.
[0105] Step S136: Based on the size of the adaptive data volume and the predicted processing time of each RPA application process node, calculate the minimum data transmission rate required for each RPA application process node, and match a data transmission channel with a real-time transmission rate not lower than the required minimum data transmission rate and whose current data transmission volume does not affect the data transmission progress of the node for each RPA application process node.
[0106] For each RPA application process node, the minimum data transfer rate required is calculated based on its applicable data volume and predicted processing time. The calculation method is to divide the applicable data volume by the predicted processing time to obtain the minimum rate requirement for completing data transfer within the specified time.
[0107] Then, channels with a real-time transmission rate no lower than the minimum data transmission rate are selected from the available data transmission channels. Based on this, the current data transmission volume of these channels is further examined, and channels whose data transmission volume has not reached channel saturation and will not affect the data transmission progress of nodes are selected. The most suitable data transmission channel is then matched to each node to ensure smooth data transmission between nodes.
[0108] Step S137: Record the specifications, quantity, and data transmission channel identifier of the computing resources allocated to each RPA application process node, mark the start time and expected release time of the computing resource allocation, and form a node resource prediction and scheduling scheme that includes node identifier, resource configuration details, transmission channel parameters, and predicted processing time.
[0109] The specifications of computing resources allocated to each RPA application process node, such as processor model, memory capacity, resource quantity, and corresponding transmission channel identification information, will be recorded in detail. Simultaneously, based on the node's expected start time for processing, the start time of the computing resource allocation will be marked, i.e., the point in time when the resources begin providing services to the node; based on the node's predicted processing duration, the expected release time of the resources will be calculated and marked, i.e., the point in time when the resources can be recycled and reused after the node's processing is completed.
[0110] This information is integrated to form a node resource prediction and scheduling scheme. This scheme includes the identifier of each node, detailed resource configuration, transmission channel parameters such as transmission rate, and the predicted processing time of each node.
[0111] Step S140: For RPA application process nodes with missing data or format deviations in the bidirectional adaptation results of the process nodes, the RPA tool is called based on the node resource prediction and scheduling scheme to perform cross-node data association derivation, supplement the missing data or correct the format deviation data, and obtain the node data after association derivation.
[0112] In the bidirectional adaptation results of process nodes, some RPA application process nodes have missing data or format deviations in their adaptation data. For these nodes, based on the computing resources and transmission channels allocated in the node resource prediction and scheduling scheme, the RPA tool is invoked to initiate a cross-node data association derivation process.
[0113] The RPA tool retrieves relevant data from associated nodes, analyzes the relationships between data, supplements missing data, corrects data with format deviations, and finally obtains processed node data after association derivation, ensuring the integrity and standardization of node data.
[0114] Step S141: Filter out RPA application process nodes with missing data or format deviations from the bidirectional adaptation results of the process nodes, extract the adaptation data corresponding to the RPA application process nodes, and determine the field name of the missing data or the specific type of the format deviation.
[0115] The bidirectional adaptation results of each process node are checked one by one to identify RPA application process nodes with data problems. These nodes either have missing fields (i.e., data missing) or the data format does not meet the node processing requirements (i.e., format deviation).
[0116] Extract the corresponding adaptation data for these nodes and perform detailed data analysis. For cases of missing data, identify the specific names of the missing fields, such as the missing "Input Invoice Authentication Date" field in the Input Tax Calculation node; for cases of format deviation, determine the specific type of deviation, such as a date field formatting "Year-Month-Day" but written as "Month / Day / Year", or a numeric field containing redundant text symbols, etc.
[0117] Step S142: According to the node resource prediction and scheduling scheme, call the RPA computing resources allocated to the RPA application process node, and start the cross-node data retrieval module of the RPA tool. The cross-node data retrieval module is used to query the adaptation data of upstream and downstream nodes that have data flow association with the current RPA application process node.
[0118] Based on the node resource prediction and scheduling scheme, obtain the RPA computing resource information allocated to the RPA application process nodes with data problems, such as the calling paths and permissions of resources like processors and memory. Then, invoke these computing resources through the resource invocation interface.
[0119] Activate the module in the RPA tool specifically designed for cross-node data retrieval. This module can query the upstream and downstream nodes in the RPA application process template that have direct or indirect data flow relationships with the current node based on the current node's identification information, and obtain the adaptation data of these related nodes.
[0120] Step S143: If data is missing, the RPA tool obtains the adaptation data of the upstream node of the current RPA application process node through the cross-node data retrieval module, analyzes the correlation between the adaptation data of the upstream node and the missing fields of the current RPA application process node, and if the adaptation data of the upstream node contains related data that can deduce the missing fields of the current RPA application process node, data deduction calculation is performed based on the correlation to obtain the supplementary data of the missing fields.
[0121] When a missing data node is detected in the current RPA application process, the cross-node data retrieval module of the RPA tool focuses on querying the upstream node's adaptation data for that node. Upstream nodes process data before the current node in the business process, and their data often has certain business relevance to the current node's data.
[0122] After obtaining the adaptation data from the upstream nodes, this data is analyzed to find information related to the missing fields of the current node. If it is found that there is related data in the adaptation data of the upstream nodes that can deduce the missing fields, such as a temporal relationship between the "invoice issuance date" field of the upstream node and the "input invoice authentication date" field missing in the current node, then data deduction and calculation are performed based on the above relationship to obtain the supplementary data for the missing fields.
[0123] Step S1431: The cross-node data retrieval module of the RPA tool queries the upstream node identifier in the RPA application process template that has a direct data flow input relationship with the current RPA application process node based on the current RPA application process node identifier, and retrieves the adaptation data of the upstream node based on the upstream node identifier.
[0124] After receiving a data retrieval request, the cross-node data retrieval module of the RPA tool queries the node relationship graph of the RPA application process template based on the unique identifier of the current RPA application process node. This node relationship graph records the data flow relationships between all nodes, and the query finds the identifier information of the upstream node that has a direct data flow input relationship with the current node.
[0125] Based on the upstream node identifier, the system retrieves the corresponding adaptation data from the system's data storage module. This data is stored in a structured form, containing multiple fields and their corresponding values.
[0126] Step S1432: Perform structural parsing on the adaptation data of the upstream node, extract key fields from the adaptation data of the upstream node, and compare the key fields with the missing fields of the current RPA application process node in terms of business logic to determine whether there is a causal relationship or a subordinate relationship between the two.
[0127] The retrieved upstream node adaptation data is structurally parsed to identify each field, determining its name, data type, and meaning. Key fields potentially related to the missing fields in the current node are then extracted; these key fields typically have a logical connection to the missing fields.
[0128] The extracted key fields are compared with the missing fields in the current RPA application process node using business logic to analyze their relationship in the business process. If a change in the value of a key field directly leads to a change in the value of a missing field, then a causal relationship is determined; if the information in the missing field is derived from the information in the key field, or belongs to the business scope represented by the key field, then a subordinate relationship is determined.
[0129] Step S1433: If a causal relationship exists, the RPA tool retrieves the corresponding causal relationship rule from the application business logic library and, according to the causal relationship rule, derives the supplementary data for the missing fields of the current RPA application process node from the corresponding fields in the adaptation data of the upstream node.
[0130] When a causal relationship is determined between a key field in an upstream node and a missing field in the current node, the RPA tool accesses the declaration business logic library. This library stores causal relationship rules between different fields in various declaration business scenarios, such as the rule that "the invoice issuance date plus the certification period equals the input invoice certification date".
[0131] Based on the determined causal relationship type, the corresponding causal relationship rule is retrieved. Then, according to the calculation method or logical relationship specified in the rule, and based on the values of key fields in the upstream node's adaptation data, supplementary data for missing fields in the current RPA reporting process node is derived. For example, based on the above rule, supplementary data for "input invoice certification date" is calculated using the upstream node's "invoice issuance date" and the preset "certification period".
[0132] Step S1434: If a subordinate relationship exists, the RPA tool extracts the relationship identification information from the corresponding field in the adaptation data of the upstream node, retrieves the corresponding business database through the relationship identification information, and obtains the data corresponding to the missing field of the current RPA application process node as supplementary data.
[0133] When a dependency relationship is identified, the RPA tool extracts relationship identification information from key fields in the upstream node's adapted data. This relationship identification information uniquely points to the data record in the business database related to the missing field. For example, the "Contract Number" field in the upstream node's data can serve as relationship identification information to retrieve detailed data related to the contract.
[0134] By accessing the corresponding business database through this association identifier information, a data retrieval operation is performed. The database is then used to locate data records containing information about the missing fields of the current node, and the data corresponding to the missing fields is extracted and used as supplementary data. For example, the "contract signing date" can be retrieved from the contract database using the "contract number" and used as supplementary data for the missing "business occurrence date" of the current node.
[0135] Step S1435: During the derivation process, the RPA tool records the upstream node identifier, key field content, and business logic rules on which the derivation is based, forming a derivation process record.
[0136] Throughout the data derivation process, the RPA tool enables logging. It records detailed information about the upstream nodes used in the derivation process to trace the data source; it records the specific content of key fields used, including field names and values; and it records the numbers or descriptions of the applied business logic rules.
[0137] These records are organized chronologically and according to the derivation steps to form a complete derivation process record. This derivation process record can be used for subsequent data auditing and problem investigation to ensure the traceability and legality of supplementary data.
[0138] Step S1436: Fill the missing fields of the current RPA application process node with the derived supplementary data, and confirm the business logic coherence between the supplementary data and the existing data of the current RPA application process node to ensure that there is no logical conflict between the supplementary data and the existing data in the application business scenario, forming a node data fragment with complete fields after association derivation.
[0139] The derived supplementary data is then filled into the missing fields of the current RPA application process node according to the field correspondence. After completion, a business logic coherence check is performed on all data in the node. This check examines whether the logical relationship between the supplementary data and existing data is reasonable in the business scenario, such as whether the order of date fields conforms to business rules, and whether the calculation relationship between numerical fields is correct.
[0140] If there is no logical conflict, the supplementary data is valid, forming a node data fragment with complete fields after association derivation; if there is a logical conflict, the derivation process needs to be checked again or other related data needs to be found for re-derivation until the data logic is coherent.
[0141] The RPA tool retrieves compatible data formats for other RPA application process nodes that have the same business logic as the current RPA application process node. Using these compatible data formats as a reference standard, it converts data with format deviations from the current RPA application process node into the reference standard format.
[0142] When a format deviation is detected in the current RPA submission process node, the RPA tool initiates a format correction mechanism. First, based on the business logic description of the current node, it determines the business process and data processing type to which it belongs. Then, it searches the RPA submission process template for other nodes with the same or highly similar business logic as the current node. These nodes have already formed standardized and adapted data formats in past processing.
[0143] Extract the adapted data formats of these reference nodes as standard templates. These templates include field naming rules, data types (such as text, numeric, date, etc.), and format specifications (such as the "year-month-day" format for dates, the number of decimal places for numeric values, etc.). Compare the data with format deviations in the current node with the standard templates to identify the specific location and type of deviations, such as incorrect date separators, redundant numerical units, or inconsistent field name spellings.
[0144] Perform format conversion operations on the deviation data of the current node according to the format requirements of the standard template. For example, convert the date format "month / day / year" to "year-month-day", remove non-numeric symbols after numeric fields, and correct spelling errors in field names. During the conversion process, ensure that the authenticity and accuracy of the data content are not affected by the format adjustment.
[0145] Step S146: After completing the data supplementation or format conversion, the RPA tool will merge the processed data with the original adapted data of the current RPA application process node to form the associated deduced node data, record the node range, data association relationship and deduction process of cross-node data retrieval, and associate it with the current RPA application process node identifier.
[0146] After completing data missing data supplementation or format deviation conversion, the RPA tool merges the processed data with the original adapted data of the current node. For data missing data supplementation, the derived supplementary data is integrated with the complete field data in the original adapted data to ensure that all fields have valid data. For format deviation conversion, the converted, standardized format data replaces the original deviated format data, preserving the original data's content information.
[0147] After merging, the resulting inferred node data contains all necessary fields for the current node and conforms to business requirements. Simultaneously, the RPA tool meticulously records the scope of nodes involved in cross-node data retrieval, including the identifiers of upstream, downstream, or reference nodes; records the relationships between data, such as causal association rule numbers and subordinate association identifiers; and records key steps and operational details in the inference process, such as data extraction methods and format conversion rules.
[0148] The above-mentioned recorded information is associated with the identifier of the current RPA application process node and stored to form a node data processing file, which facilitates the subsequent tracing and auditing of data sources and processing processes.
[0149] Step S150: Based on the business logic sequence of the RPA application process template, integrate the node data without data problems in the bidirectional adaptation result of the process nodes with the node data after association deduction to generate an application data submission package, and push the application data submission package to the target application system to obtain application submission feedback data.
[0150] After processing all node data, the node data without data issues and the node data after correlation derivation are integrated according to the business logic sequence of each node in the RPA application process template. This integration forms a complete application data chain, which is then packaged into an application data submission package according to the requirements of the target application system and pushed to the target application system. The system's returned application submission feedback data is received and recorded, completing the entire application data processing flow.
[0151] Step S151: Extract node data without data problems from the bidirectional adaptation results of the process nodes. The node data without data problems is not marked as missing data or format deviation, and has completed bidirectional adaptation with the corresponding RPA application process node.
[0152] The RPA tool filters the bidirectional adaptation results of process nodes, identifying node data that is not marked as missing or formatted incorrectly. This data has been confirmed during the adaptation process to meet the business logic requirements and data specifications of the corresponding RPA application process node. The node data without data issues is extracted, and its adaptation relationship with the corresponding node is checked to ensure the validity and completeness of the data and node correspondence.
[0153] The extracted data from nodes with no data issues will be categorized and stored according to their node identifiers to prepare data for subsequent integration steps and ensure that this data can be directly used to construct the application data link.
[0154] Step S152: Collect the associated derivation node data corresponding to all RPA application process nodes with data problems, and ensure that each RPA application process node with data problems has generated associated derivation node data. If any RPA application process node cannot solve the data problem through cross-node data association derivation, mark the node data of that RPA application process node as temporarily suspended data and exclude it from the integration scope.
[0155] The RPA tool collects the corresponding derived node data for nodes marked as having data problems in the bidirectional adaptation results of the process nodes. For each node with a data problem, it checks whether the derived node data has been successfully generated and whether the data has been verified to be free of logical conflicts and formatting issues.
[0156] If a node with data issues is found to have problems that cannot be resolved through cross-node data correlation (e.g., missing key data that cannot be derived from upstream or downstream nodes, or format deviations that cannot be corrected by reference nodes), then the data of that node will be marked as temporarily suspended. Temporarily suspended data will not be included in this integration; its node identifier and the type of unresolved issue will be recorded separately, and it will be processed later through manual processing or data supplementation.
[0157] Step S153: Refer to the business logic order of each RPA application process node in the RPA application process template, sort the node data without data problems and the node data after correlation deduction, so that the sorted node data order is consistent with the flow order of the application business.
[0158] Retrieve the business logic sequence information of each RPA application process node in the RPA application process template. This business logic sequence reflects the natural flow of the application process from start to finish, such as data verification node → input tax calculation node → output tax calculation node → tax payable calculation node → application form generation node, etc.
[0159] Based on this order, the collected node data with no data and the node data after correlation derivation are sorted. The data of each node is arranged according to its position in the business process to ensure that the sorted node data order is consistent with the actual flow order of the application business.
[0160] Step S154: Perform association field concatenation processing on the sorted node data, extract the association identifier field from each node data, and concatenate the node data of different nodes into a complete declaration data link through the association identifier field.
[0161] In the sorted node data, the association identifier field is extracted from each node. This association identifier field is typically the unified social credit code of the applicant, and it remains consistent and unique across all node data. The association identifier field establishes connections between different node data, ensuring that node data belonging to the same applicant can be accurately linked.
[0162] Following the business logic sequence and associated identification fields, the data from each node is linked together to form a complete declaration data chain. For example, the basic information verified in the data verification node is linked to the accounting data of the input tax calculation node through the unified social credit code, and then linked to the accounting data of the output tax calculation node, and so on, until all node data are linked together to form a complete data chain covering the entire declaration process.
[0163] Step S155: According to the application data organization format required by the target application system, the concatenated application data link is encapsulated into an application data submission package. The application data submission package includes package identification information, data link body and data flow record. The package identification information includes the application body identifier and submission batch code. The data flow record includes the adaptation time and correlation derivation time of each node data.
[0164] Obtain the application data organization format specification published by the target application system. This specification clarifies the structure, field naming rules, data format requirements, and encapsulation methods of the application data submission package. Based on the specification requirements, encapsulate the concatenated application data chain.
[0165] First, package identification information is generated. The applicant entity identifier is the Unified Social Credit Code from the application data, and the batch code is generated by the system based on the application date, application period, and serial number, ensuring the uniqueness of each submission package. Second, the concatenated application data link is used as the main data link, and adjusted according to standardized field naming and format requirements. Then, the adaptation time of each node's data during the adaptation process and the correlation derivation time of nodes with data problems are extracted to form a data flow record, recording the processing time of data at each node.
[0166] The package identification information, data link entity, and data flow record are combined according to the structure and order required by the specifications to form the initial content of the declaration data submission package.
[0167] Step S1551: Obtain the application data encapsulation specification published by the target application system, and parse the structural requirements, field naming rules and data format standards of the application data submission package in the application data encapsulation specification.
[0168] The RPA tool accesses the target declaration system's specification release platform via an interface to download the latest declaration data encapsulation specification file. It then parses the specification file to clarify the overall structural requirements of the declaration data submission package, such as the mandatory first-level modules (package identifier, data body, flow records, etc.) and the hierarchical relationships between modules; extracts field naming rules, including field name naming formats, prohibited characters, and fixed names for key fields; and determines data format standards, such as specific format requirements for various data types like dates, numbers, and text, data compression methods, and encryption standards.
[0169] The parsed specification content is stored as structured specification parameters to guide the subsequent data submission package encapsulation process, ensuring that the encapsulated submission package meets the requirements of the target system.
[0170] Step S1552: Generate package identifier information for the declaration data submission package, wherein the declaration subject identifier is the subject code corresponding to the associated identifier in the declaration data set, and the submission batch code is generated by the RPA system in combination with the current date and the submission sequence of the day. The submission batch code includes a date field, a time field, and a sequence field.
[0171] The applicant identifier in the package identification information directly adopts the entity code corresponding to the associated identifier in the application data set, namely the enterprise's unified social credit code, ensuring that this identifier uniquely corresponds to the applicant entity. The submission batch code is automatically generated by the RPA system, and the generation rule is to combine the current date, the submission time period of the day, and the submission sequence.
[0172] The date field uses the format "year-month-day" to represent the date the commit package was generated; the time period field divides a day into several time periods and represents them with specific codes, such as AM, PM, and PM; the sequence field indicates the generation order of the commit package within the current date and time period, represented by consecutive numbers. For example, the commit batch code of a commit package can be represented as "20250821-AM-001", where "20250821" is the date field, "AM" is the time period field, and "001" is the sequence field.
[0173] Step S1553: Take the concatenated declaration data link as the main body of the data link, and adjust the field names of the data of each node in a unified manner according to the field naming rules of the declaration data encapsulation specification.
[0174] Using the concatenated declaration data link as the main body of the data link, the field names of the data at each node are checked. By comparing with the field naming rules in the declaration data encapsulation specification, field names that do not conform to the rules are identified, such as field names containing special characters, spelling errors, or names that are too long.
[0175] Field names should be standardized according to specifications, such as removing special characters, correcting spelling errors, simplifying excessively long names, and unifying the capitalization of field names. During the adjustment process, ensure that the changes to field names do not affect the actual meaning of the field data, and that each field name is unique within the data chain to avoid field name conflicts.
[0176] Step S1554: Convert the data format of each field in the data link body according to the data format standard of the declaration data encapsulation specification.
[0177] According to the data format standards specified in the data encapsulation specifications, the data formats of each field in the main body of the data link are converted. For date fields, they are uniformly converted to the "year-month-day" or "year-month-day hour:minute:second" format required by the specifications; for numeric fields, the specified number of decimal places are retained as required by the specifications, redundant unit symbols are removed, and they are converted to a pure numeric format; for text fields, leading and trailing spaces are removed, and the character encoding format is unified.
[0178] During the conversion process, the accuracy of the data format conversion is verified, such as checking the reasonableness of dates and the precision of values, to ensure that the converted field data not only conforms to the format standards but also truly reflects the business content.
[0179] Step S1555: Extract the adaptation time of each node data in the bidirectional adaptation process of the process nodes, and the association derivation time of the RPA application process nodes with data problems. Associate the adaptation time and the association derivation time with the corresponding node identifier to form a data flow record.
[0180] The adaptation time of each node's data is extracted from the bidirectional adaptation results of the process nodes. This adaptation time records the specific moment when the node data and the corresponding RPA application process node complete the adaptation. The association derivation time of nodes with data problems is extracted from the processing records of the node data after association derivation. This association derivation time records the specific moment when the node data completes the missing data supplementation or format conversion.
[0181] The adaptation time and correlation derivation time are associated with the corresponding node identifiers to form basic entries in the data flow record. Each entry includes a node identifier and an adaptation time (if the node has no data issues) or a correlation derivation time (if the node has data issues). Following the order of nodes in the business logic sequence, these entries are integrated into a complete data flow record, clearly displaying the processing timeline of each node's data.
[0182] Step S1556: Combine the packet identification information, data link body and data flow record in the order required by the declaration data encapsulation specification to construct the initial structure of the declaration data submission packet.
[0183] According to the structural order required by the data encapsulation specifications, the packet identification information, data link body, and data flow record are combined. Typically, they are arranged in the order of packet identification information first, data link body in the middle, and data flow record last, with each part distinguished by separators or markers required by the specifications.
[0184] During the assembly process, ensure that the content of each part is complete and without omission, the structure is clear, and it meets the requirements of the specification for the overall structure of the submission package.
[0185] Step S1557: Compress the initial structure of the declaration data submission package, using a compression algorithm supported by the target declaration system to reduce the size of the declaration data submission package, and record the file size and compression efficiency before and after compression.
[0186] Compression algorithms supported by the target declaration system, such as ZIP and GZIP, are used to compress the initial structure of the declaration data submission package. During the compression process, an appropriate compression level is set to strike a balance between compression efficiency and compression time, ensuring that the size of the compressed submission package is significantly reduced, facilitating data transmission and storage.
[0187] After compression, record the original file size and the compressed file size of the submitted data package. Calculate the ratio of the two to determine the compression efficiency. Store this information as a compression record for evaluating compression effectiveness and data transmission efficiency.
[0188] Step S1558: Add a data verification identifier to the compressed declaration data submission package. The data verification identifier is generated by performing a hash operation on the entire data of the declaration data submission package. It is used by the target declaration system to confirm that the declaration data submission package has not been tampered with during transmission, thus completing the encapsulation of the declaration data submission package.
[0189] A hash operation is performed on the entire compressed declaration data submission packet using a hash algorithm supported by the target declaration system, such as SHA-256. The hash operation converts all data in the submission packet into a fixed-length hash value, which serves as the data verification identifier. This data verification identifier is unique; if any minor alteration occurs to the submission packet data during transmission, the recalculated hash value will differ from the original verification identifier.
[0190] Add the generated data verification identifier to a specified location in the compressed declaration data submission package, such as the package header or footer, to complete the final encapsulation of the declaration data submission package. The encapsulated submission package contains two parts: compressed data and a data verification identifier, ensuring that the target declaration system can verify the integrity of the data.
[0191] Step S156: Retrieve the interface communication parameters of the target declaration system. The interface communication parameters include the interface access address, data transmission encryption protocol, and identity authentication token. The RPA tool establishes an encrypted communication link with the target declaration system based on the interface communication parameters.
[0192] The RPA tool retrieves the interface communication parameters of the target declaration system from the system's configuration database. The interface access address is the network address where the target system receives declaration data submission packets, usually represented in URL form; the data transmission encryption protocol is an encryption standard that ensures secure data transmission, such as SSL / TLS; the authentication token is the credential for authentication between the RPA system and the target declaration system, containing authentication information and a validity period.
[0193] Based on the interface access address, the RPA tool initiates a connection request to the target reporting system; based on the data transmission encryption protocol, it negotiates and determines the encryption algorithm and key, and encrypts the transmission channel; by submitting an authentication token, it completes the authentication with the target reporting system. After successful authentication, a secure encrypted communication link is established.
[0194] Step S157: Transmit the declaration data submission packet to the target declaration system through an encrypted communication link. During the transmission process, the data transmission progress is tracked in real time. If the transmission is interrupted, the backup data transmission channel is reactivated based on the node resource prediction and scheduling scheme to resume transmission.
[0195] The RPA tool transmits the encapsulated declaration data submission packet to the target declaration system's interface access address via an established encrypted communication link. During transmission, a transmission progress tracking mechanism is activated to monitor the ratio of transmitted data to total data in real time, and to calculate the transmission rate and estimated remaining transmission time.
[0196] If an interruption occurs during transmission, such as a network failure or connection timeout, the RPA tool immediately detects the cause of the interruption and queries the backup data transmission channel information allocated for the current transmission task in the node resource prediction and scheduling scheme. Based on the parameters of the backup channel, an encrypted communication link is re-established, and the transmission of incomplete data continues from the point of interruption, ensuring that the declaration data submission packet can be transmitted completely to the target declaration system.
[0197] Step S158: Receive the transmission response information returned by the target declaration system. The transmission response information includes the data reception status and the integrity identifier of the received data.
[0198] After the data submission packet is transmitted, the RPA tool maintains a communication connection with the target reporting system, waiting for the system to return a transmission response. The transmission response is feedback from the target reporting system regarding the data reception status. The data reception status indicates whether the data was successfully received, such as "success" or "failure." The data integrity identifier is a hash value obtained by the target reporting system after performing a hash operation on the received data. This hash value is compared with the data verification identifier in the data submission packet to verify that the data is complete and has not been tampered with.
[0199] The RPA tool stores the received transmission response information in a local log as a preliminary record of the data transmission result.
[0200] Step S159: If the data reception status is successful, generate application submission feedback data containing packet identification information, successful reception confirmation, and subsequent processing instructions; if the data reception status is failed, record the failure type and corresponding transmission parameters, associate them with the packet identification information of the application data submission packet, and form application submission feedback data containing failure details.
[0201] When the data reception status in the transmission response information is successful, the RPA tool extracts the packet identifier information of the declaration data submission packet, combines it with the reception success confirmation information returned by the target declaration system, such as the successful reception prompt code and message, and the subsequent processing guidance provided by the system, such as the declaration data review cycle and result query method, to generate declaration submission feedback data.
[0202] When data reception fails, the RPA tool identifies the failure type from the transmission response information, such as data verification failure, format error, authentication failure, server busy, etc.; it records relevant parameters during the transmission process, such as transmission time, channel identifier used, and encryption protocol version. The failure type and transmission parameters are then associated with the packet identifier information of the declaration data submission packet to form declaration submission feedback data containing failure details, clarifying the reason for the failure and related background.
[0203] After the feedback data for the application submission is generated, it can be stored in the system's feedback data management module and simultaneously pushed to the applicant's terminal interface or associated notification channels through a preset notification mechanism. The package identification information included in the feedback data can be used to query the processing status of the application data submission package in the system later, ensuring that the applicant can promptly understand the submission status of the application data.
[0204] For the generated submission feedback data containing failure details, the RPA tool will categorize and analyze the failure type. If the failure type is data validation failure, it can further pinpoint the specific validation failure field and, combined with the validation rule description returned by the target submission system, form a detailed description of the failure reason; if it is a format error, it can clearly indicate the location of the non-compliant format and the correct format standard; if it is an identity authentication failure, it can prompt the submitter to check the validity of the identity authentication token or obtain a new token; if it is due to server overload, it can suggest resubmitting during off-peak hours.
[0205] Failed data submission packets are marked as pending and stored in a dedicated failure data buffer. Simultaneously, a retry mechanism or manual intervention is automatically triggered based on the reason for the failure. For failures caused by temporary factors such as server overload, the RPA tool will, based on the node resource prediction and scheduling scheme, re-invoke an idle data transmission channel after a preset time interval and push the data submission packet again according to the original encapsulation and transmission process. For failures requiring data modification, such as data verification failures or format errors, the failure details can be fed back to the data processing module, prompting for re-data association derivation or format correction. After the data correction is completed, the data submission packet is regenerated and pushed again.
[0206] Throughout the entire data processing flow, the system employs multiple privacy protection technologies to protect sensitive data such as the applicant's unified social credit code and the content of the application data. During the data collection phase, the application data is encrypted using an encrypted transmission protocol to prevent theft or tampering during transmission. During the data storage phase, sensitive data is anonymized, such as by replacing or masking certain fields, and access control policies restrict access to only authorized personnel. During data processing, all operations involving privacy data are performed in an encrypted environment to ensure data security during use.
[0207] Figure 2 The illustration shows exemplary hardware and software components of an RPA-based fully automated data processing system 100 for implementing the concepts of this application, provided in some embodiments of this application. For example, a processor 120 may be used in the RPA-based fully automated data processing system 100 and to perform the functions described in this application.
[0208] For example, the RPA (Real-Time Processing) fully automated data processing system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the RPA fully automated data processing system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The RPA fully automated data processing system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0209] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned RPA full-process automated reporting data processing method is realized.
[0210] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A fully automated RPA (Robotic Process Automation) method for processing application data, characterized in that, The method includes: The system receives a set of application data, which includes various application form data and supporting document data submitted by the applicant. The application form data and supporting document data are associated with each other, and the association identifier is used to uniquely associate the corresponding application data of the same applicant. Retrieve a preset RPA application process template, construct a business logic mapping matrix based on the content attributes of various types of data in the application data set, and perform bidirectional adaptation between various types of data and nodes in the RPA application process template through the business logic mapping matrix to obtain the bidirectional adaptation result of process nodes. Based on the bidirectional adaptation results of the process nodes, and combined with the historical processing data of each RPA application process node, a node processing time prediction model is constructed. Based on the node processing time prediction model, RPA computing resources and data transmission channels are allocated, and a node resource prediction scheduling scheme is generated. For RPA application process nodes with missing data or format deviations in the bidirectional adaptation results of the process nodes, the RPA tool is called based on the node resource prediction and scheduling scheme to perform cross-node data association derivation, supplement the missing data or correct the format deviation data, and obtain the node data after association derivation. Based on the business logic sequence of the RPA application process template, the node data without data problems in the bidirectional adaptation results of the process nodes are integrated with the node data after association deduction to generate an application data submission package. The application data submission package is then pushed to the target application system to obtain application submission feedback data.
2. The RPA full-process automated reporting data processing method according to claim 1, characterized in that, The process involves retrieving a preset RPA application process template, constructing a business logic mapping matrix based on the content attributes of various data types in the application data set, and bidirectionally adapting various data types to the nodes in the RPA application process template using the business logic mapping matrix to obtain bidirectional adaptation results for process nodes, including: Retrieve a preset RPA application process template. The RPA application process template contains multiple RPA application process nodes connected in series according to the application business logic. Each RPA application process node is marked with a business logic description. The business logic description includes the application business link corresponding to the node, the functional positioning of the data to be processed, and the data flow relationship with upstream and downstream nodes. The content attributes of various application form data and supporting document data in the application data set are parsed, and the business function tags of each type of data are extracted. The business function tags include the application business actions supported by the data, the business decision information they carry, and the related upstream and downstream data flows. Based on the business logic descriptions of all RPA application process nodes and the business function tags of various data, a business logic mapping matrix is constructed. The row dimension of the business logic mapping matrix is the RPA application process node identifier, the column dimension is the data business function tag, and the matrix elements are the business logic matching degree between the node and the data. The business logic matching degree between each type of data and each RPA application process node is calculated using the business logic mapping matrix, and the RPA application process node with the highest business logic matching degree is determined as the initial data adaptation node. Analyze the data flow relationship between the initial adaptation node and the upstream and downstream nodes to determine whether this type of data can be transferred to the downstream node through the initial adaptation node and meet the business logic requirements of the downstream node. If it can, the initial adaptation node is determined as the target adaptation node. If it cannot, the matching dimension of the business logic mapping matrix is readjusted, the potential business function tags of the data are added, and the matching degree is recalculated until the target adaptation node is determined or the data is marked as the data to be adapted and processed. Associate all data of the identified target adaptation nodes with the corresponding RPA application process nodes, record the business function tags of the data to be adapted and the reasons for non-adaptation, and form a two-way adaptation result of the process nodes including node identifiers, corresponding data lists and data records to be processed.
3. The RPA full-process automated application data processing method according to claim 1, characterized in that, The step involves constructing a node processing time prediction model based on the bidirectional adaptation results of the process nodes and combining it with the historical processing data of each RPA application process node. Based on this prediction model, RPA computing resources and data transmission channels are allocated, and a node resource prediction scheduling scheme is generated, including: Extract the amount of adaptation data corresponding to each RPA application process node from the bidirectional adaptation results of the process nodes, and at the same time retrieve the historical processing data of each RPA application process node. The historical processing data includes the historical adaptation data amount, corresponding processing time and resource consumption records. Feature extraction is performed on historical processing data. The extracted features are used as training samples to input into the regression model to build a node processing time prediction model. The extracted features include the amount of historical adaptive data, the number of data fields, the data association dimension, and the corresponding processing time. Input the current amount of adapted data, number of data fields, and data association dimensions of each RPA application process node into the node processing time prediction model to obtain the predicted processing time of each RPA application process node. The available computing resources ledger of the RPA system is retrieved, which includes the computing speed of idle processors, the read and write speed of memory, and the capacity of storage resources. At the same time, the real-time transmission rate and channel load records of available data transmission channels are retrieved. Based on the predicted processing time of each RPA application process node and the available computing resource ledger, allocate corresponding computing resource specifications and quantities to each RPA application process node; Based on the size of the data volume and the predicted processing time of each RPA application process node, calculate the minimum data transmission rate required for each RPA application process node, and match each RPA application process node with a data transmission channel whose real-time transmission rate is not lower than the required minimum data transmission rate and whose current data transmission volume does not affect the data transmission progress of the node. Record the specifications, quantity, and data transmission channel identifier of the computing resources allocated to each RPA application process node, mark the start time and expected release time of the computing resource allocation, and form a node resource prediction and scheduling scheme that includes node identifier, resource configuration details, transmission channel parameters, and predicted processing time.
4. The RPA full-process automated reporting data processing method according to claim 3, characterized in that, The step of extracting features from historical processing data and using the extracted features as training samples to input into a regression model to construct a node processing time prediction model includes: The historical processing data of each RPA application process node is cleaned, and the amount of historical adaptation data is extracted from the cleaned historical processing data, with the number of data bytes as the quantitative indicator of the amount of historical adaptation data; the number of data fields is extracted, and the total number of independent fields contained in each batch of historical processing data is counted; the data association dimension is extracted, and the number of association links between each batch of historical processing data and other RPA application process node data is counted. The training sample set is constructed by using the historical data volume, the number of data fields, and the data association dimension as the model input features and the corresponding historical processing time as the model output label. The training sample set is divided into a training subset and a validation subset, where the training subset is used for model training and the validation subset is used for model performance verification. The gradient boosting regression model is selected as the base model. The training subset is input into the base model for iterative training. In each iteration, the hyperparameters of the base model are adjusted according to the prediction error of the validation subset. The hyperparameters include the learning rate, the number of decision trees, and the tree depth. When the prediction error of the validation subset is less than the preset error threshold, the model training is stopped, and the final node processing time prediction model is obtained. The generalization test of the final node processing time prediction model is carried out. Historical processing data that was not used in training is selected as test samples. The test samples are input into the node processing time prediction model to obtain the predicted processing time. The deviation rate between the predicted processing time and the actual processing time is calculated. If the deviation rate is lower than the preset deviation threshold, the node processing time prediction model is confirmed to be usable. If the deviation rate is higher than the preset deviation threshold, the training samples or model structure are readjusted until the generalization of the node processing time prediction model meets the requirements.
5. The RPA fully automated reporting data processing method according to claim 1, characterized in that, For RPA application process nodes with missing data or format deviations in the bidirectional adaptation results of the process nodes, the RPA tool is invoked based on the node resource prediction and scheduling scheme to perform cross-node data association derivation, supplementing missing data or correcting format deviation data, to obtain the associated node data, including: Filter out RPA application process nodes with missing data or format deviations from the bidirectional adaptation results of the process nodes, extract the adaptation data corresponding to the RPA application process nodes, and determine the field name of the missing data or the specific type of the format deviation. According to the node resource prediction and scheduling scheme, the RPA computing resources allocated to the RPA application process node are called, and the cross-node data retrieval module of the RPA tool is started. The cross-node data retrieval module is used to query the adaptation data of upstream and downstream nodes that have data flow association with the current RPA application process node. If data is missing, the RPA tool obtains the adaptation data of the upstream node of the current RPA application process node through the cross-node data retrieval module, analyzes the correlation between the adaptation data of the upstream node and the missing fields of the current RPA application process node, and if the adaptation data of the upstream node contains related data that can deduce the missing fields of the current RPA application process node, it performs data deduction calculation based on the correlation to obtain the supplementary data for the missing fields. If the upstream node's adaptation data cannot deduce the missing fields of the current RPA application process node, the RPA tool further retrieves the preset data requirements of the downstream nodes of the current RPA application process node, and reversely deduces the reasonable value range of the missing fields of the current RPA application process node based on the preset data requirements of the downstream nodes, and determines the supplementary data for the missing fields in combination with the application business logic. If there is a format discrepancy, the RPA tool will search for the compatible data format of other RPA application process nodes that are consistent with the business logic of the current RPA application process node, and use the compatible data format as a reference standard to convert the data with the format discrepancy of the current RPA application process node into the reference standard format. After completing the data supplementation or format conversion, the RPA tool will merge the processed data with the original adapted data of the current RPA application process node to form the associated deduced node data, record the node range, data association relationship and deduction process of cross-node data retrieval, and associate it with the current RPA application process node identifier.
6. The RPA full-process automated application data processing method according to claim 5, characterized in that, If data is missing, the RPA tool obtains the adaptation data of the upstream node in the current RPA application process node through the cross-node data retrieval module, analyzes the correlation between the adaptation data of the upstream node and the missing fields of the current RPA application process node, and if the adaptation data of the upstream node contains related data that can deduce the missing fields of the current RPA application process node, data derivation calculation is performed based on the correlation to obtain supplementary data for the missing fields, including: The cross-node data retrieval module of the RPA tool queries the upstream node identifier in the RPA application process template that has a direct data flow input relationship with the current RPA application process node based on the current RPA application process node identifier, and retrieves the adaptation data of the upstream node based on the upstream node identifier. The adaptation data of the upstream node is structurally parsed, key fields are extracted from the adaptation data of the upstream node, and the key fields are compared with the missing fields of the current RPA application process node by business logic to determine whether there is a causal relationship or a subordinate relationship between the two. If a causal relationship exists, the RPA tool retrieves the corresponding causal relationship rule from the application business logic library and, according to the causal relationship rule, derives the supplementary data for the missing fields of the current RPA application process node from the corresponding fields in the adaptation data of the upstream node. If a subordinate relationship exists, the RPA tool extracts the relationship identification information from the corresponding field in the adaptation data of the upstream node, retrieves the corresponding business database through the relationship identification information, and obtains the data corresponding to the missing field of the current RPA application process node as supplementary data. During the derivation process, the RPA tool records the upstream node identifiers, key field contents, and business logic rules on which the derivation is based, forming a derivation process record; The derived supplementary data is filled into the missing fields of the current RPA application process node. The business logic coherence between the supplementary data and the existing data of the current RPA application process node is confirmed to ensure that there is no logical conflict between the supplementary data and the existing data in the application business scenario, forming a node data fragment with complete fields after association derivation.
7. The RPA full-process automated application data processing method according to claim 1, characterized in that, Based on the business logic sequence of the RPA application process template, the node data without data issues in the bidirectional adaptation results of the process nodes are integrated with the node data after association derivation to generate an application data submission package. The application data submission package is then pushed to the target application system to obtain application submission feedback data, including: Extract node data without data problems from the bidirectional adaptation results of the process nodes. The node data without data problems is not marked as missing data or format deviation, and has completed bidirectional adaptation with the corresponding RPA application process node. Collect the associated derivation node data corresponding to all RPA application process nodes with data problems, and ensure that each RPA application process node with data problems has generated associated derivation node data. If any RPA application process node cannot solve the data problem through cross-node data association derivation, mark the node data of that RPA application process node as deferred data and exclude it from the integration scope. Referring to the business logic order of each RPA application process node in the RPA application process template, sort the node data without data problems and the node data after correlation deduction, so that the sorted node data order is consistent with the flow order of the application business. Perform association field concatenation processing on the sorted node data, extract the association identifier field from each node data, and concatenate the node data of different nodes into a complete declaration data link through the association identifier field; According to the data organization format required by the target application system, the concatenated application data link is encapsulated into an application data submission package. The application data submission package includes package identification information, data link body and data flow record. The package identification information includes the application body identifier and submission batch code. The data flow record includes the adaptation time and correlation derivation time of each node data. The RPA tool retrieves the interface communication parameters of the target declaration system, which include the interface access address, data transmission encryption protocol and identity authentication token. The RPA tool establishes an encrypted communication link with the target declaration system based on the interface communication parameters. The declaration data submission packet is transmitted to the target declaration system via an encrypted communication link. The data transmission progress is tracked in real time during the transmission process. If the transmission is interrupted, the backup data transmission channel is reactivated based on the node resource prediction and scheduling scheme to restore the transmission. Receive transmission response information returned by the target declaration system, wherein the transmission response information includes data reception status and integrity identifier of received data; If the data reception status is successful, a declaration submission feedback data containing packet identification information, successful reception confirmation, and subsequent processing instructions is generated; if the data reception status is unsuccessful, the failure type and corresponding transmission parameters are recorded, and associated with the packet identification information of the declaration data submission packet to form a declaration submission feedback data containing failure details.
8. The RPA fully automated reporting data processing method according to claim 7, characterized in that, The step of encapsulating the concatenated application data links into an application data submission package according to the application data organization format required by the target application system includes: Obtain the application data encapsulation specifications published by the target application system, and analyze the structural requirements, field naming rules, and data format standards for the application data submission package in the application data encapsulation specifications. Generate the package identifier information of the declaration data submission package, where the declaration subject identifier is the subject code corresponding to the associated identifier in the declaration data set, and the submission batch code is generated by the RPA system in combination with the current date and the submission sequence of the day. The submission batch code includes the date field, time field and sequence field. The concatenated application data link is taken as the main body of the data link, and the field names of the data at each node are uniformly adjusted according to the field naming rules of the application data encapsulation specification. Convert the data format of each field in the main data link according to the data format standard of the data encapsulation specification. Extract the adaptation time of each node's data during the bidirectional adaptation process of the process nodes, and the associated derivation time of the RPA application process nodes with data problems. Associate the adaptation time and the associated derivation time with the corresponding node identifier to form a data flow record. The initial structure of the application data submission package is constructed by combining the package identification information, data link entity, and data flow record in the order required by the application data encapsulation specification. The initial structure of the declaration data submission package is compressed using a compression algorithm supported by the target declaration system to reduce the size of the declaration data submission package. The file size and compression efficiency before and after compression are recorded. A data verification identifier is added to the compressed declaration data submission package. The data verification identifier is generated by performing a hash operation on the entire data of the declaration data submission package. It is used by the target declaration system to confirm that the declaration data submission package has not been tampered with during transmission and to complete the encapsulation of the declaration data submission package.
9. The RPA fully automated reporting data processing method according to claim 1, characterized in that, The received application data set includes: A distributed data receiving cluster for deploying RPA tools, the distributed data receiving cluster containing multiple parallel receiving nodes, each of which is configured with an independent network port for simultaneously receiving data submitted by multiple claiming entities; Dedicated data receiving protocols are configured for application form data and supporting document data respectively. Application form data adopts a dedicated form data transmission protocol, which supports field-level data transmission; supporting document data adopts a file fragmentation transmission protocol, which supports uploading large file fragments. When the reporting entity initiates a data submission request, the receiving and scheduling module of the RPA tool allocates the data submission request to the parallel receiving node of the corresponding data receiving protocol according to the data type in the data submission request. Parallel receiving nodes receive application form data uploaded by the applicant, parse the field structure of the application form data, extract the preset association identifier field in the application form data, and record the content of the association identifier field and the corresponding application form data reception time; Parallel receiving nodes receive the supporting documents uploaded by the applicant, reassemble the fragmented files according to the file fragmentation transmission protocol, and extract the association identifier from the metadata of the reassembled file. The file metadata includes file naming information and file attribute information. The application form data and supporting document data corresponding to the same association identifier are grouped into a group of application data units, and each application data unit corresponds to a single application data of an applicant. All submitted data units are aggregated to form a submitted data set. A receiving summary identifier is added to the submitted data set. The receiving summary identifier includes the parallel receiving node number, the receiving completion time, and the total number of submitted data units. At the same time, the name of the applicant entity and the submitting terminal information corresponding to each submitted data unit are recorded to complete the receiving of the submitted data set.
10. An RPA (Robotic Process Automation) system for fully automated data processing of application submissions, characterized in that: The RPA full-process automated data processing system includes a processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the RPA full-process automated data processing method according to any one of claims 1-9.
Citation Information
Patent Citations
Job scheduling method and device, equipment and storage medium
CN115829263A
Distributed computing scheduling system, task processing method, equipment and storage medium
CN116360993A
Log sampling method of RPA workflow, storage medium and equipment
CN118069471A
Human resource platform management method, system and equipment based on big data
CN119067411A
Automatic tax declaration and recheck system
CN120471722A