RPA full-process automatic declaration data processing method and system

By constructing a business logic mapping matrix and a node resource prediction and scheduling scheme, the entire process of RPA application data automation was achieved, solving the problems of inaccurate and inefficient data processing in existing technologies, and improving the automation and accuracy of the application process.

CN120874803BActive Publication Date: 2025-12-09CIIC TECHNOLOGY SERVICES (CHENGDU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383360.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-09
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing RPA data processing methods are difficult to automate the entire process. They lack the ability to understand and process the overall logical relationships of the data, leading to problems such as data entry errors and data omissions, which affect the timeliness and accuracy of the application process.

Method used

By receiving the data set of applications, a business logic mapping matrix is ​​constructed for bidirectional adaptation, a node resource prediction and scheduling scheme is generated, and RPA tools are called to perform cross-node data association derivation, supplement missing data or correct data with format deviations, so as to realize automated data processing.

Benefits of technology

It has improved the automation of the application process, reduced manual intervention, enhanced the accuracy and completeness of data processing, reduced labor costs and error rates, and ensured the smooth operation of the application process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874803B_ABST
    Figure CN120874803B_ABST
Patent Text Reader

Abstract

The application provides an RPA full-process automatic declaration data processing method and system, first receives a declaration data set containing declaration form data and proof material data and having an association identifier, then calls an RPA declaration process template, constructs a business logic mapping matrix to realize bidirectional adaptation of data and process nodes, constructs a node processing time length prediction model based on the adaptation result, generates a node resource prediction scheduling scheme, calls an RPA tool for associated deduction for the node with data problems, finally integrates data to generate a declaration data submission package and pushes it to a target declaration system, and obtains declaration submission feedback data, so that full-process automation of declaration data processing is realized, and the processing efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to an RPA full-process automatic declaration data processing method and system. BACKGROUND

[0002] Under the background of rapid development of digital government affairs and commercial services, various declaration businesses are increasing, involving tax declaration, project declaration, qualification declaration and many other fields. The traditional declaration data processing method mainly relies on manual operation, and the staff needs to manually collect, sort and check various declaration form data and proof material data submitted by the declaration subject. The above-mentioned method not only has low efficiency, but also is prone to human errors, such as data entry errors, data omissions, and format irregularities, which leads to the obstruction of the declaration process and affects the timeliness and accuracy of the declaration business.

[0003] With the emergence of robot process automation (RPA) technology, some declaration businesses begin to try to introduce RPA for automatic processing. However, the existing RPA declaration data processing method is mostly limited to simple data entry and form filling, and lacks the ability to understand and process the overall logical relationship of the declaration data. For the complex business logic in the declaration data set, such as the correlation between different data and the correspondence between data and declaration process nodes, the existing method is difficult to effectively cope with. At the same time, when dealing with data missing or format deviation, manual intervention is often needed, and full-process automatic processing cannot be realized, which limits the application effect and popularization range of RPA technology in declaration data processing. SUMMARY

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide an RPA full-process automatic declaration data processing method, which comprises:

[0005] receiving a declaration data set, the declaration data set containing various declaration form data and proof material data submitted by a declaration subject, the declaration form data and the proof material data having an association identifier, the association identifier being used to uniquely associate the corresponding declaration data of the same declaration subject;

[0006] retrieving a preset RPA declaration process template, constructing a business logic mapping matrix based on the content attributes of various data in the declaration data set, and performing bidirectional adaptation of various data and nodes in the RPA declaration process template through the business logic mapping matrix to obtain a process node bidirectional adaptation result;

[0007] According to the flow node bidirectional adaptation result, a node processing time length prediction model is constructed in combination with historical processing data of each RPA declaration process node, RPA computing resources and data transmission channels are allocated based on the node processing time length prediction model, and a node resource prediction scheduling scheme is generated;

[0008] For the RPA declaration process node with data missing or format deviation in the flow node bidirectional adaptation result, cross-node data association derivation is performed based on the node resource prediction scheduling scheme to call the RPA tool, missing data is supplemented, or format deviation data is corrected, and post-association derivation node data is obtained;

[0009] According to the business logic order of the RPA declaration process template, the node data without data problems in the flow node bidirectional adaptation result and the post-association derivation node data are integrated to generate a declaration data submission package, the declaration data submission package is pushed to a target declaration system, and declaration submission feedback data is obtained.

[0010] In still another aspect, the embodiment of the present application also provides an RPA full-process automatic declaration data processing system, which comprises a processor and a machine readable storage medium, the machine readable storage medium is connected with the processor, the machine readable storage medium is used for storing programs, instructions or codes, and the processor is used for executing the programs, instructions or codes in the machine readable storage medium to realize the above-mentioned method.

[0011] Based on the above aspects, the embodiment of the present application realizes efficient bidirectional adaptation of various data and declaration process nodes by receiving a declaration data set containing declaration form data and proof material data and having a unique association identifier, calling a preset RPA declaration process template and constructing a business logic mapping matrix, accurately grasps the internal relationship between data and process, effectively solves the problem of insufficient understanding of business logic in traditional methods, constructs a node processing time length prediction model based on the adaptation result and historical processing data, and then generates a node resource prediction scheduling scheme, reasonably allocates RPA computing resources and data transmission channels, improves resource utilization efficiency, ensures smooth operation of the declaration process, calls the RPA tool to perform cross-node data association derivation for the node with data problems, realizes automatic data supplement and format correction, reduces manual intervention, improves the accuracy and integrity of data processing, finally integrates data according to the business logic order to generate a declaration data submission package and push it to a target declaration system to obtain declaration submission feedback data, completes the full-process automation of declaration data processing, significantly improves the processing efficiency and quality of declaration business, and reduces labor cost and error rate. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is an execution flow schematic diagram of the RPA full-process automatic declaration data processing method provided by the embodiment of the present application.

[0013] Figure 2 is a schematic diagram of exemplary hardware and software components of an RPA full-process automated declaration data processing system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0014] The present application will be described in detail below with reference to the accompanying drawings, Figure 1 is a flowchart of an RPA full-process automated declaration data processing method provided by an embodiment of the present application, which will be described in detail below.

[0015] Step S110: receiving a declaration data set, the declaration data set containing various types of declaration form data and proof material data submitted by a declaration subject, the declaration form data and the proof material data having an association identifier, the association identifier being used to uniquely associate the corresponding declaration data of the same declaration subject.

[0016] This embodiment takes the enterprise value-added tax monthly declaration scenario as an example, and the declaration subject is an enterprise engaged in production and business activities. The declaration data set covers various types of data submitted by the enterprise through the online declaration platform, wherein the declaration form data includes the value-added tax declaration form main table, the attached data 1-5, the value-added tax exemption declaration form, etc.; and the proof material data includes the scanned copy of the value-added tax special invoice deduction link, the input tax amount transfer proof file, and the electronic version of the contract agreement corresponding to the tax-free project, etc.

[0017] The association identifier uses the unified social credit code of the enterprise, which is recorded in the table header part of the declaration form data and the metadata information of the proof material data. Through this association identifier, all declaration form data and proof material data submitted by the same enterprise can be accurately associated, ensuring the accuracy of data attribution in the subsequent data processing process.

[0018] Step S111: deploying a distributed data receiving cluster of RPA tools, the distributed data receiving cluster containing multiple parallel receiving nodes, each parallel receiving node being configured with an independent network port for simultaneously receiving data submitted by multiple declaration subjects.

[0019] In the enterprise value-added tax monthly declaration data receiving link, the distributed data receiving cluster deployed is composed of multiple parallel receiving nodes. These parallel receiving nodes are distributed on different physical servers and work cooperatively through cluster management software. Each parallel receiving node is configured with an independent network port, which is assigned a different port identifier in the network configuration and is in an active state, capable of receiving external data transmission requests at any time.

[0020] The arrangement of multiple parallel receiving nodes enables the receiving cluster to simultaneously process data submission requests from different enterprises, avoiding congestion of a single node due to excessive data traffic during peak submission periods, and ensuring the efficiency and stability of data reception.

[0021] Step S112: Configure exclusive data receiving protocols for the declaration form data and the proof material data respectively. The declaration form data uses a form data dedicated transmission protocol, and the form data dedicated transmission protocol supports field-level data transmission. The proof material data uses a file fragmentation transmission protocol, and the file fragmentation transmission protocol supports large file fragmentation upload.

[0022] Different exclusive data receiving protocols are configured for different data types in enterprise value-added tax declaration. For declaration form data, a form data dedicated transmission protocol is used. In the data transmission process, each data field in the form is treated as an independent transmission unit, and each transmission unit contains field name, field data type, field length, and field value information. After receiving the data, the receiving end can accurately parse the content of each field according to this information, ensuring the integrity and accuracy of the form data.

[0023] For proof material data, a file fragmentation transmission protocol is used due to the large size of some proof files such as large collections of invoice scans and long-term contract scans. The file fragmentation transmission protocol splits large files into multiple file fragments according to a preset fragmentation size, and each fragment contains fragment number, total fragment number, and file identifier information. After receiving all fragments, the receiving end can recombine the fragments into a complete file according to this information, ensuring the success rate of large file transmission.

[0024] Step S113: When the declaration subject initiates a data submission request, the receiving scheduling module of the RPA tool assigns the data submission request to the parallel receiving node corresponding to the data receiving protocol according to the data type in the data submission request.

[0025] When the enterprise initiates a value-added tax declaration data submission request through the declaration terminal, the data submission request message first enters the receiving scheduling module of the RPA tool. The receiving scheduling module performs preliminary analysis on the message and extracts the data type identification field, which clearly indicates whether the data submitted this time is declaration form data or proof material data.

[0026] For example, step S1131: The receiving scheduling module of the RPA tool listens to the data submission request port of the declaration subject in real time, and when it receives a data submission request, it extracts the data type identifier from the header field of the data submission request message. The data type identifier is used to distinguish between declaration form data and proof material data.

[0027] The receiving scheduling module continuously monitors the configured data submission request port through a preset network monitoring program. When the submission request of the declarant arrives at the port, the monitoring program immediately captures the request message. The receiving scheduling module parses the message header and locates the data type identification field. The field uses a preset encoding rule, such as a specific string or a number code, to correspond to the two types of declaration form data and proof material data, thereby accurately distinguishing the data types.

[0028] Step S1132: The receiving node state monitoring account is called, which includes the current receiving task quantity, the remaining receiving capacity, and the supported data receiving protocol type of each parallel receiving node.

[0029] After determining the data type, the receiving scheduling module calls the receiving node state monitoring account through the interface. The account is maintained in real time by the cluster management system and is updated once every fixed time interval. The account records the unique identifier of each parallel receiving node, as well as the corresponding current receiving task quantity, i.e., the number of declaration data submission requests being processed by the node; the remaining receiving capacity, i.e., the maximum number of tasks that can be carried by the node; and the supported data receiving protocol type, such as a form data dedicated transmission protocol or a file fragment transmission protocol.

[0030] Step S1133: The parallel receiving nodes that support the receiving protocol corresponding to the data type in the data submission request are screened out, and the parallel receiving nodes whose current receiving task quantity does not exceed the node's regular receiving range and whose remaining receiving capacity can completely accommodate the data volume corresponding to the current declaration data submission request are preferentially selected from these parallel receiving nodes as target receiving nodes.

[0031] According to the receiving protocol corresponding to the data type identification, the receiving scheduling module screens out all parallel receiving nodes that support the protocol from the account to form a candidate node list. The nodes in the candidate node list are further screened to check whether the current receiving task quantity of each node is within its regular receiving range, which is pre-set according to the hardware performance of the node; and to check whether the remaining receiving capacity of the node is greater than or equal to the data volume corresponding to the current declaration data submission request, to ensure that the node has sufficient capacity to process the request. The nodes that meet the conditions are included in the target receiving node candidate pool.

[0032] Step S1134: If there are multiple target receiving nodes that meet the conditions, the network delay between each target receiving node and the submission terminal of the declarant is calculated, and the target receiving node with the smallest network delay is selected.

[0033] When there are two or more nodes in the target receiving node candidate pool that meet the conditions, the receiving scheduling module initiates a network delay detection procedure. The procedure sends a probe data packet to each candidate node and records the round-trip time of the data packet from the submission terminal to the node to calculate the network delay. The network delays of all candidate nodes are compared, and the node with the smallest delay value is selected as the final target receiving node to reduce the delay in the data transmission process.

[0034] Step S1135: Send a task allocation instruction to the target receiving node, the task allocation instruction containing the data submission request identifier of the declarant, the data type and the receiving parameter requirement, and updating the current receiving task quantity of the target receiving node in the receiving node state monitoring account.

[0035] The receiving scheduling module generates a task allocation instruction, which contains the unique identifier of the data submission request of the declarant for association with subsequent received data, the data type of this submission to ensure that the node receives correctly, and the receiving parameter requirement such as data transmission timeout time and verification method. After the instruction is sent to the target receiving node through the internal communication channel, the receiving scheduling module immediately updates the current receiving task quantity of the node in the receiving node state monitoring account, increasing the value by one.

[0036] Step S1136: Return the connection information of the target receiving node to the submission terminal of the declarant, the connection information containing the node IP address, port number and protocol version, guiding the declarant to establish a data transmission connection with the target receiving node.

[0037] The receiving scheduling module extracts the IP address, port number for receiving the type of data and supported protocol version from the configuration information of the target receiving node, encapsulates these information into a connection response message, and sends it to the submission terminal of the declarant. The submission terminal starts the corresponding client program according to the returned connection information, establishes a TCP / IP connection with the target receiving node, and prepares for data transmission.

[0038] Step S1137: The receiving scheduling module tracks the data receiving progress of the target receiving node in real time, and if the target receiving node has a receiving delay, the unfinished receiving task is transferred to other idle parallel receiving nodes of the same protocol.

[0039] The receiving scheduling module obtains data receiving progress information, such as the percentage of the amount of received data to the total amount of data, the receiving rate, and the like, through periodic communication with the target receiving node. The receiving progress is compared with a preset progress threshold value, and if the progress does not reach the threshold value within a specified time, it is determined that the receiving is delayed. At this time, the receiving scheduling module selects an idle node from the parallel receiving nodes of the same protocol, sends a task transfer instruction to the original target node, transfers the unfinished receiving task to the new node, updates the receiving node state monitoring account, and notifies the submission terminal of the declarant to switch to connect to the new node.

[0040] Step S1138: Record the process information allocated each time the data submission request, including the data submission request identifier, data type, target receiving node identifier, and allocation time, to form a receiving scheduling log.

[0041] After the receiving scheduling module completes the task allocation, the unique identifier of the current data submission request, the data type description, the unique identifier of the target receiving node, and the specific time of task allocation (accurate to seconds) and the like are written into the receiving scheduling log. The log is stored in a structured format, and each record contains fixed fields, facilitating subsequent log analysis, problem troubleshooting, and data statistics, and ensuring that the receiving scheduling process is traceable.

[0042] Step S114: The parallel receiving node receives the declaration form data uploaded by the declarant, parses the field structure of the declaration form data, extracts the preset association identifier field in the declaration form data, and records the content of the association identifier field and the receiving time of the corresponding declaration form data.

[0043] After the parallel receiving node receives the declaration form data uploaded by the enterprise, the form data parsing program is started. The form data parsing program parses the field structure of the form data according to the specification of the form data special transmission protocol. During the parsing process, the program identifies each field in the form, determines the position of each field in the form, and determines the data format and the like.

[0044] After completing the field structure parsing, the program will locate the preset association identifier field, i.e., the unified social credit code field of the enterprise, and extract the specific content of the field. At the same time, the system will record the receiving time of the declaration form data, and the receiving time is accurate to the millisecond level, so as to trace and manage the data receiving process in the future.

[0045] Step S115: The parallel receiving node receives the proof material data uploaded by the declarant, recombines the fragmented files according to the file fragmentation transmission protocol, and extracts the association identifier from the recombined file metadata, wherein the file metadata includes file naming information and file attribute information.

[0046] For the uploaded data of the proof materials, the parallel receiving node receives each file fragment according to the file fragment transmission protocol. During the receiving process, the parallel receiving node checks the integrity of each fragment to ensure that the received fragment data is not damaged. When all fragments are received, the node starts a file recombination program to splice the fragments into a complete proof material file according to the fragment number and the total number of fragments.

[0047] After the recombination is completed, the program reads the metadata information of the file. The file metadata includes file naming information, such as the enterprise abbreviation contained in the file name, and file attribute information, such as the creation time and the modification time. The unified social credit code of the enterprise is extracted from the above metadata as the basis for associating the proof material data with the corresponding enterprise.

[0048] Step S116: The declaration form data and the proof material data corresponding to the same association identifier are classified into a group of declaration data units, and each declaration data unit corresponds to single declaration data of a declaration subject.

[0049] After obtaining the association identifiers of the declaration form data and the proof material data, the system classifies all received data. By comparing the contents of the association identifiers, the declaration form data and the proof material data with the same unified social credit code are divided into the same group of declaration data units.

[0050] Each declaration data unit contains all the form data and proof material data submitted by the enterprise for this time of value-added tax declaration, forming a complete set of single declaration data of the enterprise.

[0051] Step S117: All declaration data units are summarized to form a declaration data set, a receiving summary identifier is added to the declaration data set, the receiving summary identifier includes the parallel receiving node number, the receiving completion time, and the total number of declaration data units, and the declaration subject name and the submission terminal information corresponding to each declaration data unit are recorded, and the receiving of the declaration data set is completed.

[0052] The system summarizes all classified declaration data units, integrates the above data units together to form the final declaration data set. During the summarizing process, a receiving summary identifier is added to the declaration data set. The parallel receiving node number in the receiving summary identifier records the identification information of each node participating in this data receiving; the receiving completion time records the time point when all data units are received; and the total number of declaration data units reflects the number of data units received this time.

[0053] Meanwhile, the system records the corresponding declaration subject name for each declaration data unit, which is extracted from the table header information of the declaration form data; and the submission terminal information includes the device identifier and network address of the submission terminal, which are obtained from the message header of the data submission request. After these operations are completed, the receiving process of the declaration data set is completed.

[0054] Step S120: The preset RPA declaration process template is called, a business logic mapping matrix is constructed based on the content attributes of various data in the declaration data set, and the various data and the nodes in the RPA declaration process template are bidirectionally adapted through the business logic mapping matrix to obtain a process node bidirectional adaptation result.

[0055] After the receiving of the declaration data set is completed, the data and process node adaptation link is entered. First, the preset enterprise value-added tax declaration RPA process template is called, which is preset according to the standard process of the value-added tax declaration business. Then, the content attributes of various data in the declaration data set are analyzed, such as the declaration items involved by the data, the data format, the data source, and the like.

[0056] Based on these content attributes, a business logic mapping matrix is constructed, which contains the association relationship information between the data attributes and the process nodes. Through the matrix, the various data and the process nodes are bidirectionally adapted to determine which process node each type of data should be matched to for processing, and which data each process node needs to process, and finally the process node bidirectional adaptation result is obtained.

[0057] Step S121: The preset RPA declaration process template is called, and the RPA declaration process template contains a plurality of RPA declaration process nodes connected in series according to the declaration business logic, each RPA declaration process node is marked with a business logic description, and the business logic description includes the declaration business link corresponding to the node, the functional positioning of the data to be processed, and the data flow relationship with the upstream and downstream nodes.

[0058] The preset enterprise value-added tax declaration RPA process template is called from the process template library of the system. The enterprise value-added tax declaration RPA process template connects a plurality of RPA declaration process nodes in series according to the business logic sequence of the value-added tax declaration, including a data verification node, an input tax amount accounting node, a sales tax amount accounting node, a taxable amount calculation node, and a declaration form generation node.

[0059] Each RPA declaration process node has a detailed business logic description. For example, the business logic description of the data verification node is: corresponding to the initial verification link of the declaration business, the data to be processed is the basic field data of all declaration forms, the function positioning is to check the standardization and integrity of the data format, the upstream node is the data receiving node, the downstream nodes are the input tax calculation node and the output tax calculation node, and the data flow relationship is to transmit the verified data to the two downstream nodes respectively.

[0060] Step S122: Content attribute analysis is performed on each type of declaration form data and proof material data in the declaration data set, and a business function tag of each type of data is extracted, the business function tag including a declaration business action supported by the data, a business decision information carried by the data, and an associated upstream and downstream data flow direction.

[0061] The content attribute of each type of data in the declaration data set is analyzed. For the declaration form data, such as the attached data of the value-added tax declaration form, the field content such as taxable sales and output tax included in the data is analyzed to determine that the declaration business action supported by the data is the statistics and declaration of output tax, the business decision information carried by the data is the calculation basis of output tax, and the associated upstream data flow direction is the enterprise sales account data and the downstream data flow direction is the taxable amount calculation node.

[0062] For the proof material data, such as the scanned copy of the special invoice for value-added tax deduction, the key information such as invoice code, invoice number, and deductible tax amount is analyzed to determine that the declaration business action supported by the data is the input tax deduction declaration, the business decision information carried by the data is the legality basis for input tax deduction, and the associated upstream data flow direction is the invoice issuer data and the downstream data flow direction is the input tax calculation node. Through the above analysis, a corresponding business function tag is extracted for each type of data.

[0063] Step S123: Based on the business logic descriptions of all RPA declaration process nodes and the business function tags of each type of data, a business logic mapping matrix is constructed, the row dimension of the business logic mapping matrix being the RPA declaration process node identifier, the column dimension being the data business function tag, and the matrix element being the business logic matching degree of the node and the data.

[0064] The business logic descriptions of all RPA declaration process nodes and the business function tags of each type of data are collected, and a business logic mapping matrix is constructed based thereon. The row dimension of the matrix sequentially arranges the identifiers of each RPA declaration process node, such as the data verification node identifier, the input tax calculation node identifier, and the like; and the column dimension sequentially arranges the business function tags of each type of data, such as the output tax declaration support tag, the input tax deduction basis tag, and the like.

[0065] Each element in the matrix represents the degree of business logic matching between the corresponding RPA declaration process node and the data business function label. The degree of matching is determined by comparing the data function positioning in the business logic description of the node with the supporting business actions, decision information, etc. in the business function label of the data. The higher the degree of coincidence, the higher the degree of matching.

[0066] Step S124: Calculate the degree of business logic matching between each type of data and each RPA declaration process node through the business logic mapping matrix, and determine the RPA declaration process node with the highest degree of business logic matching as the preliminary adaptation node for the data.

[0067] The degree of matching is calculated using the business logic mapping matrix. For each type of data, the matrix element value at the intersection of the column corresponding to its business function label and the row corresponding to the identification of each RPA declaration process node is checked, i.e. the degree of business logic matching between the data and each node is obtained.

[0068] Among all the matching degrees of the nodes, the node with the highest value is selected and determined as the preliminary adaptation node for the data. For example, the sales tax related form data has the highest degree of matching with the sales tax accounting node, so the sales tax accounting node is determined as the preliminary adaptation node for the data.

[0069] Step S125: Analyze the data flow relationship between the preliminary adaptation node and the upstream and downstream nodes, and determine whether the data can be transferred to the downstream node through the preliminary adaptation node and meet the business logic requirements of the downstream node. If it can meet the requirements, the preliminary adaptation node is determined as the target adaptation node; if it cannot meet the requirements, the matching dimension of the business logic mapping matrix is adjusted, the potential business function labels of the data are supplemented, and the degree of matching is calculated again until the target adaptation node is determined or the data is marked as adaptation pending.

[0070] For data with a determined preliminary adaptation node, the data flow relationship between the preliminary adaptation node and the upstream and downstream nodes is further analyzed. For example, the preliminary adaptation node for a certain type of data is the input tax accounting node, and the business logic requirements of the downstream node such as the taxable amount calculation node are checked to determine whether the data can meet the requirements of the data format, data content, etc. of the taxable amount calculation node after being processed by the input tax accounting node.

[0071] If the analysis result is that it can be met, the input tax amount accounting node is determined as the target adaptation node of this type of data. If it cannot be met, such as the data lacks certain derived information required for the tax amount calculation node, the matching dimension of the business logic mapping matrix is adjusted, the potential business function tags of the data are supplemented, such as the derived information generation potential tag, and then the matching degree is calculated again, the preliminary adaptation node is determined again and analyzed, until the target adaptation node that can meet the demand of the downstream node is found. If it cannot be met after multiple adjustments, the data is marked as adaptation pending data.

[0072] Step S126: associate all data determined with the target adaptation node with the corresponding RPA declaration process node, record the business function tags of the adaptation pending data and the unadapted reasons, and form the process node bidirectional adaptation result containing node identification, corresponding data list and pending data record.

[0073] The data of each type determined with the target adaptation node is associated with the corresponding RPA declaration process node to record the corresponding relationship between the data identification and the node identification in the form of an association table. For adaptation pending data, the business function tags such as data type tags, business domain tags, etc. are recorded in detail, as well as the unadapted reasons such as data format incompatible with all node requirements, data lacking key associated information, etc.

[0074] Integrate these associated information, data list and pending data record to form the process node bidirectional adaptation result. The process node bidirectional adaptation result shows the adaptation data corresponding to each RPA declaration process node and the data that has not been successfully adapted.

[0075] Step S130: According to the process node bidirectional adaptation result, combine the historical processing data of each RPA declaration process node to construct a node processing time prediction model, allocate RPA computing resources and data transmission channels based on the node processing time prediction model, and generate a node resource prediction scheduling scheme.

[0076] According to the process node bidirectional adaptation result, the data that needs to be processed by each RPA declaration process node is clear. At the same time, the historical processing data of processing similar data of each node is retrieved, and these data are used to construct a node processing time prediction model. The time required for each node to process the current data is predicted through the model, and then the predicted time and the resource status of the system are used to allocate computing resources and data transmission channels for each node, and finally a node resource prediction scheduling scheme is generated.

[0077] Step S131: Extract the adaptation data amount corresponding to each RPA declaration process node from the process node bidirectional adaptation result, and retrieve the historical processing data of each RPA declaration process node, which includes historical adaptation data amount, corresponding processing time and resource consumption record.

[0078] From the bidirectional adaptation result of the process node, the total amount of adaptation data corresponding to each RPA declaration process node is counted, that is, the adaptation data amount, which can be measured by the storage size or the number of records of the data. At the same time, the historical processing data of each RPA declaration process node is retrieved from the historical database of the system.

[0079] In the historical processing data, the historical adaptation data amount records the data amount of the past processing of the node; the corresponding processing time length records the time spent in processing the corresponding historical adaptation data amount; and the resource consumption record includes the processor resources, memory resources, network bandwidth and other information occupied in the processing process.

[0080] Step S132: Feature extraction is performed on the historical processing data, and the extracted features are input into a regression model as training samples to construct a node processing time length prediction model. The extracted features include historical adaptation data amount, data field quantity, data correlation dimension and corresponding processing time length.

[0081] Feature extraction is performed on the retrieved historical processing data. First, the historical adaptation data amount is extracted, which reflects the size of the historical processing data. Then, the total number of independent fields contained in each batch of historical processing data is counted as a data field quantity feature, which reflects the complexity of the data structure.

[0082] Next, the number of association links between each batch of historical processing data and other RPA declaration process node data is counted as a data correlation dimension feature, which reflects the association tightness between the data. The above extracted features and the corresponding historical processing time length are used to form training samples, which are input into a regression model to construct a node processing time length prediction model.

[0083] Step S1321: The historical processing data of each RPA declaration process node is cleaned, and the historical adaptation data amount is extracted from the cleaned historical processing data, with the number of data bytes as the quantitative indicator of the historical adaptation data amount; the data field quantity is extracted, and the total number of independent fields contained in each batch of historical processing data is counted; and the data correlation dimension is extracted, and the number of association links between each batch of historical processing data and other RPA declaration process node data is counted.

[0084] The historical processing data of each RPA declaration process node is cleaned to remove duplicate data, abnormal data and data with too many missing values. After cleaning, the historical adaptation data amount is extracted from the data, and the number of data bytes is used to quantitatively represent it, that is, the total number of storage bytes of each batch of historical data is counted.

[0085] When the number of data fields is extracted, the structure of each batch of historical processing data is analyzed to count the total number of independent fields contained therein, and each independent field name corresponds to a count unit. When the data association dimension is extracted, the number of association links is counted by analyzing the reference relationship, transmission relationship, etc. between each batch of historical processing data and other node data, and each independent association path is counted as an association link.

[0086] Step S1322: The historical adaptation data volume, data field quantity, and data association dimension are taken as model input features, and the corresponding historical processing time length is taken as the model output label to construct a training sample set.

[0087] The historical adaptation data volume, data field quantity, and data association dimension extracted after cleaning are determined as the input features of the model. The historical processing time length corresponding to each batch of historical processing data is taken as the output label of the model, i.e., the target value predicted by the model.

[0088] The input features and output labels are one-to-one corresponding to form a plurality of training samples. These training samples are stored in a fixed format, and each sample contains a corresponding input feature combination and output label value. For example, the historical adaptation data volume of a batch of historical processing data is a specific number of bytes, the data field quantity is a certain number, the data association dimension is a certain number, and the corresponding historical processing time length is a certain time period, which constitutes a complete training sample. All the above samples are collected together to form a training sample set.

[0089] Step S1323: The training sample set is divided into a training subset and a validation subset, wherein the training subset is used for model training, and the validation subset is used for model performance verification.

[0090] The training sample set constructed is divided. The division process is performed according to a predetermined proportion, for example, the sample set is divided into a training subset and a validation subset according to a certain proportion. When dividing, samples are selected by random sampling to ensure that the training subset and the validation subset can well represent the feature distribution of the whole sample set.

[0091] The training subset contains most of the sample data and is mainly used for parameter learning and iterative training of the model; the validation subset contains a small amount of sample data and is used to evaluate the performance of the model during the training process to timely find problems such as overfitting or underfitting of the model.

[0092] Step S1324: A gradient boosting regression model is selected as a base model, and the training subset is input into the base model for iterative training. In each iteration process, the prediction error of the validation subset is used to adjust the hyperparameters of the base model, including the learning rate, the number of decision trees, and the tree depth.

[0093] A gradient boosting regression model is selected as a base model for constructing the node processing duration prediction model. The gradient boosting regression model improves the prediction performance by constructing multiple decision trees and performing ensemble learning. The training subset is input into the base model to start the iterative training process of the model.

[0094] In each iteration training, the model constructs decision trees and updates parameters according to the input training data. At the same time, the validation subset is input into the model obtained by the current training, and the prediction error between the prediction result of the model on the validation subset and the actual output label is calculated. According to the prediction error, the hyperparameters of the model are adjusted, such as reducing the learning rate to slow down the model training speed, increasing the number of decision trees to increase the complexity of the model, or adjusting the tree depth to control the growth scale of the decision tree, until the performance of the model reaches an optimal state.

[0095] Step S1325: When the prediction error of the validation subset is less than the preset error threshold, stop the model training, and obtain the final node processing duration prediction model.

[0096] During the iterative training process of the model, the change of the prediction error of the validation subset is continuously monitored. A preset error threshold is set, which is an acceptable error range determined according to historical model training experience and actual business requirements. When the prediction error of the validation subset after a certain iteration is less than the preset error threshold, it indicates that the current performance of the model has met the requirements.

[0097] At this time, the iterative training process of the model is stopped, and the current model parameters are saved to obtain the final node processing duration prediction model. The node processing duration prediction model can accurately predict the processing duration of the RPA declaration process node based on the input feature data.

[0098] Step S1326: Generalization test is performed on the final node processing duration prediction model, historical processing data not participating in the training is selected as a test sample, the test sample is input into the node processing duration prediction model to obtain a predicted processing duration, a deviation rate of the predicted processing duration and the actual processing duration is calculated, if the deviation rate is lower than a preset deviation threshold, it is confirmed that the node processing duration prediction model is available, if the deviation rate is higher than the preset deviation threshold, the training sample or the model structure is adjusted until the generalization of the node processing duration prediction model meets the requirements.

[0099] To ensure that the node processing duration prediction model has good generalization ability and can be applied to different data scenarios, generalization test is performed on the final obtained model. Part of the data not participating in the model training and validation is selected from the historical processing data as a test sample, which also contains input features such as historical adaptation data volume, data field number, data association dimension, and corresponding actual processing duration.

[0100] The input features of the test sample are input into the node processing time prediction model to obtain the predicted processing time output by the model. The deviation rate between the predicted processing time and the actual processing time is calculated, and the deviation rate is calculated by the ratio of the difference between the two and the actual processing time. If the deviation rate is lower than the preset deviation threshold, it means that the model can also maintain good prediction performance on unseen data, and the node processing time prediction model is confirmed to be available. If the deviation rate is higher than the preset deviation threshold, the selection range of the training sample needs to be adjusted or the structure of the model needs to be modified, such as increasing the number of training samples, adjusting the hyperparameters, and then retraining and testing the model until the generalization of the model meets the requirements.

[0101] Step S133: input the current adaptive data volume, data field number, and data association dimension of each RPA declaration process node into the node processing time prediction model to obtain the predicted processing time of each RPA declaration process node.

[0102] From the process node bidirectional adaptation result, the adaptive data volume currently required by each RPA declaration process node is extracted, the number of independent data fields contained in the adaptive data is counted, and the number of association links between the data and other node data, i.e. the data association dimension, is analyzed and determined.

[0103] The three feature data are arranged according to the input format required by the node processing time prediction model, and then input into the model respectively. The model processes and calculates the input feature data, and outputs the predicted processing time corresponding to each RPA declaration process node. The predicted processing time reflects the estimated time required for the node to process the current adaptive data.

[0104] Step S134: call the available computing resource account of the RPA system, which contains the operation rate of idle processors, the read / write speed of memory, and the capacity of storage resources. At the same time, the real-time transmission rate of available data transmission channels and channel load records are called.

[0105] Through the resource management module of the RPA system, the available computing resource account is called. The available computing resource account updates the usage status and performance parameters of various computing resources in the system in real time, wherein the operation rate of idle processors reflects the number of instructions that can be processed by the processor per unit time; the read / write speed of memory indicates the speed of memory data reading and writing operation; the capacity of storage resources shows the size of the space currently available for storing data.

[0106] At the same time, the relevant information of the available data transmission channel is called from the network management module of the system, including the real-time transmission rate of each channel, i.e. the amount of data that can be transmitted per unit time, and the channel load record, which reflects the current busy degree of the channel, such as the proportion of used bandwidth to total bandwidth, etc.

[0107] Step S135: According to the predicted processing time length of each RPA declaration process node and the available computing resource account, allocate corresponding computing resource specifications and quantities for each RPA declaration process node.

[0108] Analyze the predicted processing time length of each RPA declaration process node. Nodes with longer predicted processing time length usually need to process larger data size or more complex data structure, and have relatively higher demand for computing resources. Combine the idle processor operation speed, memory read-write speed and storage resource capacity recorded in the available computing resource account to allocate appropriate computing resources for each node.

[0109] For example, for the input tax accounting node with longer predicted processing time length and more data fields, allocate processors with higher operation speed, memories with faster read-write speed, and storage resources with sufficient capacity; for the data verification node with shorter predicted processing time length, allocate relatively basic computing resource specifications. At the same time, according to the number of nodes and resource demand, reasonably determine the allocation quantity of each specification of computing resources to ensure the rationality and efficiency of resource allocation.

[0110] Step S136: According to the adaptive data size and predicted processing time length of each RPA declaration process node, calculate the minimum data transmission rate required by each RPA declaration process node, and match each RPA declaration process node with a data transmission channel whose real-time transmission rate is not lower than the required minimum data transmission rate and whose current transmission data volume does not affect the node data transmission progress.

[0111] For each RPA declaration process node, according to its adaptive data size and predicted processing time length, calculate the minimum data transmission rate required. The calculation method is to divide the adaptive data size by the predicted processing time length to obtain the minimum rate requirement for completing data transmission within a specified time.

[0112] Then, from the available data transmission channels, select channels whose real-time transmission rate is not lower than the minimum data transmission rate. On this basis, further check the current transmission data volume of these channels, select channels whose transmission data volume does not reach the saturation state and does not affect the node data transmission progress, match each node with the most suitable data transmission channel to ensure smooth data transmission between nodes.

[0113] Step S137: Record the computing resource specifications, quantities and data transmission channel identifiers allocated for each RPA declaration process node, mark the start time and estimated release time of computing resource allocation, and form a node resource prediction scheduling scheme containing node identifier, resource configuration details, transmission channel parameters and predicted processing time length.

[0114] The computing resource specifications assigned to each RPA declaration process node, such as processor model, memory capacity, resource quantity, and corresponding transmission channel identification information, are recorded in detail. At the same time, according to the predicted start processing time of the node, the start time of the computing resource allocation, that is, the time point at which the resource starts to provide services for the node, is marked; according to the predicted processing duration of the node, the predicted release time of the resource, that is, the time point at which the resource can be recycled and reused after the node processing is completed, is calculated and marked.

[0115] Integrating these information, the node resource prediction scheduling scheme is formed. The node resource prediction scheduling scheme contains the identification of each node, the specific resource configuration details, the parameters of the transmission channel such as transmission rate, and the predicted processing duration of the node.

[0116] Step S140: For the RPA declaration process nodes with data missing or format deviation in the process node bidirectional adaptation result, the RPA tool is called based on the node resource prediction scheduling scheme to perform cross-node data association derivation, supplement missing data or correct format deviation data, and obtain the associated derived node data.

[0117] In the process node bidirectional adaptation result, there are some RPA declaration process nodes whose adaptation data has data missing or format deviation problem. For these nodes, the computing resources and transmission channels allocated in the node resource prediction scheduling scheme are called to start the cross-node data association derivation process.

[0118] The RPA tool retrieves the relevant data of the associated node, analyzes the association between the data, supplements the missing data, and corrects the format deviation data, and finally obtains the processed associated derived node data, ensuring the integrity and standardization of the node data.

[0119] Step S141: From the process node bidirectional adaptation result, the RPA declaration process nodes with data missing or format deviation are screened out, the adaptation data corresponding to the RPA declaration process nodes is extracted, and the field name of the missing data or the specific type of the format deviation is determined.

[0120] The process node bidirectional adaptation result is checked one by one, and the RPA declaration process nodes with data problems in it are screened out. The adaptation data of these nodes either has the case that some fields are not filled, that is, data missing, or the data format does not meet the node processing requirements, that is, format deviation.

[0121] Extract the corresponding adaptation data of these nodes, and analyze the data in detail. For the case of data missing, the specific field name of the missing data is determined, such as the missing "purchase invoice authentication date" field in the input tax accounting node; for the case of format deviation, the specific type of deviation is determined, such as the date field should be in the format of "year-month-day" but is written as "month / day / year", or the numerical value field has extra text symbols, etc.

[0122] Step S142: According to the node resource prediction scheduling scheme, the RPA computing resources allocated for the RPA declaration process node are called, the cross-node data retrieval module of the RPA tool is started, and the cross-node data retrieval module is used to query the adaptation data of the upstream and downstream nodes associated with the current RPA declaration process node through data flow.

[0123] According to the node resource prediction scheduling scheme, the RPA computing resource information allocated for the RPA declaration process node with data problems is obtained, such as the calling path and permission of resources such as processor, memory, etc. These computing resources are called through the resource calling interface.

[0124] The module specially used for cross-node data retrieval in the RPA tool is started, which can query the upstream and downstream nodes associated with the current node directly or indirectly through data flow in the RPA declaration process template according to the identification information of the current node, and obtain the adaptation data of these associated nodes.

[0125] Step S143: If there is data missing, the RPA tool obtains the adaptation data of the upstream node of the current RPA declaration process node through the cross-node data retrieval module, analyzes the association relationship between the adaptation data of the upstream node and the missing field of the current RPA declaration process node, and if the adaptation data of the upstream node contains associated data that can derive the missing field of the current RPA declaration process node, performs data derivation calculation based on the association relationship to obtain the supplementary data of the missing field.

[0126] When it is detected that the current RPA declaration process node has data missing, the cross-node data retrieval module of the RPA tool focuses on querying the upstream node adaptation data of the node. The upstream node processes before the current node in the business process, and its data often has certain business association with the current node data.

[0127] After obtaining the adaptation data of the upstream node, these data are analyzed to find information associated with the missing field of the current node. If it is found that there is associated data in the adaptation data of the upstream node that can derive the missing field, for example, the "invoice issuing date" field of the upstream node has a time sequence association with the missing "purchase invoice authentication date" field of the current node, then based on the above association relationship, data derivation calculation is performed to obtain the supplementary data of the missing field.

[0128] Step S1431: The cross-node data retrieval module of the RPA tool queries the upstream node identifier of the RPA declaration process template that has a direct data flow input relationship with the current RPA declaration process node according to the current RPA declaration process node identifier, and retrieves the adaptation data of the upstream node based on the upstream node identifier.

[0129] After the cross-node data retrieval module of the RPA tool receives the data retrieval request, it queries the node relationship graph of the RPA declaration process template according to the unique identifier of the current RPA declaration process node. The node relationship graph records the data flow relationship between all nodes, and finds the identifier information of the upstream node that has a direct data flow input relationship with the current node through the query.

[0130] According to the upstream node identifier, the adaptation data corresponding to the upstream node is retrieved from the data storage module of the system. These data are stored in a structured form, containing multiple fields and corresponding values.

[0131] Step S1432: The adaptation data of the upstream node is structurally parsed, the key fields in the adaptation data of the upstream node are extracted, and the key fields are compared with the missing fields of the current RPA declaration process node in terms of business logic to determine whether they have a causal relationship or a subordinate relationship.

[0132] The retrieved upstream node adaptation data is structurally parsed to identify each field in the data and determine the name, data type, and meaning of each field. The key fields that may be related to the missing fields of the current node are extracted, and these key fields are usually related to the missing fields in terms of business logic.

[0133] The extracted key fields are compared with the missing fields of the current RPA declaration process node in terms of business logic to analyze their relationship in the business process. If the change in the value of the key field directly leads to the change in the value of the missing field, it is determined that they have a causal relationship; if the information of the missing field is derived based on the information of the key field, or belongs to the business category represented by the key field, it is determined that they have a subordinate relationship.

[0134] Step S1433: If there is a causal relationship, the RPA tool retrieves the corresponding causal relationship rule in the declaration business logic library and derives the supplementary data of the missing field of the current RPA declaration process node from the corresponding field in the adaptation data of the upstream node according to the causal relationship rule.

[0135] When it is determined that the key field of the upstream node has a causal relationship with the missing field of the current node, the RPA tool accesses the declaration business logic library. The declaration business logic library stores causal relationship rules between different fields in various declaration business scenarios, such as the rule "invoice issuance date plus authentication period equals input invoice authentication date".

[0136] According to the determined causal correlation type, the corresponding causal correlation rule is called, and then according to the calculation method or logical relationship specified in the rule, the supplementary data of the missing field of the current RPA declaration process node is derived based on the value of the key field in the upstream node adaptation data. For example, according to the above rule, the supplementary data of the "purchase invoice authentication date" is calculated using the "invoice issuance date" of the upstream node and the preset "authentication period".

[0137] Step S1434: If there is a dependent association, the RPA tool extracts the association identification information from the corresponding field in the adaptation data of the upstream node, retrieves the corresponding business database through the association identification information, and obtains the data corresponding to the missing field of the current RPA declaration process node as the supplementary data.

[0138] When it is determined that there is a dependent association, the RPA tool extracts the association identification information from the key field of the upstream node adaptation data, which can uniquely point to the data record related to the missing field in the business database. For example, the "contract number" field in the upstream node data can be used as the association identification information to retrieve the detailed data related to the contract.

[0139] By accessing the corresponding business database through the association identification information, the data retrieval operation is performed, and the data record containing the missing field information of the current node is found in the database, from which the data corresponding to the missing field is extracted as the supplementary data. For example, the "contract signing date" in the contract database is retrieved through the "contract number" as the supplementary data of the missing "business occurrence date" of the current node.

[0140] Step S1435: During the derivation process, the RPA tool records the upstream node identification, key field content and business logic rules relied on for derivation to form a derivation process record.

[0141] During the entire data derivation process, the RPA tool starts the log recording function. The identification information of the upstream node relied on during the derivation process is recorded in detail to trace the data source; the specific content of the key field used is recorded, including the field name and field value; and the number or specific description of the business logic rules applied is recorded.

[0142] These records are sorted in chronological order and derivation steps to form a complete derivation process record, which can be used for subsequent data auditing and problem troubleshooting to ensure the traceability and legality of the supplementary data.

[0143] Step S1436: Fill in the derived supplementary data into the missing fields of the current RPA declaration process node, and perform business logic coherence confirmation on the supplementary data and the existing data of the current RPA declaration process node to ensure that the supplementary data and the existing data have no logical conflicts in the declaration business scenario, and form an associated derived node data segment containing complete fields.

[0144] The derived supplementary data is filled into the missing fields of the current RPA declaration process node according to the field correspondence. After the filling is completed, the business logic coherence of all data of the node is checked. It is checked whether the logical relationship between the supplementary data and the existing data in the business scenario is reasonable, for example, whether the order of the date fields conforms to the business rules, whether the calculation relationship between the numerical value fields is correct, and the like.

[0145] If there is no logical conflict, it means that the supplementary data is valid, and an associated derived node data segment containing complete fields is formed; if there is a logical conflict, the derivation process needs to be rechecked or other associated data needs to be found for further derivation until the data logic is coherent.

[0146] The RPA tool retrieves the adaptive data format of other RPA declaration process nodes consistent with the business logic of the current RPA declaration process node, and converts the data of the current RPA declaration process node with format deviation into the reference standard format.

[0147] When it is detected that the current RPA declaration process node has format deviation, the RPA tool starts the format correction mechanism. First, according to the business logic description of the current node, the business link and data processing type to which it belongs are determined. Then, other nodes with the same or highly similar business logic as the current node are retrieved in the RPA declaration process template. These nodes have formed a standardized adaptive data format in the past processing.

[0148] The adaptive data format of these reference nodes is extracted as a standard template, which contains the naming rules of the fields, the data types (such as text, numerical value, date, etc.), the format specifications (such as the "year-month-day" format of the date, the number of decimal places of the numerical value, etc.). The deviated data of the current node is compared with the standard template to identify the specific position and type of the deviation, such as date separator error, numerical value unit redundancy, field name spelling inconsistency, etc.

[0149] According to the format requirements of the standard template, the format conversion operation is performed on the deviated data of the current node. For example, the date format "month / day / year" is converted to "year-month-day", the non-numerical value symbol after the numerical value field is removed, the spelling error of the field name is corrected, and the like. During the conversion process, the authenticity and accuracy of the data content are ensured not to be affected by the format adjustment.

[0150] Step S146: After completing data supplement or format conversion, the RPA tool fuses the processed data with the original adaptive data of the current RPA declaration process node to form associated deduced node data, records the node range, data association relationship and deduction process of cross-node data retrieval, and associates to the current RPA declaration process node identifier.

[0151] After completing data missing supplement or format deviation conversion, the RPA tool performs fusion operation on the processed data and the original adaptive data of the current node. For the case of data missing supplement, the supplemented data deduced is integrated with the complete field data in the original adaptive data to ensure that all fields have valid data; for the case of format deviation conversion, the converted standard format data replaces the original deviation format data, and the content information of the original data is retained.

[0152] The fusion forms associated deduced node data, which contains all necessary fields of the current node and the format meets the business requirements. At the same time, the RPA tool records in detail the node range involved in the cross-node data retrieval process, including the identifier of the upstream node, downstream node or reference node; records the association relationship between data, such as the number of causal association rules, the identifier information of dependent association, etc.; records the key steps and operation details in the deduction process, such as data extraction method, format conversion rule, etc.

[0153] The above record information is associated with the identifier of the current RPA declaration process node for storage, forming a node data processing file, which facilitates subsequent tracing and auditing of data sources and processing process.

[0154] Step S150: According to the business logic sequence of the RPA declaration process template, integrate the node data without data problems in the bidirectional adaptive result and the associated deduced node data to generate a declaration data submission package, push the declaration data submission package to the target declaration system, and obtain declaration submission feedback data.

[0155] After completing the processing of all node data, according to the business logic sequence of each node in the RPA declaration process template, the node data without data problems and the node data processed by association and deduction are integrated. Through integration, a complete declaration data link is formed, which is encapsulated into a declaration data submission package according to the requirements of the target declaration system, and is pushed to the target declaration system. The system returns the declaration submission feedback data, and the entire declaration data processing process is completed.

[0156] Step S151: Extract node data without data problems from the bidirectional adaptive result of the process node, which is not marked with data missing or format deviation and has completed bidirectional adaptation with the corresponding RPA declaration process node.

[0157] The RPA tool filters the results of the two-way adaptation of the process nodes, identifies node data that is not marked as data missing or format deviation, and confirms that these data meet the business logic requirements and data specifications of the corresponding RPA declaration process nodes in the adaptation process. Extract these node data without data problems, check if the adaptation association relationship is valid, and ensure the correspondence and integrity of the data and the nodes.

[0158] The extracted node data without data problems is classified and stored according to the node identifier, preparing the data for the subsequent integration step and ensuring that these data can directly participate in the construction of the declaration data link.

[0159] Step S152: Collect all associated derived node data corresponding to RPA declaration process nodes with data problems, ensure that each RPA declaration process node with data problems has generated associated derived node data, and if any RPA declaration process node cannot solve the data problem through cross-node data association derivation, mark the node data of the RPA declaration process node as temporary processing data and exclude it from the integration range.

[0160] The RPA tool collects the associated derived node data corresponding to the nodes marked as having data problems in the two-way adaptation results of the process nodes. Check each node with data problems to confirm whether the associated derived node data has been successfully generated and verified to have no logical conflicts or format problems.

[0161] If a node with data problems is found to be unable to solve the problem through cross-node data association derivation, such as missing key data that cannot be derived through upstream and downstream nodes, or format deviation that cannot be corrected by reference nodes, the data of the node is marked as temporary processing data. Temporary processing data is not included in this integration range, and its node identifier and unresolved problem type are recorded separately for subsequent manual processing or additional data processing.

[0162] Step S153: Refer to the business logic order of each RPA declaration process node in the RPA declaration process template, and sort the node data without data problems and the associated derived node data, so that the sorted node data order is consistent with the flow order of the declaration business.

[0163] Retrieve the business logic order information of each RPA declaration process node in the RPA declaration process template. This business logic order reflects the natural flow process of the declaration business from start to finish, such as data verification node→input tax calculation node→output tax calculation node→tax calculation node→declaration form generation node, etc.

[0164] According to the sequence, the collected node data without data problems and the associated derived node data are sorted. The data of each node is arranged according to its position in the business process, ensuring that the sorted node data sequence is consistent with the actual sequence of the declared business.

[0165] Step S154: Perform associated field concatenation processing on the sorted node data, extract the association identification field in each node data, and concatenate the node data of different nodes into a complete declaration data link through the association identification field.

[0166] In the sorted node data, the association identification field in each node data is extracted, which is usually the unified social credit code of the declaration subject and is consistent and unique in all node data. The association identification field establishes the connection between different node data, ensuring that each node data belonging to the same declaration subject can be accurately associated.

[0167] According to the business logic sequence and the association identification field, the data of each node is concatenated to form a complete declaration data link. For example, the basic information checked in the data verification node is associated to the accounting data of the input tax amount accounting node through the unified social credit code, and then to the accounting data of the output tax amount accounting node, and so on, until all node data are concatenated to form a complete data chain covering the entire declaration process.

[0168] Step S155: According to the declaration data organization format required by the target declaration system, the concatenated declaration data link is packaged into a declaration data submission package, which includes package identification information, data link subject and data flow record, the package identification information includes declaration subject identification and submission batch code, and the data flow record includes the adaptation time and association derivation time of each node data.

[0169] Get the declaration data organization format specification published by the target declaration system, which specifies the structure composition, field naming rules, data format requirements, packaging methods, etc. of the declaration data submission package. According to the specification requirements, the concatenated declaration data link is packaged.

[0170] First, the package identification information is generated, in which the declaration subject identification is the unified social credit code in the declaration data, and the submission batch code is generated by the system according to the declaration date, declaration time period and serial number, ensuring the uniqueness of each submission package. Secondly, the concatenated declaration data link is taken as the data link subject, and adjusted according to the specification field naming and format requirements. Then, the adaptation time of each node data in the adaptation process and the association derivation time of the node with data problems are extracted to form a data flow record, which records the processing time of the data in each node.

[0171] The package identification information, data link body and data flow transfer record are combined according to the structure and order required by the specification to form the initial content of the declaration data submission package.

[0172] Step S1551: Obtain the declaration data packaging specification published by the target declaration system, and parse the structure requirements, field naming rules and data format standards of the declaration data submission package in the declaration data packaging specification.

[0173] The RPA tool accesses the specification publishing platform of the target declaration system through the interface, downloads the latest declaration data packaging specification file. The specification file is parsed to determine the overall structure requirements of the declaration data submission package, such as the mandatory first-level modules (package identification, data body, flow transfer record, etc.) and the hierarchical relationship of each module; extract the field naming rules, including the naming format of the field name, the prohibited characters, the fixed name of the key field, etc.; determine the data format standards, such as the specific format requirements of various data such as date, numerical value, text, data compression method, encryption standard, etc.

[0174] The parsed specification content is stored as a structured specification parameter, which is used to guide the subsequent declaration data submission package packaging process to ensure that the packaged submission package meets the requirements of the target system.

[0175] Step S1552: Generate the package identification information of the declaration data submission package, wherein the declaration subject identification is the subject code corresponding to the association identification in the declaration data set, and the submission batch code is generated by the RPA system in combination with the current date and the submission sequence of the day. The submission batch code includes the date field, the time period field and the sequence field.

[0176] The declaration subject identification in the package identification information directly uses the subject code corresponding to the association identification in the declaration data set, i.e. the unified social credit code of the enterprise, to ensure that the identification corresponds to the uniqueness of the declaration subject. The submission batch code is automatically generated by the RPA system, and the generation rule is to combine the current date, the submission time period of the day and the submission sequence.

[0177] The date field uses the format of "year-month-day" to represent the date of generating the submission package; the time period field divides a day into several time periods and represents them with specific codes, such as morning, afternoon, evening, etc.; the sequence field represents the generation order of the submission package in the current date and time period, represented by consecutive numbers. For example, the submission batch code of a certain submission package can be represented as "20250821-AM-001", wherein "20250821" is the date field, "AM" is the time period field, and "001" is the sequence field.

[0178] Step S1553: Take the concatenated declaration data link as the data link body, and uniformly adjust the field names of each node data according to the field naming rules of the declaration data packaging specification.

[0179] Check the field name of each node data in the data link after concatenation. Identify the field names that do not conform to the naming rules in the declaration data packaging specification, such as field names containing special characters, spelling errors, excessively long names, etc.

[0180] Adjust the field names according to the specification requirements, such as removing special characters, correcting spelling errors, simplifying excessively long names, and unifying the case of field names. During the adjustment process, ensure that the adjustment of the field name does not affect the actual meaning of the field data, and that each field name is unique in the data link body to avoid field name conflicts.

[0181] Step S1554: Convert the data format of each field in the data link body according to the data format standard of the declaration data packaging specification.

[0182] According to the data format standard specified in the declaration data packaging specification, convert the data format of each field in the data link body. For date fields, convert to the "year-month-day" or "year-month-day hour: minute: second" format as required by the specification; for numerical fields, retain the specified number of decimal places, remove unnecessary unit symbols, and convert to a pure numerical format; for text fields, remove leading and trailing spaces and unify the character encoding format.

[0183] During the conversion process, check the accuracy of the data format conversion, such as date reasonableness check and numerical precision check, to ensure that the converted field data not only conforms to the format standard but also accurately reflects the business content.

[0184] Step S1555: Extract the adaptation time of each node data in the process node bidirectional adaptation process, and the associated derivation time of the RPA declaration process node with data problems, associate the adaptation time and the associated derivation time with the corresponding node identifier, and form a data flow record.

[0185] Extract the adaptation time of each node data from the process node bidirectional adaptation result, which records the specific time when the node data and the corresponding RPA declaration process node complete the adaptation; extract the associated derivation time of the node with data problems from the processing record of the associated derived node data, which records the specific time when the node data completes the missing supplement or format conversion.

[0186] The adaptation time and the correlation derivation time are associated with the corresponding node identifiers respectively to form basic entries of the data flow transfer record. Each entry contains a node identifier, an adaptation time (if the node is a node without data problems) or a correlation derivation time (if the node is a node with data problems). According to the arrangement of the nodes in the business logic sequence, the above entries are integrated into a complete data flow transfer record to clearly show the processing timeline of the data of each node.

[0187] Step S1556: The package identification information, the data link main body and the data flow transfer record are combined in the order required by the declaration data packaging specification to construct the initial structure of the declaration data submission package.

[0188] According to the structural order required by the declaration data packaging specification, the package identification information, the data link main body and the data flow transfer record are combined. Generally, they are arranged in the order of the package identification information first, the data link main body in the middle and the data flow transfer record last, and each part is distinguished by the separators or markers required by the specification.

[0189] In the combination process, it is ensured that the content of each part is complete and nothing is missed, the structure level is clear, and the requirements of the specification on the overall structure of the submission package are met.

[0190] Step S1557: The declaration data submission package with the initial structure is compressed, and a compression algorithm supported by the target declaration system is used to reduce the volume of the declaration data submission package. The file sizes before and after compression and the compression efficiency are recorded.

[0191] A compression algorithm supported by the target declaration system, such as ZIP, GZIP, etc., is used to compress the declaration data submission package with the initial structure. During the compression process, a reasonable compression level is set to balance the compression efficiency and compression time consumption, and it is ensured that the volume of the compressed submission package is significantly reduced to facilitate data transmission and storage.

[0192] After compression is completed, the original file size before compression and the file size after compression of the declaration data submission package are recorded, the compression efficiency is obtained by calculating the ratio of the two, and the above information is stored as a compression record for evaluating the compression effect and data transmission efficiency.

[0193] Step S1558: A data verification identifier is added to the compressed declaration data submission package, which is generated by performing a hash operation on the overall data of the declaration data submission package, and is used to confirm that the declaration data submission package has not been tampered with during transmission, and the packaging of the declaration data submission package is completed.

[0194] A hash operation is performed on the overall data of the compressed declaration data submission package, using a hash algorithm supported by the target declaration system, such as SHA-256, etc. The hash operation converts all data of the submission package into a fixed-length hash value, which is the data verification identifier. The data verification identifier is unique, and if any minor tampering occurs to the submission package data during transmission, the newly calculated hash value will not be consistent with the original verification identifier.

[0195] The generated data verification identifier is added to a specified position of the compressed declaration data submission package, such as the package header or the package tail, to complete the final packaging of the declaration data submission package. The packaged submission package contains both compressed data and a data verification identifier, ensuring that the target declaration system can verify the data integrity.

[0196] Step S156: The interface communication parameters of the target declaration system are retrieved, including the interface access address, data transmission encryption protocol, and identity authentication token. The RPA tool establishes an encrypted communication link with the target declaration system based on the interface communication parameters.

[0197] The RPA tool retrieves the interface communication parameters of the target declaration system from the system's configuration database. The interface access address is the network address of the target system for receiving declaration data submission packages, usually in the form of URL; the data transmission encryption protocol is an encryption standard to ensure data transmission security, such as SSL / TLS, etc.; the identity authentication token is a credential for identity verification between the RPA system and the target declaration system, including authentication information and validity period.

[0198] According to the interface access address, the RPA tool initiates a connection request to the target declaration system; based on the data transmission encryption protocol, the encryption algorithm and key are negotiated and determined to encrypt the transmission channel; the identity authentication token is submitted to complete the identity verification with the target declaration system. After verification, a secure encrypted communication link is established.

[0199] Step S157: The declaration data submission package is transmitted to the target declaration system through the encrypted communication link, and the data transmission progress is tracked in real time during transmission. If the transmission is interrupted, the standby data transmission channel is re-enabled based on the node resource prediction scheduling scheme to resume transmission.

[0200] The RPA tool transmits the packaged declaration data submission package to the interface access address of the target declaration system through the established encrypted communication link. During transmission, the transmission progress tracking mechanism is started to monitor the ratio of the amount of data transmitted to the total amount of data, and to calculate the transmission rate and the estimated remaining transmission time.

[0201] If an interruption occurs during transmission, such as network failure, connection timeout, etc., the RPA tool immediately detects the cause of the interruption and queries the backup data transmission channel information allocated for the current transmission task in the node resource prediction scheduling scheme. Based on the parameters of the backup channel, the encrypted communication link is re-established, and the transmission of the uncompleted data is continued from the breakpoint to ensure that the declaration data submission package can be transmitted completely to the target declaration system.

[0202] Step S158: Receive the transmission response information returned by the target declaration system, which contains the data reception status and the integrity identifier of the received data.

[0203] After the transmission of the declaration data submission package is completed, the RPA tool maintains the communication connection with the target declaration system and waits to receive the transmission response information returned by the system. The transmission response information is the feedback of the target declaration system on the data reception, where the data reception status indicates whether the data is successfully received, such as "success" or "failure"; the integrity identifier of the received data is the hash value obtained by the target declaration system after performing a hash operation on the received data, which is used to compare with the data verification identifier in the declaration data submission package to verify whether the data is complete and has not been tampered with.

[0204] The RPA tool stores the received transmission response information into the local log as a preliminary record of the data transmission result.

[0205] Step S159: If the data reception status is successful, generate the declaration submission feedback data containing the package identifier information, the reception success confirmation, and the subsequent processing instructions; if the data reception status is failure, record the failure type and the corresponding transmission parameters, and associate them to the package identifier information of the declaration data submission package to form the declaration submission feedback data containing the failure details.

[0206] When the data reception status in the transmission response information is successful, the RPA tool extracts the package identifier information of the declaration data submission package, combines the reception success confirmation information returned by the target declaration system, such as the prompt code and message of successful reception, and the subsequent processing instructions provided by the system, such as the declaration data audit period, the result query method, etc., to generate the declaration submission feedback data.

[0207] When the data reception status is failure, the RPA tool identifies the failure type from the transmission response information, such as data verification failure, format error, identity authentication failure, server busy, etc.; records the relevant parameters in the current transmission process, such as transmission time, channel identifier used, encryption protocol version, etc. The failure type, transmission parameters, and package identifier information of the declaration data submission package are associated to form the declaration submission feedback data containing the failure details, which clearly indicates the failure cause and related background.

[0208] After the submission feedback data is generated, it can be stored in the feedback data management module of the system, and at the same time, it is pushed to the terminal interface of the declaration subject or the associated notification channel through the preset notification mechanism. The package identification information contained in the feedback data can be used to query the processing status of the submission data package in the system subsequently, so as to ensure that the declaration subject can understand the submission situation of the declaration data in time.

[0209] For the generated submission feedback data containing failure details, the RPA tool performs classification analysis on the failure type. If the failure type is data verification failure, the specific verification failure field can be further located, and combined with the verification rule description returned by the target declaration system, a detailed failure reason description is formed; if it is format error, the format position that does not meet the requirements and the correct format standard can be clearly pointed out; if it is identity authentication failure, the declaration subject can be prompted to check the validity of the identity authentication token or reacquire the token; if it is server busy, it can be suggested to resubmit in the off-peak period.

[0210] The failed submission data package is marked as a to-be-processed state and stored in a special failed data buffer area. At the same time, the corresponding retry mechanism or manual intervention prompt is automatically triggered according to the failure reason. For the failure caused by temporary reasons such as server busy, the RPA tool will schedule according to the node resource prediction scheme, and after a preset time interval, the idle data transmission channel is called again, and the submission data package is pushed again according to the original encapsulation and transmission process; for the failure type that needs to modify the data such as data verification failure and format error, the failure details can be fed back to the data processing module to prompt to reassociate the data or correct the format, and after the data correction is completed, the submission data package is generated again and the pushing operation is performed.

[0211] In the whole declaration data processing process, the unified social credit code of the declaration subject and the private sensitive data such as the declaration data content are involved, and the system adopts multiple privacy protection technical means. In the data collection stage, the declaration data is transmitted by encryption transmission protocol to prevent data from being stolen or tampered with during transmission; in the data storage stage, the privacy sensitive data is desensitized, such as character replacement or mask processing for part of the fields, and access control strategy is adopted to limit only authorized personnel can access sensitive data; in the data processing process, all operations involving private data are carried out in an encrypted environment to ensure the security of the data in use.

[0212] Figure 2 A schematic diagram of exemplary hardware and software components of an RPA full-process automated declaration data processing system 100 that can implement the idea of the present application is shown. For example, the processor 120 can be used in the RPA full-process automated declaration data processing system 100 and used to perform the functions in the present application.

[0213] For example, the RPA full-process automatic declaration data processing system 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. Exemplarily, the RPA full-process automatic declaration data processing system 100 can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The methods of the present application can be implemented according to these program instructions. The RPA full-process automatic declaration data processing system 100 also includes an I / O interface 150 between the computer and other input and output devices.

[0214] In addition, the embodiments of the present application also provide a readable storage medium, wherein computer executable instructions are preset in the readable storage medium, and when a processor executes the computer executable instructions, the RPA full-process automatic declaration data processing method is realized.

[0215] It should be noted that, in order to simplify the description of the present application and to help the understanding of one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, drawing or description thereof.

Claims

1. A method for RPA full-process automatic declaration data processing, characterized in that, The method comprises: receiving a declaration data set, the declaration data set containing various types of declaration form data and proof material data submitted by a declaration subject, the declaration form data and the proof material data having an association identifier, the association identifier being used for uniquely associating corresponding declaration data of the same declaration subject; calling a preset RPA declaration process template, constructing a business logic mapping matrix based on content attributes of various types of data in the declaration data set, and performing bidirectional adaptation of various types of data and nodes in the RPA declaration process template through the business logic mapping matrix to obtain a process node bidirectional adaptation result; constructing a node processing time length prediction model in combination with historical processing data of each RPA declaration process node according to the process node bidirectional adaptation result, allocating RPA computing resources and data transmission channels based on the node processing time length prediction model, and generating a node resource prediction scheduling scheme; for an RPA declaration process node having data missing or format deviation in the process node bidirectional adaptation result, calling an RPA tool based on the node resource prediction scheduling scheme to perform cross-node data association deduction, supplement missing data or correct format deviation data, and obtaining post-association deduction node data; integrating node data without data problems in the process node bidirectional adaptation result and the post-association deduction node data according to a business logic sequence of the RPA declaration process template, generating a declaration data submission package, pushing the declaration data submission package to a target declaration system, and obtaining declaration submission feedback data; the calling of the preset RPA declaration process template, the construction of the business logic mapping matrix based on content attributes of various types of data in the declaration data set, the bidirectional adaptation of various types of data and nodes in the RPA declaration process template through the business logic mapping matrix, and the obtaining of the process node bidirectional adaptation result, comprising: calling a preset RPA declaration process template, the RPA declaration process template containing a plurality of RPA declaration process nodes connected in series according to declaration business logic, each RPA declaration process node being marked with a business logic description, the business logic description containing a declaration business link corresponding to the node, a functional positioning of data to be processed, and a data flow relationship with upstream and downstream nodes; performing content attribute analysis on various types of declaration form data and proof material data in the declaration data set, extracting a business function tag of each type of data, the business function tag containing a declaration business action supported by the data, business decision information carried by the data, and an associated upstream and downstream data flow direction; constructing a business logic mapping matrix based on business logic descriptions of all RPA declaration process nodes and business function tags of various types of data, the row dimension of the business logic mapping matrix being an RPA declaration process node identifier, the column dimension being a data business function tag, and the matrix element being a business logic matching degree of a node and data; calculating the business logic matching degree of each type of data and each RPA declaration process node through the business logic mapping matrix, and determining an RPA declaration process node having the highest business logic matching degree as a preliminary adaptation node of the data; analyzing a data flow relationship between the preliminary adaptation node and the upstream and downstream nodes, judging whether the data can be transferred to the downstream node through the preliminary adaptation node and meet the business logic requirements of the downstream node, if yes, determining the preliminary adaptation node as the target adaptation node, if no, re-adjusting a matching dimension of the business logic mapping matrix, supplementing potential business function tags of the data, and then calculating the matching degree again until the target adaptation node is determined or marked as adaptation to-be-processed data; associating all the data determined as the target adaptation node with the corresponding RPA declaration process node, recording the business function tags of the adaptation to-be-processed data and unadapted reasons, and forming a process node bidirectional adaptation result containing node identification, corresponding data list and to-be-processed data record.

2. The RPA end-to-end automation declaration data processing method according to claim 1, characterized in that, According to the process node bidirectional adaptation result, a node processing time length prediction model is constructed in combination with historical processing data of each RPA declaration process node, RPA computing resources and data transmission channels are allocated based on the node processing time length prediction model, and a node resource prediction scheduling scheme is generated, including: extracting the adaptation data amount corresponding to each RPA declaration process node from the process node bidirectional adaptation result, and simultaneously calling historical processing data of each RPA declaration process node, the historical processing data containing historical adaptation data amount, corresponding processing time length and resource consumption record; extracting features from the historical processing data, inputting the extracted features as training samples into a regression model, and constructing a node processing time length prediction model, the extracted features including historical adaptation data amount, data field quantity, data correlation dimension and corresponding processing time length; inputting the current adaptation data amount, data field quantity and data correlation dimension of each RPA declaration process node into the node processing time length prediction model to obtain the predicted processing time length of each RPA declaration process node; calling a usable computing resource account of the RPA system, the usable computing resource account containing the operation speed of idle processors, the read / write speed of memory and the capacity of storage resources, and simultaneously calling real-time transmission rate and channel load record of available data transmission channels; allocating corresponding computing resource specifications and quantities for each RPA declaration process node according to the predicted processing time length of each RPA declaration process node and the usable computing resource account; calculating the minimum data transmission rate required by each RPA declaration process node according to the adaptation data amount and the predicted processing time length of each RPA declaration process node, and matching a data transmission channel for each RPA declaration process node, the real-time transmission rate of which is not lower than the required minimum data transmission rate and the current transmission data amount of which does not affect the data transmission progress of the node; recording the computing resource specifications, quantities and data transmission channel identifications allocated for each RPA declaration process node, marking the start time and predicted release time of the computing resource allocation, and forming a node resource prediction scheduling scheme containing node identification, resource configuration details, transmission channel parameters and predicted processing time length. 3.The RPA end-to-end automation declaration data processing method of claim 2, wherein, The feature extraction from the historical processing data, inputting the extracted features as training samples into a regression model, and constructing a node processing time length prediction model, includes: The historical processing data of each RPA declaration process node is cleaned, the historical adaptation data volume is extracted from the cleaned historical processing data, and the data byte number is taken as the quantitative index of the historical adaptation data volume; the data field number is extracted, and the total number of independent fields contained in each batch of historical processing data is counted; the data association dimension is extracted, and the number of association links between each batch of historical processing data and other RPA declaration process node data is counted; The historical adaptation data volume, the data field number, and the data association dimension are taken as model input features, and the corresponding historical processing time is taken as a model output label to construct a training sample set; The training sample set is divided into a training subset and a validation subset, wherein the training subset is used for model training, and the validation subset is used for model performance verification; A gradient boosting regression model is selected as a base model, the training subset is input into the base model for iterative training, in each iteration process, the prediction error of the validation subset is used to adjust the hyperparameters of the base model, including the learning rate, the number of decision trees, and the tree depth; When the prediction error of the validation subset is less than a preset error threshold, the model training is stopped, and a final node processing time prediction model is obtained; The final node processing time prediction model is subjected to generalization test, historical processing data not participating in training is selected as a test sample, the test sample is input into the node processing time prediction model to obtain a predicted processing time, the deviation rate of the predicted processing time and the actual processing time is calculated, if the deviation rate is lower than a preset deviation threshold, it is confirmed that the node processing time prediction model is available, if the deviation rate is higher than the preset deviation threshold, the training sample or the model structure is re-adjusted until the generalization of the node processing time prediction model meets the requirements.

4. The RPA end-to-end automation declaration data processing method of claim 1, wherein, The RPA declaration process node with data missing or format deviation in the process node bidirectional adaptation result is based on the node resource prediction scheduling scheme to call the RPA tool to perform cross-node data association derivation, supplement the missing data or correct the format deviation data, and obtain the associated derived node data, including: The RPA declaration process node with data missing or format deviation is filtered out from the process node bidirectional adaptation result, the adaptation data corresponding to the RPA declaration process node is extracted, and the field name of the missing data or the specific type of the format deviation is determined; According to the node resource prediction scheduling scheme, the RPA computing resource allocated for the RPA declaration process node is called, and the cross-node data retrieval module of the RPA tool is started, which is used to query the adaptation data of the upstream and downstream nodes associated with the current RPA declaration process node; If there is data missing, the RPA tool obtains the adaptation data of the upstream node of the current RPA declaration process node through the cross-node data retrieval module, analyzes the association relationship between the adaptation data of the upstream node and the missing field of the current RPA declaration process node, if the adaptation data of the upstream node contains association data that can derive the missing field of the current RPA declaration process node, data derivation calculation is performed based on the association relationship to obtain the supplementary data of the missing field; If the adaptation data of the upstream node cannot derive the missing field of the current RPA declaration process node, the RPA tool further retrieves the preset data requirement of the downstream node of the current RPA declaration process node, reversely derives the reasonable value range of the missing field of the current RPA declaration process node according to the preset data requirement of the downstream node, and determines the supplementary data of the missing field in combination with the declaration business logic; If there is a format deviation, the RPA tool retrieves the adaptation data format of other RPA declaration process nodes consistent with the business logic of the current RPA declaration process node, converts the data of the current RPA declaration process node with the format deviation into the reference standard format as a reference standard, and converts the data of the current RPA declaration process node with the format deviation into the reference standard format as a reference standard; After the data supplement or format conversion is completed, the RPA tool fuses the processed data and the original adaptation data of the current RPA declaration process node to form the associated derived node data, records the node range, data association relationship and derivation process of the cross-node data retrieval, and associates to the current RPA declaration process node identifier.

5. The RPA end-to-end automation declaration data processing method according to claim 4, characterized in that, If there is data missing, the RPA tool obtains the adaptation data of the upstream node of the current RPA declaration process node through the cross-node data retrieval module, analyzes the association relationship between the adaptation data of the upstream node and the missing field of the current RPA declaration process node, and if the adaptation data of the upstream node contains associated data that can derive the missing field of the current RPA declaration process node, performs data derivation calculation based on the association relationship to obtain the supplementary data of the missing field, including: The cross-node data retrieval module of the RPA tool queries the upstream node identifier in the RPA declaration process template that has a direct data flow input relationship with the current RPA declaration process node according to the current RPA declaration process node identifier, and calls the adaptation data of the upstream node based on the upstream node identifier; The adaptation data of the upstream node is structurally analyzed, the key fields in the adaptation data of the upstream node are extracted, and the key fields are compared with the missing field of the current RPA declaration process node in terms of business logic to determine whether they have causal association or subordinate association; If there is causal association, the RPA tool calls the corresponding causal association rule in the declaration business logic library, and derives the supplementary data of the missing field of the current RPA declaration process node from the corresponding field in the adaptation data of the upstream node according to the causal association rule; If there is subordinate association, the RPA tool extracts the association identifier information from the corresponding field in the adaptation data of the upstream node, retrieves the corresponding business database through the association identifier information, and obtains the data corresponding to the missing field of the current RPA declaration process node as the supplementary data; During the derivation process, the RPA tool records the upstream node identifier, key field content and business logic rule on which the derivation is based to form a derivation process record; The derived supplementary data is filled into the missing field of the current RPA declaration process node, and the supplementary data and the existing data of the current RPA declaration process node are confirmed in terms of business logic continuity to ensure that the supplementary data and the existing data have no logical conflict in the declaration business scenario, and an associated derived node data segment containing complete fields is formed.

6. The RPA end-to-end automation declaration data processing method of claim 1, wherein, The business logic sequence according to the RPA declaration process template is integrated with the node data without data problems in the process node bidirectional adaptation result and the associated derived node data to generate a declaration data submission package, the declaration data submission package is pushed to the target declaration system, and declaration submission feedback data is obtained, including: Extracting node data without data problems from the process node bidirectional adaptation result, the node data without data problems is not marked with data missing or format deviation, and has completed bidirectional adaptation with the corresponding RPA declaration process node; Collecting associated derived node data corresponding to all RPA declaration process nodes with data problems, ensuring that associated derived node data has been generated for each RPA declaration process node with data problems, if any RPA declaration process node cannot solve the data problem through cross-node data association derivation, the node data of the RPA declaration process node is marked as temporary processing data and excluded from the integration range; Referring to the business logic sequence of each RPA declaration process node in the RPA declaration process template, the node data without data problems and the associated derived node data are sorted, so that the order of the sorted node data is consistent with the flow sequence of the declaration business; Performing associated field concatenation processing on the sorted node data, extracting the associated identification field in each node data, and concatenating the node data of different nodes into a complete declaration data link through the associated identification field; According to the declaration data organization format required by the target declaration system, the concatenated declaration data link is encapsulated into a declaration data submission package, the declaration data submission package includes package identification information, data link main body and data flow record, the package identification information includes declaration main body identification and submission batch code, and the data flow record includes the adaptation time and association derivation time of each node data; Retrieve the interface communication parameters of the target declaration system, the interface communication parameters include interface access address, data transmission encryption protocol and identity authentication token, and the RPA tool establishes an encrypted communication link with the target declaration system based on the interface communication parameters; Through the encrypted communication link, the declaration data submission package is transmitted to the target declaration system, and the data transmission progress is tracked in real time during transmission, if the transmission is interrupted, the standby data transmission channel is re-enabled based on the node resource prediction scheduling scheme to resume transmission; Receiving the transmission response information returned by the target declaration system, the transmission response information includes data reception status and integrity identification of received data; If the data reception status is successful, generate declaration submission feedback data including package identification information, reception success confirmation and subsequent processing instruction; if the data reception status is failed, record the failure type and the corresponding transmission parameter, and associate it to the package identification information of the declaration data submission package to form the declaration submission feedback data containing the failure details.

7. The RPA end-to-end automation declaration data processing method according to claim 6, characterized in that, The declaration data submission package is encapsulated according to the declaration data organization format required by the target declaration system, including: Obtain the declaration data packaging specification published by the target declaration system, analyze the structure requirements, field naming rules and data format standards of the declaration data submission package in the declaration data packaging specification; Generate the package identification information of the declaration data submission package, wherein the declaration subject identification is the subject code corresponding to the association identification in the declaration data set, and the submission batch code is generated by the RPA system in combination with the current date and the submission sequence of the day, and the submission batch code includes the date field, the time period field and the sequence field; After concatenation, the declaration data link is used as the data link subject, and the field names of each node data are uniformly adjusted according to the field naming rules of the declaration data packaging specification; According to the data format standard of the declaration data packaging specification, the data format of each field in the data link subject is converted; Extract the adaptation time of each node data in the process node bidirectional adaptation process, and the associated derivation time of the RPA declaration process node with data problems, associate the adaptation time and the associated derivation time with the corresponding node identification, and form a data flow record; Combine the package identification information, the data link subject and the data flow record in the order required by the declaration data packaging specification to construct the initial structure of the declaration data submission package; Compress the declaration data submission package of the initial structure, reduce the volume of the declaration data submission package by using the compression algorithm supported by the target declaration system, and record the file size before and after compression and the compression efficiency; Add a data verification identification to the compressed declaration data submission package, which is generated by performing a hash operation on the overall data of the declaration data submission package, and is used to confirm that the declaration data submission package has not been tampered with in the transmission process, and the packaging of the declaration data submission package is completed.

8. The RPA end-to-end automation declaration data processing method of claim 1, wherein, The received declaration data set includes: Deploy a distributed data receiving cluster of RPA tools, which includes multiple parallel receiving nodes, each of which is configured with an independent network port for simultaneously receiving data submitted by multiple declaration subjects; Configure a dedicated data receiving protocol for the declaration form data and the proof material data respectively, the declaration form data uses a form data dedicated transmission protocol, and the form data dedicated transmission protocol supports field-level data transmission; the proof material data uses a file fragment transmission protocol, and the file fragment transmission protocol supports large file fragment uploading; When the declaration subject initiates a data submission request, the receiving scheduling module of the RPA tool distributes the data submission request to the parallel receiving node of the corresponding data receiving protocol according to the data type in the data submission request; The parallel receiving node receives the declaration form data uploaded by the declaration subject, analyzes the field structure of the declaration form data, extracts the preset association identification field in the declaration form data, and records the content of the association identification field and the receiving time of the corresponding declaration form data; The parallel receiving node receives the proof material data uploaded by the declaration subject, reorganizes the fragmented file according to the file fragment transmission protocol, and extracts the association identification from the reorganized file metadata, wherein the file metadata includes file naming information and file attribute information; The declaration form data and the proof material data corresponding to the same association identifier are classified into a group of declaration data units, and each declaration data unit corresponds to single declaration data of a declaration subject; All the declaration data units are summarized to form a declaration data set, a receiving summary identifier is added to the declaration data set, the receiving summary identifier includes a parallel receiving node number, a receiving completion time and a total number of declaration data units, and the name of the declaration subject corresponding to each declaration data unit and the submission terminal information are recorded, and the receiving of the declaration data set is completed.

9. A RPA end-to-end automated declaration data processing system, characterized in that, The RPA full-process automatic declaration data processing system includes a processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to realize the RPA full-process automatic declaration data processing method in any one of claims 1-8.

Citation Information

Patent Citations

  • Human resource platform management method, system and equipment based on big data

    CN119067411A

  • Automatic tax declaration and recheck system

    CN120471722A