Semi-structured data file processing system, method and apparatus, and storage medium
By introducing data processing modules and python scripts into the RPA system, the problem that RPA cannot process mixed-format semi-structured tables is solved, efficient processing of complex data files is achieved, and the flexibility and adaptability of data processing is improved.
Patent Information
- Application Number
- PCT/CN2024/098667
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-06-12
- Publication Date
- 2025-05-22
AI Technical Summary
Existing RPA technologies are unable to effectively interpret and extract semi-structured tables in mixed formats, resulting in inefficiency in processing complex data files.
By introducing data processing modules and preset python scripts, the RPA process automation module is used to obtain semi-structured data files, and run python scripts in the data processing module to process these files and generate target data files.
The RPA process automation module has improved the flexibility and adaptability of the processing of semi-structured data files, allowing it to process data files in complex formats, and improving the efficiency and accuracy of data processing.
Smart Images

Figure CN2024098667_22052025_PF_FP_ABST
Abstract
Description
Semi-structured data file processing system, method, device and storage medium Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a semi-structured data file processing system, method, device and storage medium. Background Art
[0002] RPA (Robotic Process Automation) is an application technology used to automate business processes. Based on business logic and established rules, it uses configured software or robots to crawl and interpret applications, manipulate data, trigger responses, and communicate with other digital systems. The goal of RPA is to improve productivity, reduce costs, and minimize human error.
[0003] The growth of RPA technology is driven by several factors. First, businesses face a large number of repetitive, low-value-added tasks that consume employee time and resources. Second, the trend of digital transformation is driving businesses to adopt automation solutions to improve efficiency and competitiveness. Furthermore, advances in artificial intelligence and machine learning are enabling machines to better simulate and perform human tasks.
[0004] Currently, there are multiple RPA solutions available for enterprises to choose from and use. These solutions provide visual workflow designers, data capture and processing tools, automated script writing and execution capabilities, and the ability to integrate with other systems. They perform tasks by simulating human operations, such as filling out forms in applications, copying and pasting data, and sending emails.
[0005] However, although RPA technology has significant advantages in automating business processes, it currently only supports the simplest table styles and cannot directly interpret and extract semi-structured tables in mixed formats.
[0006] Summary of the Invention
[0007] In view of this, the present application provides a semi-structured data file processing system, method, device and storage medium to solve the problem that RPA currently only supports the simplest table style and cannot directly interpret and extract semi-structured tables in mixed formats.
[0008] In the first aspect, the present application provides a semi-structured data file processing system, including: an RPA process automation module and a data processing module; the RPA process automation module is used to obtain the semi-structured data file to be processed, and send the semi-structured data file to be processed to the data processing module; the data processing module is used to use a preset Python script to process the semi-structured data file to be processed to obtain the target data file.
[0009] The semi-structured data file processing system provided in this application processes the semi-structured data files that the RPA process automation module cannot process in the data processing module through a preset Python script, and can obtain the target data files that the RPA process automation module can continue to process, providing flexibility and adaptability for the RPA process automation module when processing semi-structured data files.
[0010] In an optional embodiment, the RPA process automation module includes: a first login submodule and an acquisition submodule; the first login submodule is used to obtain first login account information and first login password information, and log in to a preset web page based on the first login account information and the first login password information, and send a first login success instruction to the acquisition submodule; the acquisition submodule is used to obtain the semi-structured data file to be processed in the preset web page according to preset business requirements based on the first login success instruction.
[0011] In an optional embodiment, the first login submodule includes: a first acquisition unit and a capture unit; the first acquisition unit is used to obtain the first login account information and the first login password information, and send the first login account information and the first login password information to the capture unit; the capture unit is used to capture and determine the target input box in the preset web page based on the web page tag of the preset web page, and enter the first login account information and the first login password information into the target input box for login.
[0012] In an optional embodiment, the acquisition submodule includes: an access unit and a first determination unit; the access unit is used to access the target interface and send an access success instruction to the first determination unit when a first login success instruction is received; the first determination unit is used to obtain the semi-structured data file to be processed in the target interface according to preset business requirements when a access success instruction is received.
[0013] In an optional embodiment, the data processing module includes: a calling submodule, a determining submodule and a data processing submodule; the calling submodule is used to call a preset python script in a preset script library, and send the preset python script to the data processing submodule; the determining submodule is used to determine target operating parameters based on the semi-structured data file to be processed, and send the target operating parameters to the data processing submodule; the data processing submodule is used to process the semi-structured data file to be processed using the preset python script based on the target operating parameters to obtain the target data file.
[0014] This application can determine the target running parameters of the preset Python script through the semi-structured data file to be processed. Furthermore, based on the target running parameters, the preset Python script called by the data processing module is used to process the semi-structured data file to be processed. The semi-structured data file that cannot be processed by the RPA process automation module can be processed into the target data file that the RPA process automation module can continue to process, providing flexibility and adaptability for the RPA process automation module to process semi-structured data files.
[0015] In an optional embodiment, the determination submodule includes: a second determination unit, a setting unit and a third determination unit; the second determination unit is used to determine the file path and target account information based on the semi-structured data file to be processed, and send the target account information to the setting unit, and send the file path to the third determination unit; the setting unit is used to set the data type and global calculation variables based on the target account information, and send the data type and global calculation variables to the third determination unit; the third determination unit is used to determine the target operating parameters based on the file path, data type and global calculation variables.
[0016] In an optional embodiment, the data processing submodule includes: a first processing unit, a generation unit and a fourth determination unit; the first processing unit is used to process the semi-structured data file to be processed based on the target operating parameters and the preset first condition using a preset Python script to obtain a first data set, and send the first data set to the generation unit and the fourth determination unit; the generation unit is used to generate a first data file based on the first data set, and send the first data file to the fourth determination unit; the fourth determination unit is used to determine the second data set based on the first data set, and determine the target data file based on the second data set and the first data file.
[0017] In an optional embodiment, the first processing unit includes: an acquisition subunit, an adjustment subunit, a first processing subunit and a determination subunit; the acquisition subunit is used to process the semi-structured data file to be processed using a preset python script based on target operating parameters to obtain a first data subset that meets the preset first condition and a second data subset that does not meet the preset first condition, and send the first data subset to the adjustment subunit and send the second data subset to the first processing subunit; the adjustment subunit is used to adjust the data format of each data in the first data subset that does not meet the preset data format to obtain a first target data subset, and send the first target data subset to the determination subunit; the first processing subunit is used to process the second data subset using a preset placeholder to obtain a second target data subset, and send the second target data subset to the determination subunit; the determination subunit is used to determine the first data set based on the first target data subset and the second target data subset.
[0018] In an optional embodiment, the fourth determination unit includes: a second processing subunit, a generation subunit and a third processing subunit; the second processing subunit is used to obtain a second data set based on the second target data subset through a preset processing method, and send the second data set to the generation subunit; the generation subunit is used to generate a second data file based on the second data set and the first data file, and send the second data file to the third processing subunit; the third processing subunit is used to process the second data file according to preset requirements to obtain the target data file.
[0019] In an optional embodiment, the semi-structured data processing system is connected to the SPA system; the RPA process automation module is also used to receive the target data file sent by the data processing module and upload the target data file to the SPA system.
[0020] In an optional embodiment, the RPA process automation module further includes: a second login submodule and an upload submodule; the second login submodule is used to obtain second login account information and second login password information, and log in to a preset SPA webpage based on the second login account information and second login password information, and send a second login success instruction to the upload submodule; the upload submodule is used to upload the target data file to the SPA system based on the second login success instruction.
[0021] In an optional embodiment, the upload submodule includes: a second acquisition unit and an upload unit; the second acquisition unit is used to obtain the target transaction code and file configuration parameters, and send the target transaction code and file configuration parameters to the upload unit; the upload unit is used to upload the target data file to the SPA system based on the target transaction code and file configuration parameters.
[0022] In a second aspect, the present application provides a method for processing a semi-structured data file, which is used in a semi-structured data file processing system according to the first aspect or any corresponding embodiment thereof; the method comprises:
[0023] Obtain a semi-structured data file to be processed and call a preset Python script in a preset script library; determine target operating parameters based on the semi-structured data file to be processed; based on the target operating parameters, use the preset Python script to process the semi-structured data file to obtain a target data file.
[0024] The semi-structured data file processing method provided in this application uses the semi-structured data file processing system of the above-mentioned first aspect of this application or any corresponding embodiment to process semi-structured data files. Semi-structured data files that cannot be processed by the RPA process automation module can be processed into target data files that the RPA process automation module can continue to process, providing flexibility and adaptability for the RPA process automation module when processing semi-structured data files.
[0025] In a third aspect, the present application provides a semi-structured data file processing device, configured to execute the semi-structured data file processing method provided in the second aspect above; the device comprises:
[0026] The acquisition module is used to obtain the semi-structured data file to be processed and call the preset Python script in the preset script library; the determination module is used to determine the target operating parameters based on the semi-structured data file to be processed; the processing module is used to process the semi-structured data file to be processed using the preset Python script based on the target operating parameters to obtain the target data file.
[0027] In a fourth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the semi-structured data file processing method provided in the second aspect above.
[0028] In a fifth aspect, the present application provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the semi-structured data file processing method provided in the second aspect above by executing the computer instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] FIG1 is a block diagram of a semi-structured data file processing system according to an embodiment of the present application;
[0031] FIG2 is a flow chart of a method for processing a semi-structured data file according to an embodiment of the present application;
[0032] FIG3 is a structural block diagram of a semi-structured data file processing device according to an embodiment of the present application;
[0033] FIG4 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0035] Although RPA technology has significant advantages in automating business processes, it currently only supports the simplest table formats and cannot directly interpret and extract semi-structured tables in mixed formats.
[0036] For example, it's necessary to download bank transaction records from multiple banks' online banking systems based on specific dates (in the bank's designated Excel format), import the data into a pre-defined Excel template, and upload it to the SAP system. To perform this task, RPA utilizes technologies such as browser page element recognition, automated mouse clicks, keyboard input, Excel operations, and OCR recognition. The RPA robot pre-stores the required pages, along with the required actions and interaction data, to repeatedly execute the pre-set workflow.
[0037] When banks download bank statements, each bank has its own Excel format, and some versions may not conform to the Excel templates supported by the SAP financial system. Typically, RPA's built-in Excel information extraction function can be used to extract information, then reformat the data and insert it into a new Excel file according to the template. However, RPA currently only supports the simplest table format (i.e., a table with a header in the first row and data in the second and subsequent rows). Furthermore, when the downloaded Excel file format is semi-structured, RPA cannot directly interpret and extract this mixed format, i.e., semi-structured data files.
[0038] In this embodiment, a semi-structured data file processing system is provided. Figure 1 is a structural block diagram of the semi-structured data file processing system according to an embodiment of the present application. As shown in Figure 1, the semi-structured data file processing system 1 includes: an RPA process automation module 11 and a data processing module 12.
[0039] Specifically, the semi-structured data file processing system 1 is connected to the SPA system 2. The SPA system is an enterprise resource management software system.
[0040] It should be understood that the system also includes other devices and equipment.
[0041] Optionally, the RPA process automation module 11 includes: a first login submodule 111 , an acquisition submodule 112 , a second login submodule 113 and an upload submodule 114 .
[0042] The first login submodule 111 includes a first acquisition unit 1111 and a capture unit 1112 ; the acquisition submodule 112 includes an access unit 1121 and a first determination unit 1122 ; and the upload submodule 114 includes a second acquisition unit 1141 and an upload unit 1142 .
[0043] Optionally, the data processing module 12 includes a calling submodule 121 , a determining submodule 122 and a data processing submodule 123 .
[0044] The determining submodule 122 includes a second determining unit 1221 , a setting unit 1222 and a third determining unit 1223 ; the data processing submodule 123 includes a first processing unit 1231 , a generating unit 1232 and a fourth determining unit 1233 .
[0045] Furthermore, the first processing unit 1231 includes an acquisition subunit 12311 , an adjustment subunit 12312 , a first processing subunit 12313 and a determination subunit 12314 ; the fourth determination unit 1233 includes a second processing subunit 12331 , a generation subunit 12332 and a third processing subunit 12333 .
[0046] Furthermore, the functions of each device in the above system are described.
[0047] Optionally, the RPA process automation module 11 is used to obtain a semi-structured data file to be processed and send the semi-structured data file to be processed to the data processing module 12.
[0048] First, the first login account information and the first login password information are obtained in the first login submodule 111, and the user logs in to the preset web page based on the obtained first login account information and the first login password information. Optionally, after the login is successful, a first login success instruction is sent to the acquisition submodule 112.
[0049] Specifically, the first acquisition unit 1111 acquires the first login account information and the first login password information and sends them to the capture unit 1112 .
[0050] Optionally, in the capture unit 1112, based on the web page tag of the preset web page, the user name, password tag, and login button are captured from the preset web page, the input box that needs to be filled with account information, i.e., the target input box, is confirmed, and the received first login account information and first login password information are entered into the target input box to complete the login.
[0051] Secondly, after receiving the first login success instruction, the acquisition submodule 112 acquires and obtains the corresponding semi-structured data file to be processed in the preset webpage according to preset business requirements under the control of the first login success instruction.
[0052] Specifically, under the control of the first login success instruction, the access unit 1121 accesses the corresponding mold table interface and sends an access success instruction to the first determination unit 1122 .
[0053] Optionally, after receiving the first determining unit 1122 , the first determining unit 1122 obtains the corresponding semi-structured data file to be processed in the target interface according to preset business requirements.
[0054] For example, when the target interface is a bank transaction interface, the corresponding file is downloaded from the bank transaction interface according to the selected preset start and end dates. During the download, the file is downloaded in .csv format and stored in the preset working path.
[0055] Optionally, the data processing module 12 is configured to process the semi-structured data file to be processed using a preset Python script to obtain a target data file.
[0056] First, the corresponding preset Python script is called in the preset script library integrated in the calling submodule 121 , and the preset Python script is sent to the data processing submodule 123 .
[0057] Secondly, in the determination submodule 122 , target running parameters of the preset Python script are determined according to the received semi-structured data file to be processed.
[0058] Specifically, the second determining unit 1221 determines the file path and target account information corresponding to the semi-structured data file to be processed. The file path is the file save path corresponding to the semi-structured data file to be processed obtained by the first determining unit 1122 in the target interface according to the preset business requirements. The target account information is the account number involved in the semi-structured data file to be processed. For example, if the semi-structured data file to be processed is a transaction flow file of a banking system, the target account information is the bank account number involved in the file.
[0059] Optionally, the corresponding data type and global calculation variables are set in the setting unit 1222 according to the target account information. For example, the data type corresponding to the transaction flow file of the bank system may be currency, etc.; the corresponding global calculation variables may be global variables used to calculate the beginning balance of each day, etc. This embodiment does not make specific restrictions on this, and it can be determined according to the semi-structured data file to be processed.
[0060] Optionally, the data type and global calculation variables are sent to the third determining unit 1223, and the third determining unit 1223 can determine the target running parameters corresponding to the preset Python script according to the received file path, data type and global calculation variables.
[0061] For example, the target running parameters of the preset Python script corresponding to the transaction flow file of the banking system can be:
[0062] The file path of the downloaded transaction flow file;
[0063] The contents of the account-currency correspondence list corresponding to the transaction flow file.
[0064] Finally, based on the target operating parameters, the preset Python script is run in the data processing submodule 123, and the preset Python script is used to process the semi-structured data file to be processed into a target data file that can be further processed by the RPA process automation module 11.
[0065] Specifically, based on the target operating parameters, the acquisition subunit 12311 in the first processing unit 1231 processes the semi-structured data file to be processed using the preset Python script to obtain a first data subset that satisfies a preset first condition and a second data subset that does not satisfy the preset first condition. The preset first condition is used to determine whether the corresponding data can be directly obtained from the semi-structured data file to be processed.
[0066] Optionally, the first data subset is sent to the adjustment subunit 12312 , and the second data subset is sent to the first processing subunit 12313 .
[0067] Optionally, in the adjustment subunit 12312 , the data format of each data in the first data subset that does not meet the preset data format is adjusted to obtain an adjusted first target data subset. Optionally, the first target data subset is sent to the determination subunit 12314 .
[0068] Optionally, since the second data subset cannot be directly obtained from the semi-structured data file to be processed, in the first processing sub-unit 12313, the second data subset is processed using a preset placeholder to obtain a second target data subset, that is, the second target data subset is composed of the preset placeholder.
[0069] Optionally, the determining subunit 12314 may obtain a corresponding first data set according to the received first target data subset and second target data subset, and send the first data set to the generating unit 1232 .
[0070] Optionally, the generating unit 1232 generates a corresponding first data file according to the first data set, that is, the first data file lacks the data value of each data in the second target data subset, and sends the first data file to the fourth determining unit 1233.
[0071] Optionally, after receiving the first data file, the second processing subunit 12331 in the fourth determining unit 1233 obtains the data value of each data in the second target data subset through a preset processing method such as calculation, and forms a corresponding second data set.
[0072] Optionally, the second processing sub-unit 12331 sends the second data set to the generation sub-unit 12332, so that the generation sub-unit 12332 can fill the received second data set into the formed first data file, generate the corresponding second data file, and send the second data file to the third processing sub-unit 12333.
[0073] Optionally, after receiving the second data file, the third processing sub-unit 12333 processes the second data file according to preset requirements, and obtains the target data file that the final RPA process automation module 11 can continue to process. The preset requirements can be determined according to preset business needs. For example, when the semi-structured data file to be processed is a transaction flow file of a banking system, after obtaining the second data file corresponding to the transaction flow file through the above processing process, the row data in the second data file can be segmented according to the date and the bank flow file of the corresponding date, i.e., the target data file, can be obtained (if no transaction flow occurs on the corresponding date, the corresponding file will not be generated. For example, if there are 31 days in January, a maximum of 31 new independent .csv files will be generated in January, and each file only contains transaction information for that day).
[0074] Optionally, the RPA process automation module 11 is further configured to receive the target data file sent by the data processing module 12 and upload the received target data file to a corresponding SPA system connected to the semi-structured data file processing system 1 .
[0075] First, the second login account information and the second login password information are obtained in the second login submodule 113, and the user logs in to the preset SPA webpage corresponding to the SPA system according to the obtained second login account information and the second login password information to access the SPA system.
[0076] Next, after the login is successful, the second login submodule 113 sends a second login success instruction to the upload submodule 114 .
[0077] Finally, after receiving the second login success instruction, the upload submodule 114 can upload the target data file to the corresponding SPA system.
[0078] Specifically, after receiving the second successful login instruction, the second acquisition unit 1141 is used to obtain the corresponding target transaction code (for example, the target transaction code corresponding to the transaction flow file of the bank system is the corresponding electronic bank statement) and file configuration parameters, and the target transaction code and file configuration parameters are sent to the upload unit 1142.
[0079] Optionally, the uploading unit 1142 may upload the target data file to a corresponding SPA system according to the received target transaction code and file configuration parameters.
[0080] The semi-structured data file processing system provided in this embodiment can determine the target running parameters of a preset Python script based on the semi-structured data file to be processed. Optionally, based on the target running parameters, the preset Python script called by the data processing module is used to process the semi-structured data file to be processed. This system can process semi-structured data files that cannot be processed by the RPA process automation module and obtain target data files that can be continued to be processed by the RPA process automation module, thereby providing flexibility and adaptability for the RPA process automation module in processing semi-structured data files.
[0081] In one example, taking a semi-structured bank transaction flow file as an example, the processing process of the bank transaction flow file based on the semi-structured data file processing system provided above in this embodiment is described as follows:
[0082] 1. Download bank balance (RPA process automation module 11)
[0083] 1. Open the online banking website and enter your account number and password to log in.
[0084] According to the web page tags, the user name, password tags, and login button are captured from the page, the input box that needs to be filled with account information is confirmed, the preset account and password information is automatically filled in, and the mouse click is simulated to log in.
[0085] 2. Download account transaction information
[0086] Continuously simulate mouse movements and clicks, accessing the trading interface from the account, selecting the preset start and end dates (in this scenario, the start and end dates are both today), and then clicking the download button to download the file in .csv format and save it to the preset working path.
[0087] 2. Parsing the downloaded content (executed by the data processing module 12)
[0088] Specifically, the RPA process automation module 11 calls a pre-written Python script to process, process, and calculate the contents of the downloaded bank flow .csv file line by line, and then generates a new .csv file with the processed data as the upload file material that the RPA process automation module 11 submits to the SPA system 2 in the next step.
[0089] The running parameters of the Python script are:
[0090] 1. The path of the downloaded bank statement file (in the form of embedded code);
[0091] 2. The contents of the account-currency correspondence list (in the form of embedded code).
[0092] Specifically, use embedded code to set startup parameters to reduce the operation of the running steps. Before each run, you need to check whether these two parameters are set correctly and modify them in time.
[0093] The running process of the Python script, i.e. the processing process of the data processing module 12, is as follows:
[0094] First, read the bank account number involved in this file and set the corresponding currency (because these two data are in the same table and each row is consistent); at the same time, set some global variables to calculate the beginning balance of each day.
[0095] Next, the first loop begins. This loop primarily performs two tasks: 1. Fills in data that can be directly excerpted, namely, data from the first data subset (including "transaction amount," "statement date," "transaction time," and "summary"). 2. Calculates the beginning and ending balances for each date and temporarily fills in blanks with the preset placeholder "NaN" (since these two data points require calculations and cannot be directly excerpted, they cannot be directly filled in during the first loop). While excerpting the data that can be directly filled in, some of the data formats need to be adjusted. For example, the date format needs to be converted from "01 / 06 / 2023" to "20230601"; the dollar sign in front of the amount is removed and string-type numbers are converted to float format (for calculations); bank account numbers with fewer than 11 digits are padded with leading zeros; decimals with fewer than or more than two digits are rounded to two decimal places, and other data format adjustments are performed.
[0096] Then, the second loop begins to fill in the opening and closing balances for each date, split the row data by date, and generate the bank statement file for the corresponding date. (If there is no transaction flow on the corresponding date, the corresponding file will not be generated. For example, if January has 31 days, a maximum of 31 new independent .csv files will be generated for January, and each file will only contain transaction information for that day.)
[0097] Finally, a line of "Execution Completed" text is displayed on the execution terminal screen as a prompt that the Python script has completed its work.
[0098] 3. Upload to SAP system (RPA process automation module 11 execution)
[0099] 1. Log in to SAP
[0100] Specifically, similar to step 1 in step 1, simulate keyboard and mouse operations, enter the preset account information on the SAP homepage, and click the button to complete the login.
[0101] 2. Use the upload file function on the SAP web page to complete the upload operation
[0102] Continuously simulate mouse movement and click operations, find transaction code FF.5 to import the electronic bank statement, simulate mouse clicks to set the configuration parameters for uploading the file, select the final Excel file from the local computer, and finally click the Execute button. This operation requires multiple cycles, and the number of cycles depends on the number of Excel files generated in the previous step.
[0103] In this embodiment, a semi-structured data file processing method is provided, which can be used in the above-mentioned semi-structured data file processing system 1. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0104] FIG2 is a flow chart of a method for processing a semi-structured data file according to an embodiment of the present application. As shown in FIG2 , the flow includes the following steps:
[0105] Step S201: Obtain a semi-structured data file to be processed and call a preset Python script in a preset script library.
[0106] The specific implementation process refers to the above functional description of the RPA process automation module 11 and the calling sub-module 121 in the semi-structured data file processing system 1, which will not be repeated here.
[0107] Step S202: determining target operating parameters based on the semi-structured data file to be processed.
[0108] The specific implementation process refers to the above functional description of the determination submodule 122 in the semi-structured data file processing system 1, which will not be repeated here.
[0109] Step S203 : Based on the target operating parameters, the semi-structured data file to be processed is processed using a preset Python script to obtain a target data file.
[0110] The specific implementation process refers to the above functional description of the data processing submodule 123 in the semi-structured data file processing system 1, which will not be repeated here.
[0111] The semi-structured data file processing method provided in this embodiment uses the semi-structured data file processing system provided in the above embodiment of this application to process semi-structured data files. Semi-structured data files that cannot be processed by the RPA process automation module can be processed into target data files that the RPA process automation module can continue to process, providing flexibility and adaptability for the RPA process automation module when processing semi-structured data files.
[0112] In this embodiment, a semi-structured data file processing device is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments. The details already described will not be repeated here. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0113] This embodiment provides a semi-structured data file processing device, as shown in FIG3 , including:
[0114] The acquisition module 301 is used to acquire the semi-structured data file to be processed and call a preset Python script in a preset script library.
[0115] The determination module 302 is configured to determine target operating parameters based on the semi-structured data file to be processed.
[0116] The processing module 303 is used to process the semi-structured data file to be processed based on the target operating parameters using a preset Python script to obtain a target data file.
[0117] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0118] The semi-structured data file processing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0119] An embodiment of the present application also provides a computer device having the semi-structured data file processing device shown in FIG3 above.
[0120] Please refer to Figure 4, which is a structural diagram of a computer device provided by an optional embodiment of the present application. As shown in Figure 4, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 takes a processor 10 as an example.
[0121] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may also optionally include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0122] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.
[0123] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0124] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0125] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0126] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; optionally, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0127] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A semi-structured data file processing system, characterized in that: include: RPA process automation module and data processing module; The RPA process automation module is used to obtain the semi-structured data file to be processed, and send the semi-structured data file to be processed to the data processing module; The data processing module is used to process the semi-structured data file to be processed using a preset python script to obtain a target data file.
2. The system according to claim 1, characterized in that The RPA process automation module includes: a first login submodule and an acquisition submodule; The first login submodule is used to obtain first login account information and first login password information, and log in to a preset webpage based on the first login account information and the first login password information, and send a first login success instruction to the acquisition submodule; The acquisition submodule is used to acquire the to-be-processed semi-structured data file in the preset webpage according to preset business requirements based on the first successful login instruction.
3. The system according to claim 2, characterized in that The first login submodule includes: a first acquisition unit and a capture unit; The first acquisition unit is used to acquire the first login account information and the first login password information, and send the first login account information and the first login password information to the capture unit; The capture unit is used to capture and determine a target input box in the preset web page based on the web page tag of the preset web page, and input the first login account information and the first login password information into the target input box for login.
4. The system according to claim 2, characterized in that The acquisition submodule includes: an access unit and a first determination unit; The access unit is configured to access the target interface and send the access success instruction to the first determining unit when receiving the first login success instruction; The first determining unit is configured to obtain the semi-structured data file to be processed in the target interface according to the preset business requirement when receiving the access success instruction.
5. The system according to claim 1, characterized in that The data processing module includes: a calling submodule, a determining submodule and a data processing submodule; The calling submodule is used to call the preset python script in the preset script library and send the preset python script to the data processing submodule; The determination submodule is used to determine target operating parameters based on the semi-structured data file to be processed, and send the target operating parameters to the data processing submodule; The data processing submodule is used to process the semi-structured data file to be processed based on the target operating parameters using the preset Python script to obtain the target data file.
6. The system according to claim 5, characterized in that The determining submodule includes: a second determining unit, a setting unit and a third determining unit; The second determining unit is configured to determine a file path and target account information based on the semi-structured data file to be processed, and send the target account information to the setting unit, and send the file path to the third determining unit; The setting unit is used to set a data type and a global calculation variable based on the target account information, and send the data type and the global calculation variable to the third determination unit; The third determination unit is used to determine the target operation parameter based on the file path, the data type and the global calculation variable.
7. The system according to claim 5, characterized in that The data processing submodule comprises: a first processing unit, a generating unit and a fourth determining unit; The first processing unit is configured to process the semi-structured data file to be processed by using the preset Python script based on the target operating parameter and the preset first condition to obtain a first data set, and send the first data set to the generating unit and the fourth determining unit; The generating unit is configured to generate a first data file based on the first data set, and send the first data file to the fourth determining unit; The fourth determining unit is configured to determine a second data set based on the first data set, and determine the target data file based on the second data set and the first data file.
8. The system according to claim 7, characterized in that The first processing unit includes: an acquisition subunit, an adjustment subunit, a first processing subunit and a determination subunit; The acquisition subunit is used to process the semi-structured data file to be processed by using the preset Python script based on the target operation parameter to obtain a first data subset that meets the preset first condition and a second data subset that does not meet the preset first condition, and send the first data subset to the adjustment subunit, and send the second data subset to the first processing subunit; The adjusting subunit is configured to adjust the data format of each data in the first data subset that does not satisfy the preset data format to obtain a first target data subset, and send the first target data subset to the determining subunit; The first processing subunit is used to process the second data subset using a preset placeholder to obtain a second target data subset, and send the second target data subset to the determining subunit; The determining subunit is configured to determine the first data set based on the first target data subset and the second target data subset.
9. The system according to claim 8, characterized in that The fourth determining unit includes: a second processing subunit, a generating subunit and a third processing subunit; The second processing subunit is used to obtain the second data set based on the second target data subset through a preset processing method, and send the second data set to the generating subunit; The generating subunit is configured to generate a second data file based on the second data set and the first data file, and send the second data file to the third processing subunit; The third processing sub-unit is used to process the second data file according to preset requirements to obtain the target data file.
10. The system according to claim 1, characterized in that The semi-structured data processing system is connected to the SPA system; The RPA process automation module is also used to receive the target data file sent by the data processing module and upload the target data file to the SPA system.
11. The system according to claim 10, characterized in that The RPA process automation module further includes: a second login submodule and an upload submodule; The second login submodule is used to obtain the second login account information and the second login password information, and log in to the preset SPA webpage based on the second login account information and the second login password information, and send a second login success instruction to the upload submodule; The upload submodule is used to upload the target data file to the SPA system based on the second login success instruction.
12. The system according to claim 11, characterized in that The upload submodule includes: a second acquisition unit and an upload unit; The second acquisition unit is used to acquire a target transaction code and file configuration parameters, and send the target transaction code and the file configuration parameters to the upload unit; The uploading unit is used to upload the target data file to the SPA system based on the target transaction code and the file configuration parameters.
13. A method for processing semi-structured data files, used in the semi-structured data file processing system according to any one of claims 1 to 12; characterized in that: The method comprises: Get the semi-structured data file to be processed and call the preset Python script in the preset script library; Determining target operating parameters based on the semi-structured data file to be processed; Based on the target operating parameters, the preset Python script is used to process the semi-structured data file to be processed to obtain a target data file.
14. A semi-structured data file processing device, used to execute the semi-structured data file processing method according to claim 13; characterized in that: The device comprises: An acquisition module is used to acquire the semi-structured data file to be processed and call a preset Python script in a preset script library; A determination module, configured to determine target operating parameters based on the semi-structured data file to be processed; A processing module is used to process the semi-structured data file to be processed based on the target operating parameters using the preset Python script to obtain a target data file.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the semi-structured data file processing method according to claim 13.
Citation Information
Patent Citations
Bank flow processing method and device combining RPA and AI
CN112529697A
Court trial record generation method and device based on RPA and AI, equipment and medium
CN114462376A
A method, system, storage medium, and device for parsing semi-structured data.
CN114936026A
Semi-structured data file processing system, method and device and storage medium
CN117708058A
Machine learning methods and systems for extracting entities from semi-structured enterprise documents
US20230267273A1