Hierarchical file analysis system and method
By introducing hierarchical architecture and artificial intelligence technology into the file parsing system and combining cloud services, the problem of file parsing dependence on rules and high maintenance costs in the existing technology is solved, and efficient and robust file data analysis is achieved.
Patent Information
- Application Number
- CN202311810293.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2023-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology relies on rule-led methods in file parsing, which leads to the failure of the model to learn new data, high maintenance costs, and engineers need to manually find keywords, which takes a long time.
A hierarchical file parsing system is adopted, combining artificial intelligence and cloud services, and filtering files through reference strings and similarity calculations, reducing maintenance costs and improving parsing efficiency.
It realizes the rapid finding of keyword values in files, reduces development and maintenance costs, avoids system unavailability problems caused by technical updates, and improves the robustness and efficiency of file parsing.
Smart Images

Figure CN120068843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to document parsing and artificial intelligence, and particularly to a hierarchical document parsing system and method. Background Art
[0002] When developing and designing products, engineers need to study numerous documents to obtain the technical specification parameters required for design and development. Since the documents come from different customers and manufacturers, the document formats cannot be made consistent.
[0003] However, due to the variety of document formats, document interpretation highly relies on the professional capabilities of engineers, so misinterpretation is very likely to occur. The current document automatic parsing methods are mainly rule-based. Although this method is easy to build a model, the established model cannot change the parsing rules, cannot learn new data that has not been trained, and its subsequent maintenance cost is very high. Summary of the Invention
[0004] In view of this, the present invention provides a hierarchical document parsing system and method. By using a filtering model and artificial intelligence to perform preliminary document data parsing, and cooperating with a cloud service for document parsing, high-robustness document data can be obtained, and the number of calls to the cloud service can be reduced to save service costs and transmission time.
[0005] A hierarchical document parsing method according to an embodiment of the present invention includes the following steps executed by a computing device: obtaining a reference string, obtaining a plurality of first strings in a plurality of documents, where each document includes a table, the table includes a plurality of text columns and a plurality of numerical columns, and each first string is associated with the plurality of text columns; calculating a first similarity between the reference string and each first string; screening out a plurality of candidate documents from the plurality of documents, where the first similarity of each candidate document is greater than a first threshold; obtaining a plurality of second strings in the plurality of candidate documents through a second module, where each second string is associated with the plurality of text columns; calculating a second similarity between the reference string and each second string; and when the second similarity is greater than a second threshold, determining one of the plurality of numerical columns according to each second string and outputting it.
[0006] A hierarchical file parsing system according to an embodiment of the present invention includes a storage device and a computing device. The storage device is used to store a plurality of instructions, a plurality of reference strings, and a plurality of files. The computing device is electrically connected to the storage device and is used to execute the plurality of instructions to cause a plurality of operations, which include: obtaining a reference string; obtaining a plurality of first strings from the plurality of files, where each file includes a table, the table includes a plurality of text columns and a plurality of numerical columns, and each first string is associated with the plurality of text columns; calculating a first similarity between the reference string and each first string, screening out a plurality of candidate files from the plurality of files, where the first similarity of each candidate file is greater than a first threshold, obtaining a plurality of second strings from the plurality of candidate files through a second module, where each second string is associated with the plurality of text columns, calculating a second similarity between the reference string and each second string; and when the second similarity is greater than a second threshold, determining one of the plurality of numerical columns according to each second string and outputting it.
[0007] In summary, the hierarchical file parsing system and method proposed by the present invention have the following effects:
[0008] 1. Based on artificial intelligence technology, it can quickly find the keyword values in professional documents and solve the problem that engineers do not know which keywords to use, thereby reducing the overall search time.
[0009] 2. Compared with the "rule-dominated" fixed-rule file parsing method, for the new file parsing requirements of the present invention, there is no need to define new search rules, which can save the time of developers and reduce the labor cost.
[0010] 3. The present invention performs data screening based on open-source artificial intelligence models and hierarchical concepts, reduces the time for self-training artificial intelligence models, and speeds up the development speed.
[0011] 4. The design architecture of the present invention belongs to the hierarchical concept, so it is quite flexible. The models or algorithms used in each layer can be replaced, so it will not be bound by any one suite or manufacturer, avoiding the problem that the system cannot operate in the future because a certain method or suite is no longer supported or updated.
[0012] Overall, the advantages of the present invention are that it does not require developers to repeatedly develop, does not require repeated training of artificial intelligence models, is not bound by any method or suite, simplifies the original complicated steps, and improves the convenience of user operation.
[0013] The above description of the content of the present invention and the following description of the embodiments are used to demonstrate and explain the spirit and principle of the present invention, and provide a further explanation of the patent application scope of the present invention. Brief Description of the Drawings
[0014] Figure 1 It is a block architecture diagram of a hierarchical file parsing system according to an embodiment of the present invention;
[0015] Figure 2 It is a flowchart of a hierarchical file parsing method according to an embodiment of the present invention;
[0016] Figure 3 It is a schematic diagram of an example in the first stage;
[0017] Figure 4 It is Figure 2 a schematic diagram of an example of a step in;
[0018] Figure 5 It is a schematic diagram of an example in the second stage;
[0019] Figure 6 It is Figure 2 a detailed flowchart of a step in; and
[0020] Figure 7 It is Figure 6 a detailed flowchart of a step in.
[0021] Symbol Description
[0022] 10: Hierarchical file parsing system
[0023] 1: Storage device
[0024] 12: File database
[0025] 14: Table standard answer database
[0026] 16: Fuzzy comparison keyword database
[0027] 18: Artificial intelligence database
[0028] 3: Arithmetic device
[0029] A 1 ~A M : Reference string
[0030] P 1 ~P N : First string
[0031] Q 1 ~Q L : Second string
[0032] S0 - S10, S101 - S106, S1021 - S1025: Steps
[0033] TB: Candidate file Detailed implementation manners
[0034] The detailed features and characteristics of the present invention are described in detail in the following embodiments. The content is sufficient for any person skilled in the relevant art to understand the technical content of the present invention and implement it accordingly. Based on the content disclosed in this specification, the claims, and the drawings, any person skilled in the relevant art can easily understand the related concepts and characteristics of the present invention. The following embodiments further illustrate the viewpoints of the present invention in detail, but do not limit the scope of the present invention in any way.
[0035] Figure 1 It is a block architecture diagram of a hierarchical file parsing system according to an embodiment of the present invention. As Figure 1 shown. The hierarchical file parsing system 10 includes a storage device 1 and a computing device 3. The storage device 1 is used to store a plurality of files, a plurality of reference strings, a plurality of keywords, and a plurality of instructions. Among them, the plurality of files are stored in the file database 12, the plurality of reference strings are stored in the table standard answer database 14, the plurality of keywords are respectively stored in the fuzzy comparison keyword database 16 and the artificial intelligence database 18, and the plurality of instructions are stored in an area other than the above databases. The computing device 3 is electrically connected to the storage device 1 and is used to execute the plurality of instructions to cause a plurality of operations. For these operations, please refer to Figure 2 .
[0036] In one embodiment, the storage device 1 is, for example, a volatile memory and / or a non-volatile memory. The non-volatile memory includes read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), flash memory, phase-change random access memory (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), and / or ferroelectric RAM (FRAM). The volatile memory may include dynamic RAM (DRAM), static RAM (SRAM), and / or synchronous DRAM (SDRAM). In another embodiment, the storage device 1 is, for example, at least one of a hard disk drive (HDD), a solid-state drive (SSD), a compact flash (CF) card, a secure digital (SD) card, a micro SD card, a mini SD card, an extreme digital (xD) card, and a memory stick. The present invention does not limit the hardware type of the storage device 1.
[0037] In one embodiment, the arithmetic device 3 can be implemented using one or several of the following examples: a personal computer, a network server, a microcontroller (MCU), an application processor (AP), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system-on-a-chip (SOC), a deep learning accelerator, or any electronic device with similar functions. The present invention does not limit the hardware type of the arithmetic device 3.
[0038] Figure 2It is a flowchart of a hierarchical file parsing method according to an embodiment of the present invention. As Figure 2 shown, this method includes steps S0 to S10. This method is divided into three stages to perform data screening and obtain the final data. The first stage includes steps S0 to S5, the second stage includes steps S6 to S9, and the third stage includes step S10.
[0039] The purpose of the first stage is to filter out low-correlation tabular data and retain high-correlation tabular data, thereby reducing the time and cost during the operation of the second stage. Figure 3 It is a schematic diagram of an example of the first stage.
[0040] Before step S0 is executed, the user stores at least one file in the file database 12 of the storage device 1. In one embodiment, each file is, for example, a PDF file, but the present invention is not limited thereto. In one embodiment, the number of the at least one file is multiple. The file content includes one or more tables, and each table includes a text column and a numerical column.
[0041] In step S0, the computing device 3 obtains at least one reference string from the table standard answer database 14 of the storage device 1, such as Figure 3 shown A1, A2,... AM. These reference strings A1, A2,... AM are used as the basis for subsequent file screening. In one embodiment, the user pre-collects multiple tables that have been used to find data. The computing device 3 extracts the reference strings from the text columns of each table through optical character recognition and then stores them in the table standard answer database 14. The present invention does not limit the length or number of the reference strings.
[0042] In step S1, the computing device 3 runs the first module to obtain multiple first strings in multiple files. The multiple first strings are such as Figure 3 shown P1, P2,... PN. Please refer to Figure 3 , in one embodiment, the computing device 3 retrieves multiple pre-stored files from the file database 12, and the first module is an open-source table parsing tool. Free open-source parsing tools are, for example: PdfPlumber, PyPDF2, and PyMuPdf, but are not limited thereto. Taking a PDF file as an example, since the tabular data in the file is in picture format, it needs to be converted into string format through the first module (such as an open-source table parsing tool). Schematic diagrams of the above operations can be referred to Figure 3 and Figure 4 . In Figure 3 and Figure 4 , each first string (such as the first string P1) represents the text data (non-numerical data) in a table on a certain page of a certain file.
[0043] In another embodiment of step S1, the arithmetic device 3 executes a plurality of instructions in the storage device 1, converts the plurality of files from a picture format to a text format, then differentiates the text data and numerical data in the text format, and then saves the text data as the plurality of first strings.
[0044] In step S2, the arithmetic device 3 calculates a first similarity between the reference string and each first string. In one embodiment, the arithmetic device 3 runs one of the following open-source string comparison models to calculate the first similarity.
[0045] 1. Hugging Face: sentence-transformers / all-MiniLM-L6-v2;
[0046] 2. GitHub: SeanLee97 / xmnlp; and
[0047] 3. GitHub: oborchers / Fast_Sentence_Embeddings.
[0048] The string comparison model can perform a comparison of sentence similarity and generate a comparison score as the first similarity. The first similarity reflects the similarity between the reference string and the first string, and the larger the value, the more similar the two are. In another embodiment, the arithmetic device 3 can run multiple string comparison models and calculate a weighted average as the first similarity based on the multiple comparison scores generated by them.
[0049] There are two reasons for the present invention to use open-source models: First, it saves the time for users to retrain the model by themselves. Second, since the existing artificial intelligence models for sentence similarity are already quite mature, there is no need to spend manpower to rebuild them, and it is sufficient to implement step S2 with the existing open-source models.
[0050] In step S3, the arithmetic device 3 determines whether the first similarity is greater than a first threshold. In one embodiment, the first threshold is 0.95, but the present invention is not limited to this exemplary value. If the determination in step S3 is no, then go to step S4. If the determination in step S3 is yes, then retain the current file as the candidate file TB and execute step S5.
[0051] In step S4, the arithmetic device 3 discards the first strings whose first similarity is not greater than the first threshold.
[0052] Please note that in Figure 3In the exemplary schematic diagram, there are a total of N first strings P1, P2, … PN. Therefore, step S2 will be executed N times. Each execution of step S2 will generate M first similarities. In one embodiment, if at least K of the M first similarities have values greater than the first threshold, the file corresponding to this first string (e.g., P1) can be retained as a candidate file TB for use in the second stage. Otherwise, this first string is discarded. The present invention does not limit the numerical settings of M, N, and K.
[0053] In step S5, the computing device 3 filters out a plurality of candidate files TB from the plurality of files. Overall, in steps S1 to S5 of the first stage, a plurality of first strings in all the files in the storage device 1 are screened through the lightweight table parsing module in step S1 and the noise filtering mechanism in steps S2 to S3, and the files with first similarities greater than the first threshold are retained as candidate files TB. In one embodiment, the first threshold is set to 0.95, which means that 95% of the files will be filtered out in the first stage, and the remaining 5% of the candidate files TB will continue to be screened in the second stage.
[0054] If the number of candidate files TB after the first stage screening is lower than the specified number (e.g., 2), the following two methods can be adopted. The first method: The computing device 3 arranges all the first strings in descending order according to the first similarity, and then sequentially selects the files corresponding to the first strings with the highest first similarity as candidate files TB until the number of candidate files TB is equal to the specified number. The second method: The computing device 3 reduces the first threshold (e.g., from 0.95 to 0.94), and then checks whether the number of candidate files TB exceeds the threshold value. If not, the first threshold is continuously reduced until the number of candidate files TB is sufficient for the second stage.
[0055] Next, the second stage (steps S6 to S9) is carried out to further screen the candidate files TB selected in the first stage to retain the precise table. Since the amount of data to be analyzed in the second stage is significantly reduced compared to the first stage, the cost and time required for execution can be saved. Figure 5 is an exemplary schematic diagram of the second stage.
[0056] In step S6, the computing device 3 obtains a plurality of second strings from the plurality of candidate files TB through the second module, such as Figure 5 shown as Q1, Q2, … QL. Since the number of candidate files TB is less than the number of files in the first stage, the number L of the plurality of second strings is less than the number N of the plurality of first strings. Please refer to Figure 5, in one embodiment, the second module is a cloud file table parsing module. In one embodiment, the accuracy of the cloud file table parsing module in recognizing the plurality of second strings is greater than that of the open-source table parsing tool in recognizing the plurality of first strings. For example, a free open-source table parsing tool (for step S1) may have problems such as missing table rows and columns or incorrect character recognition due to incomplete table parsing. In step S6, the computing device 3 runs a cloud file table parsing module such as Azure Table Recognizer or AWS Testract to obtain the plurality of second strings. For a schematic diagram of an example of the above operation, reference can be made to Figure 5 . In Figure 5 , each second string (such as Q1) represents the text data (non-numerical data) in a table on a certain page of a certain candidate file TB.
[0057] In another embodiment of step S6, the computing device 3 executes a plurality of instructions in the storage device 1 to convert the plurality of candidate files TB from picture format to text format, then distinguish the text data and numerical data in the text format, and then save the text data as the plurality of second strings.
[0058] In step S7, the computing device 3 calculates the second similarity between the reference string and the second string, and the specific implementation manner thereof can be referred to step S3.
[0059] In step S8, the computing device 3 determines whether the second similarity is greater than the second threshold. In one embodiment, the second threshold is 0.98, but the present invention is not limited to this exemplary value. If the determination in step S8 is no, then go to step S9. If the determination in step S8 is yes, then go to step S10. The second threshold must be greater than the first threshold because in the second stage, through a more stringent threshold setting, the candidate files TB screened out in the first stage are further screened to obtain accurate tables.
[0060] In step S9, the computing device 3 discards the second strings whose second similarity is not greater than the second threshold.
[0061] Next, the third stage (step S10) is carried out. For each accurate table screened out in the second stage, candidate sub-strings are taken out for fuzzy comparison and artificial intelligence comparison, and the candidate sub-strings belonging to the correct keywords are retained and the candidate sub-strings belonging to the wrong keywords are excluded.
[0062] In step S10, the computing device 3 outputs the numerical columns in the accurate table according to each second string. Generally speaking, step S10 is to evaluate whether there are keywords matching those recorded in the fuzzy comparison keyword database 16 or the artificial intelligence database 18 in the accurate tables screened out in the second stage. If so, then continue to output the data in the numerical columns of the accurate tables. Figure 6 YesFigure 2 The detailed flowchart of step S10.
[0063] In step S101, the arithmetic device 3 selects a candidate substring from multiple substrings of the second string. Please refer to Figure 4 and Figure 5 , assuming that the second string Q1 is retained as the exact table, which includes multiple candidate substrings such as: Module, 13, 3”, 13, 3”, FHD16, 9ColorTFT… etc. In one embodiment, the arithmetic device 3 selects one candidate substring each time to execute Figure 6 the process shown, for example, selecting “Module” for the first time, “13” for the next time… and so on.
[0064] In step S102, the arithmetic device 3 compares the candidate substring with multiple fuzzy keywords in the fuzzy comparison keyword database 16 to determine the weight of the candidate substring. In one embodiment, the number of executions of step S102 is the same as the number of fuzzy keywords. Therefore, multiple executions of step S102 will generate multiple weights, and the arithmetic device 3 subsequently uses the maximum value of these weights. Figure 7 is the detailed flowchart of step S102, including steps S1021 to S1025. The method for establishing the fuzzy comparison keyword database 16 will be described in detail later.
[0065] In step S1021, the arithmetic device 3 makes a first determination based on whether the candidate substring contains a fuzzy substring. If the first determination is negative, go to step S1022; if the first determination is positive, go to step S1023.
[0066] In step S1022, the arithmetic device 3 sets the weight to a first value, and then goes to step S103.
[0067] In step S1023, the arithmetic device 3 makes a second determination based on whether the fuzzy keyword is located in the first part of the candidate substring. If the second determination is negative, go to step S1024; if the second determination is positive, go to step S1025. In one embodiment, the candidate substring consists of a first part and a second part. The first part is located in the first half, and the second part is located in the second half. For example: if the candidate substring is “power supply”, the first part is “power” and the second part is “supply”. If the candidate substring is “power”, the first part is “p” and the second part is “ower”. The present invention does not particularly limit the string lengths of the first part or the second part.
[0068] In step S1024, the arithmetic device 3 sets the weight to a second value, and then goes to step S103.
[0069] In step S1025, the arithmetic device 3 sets the weight to a third value and then proceeds to step S103. In one embodiment, the first value is less than the second value, and the second value is less than the third value. For example, the first value is 1.2, the second value is 1.5, and the third value is 1.8.
[0070] Illustrated with a practical example Figure 7 The process is as follows: Assume the candidate substring is "Power supply voltage", its first part is Power, and the second part is supply voltage; and assume the fuzzy comparison keyword database 16 stores four fuzzy comparison keywords: V, I, P, Power. Then the processes of steps S1021 to S1025 will be executed four times. The weight set for the first time is the second value because although the candidate substring Power supply voltage contains the fuzzy keyword V, the fuzzy keyword V is not in the first part Power. The weight set for the second time is the first value because the candidate substring Power supply voltage does not contain the fuzzy keyword I. The weight set for the third time is the third value because the candidate substring Power supply voltage contains the fuzzy keyword P and the fuzzy keyword P is in the first part Power. The weight set for the fourth time is the third value because the candidate substring Power supply voltage contains Power and the fuzzy keyword Power is in the first part Power. Finally, the arithmetic device 3 adopts the maximum value among these four weights, that is, the third value, as the weight described in step S102.
[0071] In step S103, the arithmetic device 3 calculates the third similarity between the candidate substring and each artificial intelligence keyword in the artificial intelligence database 18. The calculation method of the third similarity can refer to the calculation methods of the first similarity and the second similarity in steps S3 and S8. In one embodiment, multiple artificial intelligence keywords in the artificial intelligence database 18 are pre-established by the user, and then through a program or script, the first part of each artificial intelligence keyword is extracted and stored as a fuzzy comparison keyword in the fuzzy comparison keyword database 16.
[0072] In one embodiment, the number of executions of step S103 is the same as the number of artificial intelligence keywords. Therefore, multiple executions of step S103 will generate multiple third similarities, and the maximum value among these third similarities is subsequently used by the arithmetic device 3. The following is an illustration of step S103 with a practical example: Assume that the candidate substring is "Power supply voltage", and assume that the artificial intelligence keyword database 18 stores seven artificial intelligence keywords: "VLED", "ILED", "PLED", "VGG", "IGG", "PGG", "Power supply current". Then step S103 will be executed seven times. The third similarity between the artificial intelligence keyword "VLED" and the candidate substring "Power supply voltage" is calculated for the first time, the third similarity between the artificial intelligence keyword "ILED" and the candidate substring "Power supply voltage" is calculated for the second time... Eventually, seven third similarities are generated. It can be easily seen that the artificial intelligence keyword "Power supply current" is the most similar to the candidate substring "Power supply voltage" literally. Therefore, step S103 outputs the third similarity corresponding to the artificial intelligence keyword "Power supply current".
[0073] In step S104, the arithmetic device 3 calculates a comparison score based on the third similarity and the weight. In one embodiment, the comparison score is the product of the third similarity and the weight.
[0074] In step S105, the arithmetic device 3 determines whether the comparison score is greater than the third threshold. If the determination in step S105 is yes, it means that this candidate substring is the correct keyword, and the process will proceed to step S106 subsequently. If the determination in step S105 is no, it means that this candidate substring is an incorrect keyword, and this candidate substring will be discarded, and the process will return to step S101 to select the next candidate substring from the second string and repeat the process of steps S102 to S106. In one embodiment, the third threshold is, for example, 0.95, but the present invention does not limit the numerical setting of the third threshold.
[0075] In step S106, the arithmetic device 3 outputs one of the multiple value columns corresponding to the candidate substring. Specifically, in steps S2 and S7, in addition to obtaining the text in the text column of the table as the second string, the arithmetic device 3 also obtains the numerical records in the value column of the table and records the corresponding relationship between these texts and the values. Therefore, whenever the arithmetic device 3 determines a correct keyword in step S106, it can find the value corresponding to the correct keyword in the numerical records and output it according to this correct keyword and the aforementioned corresponding relationship.
[0076] In one embodiment, Figure 6The process shown can be repeatedly executed multiple times, and output numerical data that was not previously recorded in the exact document. The following is an actual example: Given the relationship V = I × P among voltage V, current I, and power P, this relationship is stored in storage device 1 in the form of an instruction. Assume that the artificial intelligence keywords stored in artificial intelligence keyword database 18 include VGG, IGG, and PGG, and the candidate substring only contains VCC and ICC but not PCC. Based on the above assumptions, after the process shown in Figure 6 is continuously executed twice, the voltage value corresponding to VCC and the current value corresponding to ICC can be output in step S106 (finding the text column VCC in the exact table through the artificial intelligence keyword VGG, and finding the text column ICC in the exact table through the artificial intelligence keyword IGG), and then the arithmetic device 3 further calculates the power value of PCC through the pre-stored instruction (V = I × P) and outputs it (PCC = VCC × ICC).
[0077] Although the data that the user wants to search for may correspond to different keywords in different documents, the differences between these keywords are not significant. For example: The voltage keyword recorded in document A is VCC, and the voltage keyword recorded in document B is VGG. Therefore, for this kind of difference, only keyword collection needs to be carried out in the initial stage, and the collected keywords are stored in the artificial intelligence keyword database 18. Even if new artificial intelligence keywords are not continuously collected later, as long as the keywords corresponding to the data to be searched do not change too much, through Figure 6 the process shown, the user can still find the required data. For example, although artificial intelligence keywords such as VCC and ICC are not stored in the artificial intelligence database 18. But as long as similar keywords such as VLED, ILED, VGG, or IGG are stored in the artificial intelligence database 18, through Figure 6 and Figure 7 the process shown, it is still possible to find the data containing VCC or ICC from the document. In contrast, the existing technology that uses a "rule-dominated" document parsing mechanism requires continuous algorithm updates. The method adopted by the present invention can greatly reduce the time for modifying the document parsing program and also reduce the time for the user to regularly update the keyword database.
[0078] In summary, the hierarchical document parsing system and method proposed by the present invention have the following effects:
[0079] 1. Based on artificial intelligence technology, it can quickly find keyword values in professional documents and solve the problem that engineers do not know which keywords to use, thereby reducing the overall search time.
[0080] 2. Compared with the "rule-based" fixed-rule document parsing method, for new document parsing requirements, the present invention does not require defining new search rules, which can save the time of developers and reduce labor costs.
[0081] 3. The present invention performs data screening based on open-source artificial intelligence models and hierarchical concepts, reducing the time for independently training artificial intelligence models and accelerating the development speed.
[0082] 4. The design architecture of the present invention belongs to the hierarchical concept, so it is quite flexible. The models or algorithms used in each layer can be replaced, so it will not be bound by any package or manufacturer, avoiding the problem that the system cannot operate in the future due to a certain method or package no longer being supported or updated.
[0083] Overall, the advantages of the present invention are that it does not require developers to repeatedly develop, does not require repeated training of artificial intelligence models, is not bound by any method or package, simplifies the original complicated steps, and improves the convenience of user operation.
[0084] Although the present invention is disclosed as above in the foregoing embodiments, it is not intended to limit the present invention. Any changes and modifications made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention. For the scope of protection defined by the present invention, please refer to the appended claims.
Claims
1. A hierarchical file parsing method, including performing by an arithmetic device: Obtain a reference string; Obtain a plurality of first strings in a plurality of files through a first module, wherein each of the files includes a table, the table includes a plurality of text columns and a plurality of numerical columns, and the first strings are associated with the text columns; Calculate a first similarity between the reference string and each of the first strings; Screen out a plurality of candidate files from the files, wherein the first similarity of each of the candidate files is greater than a first threshold; Obtain a plurality of second strings in the candidate files through a second module, wherein the second strings are associated with the text columns, and the number of the second strings is less than or equal to the number of the first strings; Calculate a second similarity between the reference string and each of the second strings; and When the second similarity is greater than a second threshold, determine and output one of the numerical columns according to each of the second strings.
2. The hierarchical file parsing method according to claim 1, wherein each of the second strings includes a plurality of sub-strings, and determining and outputting one of the numerical columns according to each of the second strings includes: Select a candidate sub-string from the sub-strings; Compare the candidate sub-string with a plurality of fuzzy keywords in a fuzzy comparison keyword database to determine the weight of the candidate sub-string; Calculate a third similarity between the candidate sub-string and each artificial intelligence keyword in an artificial intelligence database; Calculate a comparison score according to the third similarity and the weight; and When the comparison score is greater than a third threshold, output one of the numerical columns corresponding to the candidate sub-string.
3. The hierarchical file parsing method according to claim 2, wherein comparing the candidate sub-string with the fuzzy keywords in the fuzzy comparison keyword database to determine the weight of the candidate sub-string includes: Generate a first judgment based on whether the candidate sub-string contains one of the fuzzy sub-strings, wherein the candidate sub-string is composed of a first part and a second part; Generate a second judgment based on whether one of the fuzzy keywords is located in the first part; When the first judgment is "no", set the weight to a first value; When the first judgment is "yes" and the second judgment is "no", set the weight to a second value; and When the first judgment is "yes" and the second judgment is "yes", set the weight to a third value; wherein the first value is less than the second value, and the second value is less than the third value.
4. The hierarchical file parsing method according to claim 1, wherein the second threshold is greater than the first threshold.
5. The hierarchical file parsing method according to claim 1, wherein the first module operated by the arithmetic device is an open-source table parsing tool.
6. The hierarchical file parsing method according to claim 5, wherein the second module operated by the arithmetic device is a cloud file table parsing module.
7. The hierarchical file parsing method as described in claim 6, wherein the recognition accuracy of the cloud file table parsing module for the second strings is greater than the accuracy of the open-source table parsing tool for the first strings.
8. The hierarchical file parsing method as described in claim 1, further comprises: When the first similarity is not greater than the first threshold, reduce the first threshold and recalculate the first similarity between the reference string and each of the first strings.
9. The hierarchical file parsing method as described in claim 1, wherein obtaining the first strings in the files through the first module comprises: Convert the files from picture format to text format; Distinguish the text data and numerical data in the text format; and Save the text data as the first strings.
10. The hierarchical file parsing method as described in claim 1, wherein obtaining the second strings in the candidate files through the second module comprises: Convert the candidate files from picture format to text format; Distinguish the text data and numerical data in the text format; and Save the text data as the second strings.
11. A hierarchical file parsing system, comprises: A storage device for storing a plurality of instructions, a plurality of reference strings, and a plurality of files; An arithmetic device electrically connected to the storage device and used to execute the instructions to cause a plurality of operations, the operations including: Obtain one of the reference strings; Obtain a plurality of first strings in the files, wherein each of the files includes a table, the table includes a plurality of text columns and a plurality of numerical columns, and the first strings are associated with the text columns; Calculate the first similarity between the reference string and each of the first strings; Screen out a plurality of candidate files from the files, wherein the first similarity of each of the candidate files is greater than a first threshold; Obtain a plurality of second strings in the candidate files through a second module, wherein the second strings are associated with the text columns; Calculate the second similarity between the reference string and each of the second strings; and When the second similarity is greater than a second threshold, determine one of the numerical columns according to each of the second strings and output.
12. The hierarchical file parsing system as described in claim 11, wherein each of the second strings includes a plurality of sub-strings, and determining one of the numerical columns according to each of the second strings and outputting comprises: Select a candidate sub-string from the sub-strings; Compare the candidate sub-string with a plurality of fuzzy keywords in a fuzzy comparison keyword database to determine the weight of the candidate sub-string; Calculate the third similarity between the candidate sub-string and each artificial intelligence keyword in an artificial intelligence database; Calculate a comparison score according to the third similarity and the weight; and When the comparison score is greater than a third threshold, output one of the numerical columns corresponding to the candidate sub-string.
13. The hierarchical file parsing system as described in claim 12, wherein comparing the candidate sub-string with the fuzzy keywords in the fuzzy comparison keyword database to determine the weight of the candidate sub-string comprises: Generate a first determination based on whether the candidate substring contains one of the fuzzy substrings, where the candidate substring consists of a first part and a second part; Generate a second determination based on whether one of the fuzzy keywords is located in the first part; When the first determination is "no", set the weight to a first value; When the first determination is "yes" and the second determination is "no", set the weight to a second value; and When the first determination is "yes" and the second determination is "yes", set the weight to a third value; where the first value is less than the second value, and the second value is less than the third value.
14. The hierarchical document parsing system according to claim 11, wherein the second threshold is greater than the first threshold.
15. The hierarchical document parsing system according to claim 11, wherein the first module operated by the computing device is an open-source table parsing tool.
16. The hierarchical document parsing system according to claim 11, wherein the second module operated by the computing device is a cloud document table parsing module.
17. The hierarchical document parsing system according to claim 11, wherein the recognition accuracy of the cloud document table parsing module for the second strings is greater than the accuracy of the open-source table parsing tool for the first strings.
18. The hierarchical document parsing system according to claim 11, further comprising: When the first similarity is not greater than the first threshold, reduce the first threshold and recalculate the first similarity between the reference string and each of the first strings.
19. The hierarchical document parsing system according to claim 11, wherein obtaining the first strings in the files through the first module comprises: Convert the files from a picture format to a text format; Distinguish the text data and numerical data in the text format; and Save the text data as the first strings.
20. The hierarchical document parsing system according to claim 11, wherein obtaining the second strings in the candidate files through the second module comprises: Convert the candidate files from a picture format to a text format; Distinguish the text data and numerical data in the text format; and Save the text data as the second strings.