Fast grouping method and apparatus based on same group of financial profiles
By using distributed sampling and cyclic processing methods, garbled text in financial records is cleaned up, solving the problem of incorrect grouping in existing technologies. This achieves efficient and reliable financial record grouping, improving the accuracy and efficiency of financial processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 国网河北省电力有限公司高邑县供电分公司
- Filing Date
- 2023-07-11
- Publication Date
- 2026-05-19
AI Technical Summary
Existing financial record grouping methods suffer from poor grouping accuracy and high resource consumption when faced with garbled text interference, resulting in poor grouping performance and affecting subsequent financial record services and parsing processes.
By employing distributed sampling, arbitrary sampling, and cyclic processing methods, images of financial documents are extracted using scanners and OCR programs. Improved equations and cyclic processing are then used to clean up garbled text and group it, ensuring the reliability and accuracy of the grouping.
It achieves efficient and reliable financial file grouping, reduces garbled text interference, improves the accuracy and efficiency of grouping, and ensures the smooth progress of subsequent financial processing.
Smart Images

Figure CN117011869B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of rapid grouping technology for financial files of the same category, and specifically relates to a rapid grouping method and apparatus based on financial files of the same category. Background Technology
[0002] Financial records include accounting vouchers, cash journals, bank journals, subsidiary ledgers, general ledgers, financial statements including balance sheets, profit and loss statements, cash flow statements, monthly tax filings, invoices, inventory records, inventory lists, fixed asset ledgers, low-value consumables ledgers, bank statements, and balance sheets.
[0003] Financial records contain archival text, which includes financial statements and figures. Many financial records are paper documents, requiring scanning to extract images, which are then transmitted to the financial platform where OCR programs extract the relevant archival text. The archival text is then grouped, often using intelligent processing methods. Grouping parsing is based on the archival attributes of the text, grouping text objects to ensure that objects within a group are as similar as possible, while objects between groups are as different as possible. Common grouping methods include Makeshift grouping, DBSCAN grouping, and density-based grouping.
[0004] The most widely used MeanShift method aims to find l texts within a defined measurement area, grouping the document text into l groups, minimizing the span of the highest-ranking group. Within the measurement area, the minimum threshold for group similarity is defined as 2. However, in practical applications, document text grouping often suffers from garbled text interference, such as garbled characters generated during document scanning and OCR extraction. The MeanShift method is highly sensitive to garbled text, and this interference can significantly hinder the final parsing of the grouped text. Therefore, the goal is to eliminate garbled text during the grouping process; this is the essence of the garbled text grouping strategy.
[0005] Currently, while there are corresponding distributed grouping methods for grouping archive text with garbled characters, the grouping accuracy of the current methods is poor, and the resources used are quite cumbersome, resulting in poor performance in practical applications.
[0006] Therefore, the financial file grouping method based on the grouping of file text with garbled characters has also suffered from considerable side effects. Currently, due to the security and cumbersome defects of file text with garbled characters during the grouping process, the financial file grouping method based on the grouping of file text with garbled characters also has significant defects in practical application. This will result in incorrect file text in the grouped financial files, which will be detrimental to the subsequent financial file service transmission and financial file text parsing processes, and thus further hinder the processing of financial files. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention proposes a rapid grouping method and apparatus for financial documents based on the same category. A scanner scans the document image of the paper financial documents and transmits it to the financial platform. The financial platform uses an OCR program to extract the document text from the document image, stores the document text in a document text library, and then groups the document text. Through methods of distributed sampling, arbitrary sampling, and cyclic processing, not only is grouping of document text with garbled characters achieved, but this invention also boasts high reliability, guaranteed accuracy, and excellent efficiency.
[0008] The present invention employs the following technical solution.
[0009] A method for rapid grouping based on the same category of financial records includes:
[0010] Step 1: The scanner scans the paper financial documents and then transmits the images to the financial platform;
[0011] Step 2: The financial platform uses an OCR program to extract the archival text from the document image, stores the archival text in the archival text library, and then groups the archival text.
[0012] The specific methods for grouping include:
[0013] Step 2-1: Obtain the text library of files containing garbled characters to be grouped;
[0014] Step 2-2: Split the obtained archive text library and store it in a distributed manner;
[0015] Steps 2-3: On each of the distributed modules, each module performs arbitrary sampling of the archive text it stores, and uses the sampled archive text as its Chinese text library. At this time, the entire archive text library is also used as the sampling library.
[0016] Preferably, steps 2-3 specifically include:
[0017] The following equation is used as the improved objective equation:
[0018] ZD q∈Y (ZX 1≤k≤Lf(q,D k ))
[0019] In the equation, Y represents the sub-database of the obtained archive text library after removing garbled text, and Y = O^A, where O is all archive text items in the obtained archive text library, A is the text item of the removed garbled text, ^ is the symbol for removing text items, |A| = a, where a is a defined index representing the maximum critical number of garbled text items to be removed; q is the archive text within text item Y; and text item Y is divided into L groups, namely {D1, D2...D...} k ...D L}, D k It is the Chinese text of the k-th Chinese text library; f(q,D) k ) is the Chinese text D of the archive text q to the kth Chinese text library. k The interval; the number of randomly sampled archival texts is set to... Here, φ and ι are both defined indices, and ZD is relative to ZX. 1≤k≤l f(q,D k Take the highest value, ZX is the value of f(q,D) k Take the minimum amount.
[0020] Preferably, after steps 2-3, the method further includes:
[0021] Steps 2-4: In each module, perform cyclic processing on the file text library: In each iteration, randomly sample multiple file texts, and perform resampling within the sampled file texts, and add the resampled file texts to the middle text library. Then, cover the file texts in the defined range of the middle text in the middle text library, and clear the covered file texts from the sampling waiting library; after the loop is completed, obtain the final middle text library;
[0022] Steps 2-5: Obtain the Chinese language library from each module, construct the object with importance, and send the file text to the main module;
[0023] Steps 2-6: Perform garbled character grouping based on importance conditions on the main module to obtain the final multiple Chinese characters;
[0024] Step 2-7: Assign each file text in the file text library to the multiple Chinese texts obtained in Step 2-6, and remove the multiple file texts with the highest interval from the file text, thus achieving grouping of the file text library with garbled characters based on arbitrary sampling.
[0025] Preferably, steps 2-4 specifically include:
[0026] Step 2-4-1: Based on the capacity of the uncovered archive text library, using the principle of distributed sampling, randomly select multiple archive texts from the current sampling pool to obtain one arbitrary archive text.
[0027] Step 2-4-2: Next, continue to select multiple files from arbitrary file text one to obtain arbitrary file text two;
[0028] Step 2-4-3: Add any file text 2 to the current Chinese text library, and use the refreshed Chinese text library as the current Chinese text library;
[0029] Step 2-4-4: In the current Chinese text library, find the file texts that are within the defined range from the Chinese text, register them, and clear the registered file texts in the sampled pending text item;
[0030] Repeat steps 2-4-1 to 2-4-4 multiple times to finally obtain the Chinese language library.
[0031] Preferably, steps 2-4 may further include:
[0032] (1) Within the current loop, determine the capacity of the uncovered archive text library:
[0033] (2) If the number of file texts in the uncovered file text library exceeds the defined quantity (1+φ)*a, then arbitrarily select (1+φ)*a file texts from the current sampling pool as arbitrary file text one; then, continue to arbitrarily select from arbitrary file text one. Treat one file text as arbitrary file text two; add arbitrary file text two to the current text library;
[0034] (3) If the number of file texts in the uncovered file text library is not less than the defined quantity, then find a natural number s that satisfies {(1+φ)φsa} / n<|V|<{(1+φ)φ(s+1)a} / n; then, arbitrarily select from the current sampled library. Each file text is treated as arbitrary file text one; further, within arbitrary file text one, any selection can be made. Treat one file text as arbitrary file text two; add arbitrary file text two to the current Chinese text library; here, φ and ι are defined indicators, a is the number of garbled characters, |V| is the number of file texts in the current sampled library V, and n is the number of modules;
[0035] (4) After adding any file text 2 to the current Chinese text library, find the interval between the file text and the Chinese text text in the current Chinese text library that is within the interval S. pquThe archive text is registered, and the registered archive text is cleared from the current sampled pending text item; S pqu These are defined metrics;
[0036] Repeat steps (1) to (4) a total of χ*l times to finally obtain the Chinese text library; χ is a defined constant index higher than one; l is the number of Chinese texts to be extracted.
[0037] Preferably, steps 2-5 specifically include:
[0038] Use the Chinese language library Each pending text file is treated as a text file, and all file text files are grouped into the pending text file with the lowest distance from it; the importance of each text file is the number of file text files assigned to that text file. It is an operational equation and l is the number of Chinese characters to be retrieved, μ is a defined floating-point number, and the importance object is the Chinese character and its importance.
[0039] Preferably, steps 2-6 specifically include:
[0040] Using a reinforced loop, finally select l Chinese texts;
[0041] During the enhanced cycle, each selection interval is within the set range of 2*S. pqu The file text with the highest total importance value is taken as the median text; within the importance object, remove files that use this median text as their median text and have a distance of 4*S from it. pqu All archival texts covered in it; S pqu It is a defined indicator.
[0042] Preferably, the method for clearing the multiple file texts with the highest distance from the file text interval includes:
[0043] Clear the {1+φ}*a file texts with the highest interval from the current file text, where a is the number of garbled characters and φ is a defined index.
[0044] A rapid grouping device based on the same category of financial records includes:
[0045] A scanner connected to a financial platform, which includes an OCR program; the scanner scans paper financial documents and transmits the images to the financial platform, where the OCR program extracts the text from the document images and then groups the text; the unit running on the financial platform includes:
[0046] The acquisition unit is used to acquire the archive text library containing garbled characters to be grouped;
[0047] The cutting unit is used to cut the acquired archive text library and perform distributed storage.
[0048] The sampling unit is used to perform arbitrary sampling of the archive text stored in each distributed module, and to use the sampled archive text as the Chinese text library. At the same time, the entire archive text library is also used as the sampling library.
[0049] Preferably, the module running on the financial platform further includes:
[0050] The loop unit is used to perform loop processing on the file text library in each module: in each loop, multiple file texts are sampled arbitrarily, and the sampled file texts are sampled again. The sampled file texts are added to the middle text library. Then, the file texts in the defined range of the middle text library are covered, and the covered file texts are cleared from the sampling waiting library. After the loop is completed, the final middle text library is obtained.
[0051] The construction unit is used to obtain the Chinese language library on each module, construct the importance object, and transmit the file text to the main module.
[0052] Importance unit, which is used to perform garbled character grouping under importance conditions on the overall module to obtain the final multiple Chinese characters;
[0053] The clearing unit is used to assign each file text in the file text library to multiple middle texts obtained in steps 2-6, and to clear multiple file texts with the highest interval from the file text, thereby achieving grouping based on the arbitrarily sampled file text library with garbled characters.
[0054] The beneficial effects of this invention are that, compared with the prior art, the scanner of this invention scans the document image of paper financial files and then transmits it to the financial platform; the financial platform uses an OCR program to extract the file text in the document image, stores the file text in the file text library, and then groups the file text; through the methods of distributed sampling, arbitrary sampling and cyclic processing, not only is the grouping of file text with garbled characters achieved, but this invention also has high reliability, ensures accuracy and high efficiency. Attached Figure Description
[0055] Figure 1 This is a flowchart of steps 1 to 2 as described in this invention;
[0056] Figure 2 This is a partial component structure diagram of the rapid grouping device based on the same group of financial files described in this invention;
[0057] Figure 3This is a flowchart of steps 2-1 to 2-3 as described in this invention;
[0058] Figure 4 This is a flowchart of steps 2-3 to 2-7 as described in this invention;
[0059] Figure 5 This is a flowchart of steps 2-4-1 to 2-4-4 as described in this invention. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, any other embodiments obtained by those skilled in the art without creative effort are all within the protection scope of this invention.
[0061] like Figure 1 As shown, the present invention provides a method for rapid grouping of financial records based on the same category, comprising:
[0062] Step 1: The scanner scans the paper financial documents and then transmits the images to the financial platform;
[0063] Step 2: The financial platform uses an OCR program to extract the archival text from the document image (all the archival text in a document image is an archival text item), stores the archival text in the archival text library, and then groups the archival text. The archival text may contain garbled characters.
[0064] like Figure 3 As shown, the specific methods for grouping include:
[0065] Step 2-1: Obtain the text library of files containing garbled characters to be grouped;
[0066] Step 2-2: Split the archive text library obtained in Step 2-1 and store it in a distributed manner;
[0067] Steps 2-3: On each of the distributed modules (a module can be a process with a storage area for storing file text), each module performs arbitrary sampling of the file text it stores, and uses the sampled file text as its Chinese text library. At this time, the entire file text library is also used as the sampling library.
[0068] In a preferred but non-limiting embodiment of the present invention, steps 2-3 specifically include:
[0069] The following equation is used as the improved objective equation:
[0070] ZDq∈Y (ZX 1≤k≤L f(q,D k ))
[0071] In the equation, Y represents the sub-database of the obtained archive text library after removing garbled text, and Y = O^A, where O is all archive text items in the archive text library obtained in step 2-1, A is the text item of the removed garbled text, ^ is the symbol for clearing text items, |A| = a, where a is a defined index representing the highest critical number of garbled text items to be cleared; q is the archive text within text item Y; and text item Y is divided into L groups, namely {D1, D2...D...} k ...D L}, D k It is the Chinese text of the k-th Chinese text library; f(q,D) k ) is the Chinese text D of the archive text q to the kth Chinese text library. k The interval (the interval is the quantity obtained by subtracting the absolute value of the Pearson coefficient of the two texts to which the interval belongs); the number of randomly sampled archival texts is set to Here, φ and ι are both defined indices, and ZD is relative to ZX. 1≤k≤l f(q,D k Take the highest value, ZX is the value of f(q,D) k The minimum value is taken; the objective equation is used to improve the highest group span, and can find the L groups with high correlation to perform grouping on the file text, and find the file text of distant groups as garbled text to be cleared.
[0072] Through arbitrary sampling in steps 2-3, the probability of having at least one correct file text that is not garbled text is ; at the same time, the side effect of cleaning up a garbled texts during the improvement of the objective equation is .
[0073] In preferred but non-limiting embodiments of the present invention, such as Figure 4 As shown, after steps 2-3, the following is also included:
[0074] Steps 2-4: In each module, perform cyclic processing on the file text library: In each iteration, randomly sample multiple file texts, and perform resampling within the sampled file texts, and add the resampled file texts to the middle text library. Then, cover the file texts in the defined range of the middle text in the middle text library, and clear the covered file texts from the sampling waiting library; after the loop is completed, obtain the final middle text library;
[0075] In preferred but non-limiting embodiments of the present invention, such as Figure 5 As shown, steps 2-4 specifically include:
[0076] Step 2-4-1: Based on the capacity of the uncovered archive text library, using the principle of distributed sampling, randomly select multiple archive texts from the current sampling pool to obtain one arbitrary archive text.
[0077] Step 2-4-2: Next, continue to select multiple files from arbitrary file text one to obtain arbitrary file text two;
[0078] Step 2-4-3: Add any file text 2 to the current Chinese text library, and use the refreshed Chinese text library as the current Chinese text library;
[0079] Step 2-4-4: In the current Chinese text library, find the file texts that are within the defined range from the Chinese text, register them, and clear the registered file texts in the sampled pending text item;
[0080] Repeat steps 2-4-1 to 2-4-4 multiple times to finally obtain the Chinese language library.
[0081] Using the principle of distributed sampling, at least one non-garbled file text is found in each iteration. The grouped text can ensure a high degree of similarity on the sub-machine. In each iteration, the interval between each file text needs to be calculated again.
[0082] In a preferred but non-limiting embodiment of the present invention, steps 2-4 may further include:
[0083] (1) Within the current loop, determine the capacity of the uncovered archive text library:
[0084] (2) If the number of file texts in the uncovered file text library exceeds the defined quantity (1+φ)*a, then arbitrarily select (1+φ)*a file texts from the current sampling pool as arbitrary file text one; then, continue to arbitrarily select from arbitrary file text one. Treat one file text as arbitrary file text two; add arbitrary file text two to the current text library;
[0085] (3) If the number of file texts in the uncovered file text library is not less than the defined quantity, then find a natural number s that satisfies {(1+φ)φsa} / n<|V|<{(1+φ)φ(s+1)a} / n; then, arbitrarily select from the current sampled library. Each file text is treated as arbitrary file text one; further, within arbitrary file text one, any selection can be made. Treat one file text as arbitrary file text two; add arbitrary file text two to the current Chinese text library; here, φ and ι are defined indicators, a is the number of garbled characters, |V| is the number of file texts in the current sampled library V, and n is the number of modules;
[0086] (4) After adding any file text 2 to the current Chinese text library, find the interval between the file text and the Chinese text text in the current Chinese text library that is within the interval S. pqu The archive text is registered, and the registered archive text is cleared from the current sampled pending text item; S pqu These are defined metrics;
[0087] Repeat steps (1) to (4) a total of χ*l times to finally obtain the Chinese text library; χ is a defined constant index higher than one, used to control the grouping quality; the higher the value of χ, the better the grouping quality, but the longer the time required; l is the number of Chinese texts to be retrieved.
[0088] Steps 2-5: Obtain the Chinese language library from each module, construct the object with importance, and send the file text to the main module;
[0089] In a preferred but non-limiting embodiment of the present invention, steps 2-5 specifically include:
[0090] Use the Chinese language library Each pending text file is treated as a text file, and all file text files are grouped into the pending text file with the lowest distance from it; the importance of each text file is the number of file text files assigned to that text file. It is an operational equation and l is the number of Chinese characters to be retrieved, μ is a defined floating-point number higher than the defined value (that is, a floating-point number with a sufficiently high value), and the importance object is the Chinese character and its importance.
[0091] Steps 2-6: Perform garbled character grouping based on importance conditions on the main module to obtain the final multiple Chinese characters;
[0092] In a preferred but non-limiting embodiment of the present invention, steps 2-6 specifically include:
[0093] Using a reinforced loop, finally select l Chinese texts;
[0094] During the enhanced cycle, each selection interval is within the set range of 2*S. pqu The file text with the highest total importance value is taken as the median text; within the importance object, remove files that use this median text as their median text and have a distance of 4*S from it. pqu All archival texts covered in it; S pqu It is a defined indicator.
[0095] Step 2-7: Assign each file text in the file text library to the multiple Chinese texts obtained in Step 2-6, and remove the multiple file texts with the highest interval from the file text, thus achieving grouping of the file text library with garbled characters based on arbitrary sampling.
[0096] In a preferred but non-limiting embodiment of the present invention, a method for removing the plurality of file texts with the highest distance from the file text interval includes:
[0097] Clear the {1+φ}*a file texts with the highest interval from the current file text, where a is the number of garbled characters and φ is a defined index.
[0098] The time consumption of the grouping method of the present invention is positively correlated with ι. Correctly extracting l Chinese texts often yields a high probability {1-ι}) of obtaining file text groups with high similarity. The number of garbled texts cleared is {1+φ}*a. The resource consumption of the grouping method of the present invention is positively correlated with n and l.
[0099] In practical applications, the grouping method of this invention exhibits high stability and can efficiently achieve the process of grouping archive texts, demonstrating excellent efficiency.
[0100] like Figure 2 As shown, the present invention provides a rapid grouping device based on the same category of financial files, comprising:
[0101] A scanner is connected to a financial platform, which includes an OCR program; the financial platform can be a computer or a server. The scanner scans paper financial documents and transmits the images to the financial platform, where the OCR program extracts the text from the document images and then groups the text.
[0102] The units running on the financial platform include:
[0103] The acquisition unit is used to acquire the archive text library containing garbled characters to be grouped;
[0104] The cutting unit is used to cut the acquired archive text library and perform distributed storage.
[0105] The sampling unit is used to perform arbitrary sampling of the archive text stored in each distributed module, and to use the sampled archive text as the Chinese text library. At the same time, the entire archive text library is also used as the sampling library.
[0106] In a preferred but non-limiting embodiment of the present invention, the module running on the financial platform further includes:
[0107] The loop unit is used to perform loop processing on the file text library in each module: in each loop, multiple file texts are sampled arbitrarily, and the sampled file texts are sampled again. The sampled file texts are added to the middle text library. Then, the file texts in the defined range of the middle text library are covered, and the covered file texts are cleared from the sampling waiting library. After the loop is completed, the final middle text library is obtained.
[0108] The construction unit is used to obtain the Chinese language library on each module, construct the importance object, and transmit the file text to the main module.
[0109] Importance unit, which is used to perform garbled character grouping under importance conditions on the overall module to obtain the final multiple Chinese characters;
[0110] The clearing unit is used to assign each file text in the file text library to multiple middle texts obtained in steps 2-6, and to clear multiple file texts with the highest interval from the file text, thereby achieving grouping based on the arbitrarily sampled file text library with garbled characters.
[0111] The beneficial effects of this invention are that, compared with the prior art, the scanner of this invention scans the document image of paper financial files and then transmits it to the financial platform; the financial platform uses an OCR program to extract the file text in the document image, stores the file text in the file text library, and then groups the file text; through the methods of distributed sampling, arbitrary sampling and cyclic processing, not only is the grouping of file text with garbled characters achieved, but this invention also has high reliability, ensures accuracy and high efficiency.
[0112] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable appendix having computer-readable program instructions loaded thereon for causing a processor to achieve each aspect disclosed herein.
[0113] Computer-readable printed media can be tangible printed media capable of holding and printing instructions executed by a circuit. Computer-readable printed media can be—but is not limited to—electrical printed media, magnetic printed media, optical printed media, electromagnetic printed media, semiconductor printed media, or any suitable combination thereof. Further examples of computer-readable printed media (a non-exhaustive list) include: portable computer disks, hard disks, random access memory (RAM), read-only memory (RyM), erasable programmable read-only memory (EPRyM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (HD-RyM), digital multipurpose disk (DXD), memory sticks, floppy disks, mechanically encoded printed media, such as punch cards or recessed protrusions with instructions printed on them, or any suitable combination thereof. The computer-readable annotated medium used herein is not to be interpreted as the instantaneous message itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (like light pulses through power transmission cables), or electrical messages transmitted through wires.
[0114] The computer-readable program instructions expressed herein can be downloaded from computer-readable supplementary media to each computing / processing power line, or downloaded via a wireless network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external supplementary power line. The wireless network can include copper transmission cables, transmission line transmissions, wireless transmissions, routers, firewalls, switches, Wi-Fi device computers, and / or edge servers. A wireless network adapter card or wireless network port in each computing / processing power line receives the computer-readable program instructions from the wireless network and forwards the computer-readable program instructions to the computer-readable supplementary media stored in each computing / processing power line.
[0115] The computer program instructions used to execute the operations of this disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-associative instructions, microcode, firmware instructions, condition-defined values, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as ScalarQAL, H++, etc., and conventional procedural programming languages such as H' language or similar programming languages. The computer-readable program instructions can be executed entirely on a client computer, partially on a client computer, executed as a single software package, or partially on a client computer and partially on a remote computing facility. Execution can be performed on-site or entirely on a remote computer or server. In the form involving a remote computer, the remote computer can connect to the client computer via any type of wireless network—including a local area network (LAb) or a wide area network (UAb)—or can connect to an external computer (such as using an Internet service provider to connect via the Internet). In some embodiments, electronic circuitry is customized using operating condition values of computer-readable program instructions, such as programmable logic circuits, field-programmable gate arrays (processing platforms), or programmable logic arrays (PLAs), which can execute computer-readable program instructions to achieve every aspect of cost disclosure.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent updates can still be made to the specific embodiments of the present invention without departing from the spirit and scope of the present invention, and any modifications or equivalent updates should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for rapid grouping of financial records based on the same category, characterized in that, include: Step 1: The scanner scans the paper financial documents and then transmits the images to the financial platform; Step 2: The financial platform uses an OCR program to extract the archival text from the document image, stores the archival text in the archival text library, and then groups the archival text. The specific methods for grouping include: Step 2-1: Obtain the text library of files containing garbled characters to be grouped; Step 2-2: Split the obtained archive text library and store it in a distributed manner; Steps 2-3: On each of the distributed modules, each module performs arbitrary sampling of the archive text it stores, and uses the sampled archive text as its Chinese text library. At this time, the entire archive text library is also used as the sampling library. Steps 2-3 specifically include: The following equation is used as the improved objective equation: Within the equation The garbled text sub-database within the obtained archive text library was cleaned up, and , It refers to all the archival text items obtained from the archival text database. It is the text item that cleared out the garbled text. It is a symbol for clearing text items. , It is a defined metric that represents the maximum critical number of garbled texts to be removed. It is a text item The file text within; put the text items Divided into The teams are: , The first one taken Chinese text in a Chinese text library; It is an archival text To the Chinese text in the Chinese text library The interval; the number of randomly sampled archival texts is set to... ,here and These are all defined metrics. Yes Take the highest amount. Yes Take the lowest amount; After steps 2-3, it also includes: Steps 2-4: In each module, perform cyclic processing on the file text library: In each iteration, randomly sample multiple file texts, and perform resampling within the sampled file texts, and add the resampled file texts to the middle text library. Then, cover the file texts in the defined range of the middle text in the middle text library, and clear the covered file texts from the sampling waiting library; after the loop is completed, obtain the final middle text library; Steps 2-5: Obtain the Chinese language library from each module, construct the object with importance, and send the file text to the main module; Steps 2-6: Perform garbled character grouping based on importance conditions on the main module to obtain the final multiple Chinese characters; Step 2-7: Assign each file text in the file text library to the multiple Chinese texts obtained in Step 2-6, and remove the multiple file texts with the highest interval from the file text, thus achieving grouping of the file text library with garbled characters based on arbitrary sampling. Steps 2-4 specifically include: Step 2-4-1: Based on the capacity of the uncovered archive text library, using the principle of distributed sampling, randomly select multiple archive texts from the current sampling pool to obtain one arbitrary archive text. Step 2-4-2: Next, continue to select multiple files from arbitrary file text one to obtain arbitrary file text two; Step 2-4-3: Add any file text 2 to the current Chinese text library, and use the refreshed Chinese text library as the current Chinese text library; Step 2-4-4: In the current Chinese text library, find the file texts that are within the defined range from the Chinese text, register them, and clear the registered file texts in the sampled pending text item; Repeat steps 2-4-1 to 2-4-4 multiple times to finally obtain the Chinese language library. Steps 2-4 can also specifically include: (1) Within the current loop, determine the capacity of the uncovered archive text library: (2) If the number of archive texts in the uncovered archive text library exceeds the defined quantity. From the current sample pool, any sample can be selected. Each file text is treated as arbitrary file text one; then, within arbitrary file text one, further arbitrary selections are made. Treat one file text as arbitrary file text two; add arbitrary file text two to the current text library; (3) If the number of archival texts in the uncovered archival text database is not less than the defined quantity, then find the natural number. conform to Next, samples are randomly selected from the current sampling pool. Each file text is treated as arbitrary file text one; further, within arbitrary file text one, any selection can be made. Treat one file text as arbitrary file text two; add arbitrary file text two to the current Chinese text library; here. and These are all defined metrics. It's a number of random characters. It is a current sampling library. The number of file texts within. It is the number of modules; (4) After adding any file text 2 to the current Chinese text library, find the file text within the specified interval in the current Chinese text library. The archive text is registered, and the registered archive text is cleared from the current sampled text item. These are defined metrics; Repeat steps (1) to (4) in total. Next, finally obtain the Chinese language library; It is a constant indicator defined as being higher than one; This is the number of Chinese characters to be extracted; Steps 2-5 specifically include: Use the Chinese language library Each pending text file is treated as a text file, and all file text files are grouped into the pending text file with the lowest distance from it; the importance of each text file is the number of file text files assigned to that text file. It is an operational equation and , This refers to the number of Chinese characters to be extracted. It is a defined floating-point number with a higher degree of importance than defined quantity; the importance object is the Chinese text and its importance. Steps 2-6 specifically include: Using a reinforcing cycle approach, the final choice is... The Chinese text; During the enhanced cycle, each selection interval is within the set range. The file text with the highest total importance value is used as the median text; within the importance object, remove files that use this median text as their median text and have a distance from it within a set range. All archival texts included in the document; It is a defined indicator.
2. The rapid grouping method based on the same category of financial records according to claim 1, characterized in that, Methods for clearing multiple file texts with the highest interval from the current file text include: Remove the text with the highest interval from the file. A file text, here It's a number of random characters. It is a defined indicator.