Method and system for analyzing treasurers' water-sensitive words
By constructing a treasury structure tree and using OCR technology to distinguish between standard and cursive fonts, combined with partition ratio and stroke comparison, the problem of cursive signature font recognition was solved, improving the accuracy and efficiency of sensitive word analysis and supporting corporate financial decision-making and compliance operations.
Patent Information
- Application Number
- CN202510555305.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing technologies struggle to effectively identify and analyze cursive signatures, leading to omissions of key information and misjudgments during sensitive word extraction.
By constructing a treasury structure tree and using OCR text recognition technology to distinguish between standard fonts and cursive fonts, and combining partition ratio and stroke comparison methods, suitable filter words are selected to improve the recognition ability of cursive signature fonts.
It improves the ability to extract and identify sensitive words from complex signature-style transaction documents, enhances the accuracy and efficiency of sensitive word analysis, provides more reliable information support, and provides strong support for corporate financial decision-making and compliance operations.
Smart Images

Figure CN120496193B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to data processing technology, and in particular to a method and system for analyzing sensitive words in treasury flow. BACKGROUND
[0002] In the modern enterprise treasury management system, the analysis of sensitive words in treasury flow occupies a crucial position. It can accurately extract key information from massive treasury flow data, providing strong support for financial decision-making, risk control, and compliance operation of enterprises. Whether it is the strict control of financial institutions on the flow of funds or the meticulous management of manufacturing enterprises on cost expenditure, efficient and accurate sensitive word analysis technology in treasury flow is widely used in various scenarios involving financial information processing.
[0003] Currently, the existing technology has obvious shortcomings in dealing with sensitive word analysis in treasury flow bill signature. Most existing methods can only effectively process standard and neat signature fonts. Once faced with a large number of connected signature fonts in reality, they are not up to the task. Lack of effective recognition and analysis means for connected fonts, they cannot accurately match and compare connected fonts with standard fonts, leading to problems such as missing key information and misjudging sensitive words in the sensitive word extraction process.
[0004] Therefore, how to adaptively extract and identify connected signature fonts to improve the adaptability of complex signature style flow bill has become a problem to be solved. SUMMARY
[0005] The present application provides a method and system for analyzing sensitive words in treasury flow, which can adaptively extract and identify connected signature fonts to improve the adaptability of complex signature style flow bill.
[0006] In a first aspect of the present application, a method for analyzing sensitive words in treasury flow is provided, comprising:
[0007] Extracting signature sensitive words of flow bill in the flow library corresponding to each node in the treasury structure tree, identifying the signature sensitive words based on OCR text recognition, obtaining standard fonts in the signature sensitive words, and taking the remaining fonts as connected fonts;
[0008] Comparing the standard fonts with preset sensitive words in the flow library to obtain first screening words, comparing the partition proportion of connected fonts in the signature sensitive words with the corresponding standard fonts in the first screening words to obtain intermediate screening words;
[0009] Comparing the partition strokes of connected fonts with the corresponding standard fonts in the intermediate screening words to obtain adaptive screening words.
[0010] Optionally, in a possible implementation manner of the first aspect, the trey structure is initialized by the following steps, comprising:
[0011] constructing a parent node corresponding to the group company, and calling each subsidiary company of the group company to perform a summary on the flow treys to obtain a flow summary trey, and binding the flow summary trey with the parent node;
[0012] constructing a child node corresponding to each of the subsidiary companies, connecting each of the child nodes with the parent node, and binding each of the flow treys with the corresponding child node to generate the trey structure.
[0013] Optionally, in a possible implementation manner of the first aspect, the comparing the preset sensitive words in the flow trey with the standard font to obtain the primary screening words comprises:
[0014] obtaining the number of fonts in the signature sensitive words as a sensitive number, and obtaining the position of the standard font in the signature sensitive words as a standard position;
[0015] performing a one-time screening on the preset sensitive words in the flow trey based on the standard position and the standard font to obtain standard screening words;
[0016] obtaining the number of fonts of the standard screening words as a screening number, and selecting the standard screening words with the screening number equal to the sensitive number as the primary screening words.
[0017] Optionally, in a possible implementation manner of the first aspect, the comparing the continuous writing font in the signature sensitive words with the corresponding standard font in the primary screening words to obtain the intermediate screening words comprises:
[0018] uniformly dividing the font slot positions of the continuous writing font in the signature sensitive words and the corresponding standard font in the primary screening words respectively to obtain upper, middle and lower division regions corresponding to the font slot positions;
[0019] obtaining a standard partition ratio corresponding to the standard font according to the ratio of the number of font pixel points in the upper, middle and lower division regions corresponding to the standard font;
[0020] obtaining a continuous writing partition ratio corresponding to the continuous writing font according to the ratio of the number of font pixel points in the upper, middle and lower division regions corresponding to the continuous writing font;
[0021] obtaining the intermediate screening words according to the comparison between the continuous writing partition ratio and each of the standard partition ratios.
[0022] Optionally, in a possible implementation manner of the first aspect, the obtaining of the intermediate screening word according to the comparison between the continuous-stroke partition proportion and the standard partition proportion of each of the standard partitions comprises:
[0023] obtaining a first gap value according to an absolute value of a difference between the continuous-stroke partition proportion and a corresponding proportion coefficient of an upper partition in each of the standard partitions;
[0024] obtaining a second gap value according to an absolute value of a difference between the continuous-stroke partition proportion and a corresponding proportion coefficient of a middle partition in each of the standard partitions;
[0025] obtaining a third gap value according to an absolute value of a difference between the continuous-stroke partition proportion and a corresponding proportion coefficient of a lower partition in each of the standard partitions;
[0026] obtaining a total gap corresponding to each of the initial screening words according to a sum of the first gap value, the second gap value and the third gap value, and selecting an initial screening word with a total gap less than a preset gap value as the intermediate screening word.
[0027] Optionally, in a possible implementation manner of the first aspect, the obtaining of the adaptive screening word according to the comparison between the continuous-stroke font and the corresponding standard font in the intermediate screening word comprises:
[0028] obtaining a continuous-stroke region and a non-continuous-stroke region corresponding to the continuous-stroke font and the corresponding standard font in the intermediate screening word;
[0029] performing stroke comparison on the continuous-stroke region and the non-continuous-stroke region corresponding to the continuous-stroke font and the corresponding standard font in the intermediate screening word to obtain a continuous-stroke similarity and a non-continuous-stroke similarity;
[0030] obtaining a stroke similarity of each of the intermediate screening words according to an average value of the sum of the continuous-stroke similarity and the non-continuous-stroke similarity, and selecting an intermediate screening word with a stroke similarity greater than or equal to a preset similarity as the adaptive screening word.
[0031] Optionally, in a possible implementation manner of the first aspect, the obtaining of the continuous-stroke region and the non-continuous-stroke region corresponding to the continuous-stroke font and the corresponding standard font in the intermediate screening word comprises:
[0032] dividing font slots of the continuous-stroke font and the corresponding standard font in the intermediate screening word into three equal parts from top to bottom to obtain a plurality of sub-partitions corresponding to the font slots;
[0033] merging sub-partitions in which continuous-stroke strokes are located to obtain the continuous-stroke region, and taking the remaining sub-partitions as the non-continuous-stroke region.
[0034] Optionally, in a possible implementation manner of the first aspect, the stroke comparison between the corresponding standard font and the corresponding non-stroke font in the intermediate screening word includes:
[0035] The number of strokes in the corresponding standard font corresponding to the connected stroke region and the non-connected stroke region in the intermediate screening word is counted to obtain a total number of connected strokes corresponding to the connected stroke region and a total number of non-connected strokes corresponding to the non-connected stroke region;
[0036] The number of strokes of the same stroke in the connected stroke region of the connected stroke font and the corresponding standard font is obtained as a first number, and the number of strokes of the same stroke in the non-connected stroke region is obtained as a second number;
[0037] The connected stroke similarity is obtained according to the ratio of the first number to the total number of connected strokes, and the non-connected stroke similarity is obtained based on the ratio of the second number to the total number of non-connected strokes.
[0038] Optionally, in a possible implementation manner of the first aspect, the method further includes:
[0039] The historical signature data of the signature personnel corresponding to the adaptive screening word is obtained, and the historical font corresponding to the most signature number of each signature personnel in the historical signature data is determined as a reference font;
[0040] The connected stroke and the non-connected stroke in the reference font are valued based on the stroke order to obtain a reference stroke sequence corresponding to each signature personnel;
[0041] The connected stroke and the non-connected stroke in the connected stroke font are valued according to the stroke order to obtain a current stroke sequence;
[0042] The strokes in each reference stroke sequence are sequentially valued and subtracted from the stroke value in the current stroke sequence to obtain a similar stroke sequence;
[0043] The values in the similar stroke sequence are summed to obtain a similar sum value, and the similar sum value is subjected to absolute value processing to obtain a selection value, and the adaptive screening word of the signature personnel is sorted in ascending order based on the selection value to obtain an adaptive screening sequence.
[0044] In a second aspect of the embodiment of the application, a treasurers' flow sensitive word analysis system is provided, including:
[0045] The recognition module is configured to extract the signature sensitive word of the flow bill in the flow bank corresponding to each node in the treasurers' structure tree, recognize the signature sensitive word based on OCR character recognition to obtain a standard font in the signature sensitive word, and take the remaining font as a connected stroke font;
[0046] The screening module is configured to compare preset sensitive words in the flow library with the standard font to obtain primary screening words, and compare the connected writing font in the signature sensitive word with the corresponding standard font in the primary screening words to obtain intermediate screening words.
[0047] The comparison module is configured to compare the corresponding standard font in the intermediate screening words with the connected writing font in the partition stroke to obtain adaptive screening words.
[0048] In a third aspect, an electronic device is provided, including a memory, a processor and a computer program, the computer program is stored in the memory, and the processor executes the computer program to implement the method of the first aspect and various possible aspects related to the first aspect.
[0049] In a fourth aspect, a storage medium is provided, the storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect and various possible aspects related to the first aspect.
[0050] The beneficial effects of the present application are as follows:
[0051] 1、The present application can effectively process the connected writing font by sequentially extracting the signature sensitive word, using the OCR to recognize the standard font, and comparing the standard font and the connected writing font, thereby improving the extraction and recognition ability of the complex signature style flow ticket sensitive word, solving the problem of easy omission and misjudgment of sensitive word extraction when facing the connected writing font in the prior art, and improving the accuracy of the treasurers flow sensitive word analysis, providing more reliable information support for enterprise financial decision-making, risk control and compliance operation.
[0052] 2、The present application constructs a treasurers structure tree, and stores and associates the flow data of the group company and its subsidiaries in layers, which facilitates the unified management and query of the entire group business flow. This not only improves the data processing efficiency of sensitive word analysis, but also makes the analysis result better reflect the business situation of the whole group and each subsidiary, and provides strong data support for management decision-making of different levels of enterprises. Subsequently, the sensitive word search range is narrowed by standard font screening, the analysis efficiency is improved, and the similarity between the connected writing font and the standard font is measured as a whole, further screening out possible sensitive words, improving the accuracy of sensitive word recognition, and comparing the strokes of the connected writing and non-connected writing areas from the stroke level, obtaining more accurate adaptive screening words. These steps cooperate with each other to comprehensively improve the accuracy and reliability of sensitive word analysis, and reduce the analysis error caused by complex font.
[0053] 3、The application fully considers the signature habits of the signature personnel by acquiring the historical signature data of the signature personnel corresponding to the adaptive screening word, and analyzing and sorting based on the historical signature data. Through this way of sorting the adaptive screening word, the signature personnel can be determined more accurately, more valuable information is provided for the auditing personnel, personnel verification is facilitated, and the reliability and practicality of the sensitive word analysis result are further improved BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 A flow chart of a treasurers' stream sensitive word analysis method provided by the application;
[0055] Figure 2 A schematic view after font slot division provided by the application;
[0056] Figure 3 A structural schematic view of a treasurers' stream sensitive word analysis system provided by the application;
[0057] Figure 4 A hardware structural schematic view of an electronic device provided by the application. DETAILED DESCRIPTION
[0058] To make the objectives, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described below in connection with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.
[0059] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein.
[0060] It should be understood that in various embodiments of the application, the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0061] It should be understood that, in the present application, "comprising" and "having" and any variations thereof are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0062] It should be understood that, in the present application, "a plurality of" means two or more. "And / or" is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents that the front and rear associated objects are in an "or" relationship. "Including A, B and C", "including A, B, C" means that A, B and C are all included, "including A, B or C" means that one of A, B and C is included, and "including A, B and / or C" means that any one or any two or three of A, B and C is included.
[0063] It should be understood that, in the present application, "B corresponding to A", "B corresponding to A", "A corresponding to B" or "B corresponding to A" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean that B is determined only according to A, but also can be determined according to A and / or other information. The matching of A and B means that the similarity of A and B is greater than or equal to a preset threshold.
[0064] Depending on the context, "if" as used herein can be interpreted as "when" or "upon" or "in response to a determination" or "in response to detecting".
[0065] The technical solutions of the present application will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes can not be described in detail in some examples.
[0066] The present application provides a kind of treasurer flow sensitive word analysis method, as shown in figure Figure 1 The present application provides a kind of treasurer flow sensitive word analysis method, as shown in figure
[0067] S1, extract the signature sensitive word of each node in the treasurer structure tree corresponding to the flow water bill in the flow water bank, identify the signature sensitive word based on OCR character recognition, obtain the standard font in the signature sensitive word, and the rest font is regarded as a continuous writing font.
[0068] It should be noted that the existing method can only effectively process the signature font written in standard and neat, and the accuracy of recognition will be greatly reduced once facing the continuous signature font existing in large quantities in practice, thereby causing errors when checking the information in the flow water bill.
[0069] Therefore, when auditing the signature sensitive words in the flow ticket, such as auditing the personnel signature, the existing OCR text recognition is used to recognize the font of the personnel name that can be recognized, and then the connected font is analyzed, such as recognizing the Cao word in the signature corresponding to the personnel (Cao San), and the other three is a connected font. Subsequently, all the employee names with the surname Cao in the flow library corresponding to the Cao word are retrieved according to the Cao word, so as to facilitate further comparison and analysis.
[0070] The treelike data structure for organizing and managing the information related to the treasurer is a treasurer structure tree. The nodes in the tree represent different organizational levels or business units, such as a group company and its subsidiaries. Each node corresponds to a flow library for storing the flow ticket information of the level or unit. This structure facilitates hierarchical management and query of the treasurer data.
[0071] The flow library is a database for storing flow tickets. Each flow library corresponds to a node in the treasurer structure tree. The flow ticket contains detailed records of the treasurer business operations, such as the transaction time, amount, signature, and other information on the invoice.
[0072] The signature sensitive word is a word in the signature of the flow ticket that may contain a specific meaning or need to be focused on, such as a personnel handwritten signature, which is related to business risk, compliance, etc. It can be used to query whether the flow ticket of the personnel meets the authority, etc.
[0073] The standard font refers to the font that meets the standard writing specifications and can be accurately recognized by the OCR text recognition technology.
[0074] The connected font is a font that is difficult to be directly recognized by the OCR due to the connection of strokes during the signature process caused by writing habits or speed, etc.
[0075] Through the above implementation, the standard font and the connected font in the signature sensitive word can be effectively distinguished. The standard font can be directly used for subsequent comparison with the preset sensitive word, improving the accuracy and efficiency of the comparison. The separate distinction of the connected font provides a basis for subsequent special analysis of the connected font, which helps to solve the problem of sensitive word recognition caused by connected writing, thereby improving the accuracy and comprehensiveness of the entire treasurer flow sensitive word analysis.
[0076] In step S1, the treasurer structure tree is initialized by the following steps, including S11-S12:
[0077] S11, a parent node corresponding to the group company is constructed, and the flow libraries of each subsidiary of the group company are retrieved for aggregation to obtain a flow aggregation library. The flow aggregation library is bound to the parent node.
[0078] It should be noted that the business flow data of the group company and its subsidiaries has a hierarchical structure and correlation. Building the treasurer structure tree can effectively organize and manage these data, so that the subsequent sensitive word extraction and analysis work can be based on a clear data architecture. By storing and correlating the flow data of the group company and its subsidiaries in layers, the business flow of the entire group can be uniformly managed and queried, improving the efficiency and accuracy of the analysis.
[0079] The parent node is the top node in the treasurer structure tree, represents the group company, is bound with the flow summary library, and is used for storing and managing the flow information of the group company as a whole. The subsidiary company is a company subordinate to the group company.
[0080] The child node is a node in the treasurer structure tree corresponding to the subsidiary company, each child node is associated with a subsidiary company, and is connected with the parent node, and is used for storing and managing the flow information of the corresponding subsidiary company.
[0081] It is not difficult to understand that the application will construct a parent node corresponding to the group company, and summarize the flow libraries of all the subsidiary companies of the group company, thereby obtaining a flow summary library corresponding to the group company, and binding the flow summary library with the parent node.
[0082] S12, a child node corresponding to each of the subsidiary companies is constructed, the child nodes are connected with the parent node, and the flow library is bound with the corresponding child node, and a treasurer structure tree is generated.
[0083] It is not difficult to understand that the server will construct a child node corresponding to each subsidiary company, connect each child node with the parent node, and bind the flow library of the corresponding subsidiary company with the corresponding child node, thereby generating a treasurer structure tree.
[0084] By initializing the treasurer structure tree through the above steps, the flow data of the group company and its subsidiaries can be effectively organized and managed.
[0085] S2, based on the standard font, the preset sensitive words in the flow library are compared, and the initial screening words are obtained, the connected writing font in the signature sensitive word is compared with the corresponding standard font in the initial screening word, and the intermediate screening word is obtained.
[0086] It should be noted that in the treasurer flow sensitive word analysis, the existence of the standard font and the connected writing font makes the identification of the sensitive word complex. By using the standard font for preliminary screening, the search range of the sensitive word can be reduced, and the efficiency of the subsequent analysis can be improved. Due to the particularity of the connected writing font, it is difficult to directly perform accurate matching, and the method of partitioning and comparing the proportion can measure the similarity between the connected writing font and the standard font as a whole, further filter out the possible sensitive words, and improve the accuracy of the sensitive word identification.
[0087] In some embodiments, the step S2 (comparing the standard font with the preset sensitive words in the flow library to obtain the primary screening words) comprises S21-S23:
[0088] S21, obtaining the number of fonts in the signature sensitive words as the sensitive number, and the position of the standard font in the signature sensitive words as the standard position.
[0089] Specifically, the number of fonts (sensitive number) of the signature sensitive words and the position of the standard font (standard position) in the signature sensitive words are obtained, for example, the sensitive number corresponding to Cao San is 2 (two characters), and Cao is located at the first position. For another example, the sensitive number corresponding to Wang Wu Er is 3.
[0090] S22, based on the standard position and the standard font, performing a screening on the preset sensitive words in the flow library to obtain the standard screening words.
[0091] Among them, the preset sensitive words are words that need to be paid attention to in the flow library, such as personnel names, which are the target set of screening, for example, Cao San, Cao Er, Cao Yi, and Wang Wu Er, for the convenience of understanding, only simple examples are given here.
[0092] Specifically, the standard position and the standard font are used to preliminarily screen the preset sensitive words in the flow library to obtain the standard screening words. This step uses the clear characteristics of the standard font to exclude obviously inconsistent preset sensitive words, for example, by Cao and located at the first position, all personnel names with the surname Cao are screened from the database to obtain the standard screening words, and all personnel with the surname Cao are screened out.
[0093] S23, obtaining the number of fonts of the standard screening words as the screening number, and selecting the standard screening words with the screening number equal to the sensitive number as the primary screening words.
[0094] It can be understood that by comparing the number of fonts (screening number) of the standard screening words and the sensitive number, the standard screening words equal to the sensitive number are selected as the primary screening words, which further narrows the screening range. It is not difficult to understand that the 3-character, 4-character, and 3-character of the surname Cao are all screened out.
[0095] In some embodiments, the step S2 (comparing the connected font in the signature sensitive words with the corresponding standard font in the primary screening words to obtain the intermediate screening words) comprises S24-S27:
[0096] S24, the font slots of the connected font in the signature sensitive words and the corresponding standard font in the primary screening words are evenly divided respectively to obtain the upper, middle and lower division regions corresponding to the font slots.
[0097] It should be noted that the present application sets multiple font slots for the area requiring signature, and the font slot provides a clear range for accurately extracting the feature information of the font. When performing partition ratio comparison and partition stroke comparison, the setting of the font slot enables the algorithm to focus on the signature font itself, excluding the interference of irrelevant information around, thereby more accurately extracting the key features of the font and providing a reliable basis for subsequent sensitive word screening. Moreover, since a large number of signature bills are involved, each signature may have unique writing characteristics. Setting the font slot can establish a unified analysis standard for all signature fonts, so that regardless of who the signer is or how the writing style is, the same rules can be used for processing and analysis. This can ensure that the analysis results between different signatures are comparable and avoid analysis errors due to the lack of a unified standard.
[0098] Specifically, the font slots of the connected writing font and the standard font are evenly divided into upper, middle and lower divided areas, preparing for subsequent partition ratio calculation.
[0099] S25, according to the ratio of the number of font pixels in the upper, middle and lower divided areas corresponding to the standard font, the standard partition ratio corresponding to the standard font is obtained.
[0100] It should be noted that for the same character, although the connected writing font changes in form during writing, it is based on the evolution of the structure of the standard font. For example, when writing the character "guo" in connected writing, although the strokes are simplified and connected, the approximate proportion relationship of the upper, middle and lower parts is still similar to that of the standard "guo". This means that the connected writing font has an inherent relationship with the standard font in terms of regional pixel distribution, and the ratio of the pixel distribution of each region also has relative stability, providing a feasible basis for screening through partition ratio.
[0101] Specifically, the ratio of the number of font pixels in each divided area of the standard font is calculated to obtain the standard partition ratio. For example, the pixel ratio of the upper, middle and lower areas corresponding to a three-character can be 1:1:2.
[0102] S26, based on the ratio of the number of font pixels in the upper, middle and lower divided areas corresponding to the connected writing font, the connected writing partition ratio corresponding to the connected writing font is obtained.
[0103] Similarly, the ratio of the number of font pixels in each divided area of the connected writing font is calculated to obtain the connected writing partition ratio. Referring to Figure 2 For example, the pixel ratio of the upper, middle and lower areas corresponding to a three-character can be 1:1:1.8.
[0104] S27, obtaining an intermediate screening word according to comparison of the continuous writing partition proportion and each of the standard partition proportions.
[0105] It can be understood that the continuous writing partition proportion and each of the standard partition proportions are compared, and the intermediate screening word is obtained according to the comparison result. It can be understood that due to the factors of signing and continuous writing, there will inevitably be errors, and therefore, the similar primary screening words within the error allowable range will be selected as the intermediate screening words in the subsequent process.
[0106] In some embodiments, the step S27 (obtaining an intermediate screening word according to comparison of the continuous writing partition proportion and each of the standard partition proportions) includes S271-S274:
[0107] S271, obtaining a first gap value according to an absolute value of a difference value of a corresponding proportion coefficient of an upper partition in the continuous writing partition proportion and each of the standard partition proportions.
[0108] It can be understood that for the upper partition, the absolute value of the difference value of the corresponding proportion coefficient is calculated to obtain the first gap value, and the pixel point distribution difference in the upper area of the font is focused.
[0109] S272, obtaining a second gap value based on an absolute value of a difference value of a corresponding proportion coefficient of a middle partition in the continuous writing partition proportion and each of the standard partition proportions.
[0110] It can be understood that according to the same principle as S271, for the middle partition, the absolute value of the difference value of the corresponding proportion coefficient in the continuous writing partition proportion and each of the standard partition proportions is calculated to obtain the second gap value.
[0111] S273, obtaining a third gap value according to an absolute value of a difference value of a corresponding proportion coefficient of a lower partition in the continuous writing partition proportion and each of the standard partition proportions.
[0112] Specifically, similarly, the lower partition is operated to obtain the absolute value of the difference value of the corresponding proportion coefficient of the lower partition in the continuous writing partition proportion and each of the standard partition proportions, that is, the third gap value.
[0113] S274, obtaining a total gap corresponding to each of the primary screening words based on a sum of the first gap value, the second gap value and the third gap value, and selecting the primary screening word with a total gap less than a preset gap value as the intermediate screening word.
[0114] The preset gap value is a gap value artificially preset, which can be an error allowable range set according to actual conditions.
[0115] It can be understood that the first gap value, the second gap value and the third gap value obtained above are summed up to obtain a total gap corresponding to each primary screening word. This total gap comprehensively reflects the difference degree of the connected script font and the standard font in the whole font structure. A preset gap value is set as a judgment standard, and the primary screening word with a total gap less than the preset gap value is selected as an intermediate screening word. In this way, the primary screening word similar to the connected script font in the structural characteristics is selected as the intermediate screening word for subsequent analysis in a quantitative manner.
[0116] S3, performing sub-region stroke comparison on the corresponding standard font of the intermediate screening word according to the connected script font to obtain an adaptive screening word.
[0117] In some embodiments, the step S3 (performing sub-region stroke comparison on the corresponding standard font of the intermediate screening word according to the connected script font to obtain an adaptive screening word) comprises S31-S33:
[0118] S31, obtaining the connected script region and the non-connected script region corresponding to the connected script font and the corresponding standard font of the intermediate screening word.
[0119] Specifically, first, the region is divided according to the connected script font, which is divided into a connected script region and a non-connected script region, wherein the connected script region is the region where the connected strokes are located, for example, the second stroke and the third stroke of the three characters are connected, and the corresponding connected script region is the middle and lower region, and the non-connected script region is the upper region.
[0120] It is worth mentioning that when the connected script appears in the middle and lower region (for example, the lower stroke is connected with the middle stroke), if the three-part split is forcibly retained, the writing trajectory of the connected script (such as the connection of the starting and ending pen, the curvature trend) will be broken, resulting in distortion of the component features. Therefore, the middle and lower regions are merged into a whole, which respects the fact that "connected script is a coherent writing unit" and avoids comparison errors caused by mechanical splitting.
[0121] In the above manner, this region division is the basis for subsequent stroke comparison, which can more specifically analyze the stroke features of different regions.
[0122] In some embodiments, the step S31 (obtaining the connected script region and the non-connected script region corresponding to the connected script font and the corresponding standard font of the intermediate screening word) comprises S311-S312:
[0123] S311, respectively, the font slots of the connected script font and the corresponding standard font of the intermediate screening word are uniformly divided into three equal parts from top to bottom to obtain a plurality of sub-division regions corresponding to the font slots.
[0124] It can be understood that, the server divides the font slot corresponding to the standard font in the connected writing font and the intermediate filter word from top to bottom uniformly according to the principle of step S24, to obtain a plurality of sub-division zones corresponding to the font slot, that is, the upper, middle and lower division regions.
[0125] S312, merging the sub-division zones where the connected strokes in the connected writing font are located to obtain a connected region, and taking the remaining sub-division zones as non-connected regions.
[0126] It can be understood that, in the plurality of sub-division zones of the connected writing font, the sub-division zones where the connected strokes are located are identified and merged to determine the connected region. The connected strokes often span multiple sub-division zones, and by merging these regions, the special writing part of the connected writing font can be completely included. The sub-division zones not involved by the connected strokes are defined as non-connected regions. This identification and classification based on sub-division zones can clearly and accurately distinguish different stroke characteristic regions of the connected writing font, laying a foundation for subsequent targeted stroke comparison. For example, the second and third strokes of three are connected, so the middle and lower regions are merged to obtain a connected region, and the remaining upper region which is not connected is taken as a non-connected region.
[0127] It is worth mentioning that the present application can merge the sub-division zones multiple times, for example, the first and second strokes are connected, and the second and third strokes are connected, so the upper and middle regions can be merged into one region, and the middle and lower regions can be merged into one region, and the subsequent calculation is the same, that is, the similarity of the strokes corresponding to the two regions is calculated.
[0128] S32, comparing the connected region and the non-connected region corresponding to the connected writing font and the intermediate filter word with the corresponding standard font to obtain a connected similarity and a non-connected similarity.
[0129] It is worth mentioning that no matter how unique the writing form of the connected writing font is, its essence is still built on the basic structure of the standard font. Taking the Chinese character "good" as an example, even if the connected writing makes the female character side connected with the stroke of the child character, the structural relationship between the two in the standard font is still retained. This means that in the connected region, the stroke direction, start and end points, etc. are still closely related to the structure of the standard font. Similarly, the strokes in the non-connected region follow the established writing specifications and structural principles in both the connected writing font and the standard font. Therefore, comparing the strokes in the two regions respectively can accurately analyze the similarities and differences between them according to the inherent structural characteristics of the font, and provide a solid basis for sensitive word filtering.
[0130] It can be understood that the server will compare the strokes of the connected writing area and the non-connected writing area corresponding to the connected writing font and the standard font. In the connected writing area and the non-connected writing area, the main comparison is whether the strokes are similar, for example, whether the connected writing area and the non-connected writing area corresponding to the standard font and the connected writing font have the same strokes, such as horizontal strokes, vertical strokes, and points. Through comparison, the connected writing similarity and the non-connected writing similarity are obtained, and the two similarity indexes respectively reflect the similarity degree of the connected writing font and the standard font in the connected writing area and the non-connected writing area.
[0131] In some embodiments, the stroke comparison of the connected writing area and the non-connected writing area corresponding to the connected writing font and the standard font in the intermediate screening word in step S32 to obtain the connected writing similarity and the non-connected writing similarity includes S321-S323:
[0132] S321, count the number of strokes in the connected writing area and the non-connected writing area corresponding to the standard font in the intermediate screening word to obtain the total number of connected writing corresponding to the connected writing area and the total number of non-connected writing corresponding to the non-connected writing area.
[0133] It can be understood that the service area will analyze the corresponding standard font in the intermediate screening word and count the number of strokes in the connected writing area and the non-connected writing area. Through statistics, the total number of connected writing corresponding to the connected writing area and the total number of non-connected writing corresponding to the non-connected writing area are obtained, for example, the connected writing area of the lower part of the standard font is 2 strokes, and the non-connected writing area is 1 stroke, which is the number of standard strokes.
[0134] S322, obtaining the number of strokes of the same strokes in the connected writing area of the connected writing font and the corresponding standard font as the first number, and the number of strokes of the same strokes in the non-connected writing area as the second number.
[0135] It can be understood that the server will respectively obtain the number of strokes of the same strokes in the connected writing area (the first number) and the number of strokes of the same strokes in the non-connected writing area (the second number) of the connected writing font and the corresponding standard font, for example, for Cao San, Cao Yi has only one horizontal stroke in the middle and lower area, and actually has two horizontal strokes, so the first number is 1 and the second number is 0, therefore, the higher the similarity of the subsequent ratio is, the more similar the two characters are.
[0136] S323, obtaining the connected writing similarity according to the ratio of the first number and the total number of connected writing, and obtaining the non-connected writing similarity based on the ratio of the second number and the total number of non-connected writing.
[0137] It can be understood that the first quantity is compared with the total number of connected strokes to obtain a connected stroke similarity, which reflects the similarity of the connected stroke font and the standard font in the connected stroke area. The second quantity is compared with the total number of non-connected strokes to obtain a non-connected stroke similarity, which reflects the similarity of the two in the non-connected stroke area. Through the calculation of the two similarities, the similarity of the connected stroke font and the standard font is comprehensively measured from different areas.
[0138] S33, according to the average value of the connected stroke similarity and the non-connected stroke similarity, the stroke similarity of each intermediate screening word is obtained, and the intermediate screening word with a stroke similarity greater than or equal to a preset similarity is selected as an adaptive screening word.
[0139] It can be understood that the average value of the connected stroke similarity and the non-connected stroke similarity is calculated to obtain the stroke similarity of each intermediate screening word. The average value comprehensively considers the connected stroke area and the non-connected stroke area, and more comprehensively reflects the overall similarity of the connected stroke font and the standard font. Then, a preset similarity is set as a threshold, and the intermediate screening word with a stroke similarity greater than or equal to the threshold is selected as an adaptive screening word. The preset similarity is a threshold set by a person in advance, which is used to judge whether the similarity of the connected stroke font and the standard font in the stroke feature meets the requirements.
[0140] The intermediate screening words highly similar to the connected stroke font in the stroke feature are screened out in a quantitative way, which further narrows down the range of sensitive words, and the final adaptive screening word is recommended to the user.
[0141] It is not difficult to understand that after obtaining the adaptive screening word, the signature personnel screened out can be determined, and the adaptive screening word can be sorted according to the signature habit of the signature personnel from the maximum possibility to the minimum possibility, and then the result can be directly sent to the auditing personnel for personnel verification.
[0142] On the basis of the above-mentioned embodiments, A1-A5 are further included:
[0143] A1, the historical signature data of the signature personnel corresponding to the adaptive screening word is obtained, and the historical font of the connected stroke corresponding to the most signature times of each signature personnel in the historical signature data is determined as a reference font.
[0144] It can be understood that the historical signature data of the signature personnel corresponding to the adaptive screening word is obtained first, which is the basic data source for analyzing the signature habit. Then in the historical signature data, the historical font of the connected stroke corresponding to the most signature times of each signature personnel is determined as a reference font. The font corresponding to the most signature times is selected as the reference because this font can better represent the typical signature style of the signature personnel and has high representativeness and reliability. For example, the second stroke and the third stroke of the three characters are often connected by this person.
[0145] Among them, the historical signature data is the past signature records of the signers, including the signature information of the signers in different times and different business scenarios.
[0146] A2. Perform value assignment processing on the connected strokes and non-connected strokes in the reference font based on the stroke order to obtain the reference stroke sequence corresponding to each signer.
[0147] Among them, the stroke order is the order of writing the font. For example, when writing the character 'you', the first stroke is the horizontal left-falling stroke The second stroke is the right-falling stroke [[ID=By calculating the difference value, the difference between the reference font and the current font in stroke assignment can be quantified, and the similar stroke sequence records the difference value information and reflects the similarity between the two. It is not difficult to understand that the more similar the writing habits are, the closer the sum value in the corresponding similar stroke sequence is to 0.
[0154] A5, summing the values in the similar stroke sequence to obtain a similar sum value, performing absolute value processing on the similar sum value to obtain a selected value, and sorting the adaptation filtering words of the signatory in ascending order based on the selected value to obtain an adaptation filtering sequence.
[0155] It can be understood that the values in the similar stroke sequence are summed to obtain a similar sum value. The sum value comprehensively reflects the overall difference between the reference font and the current font in stroke assignment. The similar sum value is subjected to absolute value processing to obtain a selected value, which eliminates the positive and negative effects of the difference value and makes the selected value only reflect the size of the difference. Finally, the adaptation filtering words of the signatory are sorted in ascending order based on the selected value to obtain an adaptation filtering sequence. The smaller the selected value is, the more similar the stroke features of the current font and the reference font are, and the greater the possibility of the signatory is. By sorting in ascending order, the adaptation filtering words with high possibility are arranged in front, which is convenient for the auditor to verify.
[0156] Referring to Figure 3 , a structure schematic diagram of a treasurers' flow sensitive word analysis system provided by an embodiment of the present application, the treasurers' flow sensitive word analysis system comprises:
[0157] The recognition module is configured to extract the signature sensitive words of the flow tickets in the flow library corresponding to each node in the treasurers' structure tree, recognize the signature sensitive words based on OCR character recognition, obtain the standard font in the signature sensitive words, and take the remaining fonts as connected writing fonts.
[0158] The screening module is configured to compare the preset sensitive words in the flow library based on the standard font to obtain primary screening words, compare the partition proportion of the connected writing fonts in the signature sensitive words with the corresponding standard fonts in the primary screening words to obtain intermediate screening words.
[0159] The comparison module is configured to compare the partition strokes of the corresponding standard fonts in the intermediate screening words according to the connected writing fonts to obtain adaptation filtering words.
[0160] Referring to Figure 4 , a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application, the electronic device 40 comprises a processor 41, a memory 42 and a computer program; wherein
[0161] The memory 42 is used for storing the computer program, and can also be a flash memory. The computer program is, for example, an application program, a function module, etc. for implementing the above method.
[0162] The processor 41 is used for executing the computer program stored in the memory, so as to implement each step of the method performed by the device. Details can be referred to the related description in the above method embodiments.
[0163] Optionally, the memory 42 can be independent or integrated with the processor 41.
[0164] When the memory 42 is independent of the processor 41, the device can further include:
[0165] The bus 43 is used for connecting the memory 42 and the processor 41.
[0166] The present application further provides a readable storage medium, wherein the readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method provided by the various embodiments.
[0167] The readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transfer of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general or special purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0168] The present application further provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of a device can read the execution instructions from the readable storage medium, and the at least one processor executes the execution instructions to make the device implement the method provided by the various embodiments.
[0169] In the embodiments of the above apparatus, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or can also be any conventional processor, etc. The steps of the method disclosed in combination with the present application can be directly embodied as hardware processor execution, or executed by a combination of hardware and software modules in the processor.
[0170] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for analyzing sensitive words in treasury transaction records, characterized in that, include: Extract the signature sensitive words of the flow tickets in the flow reservoir corresponding to each node in the treasury structure tree, identify the signature sensitive words based on OCR text recognition, obtain the standard font of the signature sensitive words, and take the remaining font as the cursive font; The treasury structure tree is initialized using the following steps: Construct a parent node corresponding to the group company, and retrieve the transaction databases of each subsidiary of the group company for aggregation to obtain a transaction aggregation database, and bind the transaction aggregation database to the parent node; Construct child nodes corresponding to each subsidiary, connect each child node to the parent node, and bind the reservoir to the corresponding child node to generate a treasury structure tree; Based on the standardized font, the preset sensitive words in the database are compared to obtain the initial screening words. Then, the cursive font in the signature sensitive words is compared with the corresponding standard font in the initial screening words to obtain intermediate screening words, including: The number of characters in the signature sensitive words is obtained as the sensitive quantity, and the position of the standardized font in the signature sensitive words is obtained as the standardized position; Based on the standardized positions and fonts, the preset sensitive words in the reservoir are filtered once to obtain standardized filtered words; The number of characters in the standardized filter words is obtained as the filter quantity, and the standardized filter words whose filter quantity is equal to the number of sensitive words are selected as the initial filter words; The font slots of the cursive font in the signature sensitive words and the corresponding standard font in the initial screening words are evenly divided to obtain the upper division area, middle division area and lower division area corresponding to the font slots; The standard partition ratio corresponding to the standard font is obtained by the ratio of the number of font pixels in the upper, middle and lower partition regions corresponding to the standard font. Based on the ratio of the number of font pixels in the upper, middle and lower division regions corresponding to the cursive font, the cursive partition ratio corresponding to the cursive font is obtained. The intermediate filtering words are obtained by comparing the proportion of the cursive writing section with the proportion of each standard section. Based on the cursive script style, the corresponding standard fonts in the intermediate filter words are compared by stroke order to obtain suitable filter words, including: Obtain the connected and non-connected regions corresponding to the standard fonts in the connected font and the intermediate filter words; The strokes of the connected font and the corresponding standard font in the intermediate filtering words are compared to obtain the connected similarity and non-connected similarity. The stroke similarity of each intermediate filtering word is obtained by averaging the sum of the stroke similarity and the non-stroke similarity, and the intermediate filtering words with a stroke similarity greater than or equal to the preset similarity are selected as the matching filtering words. Obtain the historical signature data of the signer corresponding to the matching filter words, and determine the historical font of cursive writing corresponding to the maximum number of signatures of each signer in the historical signature data as the reference font; Based on the stroke order, the connected strokes and non-connected strokes in the reference font are assigned values to obtain the reference stroke sequence corresponding to each signatory. The connected strokes and non-connected strokes in the connected font are assigned values according to the stroke order to obtain the current stroke sequence; By sequentially calculating the difference between the stroke values assigned in each of the reference stroke sequences and the stroke values assigned in the current stroke sequence, a similar stroke sequence is obtained. The values in the similar stroke sequence are summed to obtain a similarity sum value. The absolute value of the similarity sum value is then processed to obtain a selection value. Based on the selection value, the matching filter words of the signatories are sorted in ascending order to obtain a matching filter sequence.
2. The method according to claim 1, characterized in that, The intermediate filtering words are obtained by comparing the proportion of the continuous strokes in the partition with the proportion of each of the standard partitions, including: The first difference value is obtained by the absolute value of the difference between the proportion of the continuous stroke partition and the corresponding proportion coefficient of the upper division area within each standard partition proportion; The second difference value is obtained based on the absolute value of the difference between the proportion of the continuous stroke partition and the proportion of the corresponding proportion coefficient of the middle division area within each standard partition. The third difference value is obtained by the absolute value of the difference between the proportion of the continuous stroke partition and the proportion of the lower division area within each standard partition. Based on the sum of the first gap value, the second gap value, and the third gap value, the total gap corresponding to each initial screening term is obtained, and the initial screening term with a total gap value less than the preset gap value is selected as the intermediate screening term.
3. The method according to claim 1, characterized in that, The step of obtaining the connected and non-connected regions corresponding to the standard fonts in the connected font and intermediate filter words includes: The font slots of the cursive font and the corresponding standard fonts in the intermediate filtering words are divided into three equal parts from top to bottom to obtain multiple sub-division areas corresponding to the font slots; The sub-regions containing the connected strokes in the cursive font are merged to obtain the cursive region, and the remaining sub-regions are treated as non-cursive regions.
4. The method according to claim 1, characterized in that, The step of comparing the strokes of the cursive font and the corresponding standard fonts in the intermediate selection words to obtain cursive similarity and non-cursive similarity includes: The number of strokes in the corresponding standard font of the intermediate filtering words in the connected and non-connected areas is counted to obtain the total number of connected strokes in the connected areas and the total number of non-connected strokes in the non-connected areas. The number of identical strokes in the connected stroke area of the connected stroke font and the corresponding standard font is obtained as a first number, and the number of identical strokes in the non-connected stroke area is obtained as a second number. The similarity of connected strokes is obtained based on the ratio of the first quantity to the total number of connected strokes, and the similarity of non-connected strokes is obtained based on the ratio of the second quantity to the total number of non-connected strokes.
5. A system corresponding to the treasury transaction sensitive word analysis method of claim 1, characterized in that, include: The recognition module is used to extract the signature sensitive words of the flow tickets in the flow reservoir corresponding to each node in the treasury structure tree, and to recognize the signature sensitive words based on OCR text recognition to obtain the standard font of the signature sensitive words, and to treat the remaining font as cursive font; The filtering module is used to compare the preset sensitive words in the flow database based on the standard font to obtain the initial filtered words, and to compare the proportion of cursive fonts in the signature sensitive words with the corresponding standard fonts in the initial filtered words to obtain the intermediate filtered words. The comparison module is used to perform stroke comparison of the corresponding standard fonts in the intermediate filter words according to the cursive script, and obtain the suitable filter words.
Citation Information
Patent Citations
Infrared touch frame, infrared touch screen and display equipment
CN111061391A
Character image recognition method and device based on hybrid convolution, equipment and storage medium
CN111666931A