A method and device for extracting feature data
By defining the characteristics of the feature data, extracting, verifying and cleaning of feature data, the problem of feature data identification error in the tax and financial industries is solved, and the high accuracy and efficiency extraction of feature data is achieved.
Patent Information
- Application Number
- CN202111674913.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In enterprise information management, especially in the tax and financial industries, there are identification errors after automatic collection of complex data, and manual verification is difficult, resulting in insufficient accuracy in identifying sensitive feature data.
By defining the characteristics of the feature data, the feature data is extracted, verified, cleaned and optimized, and the modules in the feature data extraction device are used to check and correct the accuracy of the feature data, including source identification, definition of features, information extraction, verification and cleaning units, and optimized with manual access.
It improves the accuracy of extracting sensitive feature data, optimizes the process of feature data, reduces the occurrence of errors, and improves work efficiency.
Smart Images

Figure CN114359567B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing and recognition, and particularly relates to a method and device for extracting feature data. Background Art
[0002] When enterprise information is centrally managed, information such as enterprise names and organization codes is required to be accurate. However, when a large amount of data is entered, there are problems such as inconsistent formats and easy errors in manual entry. Although there are many text recognition solutions available, for some relatively complex data, after being processed by automatic collection methods, there are some errors. Therefore, some methods are needed to correct errors and classify the reasons for errors every time an error occurs, so as to be used for the reprocessing of information in complex scenarios and avoid the recurrence of the same error.
[0003] There are many existing solutions that can recognize copied text and pictures, recognize text in pictures, and recognize feature data therein, such as the patent application with the application number 201710318767.2. However, for information managed in industries such as taxation and finance, the recognized information is sensitive, the digital coding is long, and errors are easy to occur. At the same time, it is difficult to perform manual verification. Therefore, a set of mechanisms are needed to improve the recognition accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for extracting feature data, which define, confirm, and continuously optimize the defined features of the feature data in the source information, so as to improve the extraction accuracy of feature data with high sensitivity requirements.
[0005] To solve the above technical problems, the present invention provides a method for extracting feature data. The method steps include:
[0006] Determine the source information, where the source information refers to a large text collection from which feature data needs to be extracted;
[0007] Determine the defined features of the feature data, where the defined features are the feature summaries of the content of the feature data that is expected to be extracted before extracting from the source information;
[0008] Extract the feature data from the source information according to the defined features of the feature data; the features of the extracted feature data are actual features;
[0009] Verify the effectiveness of the actual features of the extracted feature data, and determine whether there is an error between the actual features and the defined features;
[0010] If there is an error, perform feature data cleaning, compare the actual features of the extracted feature data with the defined features, locate the steps where the error occurs, and optimize the defined features;
[0011] After optimizing the defined features, determine the subsequent processes according to the settings, including outputting feature data, re-determining the defined features, and re-extracting the feature data.
[0012] On the other hand, the present invention also provides an apparatus for extracting feature data, including a source identification unit for generating text information that requires feature data extraction, including an image recognition conversion to text unit;
[0013] A defined feature unit for analyzing the defined features of the feature data, including an intelligent defined feature module for generalizing defined features such as length and contained content based on existing feature data;
[0014] An information extraction unit for extracting feature data from the source information in combination with the defined feature unit;
[0015] An information verification unit for inputting feature data and outputting the accuracy of the feature data through processing;
[0016] An information cleaning unit for processing feature data with errors, including deletion, storage, and analysis.
[0017] Furthermore, the apparatus for extracting feature data further includes a manual access unit and a process definition unit. The manual access unit includes a manual defined feature module, a manual corrected feature data module, and a manual corrected defined feature module, where the manual defined feature module is applied to the defined feature unit;
[0018] The process definition unit is used to determine different subsequent processes of defined feature optimization in different scenarios.
[0019] At the same time, the manual access unit can also be used to determine the use of the defined features. For example, a certain defined feature is used for extracting feature data from the source information or for verifying the feature data.
[0020] The method and apparatus for extracting feature data provided by the present invention define, extract, and then correct the feature data, continuously optimize the defined features and the entire process, accumulate and analyze the scenarios and error data, improve the extraction accuracy of sensitive feature data, and improve work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the feature data extraction method provided by the embodiment of the present invention;
[0022] Figure 2 It is a flowchart of the defined feature determination process in the feature data extraction method provided by the embodiment of the present invention;
[0023] Figure 3Another process flow chart for defining features in the feature data extraction method provided by the embodiments of the present invention;
[0024] Figure 4 The method flow chart for feature data definition optimization provided by the embodiments of the present invention;
[0025] Figure 5 It is the structural diagram of the feature data extraction device provided by the embodiments of the present invention. Detailed implementation manners
[0026] To better understand the purpose, structure and function of the present invention, the following further describes in detail a feature data extraction method and device of the present invention with reference to the accompanying drawings.
[0027] The present invention provides a method for extracting feature data from source information. As Figure 1 shown, its process includes the following steps:
[0028] S100: Determine the source information. The source information refers to a large block of text from which feature data needs to be extracted. For example, a company profile from which the company name needs to be extracted. However, in actual application scenarios, the source may be a piece of text or a picture. Therefore, determining the source information includes the effective conversion of the source information. For example, converting a picture into text, and denoising the information including irregular characters, spaces, etc., to form source information with an unlimited format but valid content.
[0029] The following is an example in this embodiment:
[0030] Source information 1 is the entire text copied from a certain system:
[0031] Taxpayer identification number: 99887766554433221X, Taxpayer name: Healthy Development Co., Ltd., Address: No. 7, 20th Floor, Building 99, Area X, Healthy Park, Nanming District, Guiyang City, Tel: 085166778899, Opening bank: Bank of China Guiyang Healthy Park Sub-branch, Account number: 556677889
[0032] Source information 2: The entire text recognized after scanning a picture from a financial information file provided by a certain staff member:
[0033] Invoice information:
[0034] Taxpayer name: Healthy Development Co., Ltd.
[0035] Taxpayer identification number: 99887766554433221X
[0036] Address: No. 7, 20th Floor, Building 99, Area X, Healthy Park, Nanming District, Guiyang City
[0037] Telephone: 0851 - 66778899
[0038] Bank of Deposit: Bank of China, Guiyang Health Park Sub - branch
[0039] Account Number: 556677889
[0040] The extractable feature data included in the above two pieces of source information are: taxpayer identification number, taxpayer name, address, telephone, bank of deposit, account number.
[0041] S110: Determine the defined features of the feature data. The defined features are the feature generalizations of the content of the feature data that one hopes to extract before extracting the source information.
[0042] The content of the defined features includes:
[0043] The relative position of the key characters of the feature data in the source information, such as the relative position of the characters or character strings that play an identifying role in the feature data starting from the first character of the source information; the relative position starting from a certain feature symbol in the source information.
[0044] The length of the key characters. For example, if the feature data is a telephone number and the key character is the area code, its length may be 3 digits or 4 digits.
[0045] The position of the key characters in the feature data, such as at the beginning, in the middle, or at the end of the character string of the feature data, and the relative position.
[0046] The length of the feature data. For example, the length of a telephone number is 11 digits.
[0047] Whether to exclude special key characters. For example, exclude I, O, Z, S, V in the taxpayer identification number.
[0048] Before determining the defined features of the feature data, first analyze and summarize the features of the source information. The source information includes multiple different or identical delimiter marks, such as full - width or half - width punctuation marks: colon (:), comma (,), space ( ), period (。), line break characters, etc. The feature data is often not limited to a single group and may be feature data with corresponding group names and contents. For example, in this embodiment, the source information includes key information names: taxpayer identification number, taxpayer name; and feature data contents, such as: 99887766554433221X, Healthy Development Co., Ltd.
[0049] Taking "taxpayer identification number" as an example, when the order of the feature data in source information 1 is stable, it is necessary to determine the defined features of the name and content respectively. The determination process is as Figure 2 shown:
[0050] S210: Starting position of the feature data name: The sixth character of the source information, with the formula as follows:
[0051] Feature data.name.starting position = Source information.position(6)
[0052] S220: Ending position of the feature data name: Feature data.name.starting position + position of the first separator after the starting position. In this embodiment, it is necessary to determine the content of the first separator after the starting position, and then determine its first position after the starting position of the feature data name in the source information. Through its position and length, the ending position of the feature data can be judged. The determination formula is as follows:
[0053] Position of the first separator = Feature data.name.starting position + Source information.Str(Feature data.name.starting position, Separator.starting position), where Source information.Str(Feature data.name.starting position, Separator.starting position) is the character length in the source information from after the starting position of the feature data name to the first separator;
[0054] Feature data.name.ending position = Position of the first separator;
[0055] Before executing this step, it is necessary to analyze the source information to determine the separator.
[0056] S230: Starting position of the feature data content: The formula is as follows:
[0057] Feature data.content.starting position = Position of the first separator + Length of the first separator + 1;
[0058] S240: Length of the feature data content; It can be determined by the separator method. In special scenarios, its content length has been determined. For example, for ID card information and the taxpayer identification number information in this embodiment, its length has been determined to be 18.
[0059] The defined features of the feature data also include characters that are included or not included. In this example, the taxpayer identification number field does not include I, O, Z, S, V. In this example, this feature is not used for feature data extraction, but can be used for accuracy verification, that is, step S250.
[0060] The present invention also provides a method for feature definition when the order of each feature data in the source information is not fixed. It can be determined by the positioning of the feature data name and the length of the feature data content, or by the separator. The method flow is as Figure 3 shown:
[0061] S310: Determine the content of the defined feature, including:
[0062] Define the name of the characteristic data: e.g., "Taxpayer Identification Number";
[0063] Define the length of the content of the characteristic data: 18
[0064] Define the method of accuracy verification: It can be verified and reviewed through data characteristics. For example, when judging the Taxpayer Identification Number, the length is 18 digits, and the field does not contain I, O, Z, S, V
[0065] S320: Extract the starting position Loc_key_name of the characteristic data name. For example: Retrieve the position where "Taxpayer Identification Number" is located in Source Information 2. The retrieval result output is: 21;
[0066] S330: Extract the length Loc_key_length of the characteristic data name, and extract the length of "Taxpayer Identification Number": 6
[0067] S340: Determine the starting position Loc_key_info of the content corresponding to the characteristic data name = 21 + 6 + 1
[0068] S341: Separation flag judgment. The character at Loc_key_info, in Source Information 2, the extraction result is: ":",
[0069] S350: Judge whether this identifier is a separation flag. If it is a separation flag, correct Loc_key_info = Loc_key_info + 1; otherwise, take the next separation flag, and the information between the next extracted separation identifiers is the key information;
[0070] S360: Locate the position Loc_next of the next separator. Take the character at Loc_key_info in the source information. In Source Information 2, this character is: "9". Take the position of the first separation marker after Loc_key_info in the source information. In Source Information 2, the next separation marker is a line feed character, and the position is 47.
[0071] S370: Extract the characteristic data according to Loc_key_info and Loc_next. The information from 21 + 6 + 1 + 1 to 47 is the content of the key information
[0072] S390: Output the extracted characteristic data and prepare for the next operation, such as accuracy verification.
[0073] S120: Extract the feature data from the source information according to the above-defined features. However, not all defined features need to be used during extraction. Therefore, this method also provides manual settings to sort the priorities or define the attributes for using all defined features in the extraction algorithm. Attributes include uses, etc. For example, in this case, the taxpayer identification number does not contain I, O, Z, S, V. The use attribute of this feature is for feature data verification, i.e., it is used during extraction, and it is used for Figure 2 the verification in step S250.
[0074] The features of the extracted feature data are actual features;
[0075] S130: Verify the effectiveness of the actual features of the extracted feature data, and determine whether there is an error between the actual features and the defined features. The verification methods include performing effectiveness verification through the defined features of the use attribute. For example, verify whether the extracted taxpayer identification number contains I, O, Z, S, V;
[0076] The verification methods also include verifying through a third-party platform.
[0077] 1. Interface call verification:
[0078] Call relevant information interfaces, such as: the invoice information platform interface to verify the information of the taxpayer name and identification number;
[0079] Obtain the organization code: Verify the information of the organization code, address, and enterprise name through the enterprise information platform interface;
[0080] For example, in this case, for the result obtained from the key information, obtain the taxpayer name in the invoice information platform, and retrieve the taxpayer name in the source information. If it exists, the verification is completed.
[0081] 2. Manual verification platform: For information that has not been verified and information that has been verified through the interface platform, it can be decided whether to enter the manual verification platform according to the requirements, which is used to verify whether the extraction of key information is accurate.
[0082] 3. Internal information library: For the verified information, its source information and key information are stored in the database for repeated use and result comparison.
[0083] S140: If there is an error, which is caused by inaccurate extracted feature data or multiple sets of feature data being generated, feature data cleaning can be performed. Compare the actual features of the extracted feature data with the defined features, establish an error library for the error between the defined features and the actual features, save the corresponding relationship established between the error and the occurrence scenario, and locate the steps where the error occurs: When the defined features are accurate and the error is caused by insurmountable factors, formulate a process for adjusting the feature data. For example: Error library screening: Establish an error library, import the source information and key information of the error into the error library, and screen the key information according to the scenario environment.
[0084] Data cleaning is to process the generated feature data, and at the same time, it can also optimize the generation process of the feature data, that is, S150: Optimize the defined features. On the basis of feature data cleaning, adjust and optimize the defined features determined in step S110.
[0085] In Figure 3 the solution of the embodiment shown, a method of locating the feature data content by using the feature data name + delimiter flag is adopted. In this embodiment, a method of optimizing the defined features after locating the link where the error occurs is provided, that is, by backtracking the link where the error occurs through the extracted feature data. As Figure 4 shown, after determining the error, perform key information feature adjustment:
[0086] S410 and S420: First, determine the corrected feature data info_key and source information info_source;
[0087] S430 and S440: Retrieve the quantity num_key of the feature data existing in the source information. In the next step, first determine its first position in the source information and perform a reverse search on the first feature data;
[0088] S440 to S490: In the example in this application, the delimiter flag length is 1. Therefore, extract the information content of 1 + the length of the feature name before this position and check if it conforms to the feature of "key information name + delimiter flag". If it does not conform, that is, this feature is determined to be abnormal;
[0089] Through Num_key = num_key - 1, perform a reverse inference judgment on the next position.
[0090] After the defined features of each key information are optimized, the processing of the steps shown in Figure 1 can be performed, or the verification steps can be redefined according to the rules.
[0091] S160: Determine the subsequent process. Define the process to be carried out next according to the scenario and requirements. For example, in the current scenario, although there are errors in the extracted feature data, after manual cleaning and adjustment, the results can be output to complete the extraction of the feature data this time. If the defined features are correct but there are defects in the source information, then after adjusting the source information, the process of S120 can be restarted; alternatively, the process of S110 can be entered to re-determine the defined features.
[0092] Through the method for extracting feature data provided by the present invention, the feature data is defined, extracted, and then corrected, continuously optimizing the defined features and the entire process, accumulating and analyzing the scenario and error data, improving the extraction accuracy of sensitive feature data, and improving work efficiency.
[0093] The present invention also provides a device for extracting feature data to cooperate with the method for extracting feature data provided by the present invention. The device for extracting feature data is mainly used to process source information, defined features, feature data, and processes, and its structure is as Figure 5 shown, mainly including the following unit modules:
[0094] The source identification unit includes an image recognition conversion module that can recognize images as text information; it also includes a text paste module that can paste large segments of text from other sources for use. After processing these text information, source information that conforms to certain specifications and can be used for feature data extraction is generated.
[0095] The defined feature unit is used to analyze the defined features of the feature data, including an intelligent defined feature module. The intelligent definition module is used to summarize the defined features such as its length and contained content according to the existing feature data. When the number of features for analysis reaches a certain amount, an artificial learning algorithm can be adopted.
[0096] The defined feature unit also includes an artificial defined feature module that provides an interface for manually defining the defined features of the feature data and retains the operation records.
[0097] The artificial defined feature module is one of the modules of the artificial access unit. In the artificial access unit, there are also included: an artificial corrected defined feature module and an artificial corrected feature data module.
[0098] The artificial corrected defined feature module is used for optimizing the defined features. Compared with the artificial defined feature module, its interface also includes a recording device for the reasons for correction, such as information on error factors between the defined features and the actual features, reverse inference steps, etc. Therefore, the artificial corrected defined feature module belongs to a complex version device of the artificial defined feature module.
[0099] The artificial corrected feature data module is for manual intervention to adjust the extracted feature data and retains the operation records.
[0100] The manual access unit can also add attributes to define features. For example, usage attributes and scenario attributes are used to determine which defined features need to be used when extracting and verifying feature data from source information. In this embodiment, for example, the length of the feature data can be used for extraction, and the characters not included in the feature data can be used for verification.
[0101] The above are the devices for preparatory work in the feature data extraction device. After the data preparation is completed, the feature data can be processed in combination with the source recognition unit and the defined feature unit. The related devices include:
[0102] An information extraction unit, which is used to extract feature data from the source information in combination with the defined feature unit;
[0103] An information verification unit, which is used to input feature data and output the accuracy of the feature data through processing;
[0104] An information cleaning unit, which is used to process the feature data with errors, including deletion, storage in the database, and analysis, and optimize the defined features;
[0105] A process definition unit, which is used to determine different subsequent processes for optimizing the defined features in different scenarios.
[0106] Through the feature data extraction device provided by the present invention, different devices are respectively deployed for the source information, the analysis and definition of the feature data to be extracted, and the extracted feature data. In the case where the algorithm for extracting feature data is not perfect, an interface for manual intervention management and a process for repeatedly optimizing the algorithm are provided, continuously improving various scenarios and the data and environment with errors, improving the extraction accuracy of sensitive feature data, and improving work efficiency.
[0107] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A method for extracting feature data, characterized in that, Including: Determine source information, where the source information refers to a large collection of text from which feature data needs to be extracted; Determine the defining features of the feature data, where the defining features are the feature generalizations of the content of the feature data that is expected to be extracted before extracting the source information; Extract the feature data from the source information according to the defining features of the feature data; the features of the extracted feature data are the actual features; Perform validity verification on the actual features of the extracted feature data to determine whether there is an error between the actual features and the defining features; If there is such an error, perform feature data cleaning, compare the actual features of the extracted feature data with the defining features, locate the steps that generate the error, and optimize the defining features; After optimizing the defining features, determine the subsequent process according to the settings, where the subsequent process includes outputting feature data, re-determining the defining features, and re-extracting the feature data.
2. The method for extracting feature data according to claim 1, wherein One set of source information includes one or more sets of feature data with corresponding names and contents, separated by multiple different or identical delimiter marks; The delimiter marks include punctuation marks that distinguish full-width and half-width characters.
3. The method for extracting characteristic data according to claim 1, wherein The determination method of the defining features includes: the relative position of the key characters of the feature data in the source information; the length of the key characters, the position of the key characters in the feature data; the length of the feature data; whether to exclude particularly critical characters.
4. The method for extracting characteristic data according to claim 1, wherein The feature data verification includes: defining feature verification, manual verification, third-party platform interface call verification, and internal information library verification.
5. The method for extracting characteristic data according to claim 4, characterized in that, The feature data cleaning includes: establishing an error library for the error between the defining features and the actual features, saving the corresponding relationship established between the error and the occurrence scenario, locating the steps that generate the error, and optimizing the defining features.
6. The method for extracting feature data according to claim 5, wherein, The optimization of the defining features is to re-determine the defining features on the basis of feature data cleaning; The optimization of the defining features includes correcting the feature data in the manual management platform.
7. The method for extracting characteristic data according to claim 6, wherein The determination of the subsequent process includes providing settings in the manual management platform.
Citation Information
Patent Citations
Text information acquisition method and device, and mobile terminal
CN106951893A
AI-based objectified attribute text automatic classification method and system
CN112966111A