Invoice detail classification method and device, electronic equipment and storage medium
By converting the format of invoice item details and applying a set of regular expressions, efficient and accurate classification of invoice details is achieved, solving the problems of low efficiency and high maintenance costs in existing technologies, and meeting the needs of insurance claims business.
Patent Information
- Application Number
- CN202511009049.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-04
AI Technical Summary
In insurance claims processing, existing technologies cannot efficiently and cost-effectively classify and review medical expense invoice details, resulting in low efficiency of manual review and high maintenance costs for fuzzy searches.
By converting non-Chinese characters in the invoice item details data into a preset format and generating a set of regular expressions using a pre-built retrieval configuration table, accurate matching and classification are achieved. Combined with multi-threaded parallel processing, the automatic classification of invoice details is realized.
It improves the efficiency of invoice detail classification, reduces maintenance costs, decreases the misjudgment rate, enhances classification accuracy and system flexibility, and adapts to changes in business needs.
Smart Images

Figure CN120892567A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to an invoice item classification method and device, electronic equipment and storage medium. BACKGROUND
[0002] In the insurance claim business, claim settlement is the core link of insurance contract performance, which is directly related to the reputation of insurance companies, risk management effectiveness and the protection of the rights and interests of the insured.
[0003] The claim settlement process usually includes key steps such as reporting, investigation and damage assessment, document review, settlement review and payment closing. Among them, the medical expense invoice submitted by the customer is one of the core contents of the document review.
[0004] Therefore, how to extract the relevant information of the medical expense invoice for classification to facilitate subsequent review and claim settlement is a technical problem that needs to be solved by those skilled in the art. SUMMARY
[0005] The present application provides an invoice item classification method, device, electronic equipment and storage medium, which greatly improves the efficiency of invoice item classification while reducing maintenance costs.
[0006] In a first aspect, the present application provides an invoice item classification method, comprising the steps of:
[0007] Converting all non-Chinese characters in the invoice item detail data into a preset format, the preset format including half-angle format and specific format;
[0008] Obtaining a corresponding regular expression set according to a pre-constructed search configuration table; the search configuration table includes at least two search rules, each search rule includes a search string, and the regular expression set includes at least two regular expressions, the regular expressions being generated according to the search string;
[0009] Classifying the invoice item detail data according to the regular expression set.
[0010] In some embodiments, the step of converting all non-Chinese characters in the invoice item detail data into a preset format includes the steps of:
[0011] Setting a corresponding target category code according to a batch of historical invoice item detail data;
[0012] Filtering target data corresponding to the target category code from the invoice item detail data.
[0013] In some embodiments, each of the search rules further comprises a rule identifier, at least one classification level field, and a status identifier; the status identifier comprises a valid flag and an invalid flag; and the obtaining a corresponding set of regular expressions from the pre-constructed search configuration table comprises the steps of:
[0014] filtering valid search rules from all of the search rules of the search configuration table, the status identifier of the valid search rules being the valid flag;
[0015] splitting the search string of the valid search rules according to a preset delimiter to obtain corresponding search keywords, and generating a corresponding positive look ahead assertion according to the search keywords, the positive look ahead assertion being the regular expression;
[0016] binding the regular expression with the corresponding rule identifier and the classification level field, and aggregating the regular expressions corresponding to all of the valid search rules to obtain the set of regular expressions.
[0017] In some embodiments, the binding the regular expression with the corresponding rule identifier and the classification level field further comprises the steps of:
[0018] if the valid search rule comprises an exclusion string, splitting the exclusion string of the valid search rule according to a preset delimiter to obtain a corresponding exclusion keyword;
[0019] generating a corresponding negative look ahead assertion according to the exclusion keyword, and obtaining the regular expression according to the negative look ahead assertion and the positive look ahead assertion.
[0020] In some embodiments, the matching and classifying of the invoice item detail data is in a multi-thread parallel manner.
[0021] In some embodiments, the matching and classifying of the invoice item detail data according to the set of regular expressions comprises the steps of:
[0022] determining whether the invoice item detail data matches the regular expression in the set of regular expressions;
[0023] if the matching result is a match, classifying the invoice item detail data into the classification level field corresponding to the matched regular expression to obtain a classification result.
[0024] In some embodiments, the matching and classifying of the invoice item detail data according to the set of regular expressions further comprises the steps of:
[0025] The classification result is verified and an accuracy rate is obtained, and if the accuracy rate does not reach a preset threshold, the classification level field and the regular expression are adjusted according to the verification result.
[0026] In a second aspect, the present application further provides a claim information processing device, the device comprising:
[0027] The conversion module converts all non-Chinese characters in the invoice item detail data into a preset format, and the preset format includes a half-angle format and a specific format.
[0028] The processing module is configured to obtain a corresponding regular expression set according to a pre-constructed search configuration table, the search configuration table includes at least two search rules, each search rule includes a search string, and the regular expression set includes at least two regular expressions, and the regular expressions are generated according to the search string.
[0029] The judgment module is configured to judge whether the invoice item detail data belongs to the claim type corresponding to the search configuration table according to the regular expression.
[0030] In a third aspect, the present application further provides an electronic device, the electronic device comprising a processor and a memory, and the processor is configured to execute a computer program stored in the memory to realize the invoice detail classification method as described in the first aspect.
[0031] In a fourth aspect, the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the invoice detail classification method as described in the first aspect.
[0032] The invoice detail classification method, device, electronic device and storage medium provided by the present application convert all non-Chinese characters in the invoice item detail data into a preset format, the preset format includes a half-angle format and a specific format; obtain a corresponding regular expression set according to a pre-constructed search configuration table, the search configuration table includes at least two search rules, each search rule includes a search string, and the regular expression set includes at least two regular expressions, and the regular expressions are generated according to the search string; and the invoice item detail data is matched and classified according to the regular expression set. The present application greatly improves the efficiency of invoice detail classification while reducing maintenance cost. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings needed in the embodiment description of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the field, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 is a flowchart of a classification method of an invoice detail provided by an embodiment of the present application.
[0035] Figure 2 is a schematic diagram of a search configuration table provided by an embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of the present application.
[0037] In the description of the embodiments of the present application, it should be understood that the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0038] The following description is given in order to enable any person skilled in the art to practice and use the present application. In the following description, details are set forth in order to explain the application. It will be apparent to those skilled in the art that the present application can be practiced without using these specific details. In other instances, well-known processes have not been described in detail in order to avoid unnecessarily obscuring the description of the embodiments of the present application. Therefore, the present application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0039] In the insurance claim business, accurately determining whether the customer has undergone a certain examination, used a certain drug, etc. is crucial for claim review. Currently, there are two main schemes for classifying claim invoice details in the insurance claim industry. One is the manual review scheme, in which the claim reviewer classifies and determines by visually checking the invoice details during manual review. This approach has obvious drawbacks, as the number of invoice details is large and complex, and manual review takes a lot of time, with extremely low work efficiency. The second is the fuzzy search scheme, which uses like fuzzy matching of keywords for judgment. Compared to manual review, this scheme has some improvements, but it requires a lot of manpower for maintenance, and when new or adjusted search keywords or exclusion keywords are needed, developers must manually update the code to re-accurately determine, which greatly increases the maintenance cost and time cost. Therefore, both of the above two schemes have serious deficiencies in efficiency and maintenance convenience.
[0040] Before presenting the embodiments, the following terms are explained.
[0041] Claim settlement is the core link of insurance business, which not only embodies the legal nature of contract performance, but also reflects the functional nature of risk management. Its process usually includes reporting, investigation, document review, review and settlement, etc. The insured or beneficiary should pay attention to timely reporting, keep complete evidence, and ensure the authenticity of information to protect their rights and interests.
[0042] Invoice item details can restore the real scene of the accident or treatment by recording specific medical items, drug names, unit prices and quantities, prevent fictitious expenses or exaggerated losses, and play multiple roles in insurance claim settlement, such as authenticity verification, responsibility definition, amount calculation, risk prevention and control. It is the core data link connecting insurance contracts and actual losses.
[0043] Regular expressions are a powerful text matching tool that describes the matching pattern of a string through specific syntax rules, which can be used for complex search, replacement and extraction operations in text. In this patent, it is used for accurate retrieval of keywords in invoice item detail data.
[0044] Semi-structured text information refers to text data with a certain structure but relatively flexible format. Invoice item detail data usually contains fixed field categories, but the specific content format may be diverse, such as examination item names, drug names, etc. The values of these fields may be of different lengths and formats, so they belong to semi-structured text information.
[0045] The full-width half-width conversion function is a self-defined function for converting full-width characters into half-width characters. Full-width characters refer to characters that occupy two standard character widths in a computer system, while half-width characters occupy one standard character width. Since the invoice item detail data is mostly of the full-width type, in order to unify the data format and facilitate subsequent processing, conversion from full-width to half-width is required.
[0046] The high-frequency non-semantic element is an element that frequently appears in a certain data set but does not have an explicit semantic value itself, such as some punctuation marks, stop words (such as "of", "is", "and", etc.) or certain formatting symbols (such as HTML tags) in the text.
[0047] The following describes the invoice item classification method, device, electronic equipment and storage medium of the present application in conjunction with the accompanying drawings of the specification to solve the above problems.
[0048] Referring to Figure 1 , Fig. 1 is a flow diagram of an invoice item classification method provided by an embodiment of the present application, as Figure 1 shown, the invoice item classification method includes the following steps: Figure 1
[0049] S100, convert all non-Chinese characters in the invoice item detail data into a preset format, the preset format including half-width format and specific format.
[0050] Specifically, the data source of the invoice item detail data includes an invoice item list (such as an invoice item name, an inspection item name, etc.) issued by a legal registered institution providing goods or services, and the legal registered institution providing goods or services includes a medical institution (such as a physical examination center, a hospital, etc.), a vehicle service institution (such as an automobile repair shop / 4S shop, a vehicle detection / evaluation institution, etc.), and a travel service institution (such as an airline, a railway company, a hotel, etc.). Among them, the data characteristics of the invoice item detail data are semi-structured text, and the invoice item detail data of the semi-structured text includes Chinese characters and non-Chinese characters, the Chinese characters include Chinese and Chinese symbols (such as Chinese quotation marks), and the non-Chinese characters include English, numbers, symbols (such as commas, parentheses, etc.), etc. All non-Chinese characters in the invoice item detail data are converted from full-width format to half-width format by using a full-width and half-width conversion function. For example, full-width numbers / letters / symbols are converted to half-width numbers / letters / symbols, for example, full-width comma “,” is converted to half-width comma “,”. In addition, the non-Chinese characters are converted from the original format to a specific format, and the specific format can be in lowercase format or in uppercase format. According to the needs, the non-Chinese characters are uniformly converted from the original uppercase format to the lowercase format, or the non-Chinese characters are uniformly converted from the original lowercase format to the uppercase format. For example, in the medical insurance claim settlement scenario, “CT” is converted to “ct”. Of course, the first space of the invoice item detail data can also be removed. In this way, converting all non-Chinese characters in the invoice item detail data to a preset format can unify the format of the non-Chinese characters in the invoice item detail data, eliminate text noise, and improve the accuracy of regular matching.
[0051] Since the invoice item detail data may contain interference information, it leads to misjudgment when matching and classifying. In order to solve the above problem, before S100, that is, before converting all non-Chinese characters in the invoice item detail data to a preset format, the following steps are included:
[0052] S010, setting a corresponding target category code according to batch historical invoice item detail data;
[0053] S020, screening target data corresponding to the target category code from the invoice item detail data.
[0054] Specifically, assuming that the invoice item detail data is a detail list, the detail list includes multiple pieces of detail data, each piece of detail data is a dictionary containing a detail description and a corresponding max_code, and the max_code is a field name representing the category code of the invoice item detail, which is used to distinguish different types of fees or services. The target category code is usually determined according to business needs or consultation of claim experts, or representative samples (recommended 5000+ historical invoice item detail data) can be collected for statistical analysis to identify the high-frequency semantic elements corresponding to the detail category code involved in the classification decision as the target category code. Of course, the target category code can also be determined according to any combination of business needs, consultation of claim experts, and statistical analysis. After determining the target category code, a filtering function corresponding to the target category code can be created, and the filtering function can be used to traverse the invoice item detail data to filter out the target data corresponding to the target category code from the invoice item detail data, so as to filter out the noise data corresponding to the noise category in the invoice item detail data and retain the target data corresponding to the target category code.
[0055] Of course, the corresponding target category code and noise category can also be set according to the batch of historical invoice item detail data, and the noise data corresponding to the noise category can be filtered from the invoice item detail data, and the target data corresponding to the target category code can be screened out. Among them, the noise category can be determined according to business needs or consultation of claim experts, or statistical analysis or correlation analysis can be performed on representative samples. The statistical analysis is to count the frequency of occurrence of each max_code, and filter out the several max_code with the lowest frequency of occurrence as the noise category, or to count the high-frequency non-semantic elements (such as punctuation marks, parentheses and their contents in parentheses, price units, check numbers, medical advice notes, extra spaces, etc.) that do not participate in classification decision as the noise category. The correlation analysis is to calculate the correlation of each max_code with the target category code, and filter out the several max_code with the lowest correlation with the target category code as the noise category. Of course, the noise category can also be determined according to any combination of business needs, consultation of claim experts, statistical analysis, and correlation analysis. The filtering function (such as regular expression) is generated according to the noise type to filter out the noise data corresponding to the noise category in the invoice item detail data.
[0056] For example, assume that the invoice item detail data of a medical institution includes a surgery fee (7), a medicine fee (10), an examination fee (15), a material fee (22), and the like. For example, if the target category code includes the surgery fee (7) and the examination fee (15), the invoice item detail data can be traversed using a filter function corresponding to the target detail category code, and it is checked whether each detail is in the target category code. If it is in the target category code, the detail data is retained; otherwise, the detail data is filtered out. Of course, the noise data corresponding to the noise category in the invoice item detail data can also be screened out using a filter function generated according to the noise type in the retained detail data, for example, the invoice item detail data retained after the preliminary screening in accordance with the target category code is "Painless colonoscopy (including anesthetic)
backup
backup
[0057] Before the non-Chinese characters in the invoice item detail data are converted, the noise data corresponding to the noise category in the invoice item detail data is screened out by using the set target category code, and the target data corresponding to the target category code is retained, so that unnecessary processing of irrelevant data is avoided, thereby improving the overall processing efficiency and reducing unnecessary consumption of computing resources. In addition, the interference of noise data on the classification result can also be reduced, the accuracy of regular expression retrieval is improved, and the misjudgment rate of the matching classification result is greatly reduced, thereby improving the accuracy and reliability of the classification matching. When it is judged whether a gastroscopic examination is performed, only the invoice details related to the gastroscopic examination are included in the screened data, and the interference of other irrelevant examinations or medicines is reduced.
[0058] S200, obtaining a corresponding regular expression set according to a pre-constructed search configuration table; the search configuration table includes at least two search rules, each search rule includes a search string, and the regular expression set includes at least two regular expressions, and the regular expressions are generated according to the search strings.
[0059] Specifically, as shown in Figure 2 Figure 2 An example diagram of a retrieval configuration table. Before generating a set of regular expressions used to match classifications in this application, a retrieval configuration table needs to be created in advance. Among them, the retrieval configuration table includes at least two retrieval rules, and each retrieval rule includes a retrieval string, and the retrieval string includes retrieval keywords connected by a preset delimiter. For example, a retrieval string is "inner_|_scope", where "_|_" is the preset delimiter, and "inner" and "scope" are the retrieval keywords. Among them, the regular expression is generated according to the retrieval string in the current retrieval rule, and one retrieval rule can correspond to generating one regular expression. This application can generate a corresponding set of regular expressions according to the retrieval rules in the pre-constructed retrieval configuration table. Among them, the number of regular expressions in the set of regular expressions is the same as the number of retrieval rules with a valid flag as the status identifier in the following embodiments. That is, if there are n retrieval rules with a valid flag as the status identifier in the retrieval configuration table, then the set of regular expressions corresponds to n regular expressions.
[0060] In some embodiments, each of the retrieval rules further includes a rule identifier, at least one classification level field, and a status identifier; the status identifier includes a valid flag and an invalid flag; S200 includes, that is, the obtaining the corresponding set of regular expressions according to the pre-constructed retrieval configuration table includes the steps:
[0061] S210. Screen out valid retrieval rules from all the retrieval rules in the retrieval configuration table, and the status identifier of the valid retrieval rule is the valid flag;
[0062] S220. Split the retrieval string of the valid retrieval rule according to the preset delimiter to obtain the corresponding retrieval keywords, generate the corresponding positive lookahead assertion according to the retrieval keywords, and determine the positive lookahead assertion as the regular expression;
[0063] S250. Bind the regular expression to the corresponding rule identifier and classification level field, and summarize the regular expressions corresponding to all the valid retrieval rules to obtain the set of regular expressions.
[0064] Specifically, as Figure 2 shown, each retrieval rule, in addition to including a retrieval string (keyword), further includes a rule identifier (keyword_id), at least one classification level field (type N), and a status identifier. Among them, the rule identifier is used to uniquely identify each retrieval rule. Among them, the classification level field is used to define different levels of classification, and the classification level field can help classify the invoice item detail data to be classified into different levels in the subsequent process, which is convenient for subsequent statistics and analysis. For example, as Figure 2As shown, type1 is a large category, type2 is a small category, and type3 is a more specific level, and more levels such as type4, type5, etc. can also be included. Among them, the state identifier (rule_status) is a flag used to mark whether the search rule is valid, and the state identifier includes a valid flag (for example, A represents validity) and an invalid flag (for example, T represents invalidity). After obtaining the search configuration table, the search configuration table is traversed, and for each search rule, the state identifier is first checked. If the state identifier is an invalid flag, the search rule is directly skipped, and no corresponding regular expression is generated. If the state identifier is a valid flag, it is determined that the search rule is a valid search rule. The set of regular expressions is dynamically generated in real time through the search rule in the configuration table, the business rule is converted into a configurable text matching logic, and efficient and accurate classification of invoice item detail data can be achieved.
[0065] In actual use, since the search string may have multiple search keywords spliced, the split function needs to be used for cutting parameter passing. The split function is usually used to split a string according to a specified preset delimiter, wherein the preset delimiter avoids conflict with common symbols in the text (such as comma, semicolon, and " / "), and an uncommon symbol such as " _|_ " is used. The key words required for judgment for a certain type of check are cut by the preset delimiter " _|_ ", and " _|_ " represents the relationship of "and", that is, multiple key word conditions need to be met at the same time. The exclusion string in the following embodiment also uses this usage. The search string in a valid search rule is split into multiple individual search keywords using the split function, and the processed search keywords are combined to form the positive lookahead assertion corresponding to the search string in the valid search rule. If the exclusion string is not included in the valid search rule, a regular expression is generated according to the positive lookahead assertion. For example, the positive lookahead assertion according to the search string part is: (?=.*A)(?=.*B), and the regular expression generated according to the positive lookahead assertion of the search string part is: (?i)^(?=.*A)(?=.*B).
[0066] Among them, (?i) is an inline modifier, indicating that the following match is case-insensitive. ^ is an anchor character to indicate the start of the matched string. (?=.*A) is a positive lookahead assertion that ensures that the search keyword "A" (case-insensitive) exists in the string. (?=.*B) is another positive lookahead assertion that ensures that the search keyword "B" (case-insensitive) exists in the string. This positive lookahead assertion checks the entire string to ensure that it contains "A" and "B" regardless of case.
[0067] Subsequently, the regular expression generated according to a valid search rule is bound with the rule identifier and the classification level field of the valid search rule, so that the regular expression, the rule identifier and the classification level field are associated with each other, and the original search rule corresponding to each regular expression can be quickly traced back. Thus, when the business requirement changes and the search rule needs to be modified or added, the corresponding regular expression can be quickly found through the rule identifier for corresponding adjustment. In this way, manual search and modification of the regular expression in a large amount of code can be avoided, and the possibility of error can be reduced. In addition, during the verification, debugging and testing process of the following embodiment, the specific regular expression can be quickly located through the rule identifier and the classification level field to check whether it correctly matches the expected data, which helps to quickly find and repair problems. For example, if "nasal endoscope" is misjudged as a gastrointestinal endoscope, the corresponding search rule can be quickly located through the rule identifier, and the exclusion string can be added or modified. Also, for example, when it is found that the "internal_|_mirror" rule misjudges an arthroscope, the status identifier can be immediately changed from the valid flag to the invalid flag, and the optimized rule can be added without the intervention of development.
[0068] The present application filters valid search rules with the status identifier being the valid flag from all search rules of the search configuration table, so that the subsequent search process only processes the rules that are actually needed to take effect, avoiding invalid operation on invalid rules and improving search efficiency. The search string of the valid search rule is divided into search keywords according to a preset delimiter, and the corresponding regular expression is generated, and then the regular expression is bound with the rule identifier and the classification level field, so that the invoice item detail data can be quickly and accurately classified, and the change of business requirement can be flexibly adapted, greatly reducing the maintenance cost and development difficulty.
[0069] In some embodiments, before S250, i.e., before the regular expression is bound with the corresponding rule identifier and classification level field, the regular expressions corresponding to all valid search rules are summarized to obtain the regular expression set, the method further includes the following steps:
[0070] S230, if the valid search rule includes an exclusion string, the exclusion string of the valid search rule is divided into corresponding exclusion keywords according to a preset delimiter;
[0071] S240, the corresponding negative look-ahead assertion is generated according to the exclusion keyword, and the regular expression is obtained according to the negative look-ahead assertion and the positive look-ahead assertion.
[0072] Specifically, the search rule in the search configuration table can or can not include an exclude keyword, which includes exclude keywords connected by a preset separator. Therefore, after determining the valid search rule and obtaining the positive lookahead assertion corresponding to the search string in the above manner, it can be determined whether the valid search rule includes an exclude keyword. If the valid search rule does not include an exclude keyword, a regular expression is directly obtained according to the search string according to the above embodiment. Conversely, if the valid search rule includes an exclude keyword, the exclude keyword in the valid search rule is also split into multiple individual exclude keywords using the split function, and the processed exclude keywords are combined. In this way, a negative lookahead assertion corresponding to the exclude keyword in the valid search rule is formed. Then, the positive lookahead assertion of the search keyword part and the negative lookahead assertion of the exclude keyword part are combined to form a complete regular expression. For example, the negative lookahead assertion generated according to the exclude keyword part is: (?!.*C)(?!.*D).
[0073] (?!.*C) is a negative lookahead assertion that ensures that the search keyword "C" (case-insensitive) exists in the string. (?!.*D) is another negative lookahead assertion that ensures that the search keyword "D" (case-insensitive) exists in the string. This negative lookahead assertion checks the entire string to ensure that it does not contain "C" or "D", regardless of case.
[0074] The present application splits the exclude keyword in the search rule according to the preset separator, thereby obtaining the specific exclude keyword, and generates the corresponding negative lookahead assertion according to the exclude keyword. The regular expression is obtained according to the negative lookahead assertion and the positive lookahead assertion, which can more accurately identify and exclude invoices that do not meet the conditions in the regular expression matching process, thereby improving the accuracy of the classification result. In addition, since the processing logic of the search string and the exclude keyword is completely controlled by the search configuration table, the claims officer can adjust the search rule at any time according to the actual demand without the need for the developer to modify the code. This configuration maintenance method greatly reduces the maintenance cost and can flexibly cope with various complex business demands.
[0075] S300, matching and classifying the invoice item detail data according to the regular expression set.
[0076] In some embodiments, the matching and classifying of the invoice item detail data is in a multi-thread parallel manner.
[0077] Specifically, a thread pool is used to manage threads to avoid the overhead caused by frequent creation and destruction of threads. The thread pool can create a certain number of threads in advance, and assign different matching tasks to these threads according to the regular expression set, that is, assign a thread to each matching task, and each thread is responsible for matching a regular expression with the invoice item detail data, wherein the regular expressions corresponding to different threads are different. Different threads store their matching results in a shared data structure, and after all threads complete the task, the main thread aggregates the matching results of each thread to generate the final classification result. The task of each thread of the present application is very clear, and only responsible for the matching of one regular expression, which reduces the task coordination complexity between threads, facilitates management and debugging, because the function of each thread is single, and when a problem occurs, it can be quickly located, and the matching of the invoice item detail data with multiple regular expressions can be processed in parallel, which significantly improves the classification matching speed of the invoice item detail data, while maintaining the scalability and maintainability of the system.
[0078] In some embodiments, S300 includes, that is, the matching classification of the invoice item detail data according to the regular expression set includes the steps of:
[0079] S310, judging whether the invoice item detail data matches the regular expression in the regular expression set;
[0080] S320, if the matching result is matched, classifying the invoice item detail data into the classification level field corresponding to the matched regular expression to obtain the classification result.
[0081] Specifically, the invoice item detail data is matched with a certain regular expression in the regular expression rule set and the matching result is output, and the data is classified according to the matching result to output the classification result. Wherein, the matching result includes that a certain regular expression does not match the invoice item detail data, and a certain regular expression matches the invoice item detail data. If a certain regular expression does not match the invoice item detail data, the classification result is that the invoice item detail data does not conform to the regular expression and the corresponding classification level field, and the invoice item detail data is not classified into the classification level field corresponding to the regular expression participating in the matching. Of course, if the matching result is matched, that is, a certain regular expression matches the invoice item detail data, the invoice item detail data is classified into the classification level field corresponding to the regular expression participating in the matching.
[0082] If the invoice item detail data meets the classification level fields of multiple search rules corresponding to multiple regular expressions at the same time, the processing mode of the classification result depends on the specific design and business requirements of the system.
[0083] One case is multiple classification results coexist, i.e. if business requirement allows, invoice line item data can be classified into multiple classification hierarchy fields, in this case, the classification results will contain all matched classification hierarchy fields. The same classification (e.g. gastroenteroscopy) in this application supports multiple search rules, and the invoice line item data details only need to meet any search rule to be classified into the corresponding classification hierarchy field.
[0084] Another case is the highest priority classification hierarchy field, i.e. if the business requirement requires that the invoice line item data can only be classified into one classification hierarchy field, the priority of each classification hierarchy field can be set, and the corresponding classification result will be matched according to the highest priority matching rule. For example, assuming that the invoice line item data is "patient undergoes gastroscopy and colonoscopy", the search rules include search rule 1: search string: stomach | _mirror, classification hierarchy field: type1 = endoscope, type2 = gastroenteroscopy, type3 = gastroscope, priority: high. Search rule 2: search string: intestine | _mirror, classification hierarchy field: type1 = endoscope, type2 = gastroenteroscopy, type3 = colonoscopy, priority: low. In this case, the classification will be performed according to the highest priority rule, and the classification result output is: type1 = endoscope, type2 = gastroenteroscopy, type3 = gastroscope.
[0085] Another case is the most fine-grained classification hierarchy field, if the business requirement needs the most fine-grained classification, the system can select the most fine-grained classification hierarchy field as the final classification result. For example, assuming that the invoice line item data is "patient undergoes gastroscopy and colonoscopy", the search rules include search rule 1: search string: stomach | _mirror, classification hierarchy field: type1 = endoscope, type2 = gastroenteroscopy, type3 = gastroscope. Search rule 2: search string: intestine | _mirror, classification hierarchy field: type1 = endoscope, type2 = gastroenteroscopy, type3 = colonoscopy. In this case, the most fine-grained classification hierarchy field will be selected for classification, and the classification result output is: type3 = colonoscopy, type3 = gastroscope.
[0086] The application matches the invoice item detail data with the regular expression in the regular expression set, judges whether the invoice item detail data matches the regular expression, and if the matching result is that the invoice item detail data meets the regular expression, the classification result is obtained by classifying the matched regular expression corresponding to the classification level field. Based on the automatic retrieval and classification method of the regular expression, the application can efficiently complete the matching and classification of the invoice item detail data, not only improves the processing efficiency, but also reduces the maintenance cost, and can flexibly adapt to the changes of business requirements. Finally, the system can quickly and accurately identify whether the invoice item detail data meets the classification result set by the insurance claim item, and provide strong support for subsequent insurance claim review.
[0087] The application can avoid retrieval errors caused by inconsistent data formats by converting all non-Chinese characters in the invoice item detail data into a preset format to ensure data consistency. The standardized data reduces the complexity caused by format differences, making the regular expression matching more efficient and improving the retrieval efficiency. In addition, after unifying the format, it can avoid misjudgment caused by case or full-angle / half-angle difference, and improve the accuracy of subsequent classification. The application defines the retrieval rules by retrieving the configuration table, and business personnel can adjust the retrieval keywords and exclusion keywords in the retrieval rules at any time according to actual needs, without the need for developers to modify the code, which greatly improves the flexibility and scalability of the system. The application matches and classifies the invoice item detail data according to the regular expression set, which can quickly, accurately and automatically process a large amount of data to classify the invoice details, reduce the workload of manual review, and reduce the labor cost. Compared with manual review or simple fuzzy matching, the classification efficiency is greatly improved, the claim review process is accelerated, the processing speed of insurance business is improved, and the actual needs of insurance claim business are met.
[0088] For example, if a search rule includes both a search string and an exclusion string, the regular expression generated according to the search rule is used to match the invoice item detail data. If the invoice item detail data matches the search keyword corresponding to the regular expression and does not contain the exclusion keyword corresponding to the regular expression, it is determined that the invoice item detail data meets the classification. For example, when determining whether a customer has undergone a gastroscopy, the search configuration table is used to search for whether the invoice item detail data contains both
A
B
C
D
E
F
G
E_|_F_|_G
[0089] In actual applications, regular expression matching may be misjudged due to the complexity of the rules or the diversity of the data. To solve the above problems, in some embodiments, the matching and classification of the invoice item detail data according to the set of regular expressions includes the following steps:
[0090] The classification results are verified and the accuracy is obtained. If the accuracy does not reach a preset threshold, the classification hierarchy field and the regular expression are adjusted according to the verification results.
[0091] Specifically, manual sampling inspection verification can be performed, that is, a certain number of samples are randomly selected from each classification result, and it is manually checked whether these samples are correctly classified, and the number of correct classifications is calculated. The accuracy is calculated by the number of correct classifications and the total number of classification results. Of course, cross-validation methods can also be used for automatic verification, that is, the data set is divided into a training set and a test set, classification rules are generated on the training set, and then the accuracy of the classification results is verified on the test set. If the accuracy does not reach a preset threshold, either one or both of the regular expression and the classification hierarchy field are adjusted according to the verification results to improve the accuracy of the classification. For example, it may be necessary to add or modify search keywords and exclusion keywords, or adjust the classification hierarchy field (such as type1, type2, type10) according to business requirements to more accurately reflect the classification of the data. Of course, if there is a user interaction interface, user feedback on the classification results can be collected and adjusted according to the user feedback. Classification rules can also be reviewed and updated regularly to adapt to changes in business requirements.
[0092] This application uses a verification step to check whether the classification results meet expectations. If the verification results show that the accuracy of the classification results does not reach the preset threshold, the classification logic can be adjusted based on the verification results. For example, the regular expressions can be modified or the classification hierarchy fields can be adjusted. This can reduce misjudgments caused by improper regular expression design and improve classification accuracy. Moreover, if certain regular expressions cause performance issues (e.g., excessive nesting leading to decreased matching speed), these problems can be identified through the verification step, and the regular expressions can be optimized or the matching logic adjusted to improve the system's performance and reliability. Regularly verifying the classification results ensures that the system's classification logic always meets business requirements, and based on the verification results, the classification logic can be continuously optimized to better adapt to changes in data and business needs.
[0093] For example, the method of this invention was fully applied in the Smart Claim III project. The factors in the hospital behavior model of this project included the judgment of two categories of factors: the rationality of examinations and tests, and the rationality of treatments. These included, but were not limited to, judgments on various tests such as gastroscopy, colonoscopy, cervical spine MRI, lumbar spine MRI, cranial MRI, blood typing, glycated hemoglobin, echocardiography, chest CT, rabies vaccination, eight infectious disease tests, myocardial enzymes, thyroid function tests, specific expensive drugs, specific expensive medical devices, and specific expensive treatments. All of these judgments were implemented using the regular expressions of this invention, enabling the successful development of these factors and their important role in actual modeling. One factor is listed here for detailed explanation.
[0094] Factor: [Number of hospital visits involving gastroscopy and colonoscopy within xxx days].
[0095] Factor logic explanation: This factor aims to determine the percentage of visits involving gastroscopy and colonoscopy within a certain number of days at each hospital. The determination of whether a hospital performed gastroscopy or colonoscopy requires checking if the invoice item details contain the search keywords "*gastroscopy*", "*gastroscop*", "*colonoscop*", "*enteroscop*", or "Endoscope". If these keywords are present, the hospital is considered to have performed gastroscopy or colonoscopy.
[0096] Claim users can set corresponding search keywords and exclusion keywords for different examination items (such as gastroscopy, colonoscopy, cervical spine MRI, etc.) in the detailed search configuration table according to their business needs. For example, for gastroscopy, the search keywords can be set as "gastroscopy" and "endoscopy," and exclusion keywords (if any) can be set according to the actual situation. Business notes should also be added to clarify the classification criteria. See below for the search strings composed of search keywords for gastroscopy in the search configuration table. Figure 2The system converts all full-width characters in the invoice item detail data into half-width characters according to a self-defined full-width and half-width conversion function after obtaining the invoice item detail data. Then, the system searches the converted invoice item detail data according to a regular expression generated according to a search string set in a search configuration table. For example, when determining whether a piece of invoice item detail data belongs to the category of gastroscopy, the system searches whether the piece of invoice item detail data contains both the keyword
stomach
scope
stomach
scope
[0097] The present application also provides a claim information processing device, which comprises:
[0098] A conversion module configured to convert all non-Chinese characters in the invoice item detail data into a preset format, wherein the preset format comprises a half-width format and a specific format;
[0099] A processing module configured to obtain a corresponding set of regular expressions according to a pre-constructed search configuration table, wherein the search configuration table comprises at least two search rules, each search rule comprises a search string, and the set of regular expressions comprises at least two regular expressions, which are generated according to the search string;
[0100] A judgment module configured to determine whether the invoice item detail data belongs to a claim type corresponding to the search configuration table according to the regular expression.
[0101] The present application also provides an electronic device comprising a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the classification method of invoice details as described in the above embodiments.
[0102] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the classification method of invoice details as described in the above embodiments.
[0103] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments. It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, the specific working process of the data acquisition system and the corresponding unit described above and the beneficial effects brought by the same can be referred to the description of the manufacturing method of the cylindrical array ultrasonic transducer in the above embodiments, and will not be described here in detail.
[0104] The classification method and device of the invoice details, the electronic device and the storage medium provided by the embodiments of the present application are described in detail above, the principle and implementation manner of the present application are described by applying specific examples, and the above embodiment is only used to help understand the method and core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will be changed, and the above description should not be understood as the limitation of the present application.
Claims
1. A method for classifying invoice details, characterized in that, Including the following steps: Convert all non-Chinese characters in the invoice item details data into a preset format, which includes half-width format and specific format; The corresponding set of regular expressions is obtained based on the pre-built retrieval configuration table; The search configuration table includes at least two search rules, each search rule includes a search string, and the regular expression set includes at least two regular expressions, which are generated based on the search string. The invoice item details data are matched and categorized based on the set of regular expressions.
2. The method for classifying invoice details according to claim 1, characterized in that, Before converting all non-Chinese characters in the invoice item detail data to a preset format, the following steps are included: Set the corresponding target category code based on the batch of historical invoice item details data; Filter the target data corresponding to the target category code from the invoice item details data.
3. The method for classifying invoice details according to claim 1, characterized in that, Each of the aforementioned search rules further includes a rule identifier, at least one category level field, and a status identifier; the status identifier includes a valid flag and an invalid flag; obtaining the corresponding set of regular expressions based on the pre-built search configuration table includes the following steps: Valid search rules are selected from all the search rules in the search configuration table, and the status identifier of the valid search rule is the valid flag; The search string of the effective search rule is segmented according to the preset delimiter to obtain the corresponding search keywords. A corresponding positive look-ahead assertion is generated according to the search keywords, and the positive look-ahead assertion is determined to be the regular expression. The regular expression is bound to the corresponding rule identifier and the classification level field, and the regular expressions corresponding to all the valid retrieval rules are summarized to obtain the regular expression set.
4. The method for classifying invoice details according to claim 3, characterized in that, Before binding the regular expression with the corresponding rule identifier and the classification hierarchy field, the method further includes the following steps: If the valid retrieval rule includes an exclusion string, the exclusion string of the valid retrieval rule is segmented according to a preset delimiter to obtain the corresponding exclusion keywords; The corresponding negative lookahead assertion is generated based on the exclusion keyword, and the regular expression is obtained based on the negative lookahead assertion and the positive lookahead assertion.
5. The method for classifying invoice details according to claim 1, characterized in that, The method for matching and classifying the detailed data of the invoice items is a multi-threaded parallel approach.
6. The method for classifying invoice details according to any one of claims 1 to 5, characterized in that, The step of matching and classifying the invoice item detail data according to the set of regular expressions includes the following steps: Determine whether the invoice item details data matches the regular expression in the regular expression set; If the matching result is a match, the detailed data of the invoice items will be classified into the category level field corresponding to the matching regular expression to obtain the classification result.
7. The method for classifying invoice details according to claim 6, characterized in that, The step of matching and classifying the invoice item detail data according to the set of regular expressions includes the following steps: The classification results are verified and the accuracy is obtained. If the accuracy does not reach a preset threshold, the classification level field and the regular expression are adjusted according to the verification results.
8. A device for classifying invoice details, characterized in that, The device includes: The conversion module converts all non-Chinese characters in the invoice item details data into a preset format, which includes half-width format and specific format. The processing module is used to obtain a set of regular expressions based on a pre-built search configuration table; the search configuration table includes at least two search rules, each search rule includes a search string, and the set of regular expressions includes at least two regular expressions, which are generated based on the search string; The judgment module is used to determine, based on the regular expression, whether the invoice item details data belongs to the claim type corresponding to the retrieval configuration table.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the processor executing a computer program stored in the memory to implement the invoice details classification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the method for classifying invoice details as described in any one of claims 1 to 7.