Intelligent identification method for wine reimbursement violation based on rule and large model fusion
By employing a layered collaborative architecture of rule engine and large language model, the problems of limited scenario coverage, low processing efficiency, and lack of evidence in the price compliance detection of liquor invoices are solved. This achieves efficient and accurate compliance detection and traceable conclusions, making it suitable for market supervision and corporate financial auditing.
Patent Information
- Application Number
- CN202511610003.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies for detecting compliance with liquor invoice pricing have limitations in terms of scenario coverage, processing efficiency, and the potential for misjudgments and unsubstantiated conclusions in feature extraction, making it difficult to meet the needs of market supervision and corporate financial auditing.
It adopts a layered collaborative architecture based on a rule engine and a large language model. The rule engine quickly filters simple scenarios, while the large language model handles complex scenarios. It constructs multi-dimensional compliance judgment logic and combines data preprocessing and feature correction strategies to achieve fault-tolerant processing of abnormal invoice data and category-differentiated judgment.
It significantly improves the scope of scenarios covered by the compliance detection of liquor invoice prices, the accuracy of statistical judgment, and the efficiency of data processing, and achieves the traceability of compliance conclusions, adapting to the actual needs of tax supervision and corporate financial auditing.
Smart Images

Figure CN121330705A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the cross field of artificial intelligence and commodity price compliance detection, in particular to a wine price compliance intelligent judgment method based on "rule engine + large language model". BACKGROUND
[0002] With the multi-dimensional extension of wine commodity circulation scenarios to online sales, offline retail and inter-enterprise bulk trade, invoices as the core evidence for transaction compliance tracing, the "chaotic nature" of the recorded information and the "multi-scenario checking needs" have put higher requirements on the statistical analysis efficiency and scene adaptability of price compliance detection. On the one hand, the wine information in the invoice often presents a state of redundancy information and key feature information mixed together, among which the redundancy information covers irrelevant descriptions, emotional expressions, etc., while the key feature information includes core contents such as product specifications, product quantities, etc. At the same time, the non-standardized expression caused by manual input difference also significantly increases the difficulty of extracting useful information - the traditional data processing method lacks efficient statistical screening logic, making it difficult to effectively eliminate data noise and accurately focus on core variables, which directly leads to limited sample quality and processing efficiency of subsequent compliance analysis. On the other hand, invoice price compliance needs to complete "single ticket, single product, single price" and "bulk discount single price" double-dimensional statistical verification. Although traditional technology can complete such statistical calculation through manual verification or single rule, it has the defect of low sample processing efficiency, which is difficult to adapt to the batch detection needs of large sample invoices, and the scene adaptability of statistical verification logic is weak, which cannot flexibly cope with the statistical caliber differences of different categories and different packaging specifications, ultimately making it difficult to meet the core needs of statistical analysis timeliness and accuracy for efficient compliance detection.
[0003] Current judgment techniques for invoice wine price compliance fall into two categories. The first is the traditional rule-driven scheme, which relies on manual preset static threshold and regular matching to extract invoice information. Although it has basic efficiency, it has three bottlenecks: (1) it can only handle simple scenarios with explicit feature annotation and unit standardization of invoice information, and cannot handle implicit semantics or non-standard units; (2) it does not establish a differentiated judgment system for wine sub-categories, which is prone to misjudgment due to differences in category benchmarks; (3) the rule maintenance cost grows exponentially with the expansion of invoice formats and wine categories, and has poor scalability. The second is the pure large language model-driven scheme, which directly calls the model to parse invoices, extract features and judge compliance. However, it has obvious shortcomings: the time consumption of a single call is much higher than that of the rule scheme, which cannot meet the batch audit needs of thousands of invoices per day; it is prone to generate false features that do not match the actual records of the invoice due to "hallucination", and lacks structured parsing ability, making it difficult to associate the "product name-amount" correspondence; it only outputs conclusions and cannot trace the judgment basis, which does not meet the requirements of "traceability" of tax supervision and enterprise risk attribution.
[0004] In addition, both solutions lack the ability to handle invoice data anomalies, and are prone to triggering system anomalies when faced with non-numerical amounts, invalid units, etc. Furthermore, they have not constructed a closed loop for the entire process of "invoice data governance - feature extraction - compliance verification - conclusion interpretation". They have neither taken into account the characteristics of concise and diverse invoice information, nor formed a hierarchical collaborative judgment logic, which makes it difficult to meet the actual requirements in terms of accuracy, efficiency and interpretability. Summary of the Invention
[0005] Regarding the application of existing technologies in the compliance detection of alcohol prices, rule-driven solutions can only cover simple scenarios with explicit information and standardized units, making it difficult to adapt to the statistical analysis needs of non-standard information; pure model-driven solutions suffer from low processing efficiency and the tendency for feature extraction to produce "illusions." Furthermore, both types of solutions have the problem of traceability defects due to the lack of statistical evidence to support compliance conclusions. This invention aims to provide an intelligent identification method for alcohol reimbursement violations based on the fusion of rules and large models. By using a hierarchical collaborative intelligent judgment of alcohol price compliance based on a "rule engine + large language model," it significantly improves the scenario coverage, statistical judgment accuracy, data processing efficiency, and conclusion traceability of invoice alcohol price compliance detection, thereby meeting the core needs of market supervision departments and corporate financial audits for statistical compliance and judgment accuracy.
[0006] This invention first uses a layered collaborative architecture to enable the rule engine to quickly filter simple scenarios and then uses a large language model to fill in gaps for in-depth analysis of complex scenarios. Combined with a systematic data preprocessing mechanism, it achieves statistical fault tolerance for abnormal invoice data. Relying on a dynamic feature correction strategy, it adapts the statistical logic of the correlation between units and quantities. Then, through a multi-dimensional compliance verification system, it forms statistical standards with category differentiation. Ultimately, it solves the core problems of traditional rule solutions having limited scenario coverage, pure model solutions having low processing efficiency, and insufficient reliability of results. This enables intelligent, statistically accurate, and traceable judgment of liquor prices on invoices.
[0007] Based on the above problems, the present invention provides an intelligent identification method for violations of alcohol reimbursement rules based on the fusion of rules and large models, which includes the following main steps.
[0008] S1: Receive and process invoice data: Based on OCR or multimodal large language model, perform structured parsing on the invoice content uploaded by the user to generate structured data containing information such as product name, model specifications, quantity, amount, and unit of goods.
[0009] S2: Standardization Processing to Generate Unit Price and Extract Capacity: Based on the structured data in S1, the quantity and amount fields are first standardized to calculate and generate the unit price of the product; then, the product capacity is extracted through structured text parsing technology. If there is no explicit capacity identifier, the industry-standard benchmark value of the liquor industry is used to supplement it, and the standardized unit price and capacity data are output.
[0010] S3: Establish compliance determination criteria: According to the "Provisions for the Management of Business Entertainment of State-owned Enterprises", two types of compliance determination criteria are established - ① Low price and regular scene criteria: the core basis is the unit price of the commodity and the unit of the commodity; ② Wine commodity volume and category price criteria: the core basis is the unit price of the commodity, volume and wine category (wine category is obtained from the commodity name, model specification).
[0011] S4: Output preliminary compliance results: Based on the determination criteria of S3, combined with the unit price and volume data generated in S2, determine: if any of the criteria are met, output the determination status as normal and provide compliance basis; if none of the criteria are met, proceed to the next step.
[0012] S5: Build core feature extraction model: Based on the prompt engineering and public large language model, generate high-quality training corpus, train the core feature extraction model, and use it to extract core feature data such as product type, quantity, and actual volume from the structured data obtained in S1.
[0013] S6: Compliance determination for complex scenarios: For each invoice, first obtain the core features extracted by the S5 model, such as product type, actual quantity, and actual volume. Then, for complex scenarios such as unit abnormalities, name and unit mismatches, and non-standard units, build a multi-dimensional compliance determination logic. Substitute the extracted core features into the multi-dimensional compliance determination logic for determination, and finally output the determination status (normal / over standard / suspected) and determination basis of the invoice.
[0014] Further, in the S1 step, the main content is to receive the user uploaded wine invoice, through format unification, image processing, text recognition and other operations, the invoice information is parsed into structured data such as commodity name, model specification, quantity, amount, unit, etc. to provide basic data support for subsequent standardized processing and compliance determination, which includes the following key sub-steps: S11: Invoice reception and format unification: Receive user uploaded wine invoices, receive formats including pictures, PDF, etc. Classify and organize invoices of different formats, split multi-page PDF invoices into single pages, ensure that each invoice corresponds to an independent processing object, and store them in the invoice temporary folder and mark them with a unique number; S12: Invoice image preprocessing: For invoice images in picture and PDF format, first correct the tilted invoice to horizontal posture through perspective transformation, and adjust different resolution images to the preset standard size; then use Gaussian filter to remove watermarks, printing stains and other noise, crop the edge blur area, and output clear standardized invoice images; S13: Key text optical character extraction: Based on the standardized image output by S12, use OCR technology to extract the "product name", "model specification", "quantity", "amount", and "unit" five types of key text in the invoice. The extraction range covers the invoice commodity detail column and the summary column. The extracted raw text is temporarily stored according to the invoice number; S14: Text matching with preset fields: The five fixed fields of "product name", "model specification", "quantity", "amount", and "unit" are preset. According to the semantic and positional rules of the invoice text, the raw text extracted by S13 is matched to the corresponding field one by one to form the initial field data; S15: Preliminary cleaning of field data: Basic cleaning of matched field data, stripping non-numeric characters from text data in "quantity" and "amount" fields and converting them to numeric type; output the preliminary cleaned field data; S16: Structured data output: The field data cleaned initially is integrated into structured data containing "product name", "model specification", "quantity", "amount", "unit", and "invoice number". It supports JSON format storage and is delivered to S2 step.
[0015] Further, in S2 step, the main content is based on the structured data output by S1 to calculate the unit price and extract standardized capacity data, providing accurate numerical support for subsequent compliance judgment, which includes the following key sub-steps: S21: Basic data validity check: Check the "quantity", "amount", and "model specification" fields output by S1. Quantity must be a positive integer, amount must be a non-negative number, and model specification must contain extractable capacity information or conform to industry conventions; S22: Standardized calculation of unit price: The quantity and amount of valid data are uniformly rounded to two decimal places, and the formula "unit price = amount ÷ quantity" is used for calculation; S23: Directional extraction of capacity information: Extract the capacity values containing ml and liter indicators from "model specification" or "product name" in S1. Output the original capacity data in integer form; S24: Missing capacity information supplement: For goods without extracted capacity, supplement according to the industry benchmark values of 500ml for white wine and 750ml for red wine. Output standardized capacity data (unit: ml); S25: Standardized data integration output: Integrate unit price, capacity, quantity, and amount to form a standardized data set. Synchronize to S3 step.
[0016] Further, in S3 step, the main content is to clarify the compliance judgment standards for low price and regular scenarios, and wine products according to the "Provisions on the Management of Business Entertainment of State-owned Enterprises", forming a structured judgment basis and providing a unified standard for subsequent compliance result output, which includes the following key sub-steps: S31: Comply with file clause extraction: Sort out the clauses related to the price, volume, and category of wine products in the "Provisions on the Management of Business Entertainment of State-owned Enterprises", extract the specific content of the low-price standard, the conventional unit requirement, the upper limit of the unit price of different wine categories, and the corresponding volume reference, and form a clause summary document; S32: Low-price and conventional scene standard refinement: Based on the extracted clause summary document, the unit price threshold of the low-price scene and the product unit range (such as bottle, box) of the conventional scene are determined; the condition of "unit price not exceeding the threshold and unit within the conventional range" is determined as the compliance condition of the scene, and the judgment rule is formed which can be directly called; S33: Wine product standard refinement: According to the clause summary document, the categories of white wine and red wine are distinguished, and the upper limit of the unit price and the corresponding volume reference of each category are determined; the condition of "unit price not exceeding the corresponding category limit and volume meeting the reference" is determined as the compliance condition of wine products, and the specific judgment values of different categories are refined; S34: Standard structured storage: The low-price and conventional scene standards, and the wine product standards are arranged into a structured data table containing "scene type, judgment condition, specific value, and clause basis", which supports quick query by category and scene, and is synchronized to the judgment module for S4 calling.
[0017] Further, in the S4 step, the main content is to combine the unit price and volume data output by S2 with the compliance judgment standards determined by S3 to complete the preliminary compliance judgment and output the results. If it meets the standards, the compliance basis is given, otherwise it is transferred to the subsequent link. The specific steps include the following key sub-steps: S41: Judgment data association matching: The "unit price, volume, product unit, and product type" data output by S2 are matched with the "low-price and conventional scene standards, and wine product standards" of S3 according to the product type and product unit, ensuring that each product can be associated with the corresponding judgment standard; S42: Low-price and conventional scene compliance judgment: According to the matched low-price and conventional scene standards, if the unit price of the product does not exceed the threshold and the product unit is within the conventional range, the judgment state is directly output as normal, and the compliance basis is also marked; S43: Wine scene compliance judgment: If the low-price and conventional scene standards are not met, the wine product standards are used for judgment. If the unit price does not exceed the corresponding category limit and the volume meets the reference, the judgment state is output as normal, and the compliance basis is marked; S44: Non-compliance data marking and transfer: If neither the low-price and conventional scene standards nor the wine scene standards are met, the product data is marked as "further verification required", and the original data such as unit price, volume, and product unit are associated, and are simultaneously transferred to S5.
[0018] Further, in step S5, the main content is to build a core feature extraction model: first, define the extraction target (product type, actual quantity, actual capacity) and data format through demand definition, provide training data through corpus preparation, then adapt the prompt engineering design instructions, complete the training fine-tuning and performance verification with the large language model, and finally extract the structured core feature data accurately through the deployment of the product name and model specification text in the "further verification" data transferred from S4, and synchronize to S6 to provide data support for complex scenario compliance judgment, including the following key sub-steps: S51: Model requirement definition: determine the core features that the model needs to extract as product type, actual quantity, and actual capacity, and clearly define the input data format and output data format that the model needs to adapt, and set the performance index requirements for model accuracy and recall rate; S52: Training corpus preparation: collect product name and model specification text data in historical wine invoices, organize them in "input text-feature annotation" format, and generate high-quality annotated corpus; remove duplicates, clean up, and exclude invalid text, and divide the corpus into training set, validation set, and test set; S53: Prompt engineering design: design prompt word templates according to feature extraction requirements, and clearly require the model to identify product type, analyze actual quantity and unit, extract actual capacity value, and judge product unit, to ensure that the prompt word instructions are clear and unambiguous; S54: Model training and fine-tuning: based on the basic large language model, combine individual samples in the training set with the prompt word templates designed in S53, input the large language model to get the feature extraction results of the sample; compare the extraction results with the sample annotation features, use the combination loss function of cross-entropy loss and mean square error, calculate the loss value; adjust the model parameters through back propagation to adapt to the feature extraction task, and simultaneously monitor the model performance in real time with the validation set to optimize the fine-tuning strategy (such as adjusting the learning rate) until the indicators meet the requirements; S55: Model performance evaluation: use the test set to evaluate the trained model, detect whether the accuracy and recall rate of core feature extraction meet the preset indicators; analyze the extraction errors, supplement the corpus and fine-tune the model again; S56: Model deployment: deploy the qualified model to the data processing platform, configure the interface to receive the product name and model specification text in the "further verification" data transferred from S4, output structured core feature data, and synchronize to S6.
[0019] Further, in the S6 step, the main content is to complete the complex scene compliance judgment based on the core feature data output in S5 and the original data (data that neither meets the low price and regular scene standard nor meets the wine scene standard) transferred in S4, generate a judgment result containing state and reason, and specifically includes the following key sub-steps: S61: Data integration and format processing: associate the original data transferred in S4 with the core feature data output in S5 according to the invoice number, convert the quantity in the original data into an integer and the amount into a floating point number to ensure that the data format is uniform and computable; S62: Unit price secondary calculation: calculate according to the formula unit price = amount ÷ quantity, if the quantity is empty, 0 or a non-numeric value, set the quantity to the default quantity 1, the calculation result is rounded to the standard decimal place, and is used as the basis data for judgment; S63: Feature data association and standard matching: extract the type, quantity, and capacity fields in the core feature data, and match the preset wine standards (including the capacity benchmarks of different types of wine and the upper limit of unit price) and price policies (the highest price per milliliter, the highest price for a whole box, etc.); S64: Complex scene classification judgment: execute judgment for different complex scenes: if the actual quantity is 1 when the unit is empty, mark as suspected; if the unit is bottle when the commodity name contains whole box or box, mark as suspected; if the unit is box and the actual quantity is 1, adjust to 2 by default and mark the adjustment identifier; if the unit is not box or bottle and the unit price exceeds the standard, mark as suspected; S65: Judgment state and single bottle price determination: compare the calculated unit price with the matching standard: if it does not exceed the standard, mark as normal; if it exceeds the standard and the unit is bottle or box (quantity is clear), mark as exceeding the standard; if it cannot be determined due to abnormal unit or questionable quantity, mark as suspected; simultaneously calculate the single bottle price (total amount ÷ actual quantity); S66: Commodity compliance result output: generate the result in the state + reason format, and the normal needs to explain that the unit price does not exceed the standard, and the exceeding the standard and suspected states need to explain the abnormal reasons.
[0020] The application also provides a server, characterized by comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program comprises instructions for executing the above method.
[0021] The application also provides a computer readable storage medium having a computer program stored thereon, characterized in that the computer program is executed by a processor to implement the above method.
[0022] Compared with the prior art, the application has the following positive effects: Compared with existing technologies, the core strategic innovation of this invention lies in adopting a layered collaborative architecture of "rule engine + large language model": first, the rule engine quickly filters simple scenarios such as low prices and regular units, and then the large language model fills in the gaps to handle complex scenarios such as non-standard units and name-unit mismatches. This not only breaks through the limitation of traditional rule-driven solutions that can only handle explicit standardized information, but also avoids the time-consuming batch processing of pure large language model-driven solutions, significantly improving the breadth of scenario coverage and data processing efficiency. In the feature extraction stage, by prompting clear feature extraction instructions in the engineering design, combined with a loss function, and combining historical invoice corpus annotation and secondary fine-tuning to train the core feature extraction model, the cost of pure large language model-driven solutions is effectively reduced. The system addresses the "illusion" risk of language models and establishes differentiated compliance judgment rules for specific categories such as baijiu and red wine based on the "Regulations on the Management of Business Entertainment in State-owned Enterprises." This avoids misjudgments caused by traditional rules that do not differentiate between categories and allows for flexible adjustments as policies are updated or categories expand. It also overcomes the limitation that the maintenance cost of traditional rules increases exponentially with the growth of categories. Furthermore, a full-process data anomaly handling strategy is constructed through data cleaning, supplementation of missing capacity benchmark values, and marking of "suspected" status in questionable scenarios. This avoids invalid data triggering system anomalies, enhances system robustness, and records the rules on which the judgment is based from the initial judgment to the complex judgment, making compliance conclusions traceable and fully adaptable to the needs of tax supervision and enterprise risk attribution. This innovative strategy ultimately solves the problems of narrow scenario coverage in traditional rule-driven solutions, low efficiency and susceptibility to "illusions" in pure large language model-driven solutions, and the lack of evidence for conclusions, poor tolerance for data anomalies, and high maintenance costs in both types of solutions. It achieves the goals of broad scenario coverage, high processing efficiency, accurate judgment, traceable conclusions, strong system robustness, and scalability for liquor invoice price compliance detection, and can be directly adapted to the actual needs of corporate financial audits and regulatory compliance detection. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the specific implementation process. Detailed Implementation
[0024] To further illustrate the technical solution of this invention, the invention will be described in detail below with reference to specific implementation processes. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0025] The detailed implementation flowchart of the intelligent judgment method for the compliance of liquor prices in invoice information based on the hierarchical collaboration of "rule engine + large language model" described in this invention is attached. Figure 1 As shown, attached Figure 1The specific implementation of each step will be described in detail below in combination with the flow and actual invoice data processing scenarios.
[0026] S1: Receiving and processing invoice data: based on OCR or multi-modal large language model, the invoice content uploaded by the user is structured and parsed to generate structured data containing information such as product name, model specification, quantity, amount, unit, etc. The specific implementation includes the following key sub-steps.
[0027] S11: Invoice receiving and format unification: receiving the user uploaded wine invoice, the receiving format covers types such as pictures and PDFs. Classify and organize invoices of different formats, split multi-page PDF invoices into single pages, ensure that each invoice corresponds to an independent processing object, and store them in the invoice temporary folder and mark them with a unique number.
[0028] S12: Invoice image preprocessing: for the invoice images converted from pictures and PDFs, first correct the inclined invoice to a horizontal posture through perspective transformation, and adjust different resolution images to the preset standard size; then use Gaussian filtering to remove watermarks, printing stains and other noise, crop the edge blur area, and output clear standardized invoice images.
[0029] (1) Key text optical character extraction: based on the standardized images output by S12, use OCR technology to extract the "product name", "model specification", "quantity", "amount", and "unit" five types of key text in the invoice, the extraction range covers the invoice commodity detail column and the summary column, and the extracted original text is temporarily stored according to the invoice number.
[0030] (2) Text and preset field matching: preset five fixed fields of "product name, model specification, quantity, amount, and unit", according to the semantic and position rules of the invoice text, match the original text extracted by S13 to the corresponding field one by one to form the initial field data.
[0031] (3) Field data preliminary cleaning: basic cleaning of the matched field data, stripping non-numeric characters in the "quantity" and "amount" fields and converting them to numeric type; output the preliminary cleaned field data.
[0032] (4) Structured data output: integrate the field data cleaned initially into structured data containing "product name, model specification, quantity, amount, unit, and invoice number", support JSON format storage, and pass to S2 step.
[0033] S2: Standardization processing: Generate unit price and extract capacity: Based on the structured data of S1, first standardize the quantity and amount fields, calculate the unit price of the product, and then extract the product capacity through structured text analysis technology. If there is no clear capacity identifier, use the industry general benchmark value to supplement, output the standardized unit price and product capacity data. Specifically, it includes the following key sub-steps: (1) Basic data validity check: Check the "quantity", "amount", "model specification" fields output by S1; Quantity must be a positive integer, amount must be a non-negative number, and model specification must contain extractable capacity information or conform to industry conventions; (2) Unit price standardization calculation: The number and amount of valid data are unified to 2 decimal places, and calculated according to the formula "unit price = amount ÷ quantity"; (3) Capacity information directional extraction: Extract the capacity value containing ml, liter identifier from "model specification" or "product name" in S1. Output the original capacity data in integer form; (4) Missing capacity information supplement: For products without extracted product capacity, supplement with the industry benchmark value of 500ml for white wine and 750ml for red wine, and output the standardized capacity data (unit: ml); (5) Standardized data integration output: Integrate unit price, product capacity, quantity, and amount to form a standardized data set. Synchronize to S3 step.
[0034] S3: Develop compliance determination standards: According to the "Provisions for the Management of Business Entertainment of State-owned Enterprises", establish two types of compliance determination standards - ① Low price and conventional scene standard: The core basis is the unit price of the product and the unit of the product; ② Wine product capacity and category price standard: The core basis is the unit price of the product, capacity and wine category. Specifically, it includes the following key steps: (1) Extracting file clauses for compliance: Analyze the clauses related to wine product price, capacity, and category in the "Provisions for the Management of Business Entertainment of State-owned Enterprises", extract the specific content related to low price standard, conventional unit requirement, different wine price ceiling, and capacity corresponding benchmark, and form a clause summary document; (2) Low price and conventional scene standard refinement: Based on the extracted clause summary document, determine the unit price threshold for low price scenarios and the product unit range (such as bottles, boxes) for conventional scenarios; determine that "the unit price does not exceed the threshold and the unit is within the conventional range" as the compliance condition for this low price and conventional scenario, and form a callable low price and conventional scenario standard; (3) Wine product standard refinement: According to the clause summary document, distinguish between white wine and red wine categories, and clearly define the unit price ceiling and corresponding capacity benchmark for each category. Determine that "the unit price does not exceed the corresponding category ceiling and the capacity meets the benchmark" as the compliance condition for wine products, and refine the specific determination values for different categories; (4) Standard structured storage: The low price and regular scene standards, wine product capacity and product category price standards are sorted into a structured data table containing "scene type, judgment condition, specific value, and basis clause". It supports quick query by product category and scene and is synchronized to the judgment module for S4 calling.
[0035] S4: Output preliminary compliance results: According to the judgment standards of S3, combined with the unit price and capacity data generated by S2, the judgment is made: if any standard is met, the judgment state is normal and the compliance basis is given; if none is met, the next step is entered. It includes the following key sub-steps: (1) Judgment data association matching: The "unit price, product capacity, product unit, product type" data output by S2 are matched with the "low price and regular scene standards, wine product standards" according to product type and product unit, to ensure that each product can be associated with the corresponding judgment standard; (2) Low price and regular scene compliance judgment: According to the matched low price and regular scene standards, if the unit price of the product does not exceed the threshold and the product unit is within the regular range, the judgment state is directly output as normal, and the compliance basis is marked synchronously; (3) Wine scene compliance judgment: If the low price and regular scene standards are not met, the wine product standards are used for judgment. If the unit price does not exceed the upper limit of the corresponding category and the capacity meets the benchmark, the judgment state is output as normal and the compliance basis is marked; (4) Non-compliance data marking and transfer: If neither the low price and regular scene standards nor the wine scene standards are met, the product data is marked as "further verification required" and associated with the original data such as unit price, product capacity, and product unit, and is transferred to S5.
[0036] S5: Build core feature extraction model: Based on the prompt engineering and public large language model, generate high-quality training corpus, train core feature extraction model, and use it to extract core feature data such as product type, quantity, and actual capacity from the product name and model specifications in the structured data obtained from S1. It includes the following key sub-steps: (1) Model requirement clarification: Determine the core features that the model needs to extract, which are product type, actual quantity, and actual capacity. Clarify the input data format and output data format that the model needs to adapt, and set the performance index requirements for model accuracy and recall rate; (2) Training corpus preparation: Collect product name and model specification text data from historical wine invoices, organize them in "input text-feature annotation" format, and generate high-quality annotated corpus; remove invalid texts by de-duplicating and cleaning the corpus, and divide it into training set, validation set, and test set; (3) Prompt engineering design: Design prompt word templates according to feature extraction requirements, and clearly require the model to identify product types, analyze actual quantities and units, extract actual capacity values, and judge product units to ensure that prompt word instructions are clear and unambiguous.
[0037] The core feature prompt instructions for wine invoices are as follows: Role: You are an expert in the wine industry Task: Extract the information in the name according to the following requirements: #Task Description Based on the provided name, determine the type of wine (white wine, red wine), and extract the capacity and quantity of the wine #Explanation of each field ▼Type of wine - Red wine: Clearly mentions red wine, grape wine, dry red, etc. - White wine: Includes white wine, beer, grain wine, and foreign wine, and those that do not clearly indicate the type of wine are classified as white wine ▼Wine capacity - Refers to the capacity of each bottle of wine, with units of milliliters (ml) or liters (L) If the unit is not explicitly mentioned, it is assumed to be ml If the capacity is not explicitly mentioned, white wine is assumed to be 500 ml and red wine is assumed to be 750 ml
[0038] ▼Quantity of wine - Refers to the number of bottles of wine in a box, for example, "500ML*6" indicates 6 bottles; "500ml*2 bottles*3 boxes" indicates 2*3=6 bottles If a box is mentioned but the quantity is not, it is assumed to be 1 bottle If the quantity is not mentioned, it is assumed to be 1 bottle #Requirements 1. The unit of capacity is converted to milliliters (ml) and only the numerical part is output 2. The quantity is output as an integer, representing the total number of bottles, and 0 is also considered an integer #Output Requirements ```json {{ "Type": "Red wine or white wine", "Capacity": "Number", "Quantity": "Number" }} ``` #Special Reminders - Ensure that the JSON format is absolutely standard and has no syntax errors #User Input: {content} / no_think.
[0039] (4) Model training and fine-tuning: Based on the base large language model, combine the single sample in the training set with the prompt word template designed by S53, input the large language model to get the feature extraction result of the sample; compare the extraction result with the sample labeled features, use the cross-entropy loss and mean square error combined loss function to calculate the loss value; adjust the model parameters through back propagation to adapt to the feature extraction task, and at the same time use the validation set to monitor the model performance in real time, optimize the fine-tuning strategy (such as adjusting the learning rate) until the indicators meet the standards.
[0040] (5) Model performance evaluation: Use the test set to evaluate the trained model, detect whether the accuracy and recall rate of core feature extraction meet the preset indicators; analyze the extraction errors of the samples, supplement the corpus and fine-tune the model again.
[0041] (6) Model deployment: Deploy the qualified model to the data processing platform, configure the interface to receive the product name and model specification text in the "further verification" data transferred by S4, output the structured core feature data, and synchronize to S6.
[0042] S6: Complex scenario compliance judgment: For each invoice, first obtain the core features of the product type, actual quantity, and actual capacity extracted by S5 model, and then for complex scenarios such as unit abnormality, name and unit mismatch, and non-standard unit, construct multi-dimensional compliance judgment logic, input the extracted core features into the logic for judgment, and finally output the judgment state (normal / over-standard / suspected) and judgment basis of the invoice. Specifically, it includes the following steps: (1) Data integration and format processing: Associate the original data transferred by S4 with the core feature data output by S5 according to the invoice number, convert the quantity in the original data to an integer and the amount to a floating point number to ensure uniform data format and calculation; (2) Unit price secondary calculation: Calculate the unit price according to the formula unit price = amount ÷ quantity, if the quantity is empty, 0 or non-numeric, default the quantity to 1, and the calculation result is rounded to the standard decimal place, which is used as the basis for judgment; (3) Feature data association and standard matching: Extract the type, quantity, and capacity fields in the core feature data, and match the preset wine standards (including capacity benchmarks for different types of wine and unit price upper limits) and price policies (single milliliter maximum price, whole box maximum price, etc.); (4) Complex scenario classification judgment: For different complex scenarios, if the unit is empty and the actual quantity is 1, mark it as suspected; if the product name contains whole box or box but the unit is bottle, mark it as suspected; if the unit is box and the actual quantity is 1, default it to 2 and mark it as adjusted; if the unit is not box or bottle and the unit price exceeds the standard, mark it as suspected; (5) Determine the state and single bottle price: compare the calculated unit price with the matching standard: if it is not over the standard, mark it as normal; if it is over the standard and the unit is bottle or box (quantity is clear), mark it as over the standard; if it cannot be determined due to abnormal unit or questionable quantity, mark it as suspected; calculate the single bottle price (total amount ÷ actual quantity) at the same time; (6) Commodity compliance result output: generate results in state + reason format, normal needs to explain that the unit price is not over the standard, over the standard and suspected state needs to explain the abnormal reason.
[0043] Although the specific embodiments of the present application are disclosed for the purpose of illustrating the present application, the purpose is to help understand the content of the present application and to implement it, those skilled in the art can understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present application and the appended claims. Therefore, the present application should not be limited to the disclosed content of the best embodiment, the scope of the present application claimed is the scope defined by the claims.
Claims
1. A rule and large model fusion-based intelligent identification method for liquor reimbursement violations, comprising the following steps: 1) structurally analyzing the content of an invoice uploaded by a user to generate structured data of the invoice; The structured data includes commodity name, model specification, quantity, amount, and commodity unit; 2) after standardizing the quantity and amount in the structured data, calculating the unit price of the commodity; extracting the commodity capacity from the invoice, if the commodity capacity cannot be extracted from the invoice, supplementing the commodity capacity of the invoice with the general benchmark value of the liquor industry, and outputting the standardized unit price and capacity data; 3) establishing a number of compliance judgment standards according to the regulations; 4) judging the invoice according to the compliance judgment standards and the unit price and capacity data generated in step 2): if the invoice meets any of the compliance judgment standards, the status of the invoice is determined to be normal and the compliance basis is given; if none of the compliance judgment standards is met, step 5) is entered; 5) extracting core feature data from the commodity name and model specification of the structured data; the core feature data includes commodity type, quantity, and actual capacity; 6) using the set multi-dimensional compliance judgment logic to judge the core feature data, and determining the status of the invoice according to the judgment result and giving the judgment basis.
2. The method of claim 1, wherein, Based on OCR or multi-modal large language model, the content of the invoice is structurally analyzed to generate the structured data of the invoice.
3. The method of claim 1, wherein, The method for extracting the commodity capacity from the invoice is to extract the capacity value containing capacity identification information from the model specification or commodity name of the structured data as the commodity capacity; the capacity identification information is ml or liter.
4. The method of claim 1, wherein, The compliance judgment standards include: ① low price and regular scene standard, ② liquor commodity capacity and category price standard; wherein the compliance basis of the low price and regular scene standard is the unit price of the commodity and the commodity unit; the compliance basis of the liquor commodity capacity and category price standard is the unit price of the commodity, the capacity and the liquor category.
5. The method of claim 4, wherein, The method for establishing the compliance judgment standards is: 31) combing the provisions related to the price, capacity and category of liquor commodities in the regulations, extracting the specific contents related to the low price standard, the regular unit requirement, the upper limit of the unit price of different liquors, and the corresponding capacity benchmark, and forming a provision abstract document; 32) setting the unit price threshold of the low price scene and the commodity unit range of the regular scene based on the provision abstract document, determining that "the unit price does not exceed the threshold and the unit is within the regular range" is the compliance condition of the low price and regular scene, and generating the low price and regular scene standard; 33) distinguishing between baijiu and red wine categories according to the provision abstract document, clearly defining the upper limit of the unit price and the corresponding capacity benchmark for each category, determining that "the unit price does not exceed the corresponding category upper limit and the capacity meets the benchmark" is the compliance condition of liquor commodities, refining the specific judgment values of different categories, and generating the liquor commodity capacity and category price standard.
6. The method of claim 5, wherein, The method for judging whether the invoice meets the compliance judgment standards is: 41) The unit price, commodity volume, commodity unit, and commodity type obtained from the invoice are matched with the low price and regular scenario standard, the wine commodity volume and category price standard according to commodity type and commodity unit; 42) First, the low price and regular scenario standard is judged. If the commodity unit price does not exceed the threshold value and the commodity unit is within the regular range, the determination state is output as normal and the compliance basis is marked; 43) If the low price and regular scenario standard is not met, the wine commodity volume and category price standard is judged. If the unit price does not exceed the upper limit of the corresponding category and the volume meets the standard, the determination state is output as normal and the compliance basis is marked.
7. The method of claim 1, wherein, The method for extracting the core feature data from the commodity name and model specification of the structured data is: 51) The core feature extraction model needs to extract the core features of commodity type, actual quantity, and actual volume; 52) Collect the commodity name and model specification text data in the historical wine invoice, organize them in the "input text-feature marking" format, and generate training samples; 53) Design a prompt word template according to the feature extraction requirements, clearly require the core feature extraction model to identify the commodity type, analyze the actual quantity and unit, extract the actual volume value, and judge the commodity unit, to ensure that the prompt word instruction is clear and unambiguous; 54) Based on the basic large language model, combine the training samples with the prompt word template, input the large language model to get the feature extraction result of the training sample, compare the extraction result with the labeled features of the training sample, use the cross-entropy loss and mean square error combined loss function to calculate the loss value, and adjust the parameters of the core feature extraction model to adapt to the feature extraction task through back propagation; 55) Use the trained core feature extraction model to extract the core feature data from the commodity name and model specification of the structured data.
8. The method of claim 1, wherein, The method for judging the core feature data using the set multi-dimensional compliance judgment logic is: 61) When the commodity unit in the structured data is empty, if the actual quantity in the core feature data is 1, it is marked as suspicious; when the commodity name in the structured data contains whole box or box, but the commodity unit is bottle, it is marked as suspicious; when the commodity unit in the structured data is box, if the actual quantity in the core feature data is 1, the commodity unit is adjusted to 2 and an adjustment mark is marked; when the unit is not box or bottle, if the unit price exceeds the standard, it is marked as suspicious; 62) Compare the calculated unit price with the compliance judgment standard. If it does not exceed the compliance judgment standard, it is marked as normal; if it exceeds the compliance judgment standard and the unit is bottle or box, it is marked as exceeding the standard; if it cannot be clearly judged due to unit abnormalities and quantity doubts, it is marked as suspicious; the single bottle price is calculated synchronously.
9. A server, characterized by A computer program is stored in a memory and executed by a processor, the computer program comprising instructions for executing the method of any one of claims 1 to 8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 8.