Dynamic layout adaptability form identification method based on YOLO

By combining the YOLO algorithm and optical character recognition with risk detection technology, the accuracy and risk assessment problems of traditional form recognition methods in complex scenarios have been solved. This has enabled precise location and risk assessment of fee fields, improving form processing efficiency and user interaction experience.

CN120954029APending Publication Date: 2025-11-14BEIJING HONGSHAN INFORMATION TECH RES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511266250.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional form recognition methods cannot accurately identify field locations and content when faced with complex and ever-changing real-world scenarios. They lack the ability to understand the semantics of form content and assess risks, leading to erroneous data entering subsequent business processes and affecting the accuracy of decision-making.

Method used

The YOLO algorithm is used to detect the location of the fee field, combined with optical character recognition to process digital information, and the digital information is verified by calculation formula. Risk detection is initiated, the risk level is assessed using a preset risk pattern library, and the form structure is restored through layout reconstruction algorithm. User operation behavior data is analyzed to optimize the interaction method.

Benefits of technology

It enables accurate identification and risk assessment of complex forms, improves data accuracy and user interaction experience, and is suitable for intelligent processing of diverse layouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954029A_ABST
    Figure CN120954029A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLO-based dynamic layout adaptive form identification method, which comprises the following steps of: acquiring form image data, and detecting a cost field position through a YOLO algorithm to obtain a field coordinate; digital information is extracted according to the field coordinates, optical character recognition is adopted to process the digital information, and a cost item classification and calculation formula is determined; verifying the digital information through a calculation formula, judging risk factors, obtaining a corresponding abnormal data mode from a preset risk mode library, and obtaining a risk level evaluation result; processing the abnormal layout by adopting a layout reconstruction algorithm, and determining a recovered form layout structure; and through the recovered form layout structure, obtaining user operation behavior data, judging a behavior mode, and optimizing an interaction mode by adopting a self-adaptive layout generation method to obtain a final acceptance quality evaluation report. According to the method, the form processing efficiency, the data accuracy and the user interaction experience are remarkably improved, and the method is suitable for intelligent processing of complex form scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent form recognition and processing technology, and particularly relates to a dynamic layout adaptive form recognition method based on YOLO. Background Technology

[0002] In the digital age, forms, as a core tool for information collection and processing, are widely used in various fields such as finance, administration, healthcare, human resources, and logistics. For example, in finance, forms are used to record expense reimbursements and invoice information; in healthcare, forms are used for patient information registration and medical record keeping. The efficiency and accuracy of form processing directly affect the smoothness of business processes and the reliability of decision-making. With the increasing demand for automation and intelligence from enterprises, AI-based form recognition technology has become a key means to improve efficiency and reduce labor costs.

[0003] However, traditional form recognition methods often exhibit significant limitations when dealing with complex and ever-changing real-world scenarios. Existing form recognition technologies primarily rely on template matching or rule engines. These methods perform well when handling forms with fixed formats and simple layouts, but they often fail to accurately identify field positions and content when faced with diverse layouts and dynamically changing forms.

[0004] Furthermore, traditional methods lack the ability to semantically understand and assess the risks associated with form content. If anomalies or errors occur in the form data, traditional methods cannot detect and handle them in a timely manner, potentially leading to erroneous data entering subsequent business processes and impacting the accuracy of decision-making.

[0005] Therefore, researching a form recognition method that can adapt to diverse layouts and possess intelligent recognition and risk assessment capabilities has become an urgent need in the current technological field. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a method for dynamically adaptable form recognition based on YOLO, comprising the following steps: Obtain the form image data, use the YOLO algorithm to detect the location of the fee field, and obtain the field coordinates; Numerical information is extracted from field coordinates, and optical character recognition is used to process the numerical information to determine the cost item classification and calculation formula. The numerical information is verified by calculation formulas. If the verification result shows an anomaly, risk detection is initiated to identify risk factors. For the identified risk factors, the corresponding abnormal data patterns are retrieved from the preset risk pattern library to obtain the risk level assessment results; Based on the risk level assessment results, a layout reconstruction algorithm is used to process abnormal layouts and determine the restored form layout structure. By reconstructing the form layout, user behavior data can be obtained to determine behavior patterns. Based on the identified behavior patterns, an adaptive layout generation method is used to optimize the interaction methods, resulting in the final acceptance quality assessment report.

[0007] Optionally, the step of obtaining form image data and detecting the location of the fee field using the YOLO algorithm to obtain the field coordinates includes: Obtain form image data, process the image using a pre-trained YOLO algorithm model, and determine the bounding box and coordinate information of the cost field; The cost field region is extracted from the detected bounding box, and the text within the region is parsed using optical character recognition technology to obtain the preliminary digital content; The initial numerical content is cleaned using regular expressions. Based on the cleaned numerical information and coordinate information, structured field data is generated, including cost values ​​and corresponding locations. The integrity of the structured field data is judged by a preset threshold. If the integrity is lower than the threshold, the optical character recognition technology is re-executed for parsing. Retrieve the re-parsed digital content, update the structured field data, and determine the final cost field information; Based on the final cost field information, generate mapping data including field coordinates and numerical content, and output it in a structured format.

[0008] Optionally, the step of extracting numerical information based on field coordinates, processing the numerical information using optical character recognition, and determining the cost item classification and calculation formula includes: Based on the digital information and combined with the preset expense item classification rules, the digital information is matched with the expense item category to generate the classification result; The attributes of the cost items are extracted from the classification results, and the cost calculation formula is generated using the decision tree algorithm to obtain the formula structure. By combining the formula structure with field coordinates, structured data including cost items and calculation formulas is generated to determine the cost mapping relationship; If the completeness of the cost mapping relationship is lower than the preset threshold, the text content is re-parsed, the classification results and formula structure are updated, and the final cost information is obtained. Based on the final cost information, generate structured output including field coordinates, cost items, and calculation formulas to determine the final mapping data.

[0009] Optionally, the step of verifying numerical information through a calculation formula, and initiating risk detection to determine risk factors if the verification result shows an anomaly, includes: The digital information is verified using a preset calculation formula to obtain the verification result; If the verification result deviates from the preset threshold, abnormal data is extracted from the verification result, and risk detection is initiated. The k-means clustering algorithm is used to group the outlier data to obtain a set of data features; By matching the data feature set with a pre-set risk element database, it is determined whether a risk element exists, and the risk element judgment result is obtained.

[0010] Optionally, the step of obtaining corresponding abnormal data patterns from a preset risk pattern library for the identified risk factors to obtain risk level assessment results includes: For the identified risk factors, a regular expression matching method is used to determine whether the risk factors meet the preset rules, and a preliminary set of risk factors is obtained. If the initial risk factor set includes at least one anomaly identifier, then the corresponding anomaly data pattern is obtained from the preset pattern library to determine the anomaly data pattern set. The decision tree algorithm is used to classify the set of abnormal data patterns, determine the risk level of each abnormal data pattern, and obtain the risk level distribution. Based on the risk level distribution, contextual data matching the risk level is extracted from historical data records, and the Naive Bayes algorithm is used to calculate the matching probability between the contextual data and the risk level to obtain the probability distribution results. If the probability distribution result exceeds the preset threshold, additional verification data is extracted from the digital information, and the consistency between the additional verification data and the abnormal data pattern is judged by the data comparison method to determine the final risk level assessment result.

[0011] Optionally, the step of processing abnormal layouts using a layout reconstruction algorithm based on the risk level assessment results to determine the restored form layout structure includes: If the risk level exceeds the preset threshold, key node data is extracted from the abnormal layout, and the distance between nodes is calculated using a grid adjustment algorithm to determine the preliminary adjusted layout framework. Based on the initially adjusted layout framework, the data distribution characteristics in the abnormal layout are obtained, and the K-means algorithm is used to cluster the characteristics to obtain the clustered data groups. If the clustered data group includes at least one anomaly identifier, the corresponding layout optimization rule is obtained from the preset rule base to determine the optimized layout parameters; Based on the optimized layout parameters, structural correlation data is extracted from the form layout, and the consistency between the correlation data and the layout parameters is judged by the data comparison method to obtain the consistency verification result. If the consistency verification result meets the preset standard, the grid parameters are adjusted according to the verification result to determine the final reconstructed form layout structure; Based on the final reconstructed form layout structure, structural stability data is obtained, and the matching probability of the stability data is calculated using the Naive Bayes algorithm to obtain the stability evaluation result. If the stability assessment result reaches the preset threshold, additional validation data is extracted from the form layout, the degree of matching between the additional validation data and the reconstructed structure is judged, and the final layout optimization result is determined.

[0012] Optionally, the step of optimizing the interaction method using an adaptive layout generation method based on the determined behavior pattern includes: Based on the behavioral pattern feature set, a hidden Markov model is used to analyze the temporal distribution characteristics of the operation sequence and determine the operation time distribution pattern. If the distribution pattern of operation time deviates from the preset time threshold, click frequency statistics are extracted from the operation behavior data, and statistical analysis methods are used to determine whether the click frequency conforms to the normal interaction mode, so as to obtain the click frequency evaluation result. Based on the click frequency evaluation results, the input content patterns are extracted from user operation behavior data, and the consistency between the input content and the preset rules is compared using pattern matching methods to determine the compliance results of the input content. If the compliance result of the input content meets the preset standard, the operation path trajectory is extracted from the operation behavior data, and the smoothness of the path trajectory is calculated using the trajectory analysis method to obtain the path trajectory smoothness data. Based on the path trajectory smoothness data, the interface response speed is extracted from the form layout structure. The response time analysis method is used to determine the degree of matching between the response speed and the operation path, and to determine the need to adjust the interaction method. If the need to adjust the interaction method reaches a preset threshold, then the interaction optimization rules are obtained from the preset rule base, and the interaction parameters of the form layout are adjusted using the rule application method to obtain the optimized interaction method configuration.

[0013] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method.

[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method.

[0016] Compared with the prior art, the present invention has the following advantages and technical effects: This invention discloses a method that addresses issues such as inaccurate identification of fee fields, data anomalies, and poor user interaction in business scenarios involving form images. By integrating the YOLO algorithm, optical character recognition (OCR), risk detection, and adaptive layout generation technology, it automates the entire process from fee field location to interaction optimization. First, the YOLO algorithm is used to accurately locate the coordinates of the fee field and initially extract numerical information. OCR is then used to further confirm the fee item classification and calculation formula, resolving the issue of inaccurate field identification. For abnormal data, a risk detection module is activated, using a preset risk pattern library to assess the risk level and resolve the data anomaly issue. Based on the risk assessment results, a layout reconstruction algorithm is used to restore the form structure. Furthermore, by analyzing user operation behavior data, behavior patterns are determined, and an adaptive layout generation method is used to optimize the interaction, resolving the issue of poor interaction. This invention significantly improves form processing efficiency, data accuracy, and user interaction experience, and is suitable for intelligent processing of complex form scenarios. Attached Figure Description

[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the risk factor judgment process according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the optimization of interaction parameters according to an embodiment of the present invention. Detailed Implementation

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0020] Example 1 This embodiment provides a method for recognizing dynamically layout-adaptive forms based on YOLO, including the following steps: Obtain the form image data, use the YOLO algorithm to detect the location of the fee field, and obtain the field coordinates; Numerical information is extracted from field coordinates, and optical character recognition is used to process the numerical information to determine the cost item classification and calculation formula. The numerical information is verified by calculation formulas. If the verification result shows an anomaly, risk detection is initiated to identify risk factors. For the identified risk factors, the corresponding abnormal data patterns are retrieved from the preset risk pattern library to obtain the risk level assessment results; Based on the risk level assessment results, a layout reconstruction algorithm is used to process abnormal layouts and determine the restored form layout structure. By reconstructing the form layout, user behavior data can be obtained to determine behavior patterns. Based on the identified behavior patterns, an adaptive layout generation method is used to optimize the interaction methods, resulting in the final acceptance quality assessment report.

[0021] like Figure 1 As shown, as an optional implementation, the specific steps include: S101. Obtain the form image data, detect the location of the fee field using the YOLO algorithm, and obtain the field coordinates and preliminary numerical information: The process begins by acquiring form image data and processing it using a pre-trained YOLO algorithm model to determine the bounding box and coordinates of the cost field. The cost field region is extracted from the detected bounding box, and optical character recognition (OCR) is used to parse the text within the region, yielding preliminary numerical content. If the preliminary numerical content contains non-numeric characters, regular expressions are used to clean the data, identifying only pure numeric information. Based on the cleaned numerical information and coordinates, structured field data containing the cost value and its corresponding location is generated. The integrity of the structured field data is assessed using a preset threshold; if the integrity falls below the threshold, OCR parsing is re-executed. The re-parsed numerical content is then acquired, updating the structured field data and determining the final cost field information. Finally, based on the final cost field information, mapping data containing coordinates and numerical content is generated and output in a structured format.

[0022] For example, when processing image data from medical expense reimbursement documents, a pre-trained YOLO algorithm model can be used to accurately locate the expense field. The YOLO algorithm uses a convolutional neural network to segment the image into a grid and predicts the bounding box and class probability of objects in each grid.

[0023] In one possible implementation, assuming a hospital invoice image is being processed, the YOLO model scans the image, detects the cost field region, and outputs bounding box coordinates such as the top-left corner (100, 150) and bottom-right corner (300, 200). These coordinates define the rectangular area of ​​the cost field, providing accurate spatial information for subsequent processing. The advantage of using YOLO lies in its efficiency and accuracy, enabling rapid location of target regions in complex backgrounds and reducing false detections.

[0024] Specifically, after extracting the cost field area, optical character recognition (OCR) technology is used to parse the text content. Assuming the extracted area contains the text "Total Cost: 1234.56 yuan", OCR technology converts the image into text by recognizing the shape and arrangement of characters. Common OCR tools such as Tesseract scan the pixel distribution to generate the initial numerical content "1234.56 yuan". However, sometimes the recognition result may contain noise, such as "1234,56 yuan" or "1234.56 yuan+". Therefore, regular expressions are used to clean the data, ultimately obtaining the pure number "1234.56". The advantage of regular expressions lies in their flexibility and efficiency, ensuring the accuracy of the numerical content.

[0025] In one embodiment, when generating structured field data, the cleaned number "1234.56" is combined with the bounding box coordinates (100, 150, 300, 200), and the output format is as follows: {"Cost": 1234.56, "Coordinates": {"Top Left": [100, 150], "Bottom Right": [300, 200]}}. This structured data is convenient for system storage and retrieval.

[0026] It should be noted that, to ensure data integrity, a threshold of 90% is set, and integrity is determined by checking whether the field contains complete numbers and valid coordinates. If the recognition result is only "1234.", and the integrity is below the threshold, optical character recognition is re-executed. This may involve adjusting the image contrast and re-analyzing to obtain the correct result "1234.56". Re-analysis effectively corrects errors from the initial recognition and improves data reliability.

[0027] For example, after the final cost field information is generated, the mapped data can be output as {"Cost": 1234.56, "Location": {"Top Left": [100, 150], "Bottom Right": [300, 200]}}. This mapped data not only retains the cost value but also records its specific location in the image, facilitating subsequent verification or display. The entire process, through YOLO positioning, optical character recognition parsing, regular expression cleaning, and threshold judgment, achieves automated conversion from image to structured data. This method significantly improves the efficiency and accuracy of cost information extraction, reduces manual intervention, and is suitable for large-scale document processing scenarios, such as hospital financial systems or insurance reimbursement processes, ensuring data consistency and traceability.

[0028] S102. Extract numerical information based on field coordinates, process the numerical information using optical character recognition (OCR), and determine the cost item classification and calculation formula: Based on the numerical information and pre-defined expense item classification rules, the numerical information is matched with the expense item categories to generate classification results. Attributes of the expense items are extracted from the classification results, and a decision tree algorithm is used to generate expense calculation formulas, resulting in a formula structure. Using this formula structure and field coordinates, structured data containing expense items and calculation formulas is generated to determine the expense mapping relationship. If the completeness of the expense mapping relationship is below a pre-defined threshold, the text content is re-parsed, the classification results and formula structure are updated, and the final expense information is obtained. Based on the final expense information, a structured output containing field coordinates, expense items, and calculation formulas is generated to determine the final mapping data.

[0029] For example, when processing image data from medical expense reimbursement documents, the extracted numerical information needs to be matched with the expense item categories to generate classification results. Suppose a hospital invoice contains the extracted numerical information "Hospitalization fee: 5000.00 yuan, Medication fee: 1234.56 yuan". The preset classification rule can be defined as: hospitalization fee is classified as "fixed cost", and medication fee is classified as "variable cost". Through comparison with the rule, the numerical information "5000.00" matches as a fixed cost, and "1234.56" matches as a variable cost. This classification method is achieved through a predefined rule table, which associates numerical information with categories based on the name keywords of the expense items, such as "hospitalization" and "medication". After the classification results are generated, the attributes of the expense items are extracted. For example, the attributes of fixed costs include "fixed amount, no discount", and the attributes of variable costs include "adjustable, requires verification".

[0030] In one possible implementation, a decision tree algorithm is used to generate the cost calculation formula. The decision tree constructs logical branches based on the attributes of the cost items.

[0031] For example, the formula for fixed costs is "Total Cost = Fixed Amount", and the formula for variable costs is "Total Cost = Drug Cost × Verification Coefficient". Assuming the verification coefficient is 0.9, the formula for calculating drug costs of 1234.56 yuan is "1234.56 × 0.9 = 1111.10 yuan". The decision tree uses attribute judgments, such as "whether it is adjustable", to generate the corresponding formula structure, ensuring that the formula matches the cost category. The formula structure, combined with field coordinates, such as hospitalization fees with coordinates of top left (50, 100) and bottom right (250, 150), and medication fees with coordinates of top left (100, 200) and bottom right (300, 250), generates structured data as follows: {"Hospitalization Fee":{"Amount":5000.00,"Formula":"Fixed Amount","Coordinates":{"Top Left":[50, 100]","Bottom Right":[250, 150]}}, "Medication Fee":{"Amount":1111.10,"Formula":"Amount × 0.9","Coordinates":{"Top Left":[100, 200]","Bottom Right":[300, 250]}}}.

[0032] It should be noted that if the completeness of the cost mapping relationship is below a preset threshold, such as 90%, the text content will be re-parsed. For example, if the initial classification misidentifies "1234.56" as "1234", the completeness is below the threshold due to missing decimal parts. Re-parse can be performed by adjusting the image brightness to enhance text clarity, re-extracting "1234.56", updating the classification result to variable cost, and adjusting the formula structure to "1234.56 × 0.9 = 1111.10". The final output is a structured list containing field coordinates, expense items, and calculation formulas: {"Hospitalization Fee":{"Amount":5000.00,"Category":"Fixed Fee","Formula":"Fixed Amount","Coordinates":{"Top Left":[50,100]","Bottom Right":[250,150]}}, "Medication Fee":{"Amount":1111.10,"Category":"Variable Fee","Formula":"Amount × 0.9","Coordinates":{"Top Left":[100,200]","Bottom Right":[300,250]}}}. This structured data ensures the correlation between expense information and spatial location, facilitating subsequent verification. Re-parsing and threshold judgment improve data accuracy, while classification and formula generation enhance the automation of expense processing, making it suitable for medical reimbursement scenarios.

[0033] S103. Verify the numerical information using the calculation formula. If the verification result shows an anomaly, activate the risk detection module to determine if any risk factors exist. like Figure 2 As shown, the numerical information is verified using a preset calculation formula to obtain the verification result. If the verification result deviates from the preset threshold, abnormal data is extracted from the verification result, and the risk detection module is activated. The k-means clustering algorithm is used to group the abnormal data to obtain a data feature set. The data feature set is matched with a preset risk element database to determine whether a risk element exists, and the risk element judgment result is obtained.

[0034] For example, in medical expense reimbursement scenarios, the extracted numerical information is verified using a preset calculation formula to ensure the accuracy and compliance of the data. The preset formula may be based on the category attributes of the expense item; for example, the formula for fixed expenses is "Total Expenses = Extracted Amount," and the formula for variable expenses is "Total Expenses = Extracted Amount × Verification Coefficient." Suppose a hospital invoice contains "Hospitalization Fee: 5000.00 yuan, Medication Fee: 1234.56 yuan." During verification, the hospitalization fee is directly compared to the extracted amount of 5000.00 yuan to see if it matches the formula, while the medication fee needs to be multiplied by a verification coefficient of 0.9, with an expected result of 1111.10 yuan. If the verification result is 1111.00 yuan, deviating from the expected amount by 0.10 yuan and exceeding the preset threshold by 0.05 yuan, it is determined to be abnormal data.

[0035] In one possible implementation, outlier data is extracted by comparing the difference between the verification result and the expected value. For the aforementioned drug expense, the system records a difference of 0.10 yuan and marks the record containing the amount of 1234.56 yuan and the coordinates (top left, 100, 200) and (bottom right, 300, 250) as outlier. Risk detection is then performed on the outlier data, initiating further analysis. The core of risk detection lies in identifying potential errors or fraudulent activities, such as amount tampering or false alarms.

[0036] Specifically, the k-means clustering algorithm is used to group outlier data to discover patterns. Suppose the system collects multiple outlier data entries, such as drug expense amounts of 1234.56 yuan, 1235.00 yuan, and 1200.00 yuan respectively. The clustering algorithm divides the data into two groups based on the amount difference and coordinate distribution: one group has small deviations, possibly due to data entry errors; the other group has larger deviations, possibly due to intentional tampering. After grouping, each group generates a set of data features. For example, the features of the small deviation group are "amount difference less than 0.5 yuan, coordinates concentrated at the bottom of the invoice," while the features of the larger deviation group are "amount difference greater than 1 yuan, coordinates dispersed."

[0037] For example, when matching the data feature set with a preset risk element database, the risk element database may contain rules such as "amount difference greater than 1 yuan and dispersed coordinates may indicate fraudulent behavior." Small deviation groups match the "entry error" element, indicating lower risk; larger deviation groups match the "fraud suspicion" element, requiring further verification. The matching process is achieved by comparing the attributes of the feature set with the keywords of the risk element database, ensuring the accuracy of the judgment results. After the risk element judgment results are generated, the system can output structured data, such as {"Drug Costs":{"Amount":1234.56, "Abnormal Difference":0.10, "Risk Element":"Entry Error", "Coordinates":{"Top Left":[100,200], "Bottom Right":[300,250]}}}.

[0038] In one possible implementation, risk detection can be optimized by combining historical data.

[0039] For example, if the system detects that a hospital's invoices frequently have minor discrepancies due to blurry printing, they are prioritized as low-risk to avoid misjudgments. This approach, combining clustering and matching, improves the accuracy of anomaly data processing while reducing the workload of manual verification.

[0040] S104. For the identified risk factors, retrieve the corresponding abnormal data patterns from the preset risk pattern library to obtain the risk level assessment results: Risk elements are extracted from digital information. Regular expression matching is used to determine if these risk elements conform to preset rules, resulting in a preliminary risk element set. If the preliminary risk element set contains at least one anomaly identifier, the corresponding anomaly data pattern is retrieved from a preset pattern library to determine the anomaly data pattern set. A decision tree algorithm is used to classify the anomaly data pattern set, determining the risk level of each anomaly data pattern and obtaining a risk level distribution. Based on the risk level distribution, contextual data matching the risk level is extracted from historical data records. A Naive Bayes algorithm is used to calculate the matching probability between the contextual data and the risk level, obtaining a probability distribution result. If the probability distribution result exceeds a preset threshold, additional verification data is extracted from the digital information. A data comparison method is used to determine the consistency between the additional verification data and the anomaly data pattern, determining the final risk level assessment result.

[0041] For example, in a business scenario of digital information risk assessment, suppose a financial institution needs to identify potential risks from transaction data. Regular expression matching methods can be used to extract risk elements. Predefined regular expression rules can be used to scan key fields in transaction data, such as transaction amount, transaction time, and account identifier, to determine whether they match abnormal patterns.

[0042] For example, a regular expression rule defines transactions exceeding 100,000 yuan occurring between midnight and 4 AM as potential risk factors. A scan reveals a transaction of 150,000 yuan occurring at 2 AM, which meets the rule and is included in the initial risk factor set. This method efficiently filters anomalies and reduces the false positive rate.

[0043] In one possible implementation, if the initial risk factor set contains anomaly markers, such as the aforementioned high-value nighttime transactions, the system retrieves anomaly data patterns from a pre-defined pattern library. This pattern library might include patterns such as "high-value transactions + abnormal time" or "multiple small transfers + frequent operations." Assuming a "high-value transactions + abnormal time" pattern is retrieved, the system categorizes it into the anomaly data pattern set. This approach ensures comprehensive coverage of anomaly patterns, facilitating subsequent classification.

[0044] Specifically, decision tree algorithms are used to classify sets of abnormal data patterns and determine their risk levels. Decision trees generate classification paths based on features such as transaction amount, time, and account history.

[0045] For example, a transaction exceeding 100,000 yuan from a new account might be classified as high-risk; a transaction from an established account would be classified as medium-risk. Assuming a 150,000 yuan transaction originates from a new account, the decision tree would output a high-risk rating. This method ensures accurate risk level classification through multi-dimensional feature analysis.

[0046] For example, based on risk level distribution, the system extracts contextual data from historical records. Suppose a high-risk transaction involves an account, and historical records show that this account has recently had multiple unusual logins. The Naive Bayes algorithm calculates the matching probability between the contextual data and the risk level.

[0047] For example, the correlation between abnormal logins and high risk is 0.9, and the calculated matching probability is 0.85, exceeding the preset threshold of 0.8. This method uses probability analysis to enhance the reliability of the judgment.

[0048] In one possible implementation, if the probability distribution exceeds a threshold, the system extracts additional verification data, such as transaction IP addresses or device fingerprints. The data comparison method checks the consistency of this data with abnormal data patterns.

[0049] For example, if a transaction IP originates from a high-risk area and matches the "abnormal area transaction" pattern in the pattern library, it is ultimately identified as high-risk. This multi-layered verification improves the accuracy of the assessment.

[0050] For example, the above process can be applied to anti-money laundering scenarios. A large overnight transfer from an account triggers regular expression matching, the decision tree classifies it as high-risk, a Bayesian algorithm combined with historical abnormal logins confirms the high probability, and finally, IP comparison pinpoints the risk. This multi-faceted analysis ensures the comprehensiveness and accuracy of risk identification and is suitable for single financial risk assessment scenarios.

[0051] S105. Based on the risk level assessment results, use a layout reconstruction algorithm to process abnormal layouts and determine the restored form layout structure: If the risk level exceeds a preset threshold, key node data is extracted from the abnormal layout, and a grid adjustment algorithm is used to calculate the distance between nodes to determine the preliminary adjusted layout framework. Based on the preliminary adjusted layout framework, the data distribution characteristics in the abnormal layout are obtained, and the K-means algorithm is used to cluster the features to obtain clustered data groups. If the clustered data groups contain at least one abnormal identifier, the corresponding layout optimization rule is obtained from the preset rule base to determine the optimized layout parameters. Based on the optimized layout parameters, structural correlation data is extracted from the form layout, and a data comparison method is used to determine the consistency between the correlation data and the layout parameters to obtain a consistency verification result. If the consistency verification result meets the preset standard, the grid parameters are adjusted according to the verification result to determine the final reconstructed form layout structure. Based on the final reconstructed form layout structure, structural stability data is obtained, and the Naive Bayes algorithm is used to calculate the matching probability of the stability data to obtain a stability evaluation result. If the stability evaluation result reaches a preset threshold, additional verification data is extracted from the form layout, and the degree of matching between the additional verification data and the reconstructed structure is determined to determine the final layout optimization result.

[0052] Specifically, in the business scenario of digital information risk assessment, financial institutions need to identify potential risks from transaction data and optimize data layout to improve assessment efficiency.

[0053] For example, a bank needs to process massive amounts of transaction form data, involving fields such as transaction amount, time, and account. It must ensure the data layout is reasonable to support risk analysis. The following analysis and examples focus on key technical topics.

[0054] For example, when extracting key node data from an abnormal layout, core fields such as transaction amount and account identifier can be selected as nodes. Suppose a transaction contains an amount of 200,000 yuan, the account is a newly registered account, and the time is 3 AM. This node data is extracted through field filtering. The grid adjustment algorithm calculates the distance between nodes and, based on the strength of the correlation between fields, such as the correlation between amount and time, generates a preliminary adjusted layout framework. Assuming a high correlation between amount and time, the framework places these two elements in adjacent positions for easier subsequent analysis.

[0055] In one possible implementation, when acquiring data distribution characteristics, the system analyzes the distribution range of transaction amounts and the concentration over time periods.

[0056] For example, data distribution shows that 80% of transactions are below 10,000 yuan, with only 5% occurring in the early morning, indicating that abnormal transactions are concentrated in high-value nighttime transactions. The K-means algorithm clusters these features, assuming they are divided into three groups: normal transactions, low-risk anomalies, and high-risk anomalies. A transaction of 200,000 yuan occurring in the early morning is classified as a high-risk anomaly because it contains an anomaly identifier, triggering the retrieval of layout optimization rules from the rule base.

[0057] For example, the rule base might contain rules such as "prioritize high-value transactions" and "concentrate display during abnormal time periods." Based on this, the optimized layout parameters place high-value transactions at the top of the layout for emphasis. The system extracts structural correlation data from the form layout, such as the relationship between account history and transaction amount, and judges its consistency with the optimized parameters. Suppose an account has recently been making frequent transactions with abnormally high transaction amounts, which matches the "prioritize high-value transactions" rule, and the verification passes.

[0058] In one possible implementation, when adjusting grid parameters, the system optimizes field spacing and display order based on verification results.

[0059] For example, the spacing between fields for high-risk transactions is reduced, and they are highlighted at the top of the form, forming the final reconstructed layout structure. When obtaining structural stability data, the logical consistency between fields is analyzed, such as the matching degree between transaction amounts and account history. The Naive Bayes algorithm calculates the stability matching probability; assuming a probability of 0.9, exceeding the threshold of 0.8 indicates a stable layout.

[0060] For example, additional validation data, such as transaction IP addresses, can be extracted from the form layout to determine their match with the reconstructed structure. Assuming the IP addresses originate from high-risk areas and align with the highlighted high-risk transactions in the layout, the optimization result is ultimately confirmed. This multi-layered layout optimization ensures clear data presentation, making risk points readily apparent and significantly improving risk assessment efficiency.

[0061] S106. Obtain user operation behavior data through the restored form layout structure, and determine whether the interaction method needs to be adjusted based on the behavior pattern: like Figure 3 As shown, user operation behavior data is obtained through the restored form layout structure. Sequence analysis is used to extract behavior sequence features, resulting in a set of behavior pattern features. Based on this feature set, a Hidden Markov Model (HMM) is used to analyze the time distribution characteristics of the operation sequence and determine the operation time distribution pattern. If the operation time distribution pattern deviates from a preset time threshold, click frequency statistics are extracted from the operation behavior data. Statistical analysis is used to determine whether the click frequency conforms to the normal interaction pattern, resulting in a click frequency evaluation result. Based on the click frequency evaluation result, input content patterns are extracted from the user operation behavior data. Pattern matching is used to compare the consistency of the input content with preset rules, determining the input content compliance result. If the input content compliance result meets the preset standard, the operation path trajectory is extracted from the operation behavior data. Trajectory analysis is used to calculate the smoothness of the path trajectory, obtaining path trajectory smoothness data. Based on the path trajectory smoothness data, the interface response speed is extracted from the form layout structure. Response time analysis is used to determine the degree of matching between the response speed and the operation path, determining the interaction method adjustment requirements. If the interaction method adjustment requirements reach a preset threshold, interaction optimization rules are obtained from a preset rule base. Rule application methods are used to adjust the interaction parameters of the form layout, resulting in an optimized interaction method configuration.

[0062] Specifically, in the business scenario of digital information risk assessment, financial institutions need to extract key features from user behavior data to optimize the interaction of form layouts and improve risk identification efficiency. The following analysis and examples address various technical topics, focusing on the scenario of bank transaction form data processing, ensuring consistency with historical information domains.

[0063] For example, when acquiring user behavior data, the system records user actions on transaction forms, such as clicks, input, and page switching. Suppose a user clicks the amount input box multiple times while filling in the transaction amount field, and then quickly switches to the account field. Sequence analysis methods extract behavioral sequence features, identifying the order and time intervals of user actions, forming a set of behavioral pattern features.

[0064] For example, if the system detects that a user completes the amount input, account selection, and submission within 5 seconds, it generates a feature set that includes the "quick input-switch-submit" pattern.

[0065] In one possible implementation, a Hidden Markov Model (HMM) is used to analyze the temporal distribution characteristics of the operation sequence. The system trains the model based on historical data, assuming that the normal operation time distribution is an average of 2 seconds per step. A user's operation time interval is 0.5 seconds, deviating from the preset threshold by 1 second, indicating a potential anomaly. Further model analysis reveals that the user's operations are concentrated in the early morning, exhibiting an abnormal time distribution pattern, triggering subsequent analysis.

[0066] For example, if the operation time distribution deviates from a threshold, the system extracts click frequency statistics. Suppose a user clicks the amount field 8 times within 10 seconds, far exceeding the normal interaction average of 3 times. Statistical analysis compares this to historical data, determines the click frequency is abnormal, and concludes with an assessment result of "high-frequency clicking," potentially indicating automated operation or abnormal behavior.

[0067] In one possible implementation, the system extracts patterns from the input content based on click frequency evaluation results. Assume a user enters a transaction amount of 100,000 yuan, and repeatedly enters the same amount multiple times. A pattern matching method compares the input content with preset rules, such as "normal transaction amount variation is less than 50%". If the input content shows high consistency, the compliance result is passed; otherwise, further analysis is triggered.

[0068] For example, after compliance is approved, the system extracts the user's operational path. Assuming the user navigates from the amount field to the account field and then to the submit button, the path is a linear one. The path analysis method calculates fluency based on mouse movement distance and operation time, resulting in a "high fluency" score, indicating the user is proficient in the operation.

[0069] In one possible implementation, the system extracts the interface response time from the form layout. Assuming the system responds in 0.3 seconds after the amount field is entered, which is below the average of 0.5 seconds, the response time analysis method determines that the response time matches the operation path, indicating that no adjustment to the interaction method is needed. If the response time is 1 second, exceeding the threshold of 0.8 seconds, then an adjustment to the interaction method is triggered.

[0070] For example, if the interaction method needs to be adjusted, the system retrieves optimization rules from the preset rule library, such as "prioritize responses to frequently clicked fields." The rule application method adjusts the form layout parameters, increases the response priority of the amount field, and generates an optimized interaction configuration to ensure smoother user operation and improve the interaction efficiency of risk assessment.

[0071] S107. Based on the identified behavior patterns, an adaptive layout generation method is used to optimize the interaction method, resulting in the final acceptance quality assessment report: By optimizing the interactive configuration, acceptance quality data is obtained, and quality analysis methods are used to determine the compliance of acceptance quality, thus obtaining the final acceptance quality assessment data.

[0072] As an alternative implementation method, specifically in the scenario of processing bank transaction form data, behavioral pattern features are extracted through user operation sequences, and the interaction method is optimized to improve the efficiency of risk assessment.

[0073] For example, the system records the user's action sequence on a transaction form, such as clicking, inputting, and switching pages. Suppose a user, while filling out a transaction form, first clicks the amount input box, enters 10,000 yuan, switches to the receiving account field, and then clicks the submit button. Cluster analysis methods classify the action sequences, generating a set of behavioral pattern categories.

[0074] In one embodiment, the system trains a K-means clustering model based on historical data to classify user operation sequences into three categories: "fast operation," "cautious operation," and "abnormal operation." Assuming a user operation sequence of "click amount - input - switch account - submit," taking 3 seconds, the clustering result is classified as "fast operation," meeting the preset classification threshold of "normal operation taking less than 5 seconds." Based on the behavioral pattern classification set, the system extracts user interaction preferences.

[0075] For example, if users prefer quick input with a fixed path, classification algorithms such as decision trees can generate a set of interaction preferences, labeled as "efficient input preferences".

[0076] In one embodiment, the system analyzes multiple user operation data and finds that 90% of the operations are concentrated in the amount and account fields, generating a preference set "Prioritize Amount and Account Interactions". Based on the interaction preference set, the system obtains interface layout adjustment rules from a preset rule base, such as "Efficient Input Preference Corresponds to Simplified Field Layout". The rule matching method adjusts the form layout parameters, placing the amount and account fields at the top of the page to reduce navigation steps. Through the adjusted interface layout parameters, the system obtains interaction response speed data.

[0077] For example, the adjusted response time for the amount field input is 0.2 seconds, lower than the average of 0.4 seconds. The time analysis method determines the response speed is compliant and generates a response speed assessment data of "high response". If the response speed assessment data meets the preset speed threshold of "less than 0.3 seconds", the system further extracts path trajectory data. Assuming the user's mouse movement path from the amount field to the account field is a straight line, the trajectory analysis method calculates a smoothness metric, yielding "high smoothness" assessment data, indicating smooth operation. Based on the smoothness assessment data, the system retrieves interaction optimization rules from the rule base, such as "high smoothness prioritizes maintaining the existing layout". The rule application method generates optimized interaction configurations, keeping the priority of the amount and account fields unchanged.

[0078] In one embodiment, the system adds support for keyboard shortcuts to further improve operational efficiency. By optimizing the interaction configuration, the system obtains acceptance quality data.

[0079] For example, the error rate for users completing transaction forms is 1%, lower than the average of 3%. Quality analysis methods determine that the acceptance quality is compliant, generating "high-quality" final acceptance quality assessment data to ensure that the interaction method is efficient and stable.

[0080] For example, in high-frequency trading scenarios, the system detected that users prefer to switch fields quickly. After adjusting the layout, the response speed improved by 20% and the smoothness of operation improved by 15%.

[0081] Understandably, the above methods significantly improve user operation efficiency and the accuracy of risk assessment through behavioral pattern analysis and layout optimization.

[0082] Preferably, the system updates the rule base regularly to ensure it adapts to different user behavior patterns and maintains dynamic optimization of the interaction method.

[0083] Example 2 This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in Embodiment 1.

[0084] Example 3 This embodiment also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1.

[0085] Example 4 This embodiment also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in Embodiment 1.

[0086] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for dynamically layout-adaptive form recognition based on YOLO, characterized in that, Includes the following steps: Obtain the form image data, use the YOLO algorithm to detect the location of the fee field, and obtain the field coordinates; Numerical information is extracted from field coordinates, and optical character recognition is used to process the numerical information to determine the cost item classification and calculation formula. The numerical information is verified by calculation formulas. If the verification result shows an anomaly, risk detection is initiated to identify risk factors. For the identified risk factors, the corresponding abnormal data patterns are retrieved from the preset risk pattern library to obtain the risk level assessment results; Based on the risk level assessment results, a layout reconstruction algorithm is used to process abnormal layouts and determine the restored form layout structure. By reconstructing the form layout, user behavior data can be obtained to determine behavior patterns. Based on the identified behavior patterns, an adaptive layout generation method is used to optimize the interaction methods, resulting in the final acceptance quality assessment report.

2. The method according to claim 1, characterized in that, The process of acquiring form image data and detecting the location of the fee field using the YOLO algorithm to obtain the field coordinates includes: Obtain form image data, process the image using a pre-trained YOLO algorithm model, and determine the bounding box and coordinate information of the cost field; The cost field region is extracted from the detected bounding box, and the text within the region is parsed using optical character recognition technology to obtain the preliminary digital content; The initial numerical content is cleaned using regular expressions. Based on the cleaned numerical information and coordinate information, structured field data is generated, including cost values ​​and corresponding locations. The integrity of the structured field data is judged by a preset threshold. If the integrity is lower than the threshold, the optical character recognition technology is re-executed for parsing. Retrieve the re-parsed digital content, update the structured field data, and determine the final cost field information; Based on the final cost field information, generate mapping data including field coordinates and numerical content, and output it in a structured format.

3. The method according to claim 1, characterized in that, The process of extracting numerical information based on field coordinates, processing the numerical information using optical character recognition, and determining the expense item classification and calculation formula includes: Based on the digital information and combined with the preset expense item classification rules, the digital information is matched with the expense item category to generate the classification result; The attributes of the cost items are extracted from the classification results, and the cost calculation formula is generated using the decision tree algorithm to obtain the formula structure. By combining the formula structure with field coordinates, structured data including cost items and calculation formulas is generated to determine the cost mapping relationship; If the completeness of the cost mapping relationship is lower than the preset threshold, the text content is re-parsed, the classification results and formula structure are updated, and the final cost information is obtained. Based on the final cost information, generate structured output including field coordinates, cost items, and calculation formulas to determine the final mapping data.

4. The method according to claim 1, characterized in that, The process of verifying digital information through calculation formulas, and initiating risk detection if the verification result shows an anomaly, involves identifying risk factors, including: The digital information is verified using a preset calculation formula to obtain the verification result; If the verification result deviates from the preset threshold, abnormal data is extracted from the verification result, and risk detection is initiated. The k-means clustering algorithm is used to group the outlier data to obtain a set of data features; By matching the data feature set with a pre-set risk element database, it is determined whether a risk element exists, and the risk element judgment result is obtained.

5. The method according to claim 1, characterized in that, For the identified risk factors, the corresponding abnormal data patterns are retrieved from a preset risk pattern library to obtain the risk level assessment result, including: For the identified risk factors, a regular expression matching method is used to determine whether the risk factors meet the preset rules, and a preliminary set of risk factors is obtained. If the initial risk factor set includes at least one anomaly identifier, then the corresponding anomaly data pattern is obtained from the preset pattern library to determine the anomaly data pattern set. The decision tree algorithm is used to classify the set of abnormal data patterns, determine the risk level of each abnormal data pattern, and obtain the risk level distribution. Based on the risk level distribution, contextual data matching the risk level is extracted from historical data records, and the Naive Bayes algorithm is used to calculate the matching probability between the contextual data and the risk level to obtain the probability distribution results. If the probability distribution result exceeds the preset threshold, additional verification data is extracted from the digital information, and the consistency between the additional verification data and the abnormal data pattern is judged by the data comparison method to determine the final risk level assessment result.

6. The method according to claim 1, characterized in that, The step of processing abnormal layouts using a layout reconstruction algorithm based on the risk level assessment results and determining the restored form layout structure includes: If the risk level exceeds the preset threshold, key node data is extracted from the abnormal layout, and the distance between nodes is calculated using a grid adjustment algorithm to determine the preliminary adjusted layout framework. Based on the initially adjusted layout framework, the data distribution characteristics in the abnormal layout are obtained, and the K-means algorithm is used to cluster the characteristics to obtain the clustered data groups. If the clustered data group includes at least one anomaly identifier, the corresponding layout optimization rule is obtained from the preset rule base to determine the optimized layout parameters; Based on the optimized layout parameters, structural correlation data is extracted from the form layout, and the consistency between the correlation data and the layout parameters is judged by the data comparison method to obtain the consistency verification result. If the consistency verification result meets the preset standard, the grid parameters are adjusted according to the verification result to determine the final reconstructed form layout structure; Based on the final reconstructed form layout structure, structural stability data is obtained, and the matching probability of the stability data is calculated using the Naive Bayes algorithm to obtain the stability evaluation result. If the stability assessment result reaches the preset threshold, additional validation data is extracted from the form layout, the degree of matching between the additional validation data and the reconstructed structure is judged, and the final layout optimization result is determined.

7. The method according to claim 1, characterized in that, The method of optimizing interaction by using an adaptive layout generation approach for the identified behavior patterns includes: Based on the behavioral pattern feature set, a hidden Markov model is used to analyze the temporal distribution characteristics of the operation sequence and determine the operation time distribution pattern. If the distribution pattern of operation time deviates from the preset time threshold, click frequency statistics are extracted from the operation behavior data, and statistical analysis methods are used to determine whether the click frequency conforms to the normal interaction mode, so as to obtain the click frequency evaluation result. Based on the click frequency evaluation results, the input content patterns are extracted from user operation behavior data, and the consistency between the input content and the preset rules is compared using pattern matching methods to determine the compliance results of the input content. If the compliance result of the input content meets the preset standard, the operation path trajectory is extracted from the operation behavior data, and the smoothness of the path trajectory is calculated using the trajectory analysis method to obtain the path trajectory smoothness data. Based on the path trajectory smoothness data, the interface response speed is extracted from the form layout structure. The response time analysis method is used to determine the degree of matching between the response speed and the operation path, and to determine the need to adjust the interaction method. If the need to adjust the interaction method reaches a preset threshold, then the interaction optimization rules are obtained from the preset rule base, and the interaction parameters of the form layout are adjusted using the rule application method to obtain the optimized interaction method configuration.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-7.