Receipt interface type generation method and system based on target data

By screening high-quality data sources and dynamically generating document interface templates, the problems of insufficient flexibility and uneven data quality in traditional document interface generation methods are solved, the accuracy and reliability of the document interface are improved, and it adapts to complex and changing business environments.

CN120687090AActive Publication Date: 2025-09-23SHENZHEN TEWEI KECHUANG INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510799725.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-23
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Traditional document interface generation methods rely on manual design and fixed templates, which lack flexibility and are difficult to adapt to diverse and dynamically changing business needs. In addition, data quality varies, affecting the accuracy and reliability of the document interface.

Method used

By extracting the amount of erroneous data in the target data and its copies, comparing it with the preset parameter threshold, screening high-quality data sources, and using dynamic metadata technology to analyze business logic characteristics, calculate multi-dimensional similarity and weight distribution, and dynamically generate document interface templates.

Benefits of technology

It improves the accuracy and completeness of the document interface, reduces business errors and decision-making errors caused by data quality issues, enhances user experience and system flexibility, and adapts to different business needs and data structure changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687090A_ABST
    Figure CN120687090A_ABST
Patent Text Reader

Abstract

The invention discloses a receipt interface type generation method and system based on target data, and belongs to the technical field of receipt processing. The problems that an existing method is insufficient in flexibility and uneven in data quality are solved, the high-quality data source can be screened out by extracting the target data and the error data size in the copy of the target data and comparing the error data size with the preset parameter threshold value, the accuracy and integrity of a receipt interface are ensured, and the data quality is improved. Service errors and decision errors caused by data quality problems are reduced, so that the use experience of a user is improved; by calculating the multi-dimensional similarity and combining the weight distribution, the system can preferentially consider some key features, so that the requirements in a specific service scene are better met, the receipt interface template which most conforms to the target data service logic is found, and the receipt interface generation accuracy and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document processing, and in particular to a method and system for generating document interface types based on target data. Background Art

[0002] In modern enterprise information systems, document management is a core component of business processes. Traditional document interface generation methods rely primarily on manual design and fixed templates. This approach presents the following problems when faced with diverse and dynamically changing business needs: 1. Lack of flexibility: Fixed templates are difficult to adapt to changes in business scenarios. Every business adjustment requires redesigning the document interface, which cannot quickly respond to changes in business needs. This is time-consuming and labor-intensive, affecting the company's operational efficiency.

[0003] 2. Data quality varies: Data from different sources may have quality issues, such as incorrect data, duplicate data, etc., which affects the accuracy and reliability of the document interface. Therefore, it does not meet the existing needs, so we propose a document interface type generation method and system based on target data. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for generating document interface types based on target data. By extracting the amount of erroneous data in the target data and its copies and comparing it with a preset parameter threshold, it can screen out high-quality data sources, ensure the accuracy and completeness of the document interface, reduce business errors and decision-making errors caused by data quality problems, and improve the user experience; by calculating multi-dimensional similarity and combining weight distribution, the system can give priority to certain key features, so as to better meet the needs of specific business scenarios, ensure that the document interface template that best meets the business logic of the target data is found, thereby improving the accuracy and reliability of the document interface generation, and solving the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: A method for generating a document interface type based on target data, the method comprising the following steps: Obtain historical data and design various document interface templates based on different business scenarios and document types in the historical data; Analyze historical data based on dynamic metadata technology and define business logic features for each document interface template. Business logic features include: field type, number of fields, and business logic mode. The defined document interface templates are stored in the template matching library. The collected target data is mapped, cleaned, and quality-assessed, and then stored to form a global metadata directory of the target data. The quality assessment process includes: counting the amount of erroneous data in the target data and its copies, and performing a graded data quality assessment using field type entropy analysis and a dynamic threshold adjustment algorithm. Analyze the metadata of target data based on dynamic metadata technology and extract the business logic characteristics of target data; Based on the feature matching algorithm, the best matching document interface template is found in the template matching library for the business logic features of the target data; According to the matching document interface template, map the target data fields to the form components defined in the document interface template; The document interface is dynamically generated based on the field mapping relationship and the document interface template layout.

[0006] Furthermore, when mapping and cleaning the collected target data, it also includes: Extracting error data and corresponding quantities from the target data as a first data amount; Acquire data related to target data from multiple channels as copies of the target data; Extracting the error data and the corresponding quantity in the target data copy as the second data volume; Presetting a parameter threshold, comparing the first data volume and the second data volume with the preset parameter threshold respectively, to obtain a comparison result; Based on the comparison result, the collection quality of the target data and the target data copy is evaluated to obtain a data quality evaluation result; Based on the evaluation results, select the target data for generating the document interface.

[0007] Furthermore, parameter thresholds are preset, including: Extract the total quantity of fields in the document interface; Classify the fields in the document interface and obtain the corresponding field types in the document interface; For each field type, extract the ratio of the number of fields corresponding to the field type to the total number of fields; The Shannon entropy value corresponding to each field type is obtained by using the ratio of the number of fields corresponding to each field type to the total number of fields; Normalizing the Shannon entropy value to obtain the normalized Shannon entropy value; Comparing the first data volume with the second data volume to obtain a data volume difference between the first data volume and the second data volume; Performing summing and averaging processing on the first data volume and the second data volume to obtain an average data volume corresponding to the first data volume and the second data volume as a whole; The data volume difference is compared with the data volume average value to obtain a data volume ratio coefficient; The parameter threshold is set using the data volume ratio coefficient combined with the normalized Shannon entropy value corresponding to the field type.

[0008] Furthermore, the parameter threshold is set by using the data volume ratio coefficient in combination with the normalized Shannon entropy value corresponding to the field type, including: Retrieve the normalized Shannon entropy value corresponding to the data volume ratio coefficient and field type; Comparing the data volume ratio coefficient with the normalized Shannon entropy value corresponding to the field type; When the data volume ratio coefficient does not exceed the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database as a parameter threshold for subsequent collection quality assessment; When the data volume ratio coefficient exceeds the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database and the benchmark parameter threshold is adjusted; The parameter threshold obtained after adjustment is used as the parameter threshold for subsequent acquisition quality assessment.

[0009] Furthermore, based on the feature matching algorithm, the best matching document interface template is found in the template matching library for the business logic features of the target data, specifically: Standardize the business logic features of each document interface template in the template matching library; Compare the field type similarity, field quantity similarity, and business logic model similarity between the target data and the template to obtain multi-dimensional similarity; Assign different weights to field types, number of fields, and business logic patterns based on business needs; Multiply the similarity of each dimension by the corresponding weight and sum them up to get the comprehensive similarity; Select the template with the highest overall similarity as the most matching document interface template; If there are multiple templates with the same similarity, the selection is made based on other factors.

[0010] Furthermore, the metadata of the target data is analyzed based on dynamic metadata technology to extract the business logic features of the target data, specifically: Determine the specific goals of metadata analysis and the source of target data, and use metadata management tools to automatically collect metadata of target data, including technical metadata and business metadata; Identify the field types in the target data by analyzing the metadata of the target data, and count the total number of fields in the target data and the number of each field type; Analyze the dependencies between fields, extract the validation rules of the fields, and identify the calculation logic of the fields to obtain the business logic characteristics of the target data.

[0011] Furthermore, based on the comparison results, the collection quality of the target data and the target data copy is evaluated, specifically: If the comparison result shows that the first data volume exceeds the parameter threshold, and the second data volume does not exceed the parameter threshold, it indicates that the collection quality of the target data corresponding to the first data volume is low, and the target data copy corresponding to the second data volume will be used first as the target data for the current document generation interface; If the comparison result shows that the first data volume does not exceed the parameter threshold, but the second data volume exceeds the parameter threshold, it indicates that the collection quality of the target data copy corresponding to the second data volume is low. In this case, the target data corresponding to the first data volume will be used as the target data for the current document generation interface. If the comparison result shows that both the first and second data volumes exceed the parameter threshold, the acquisition channels of the target data and the target data copy are verified to determine whether there are any omissions in the transmission process of the target data and the target data copy. If the transmission is verified to be consistent, the data with the lower percentage exceeding the parameter threshold is preferentially used as the target data for the current document generation interface. If the comparison result shows that both the first data volume and the second data volume do not exceed the parameter threshold, the target data corresponding to the first data volume is preferentially used as the target data for the current document generation interface.

[0012] Furthermore, after the defined document interface template is stored in the template matching library, it includes: Establish a fixed review cycle and review plan to review document interface templates and business logic features; Establish a user feedback mechanism to regularly collect user feedback on document interface templates and business logic features; Based on the review results and user feedback, formulate improvement measures and make timely adjustments and optimizations.

[0013] A document interface type generation system based on target data, the system comprising: a data acquisition unit and a data processing unit, the data processing unit comprising: a quality assessment module, a feature extraction module and a feature matching module; a data acquisition unit configured to obtain target data from a relational database and obtain copies of the target data from multiple channels; a quality assessment module configured to obtain the amount of data corresponding to erroneous data in the target data and the target data copy, evaluate the data quality of the target data and the target data copy based on a preset parameter threshold, and select the optimal data source based on the evaluation result; A feature extraction module is configured to analyze the metadata of the target data based on dynamic metadata technology, identify the dependencies between fields, verification rules and calculation logic, and obtain the business logic features of the target data; A feature matching module configured to use a feature matching algorithm to find the best matching document interface template in a template matching library; The template construction module is configured to design multiple document interface templates based on historical data and business scenarios, and store the multiple document interface templates to form a template matching library; An interface generation module is configured to associate the fields of the target data with the form components defined by the matching document interface template, and to construct the document interface based on the association results of the fields and the layout of the document interface template; A user feedback module is configured to regularly collect user feedback on document interface templates and business logic features; The review and optimization module is configured to regularly review document interface templates and business logic features, and formulate improvement measures based on user feedback.

[0014] Furthermore, the data acquisition unit includes: a data cleaning module configured to map and cleanse target data and its metadata and a copy of the target data; The data storage module is configured to store metadata of target data to form a global metadata directory.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention can screen out high-quality data sources by extracting the amount of erroneous data from the target data and its copies and comparing it with preset parameter thresholds. This can ensure the accuracy and completeness of the document interface, reduce business errors and decision-making errors caused by data quality issues, and thus improve the user experience. When using the document interface, users can obtain accurate information, reduce operational difficulties caused by data errors, and improve work efficiency.

[0016] 2. The present invention calculates multi-dimensional similarity and combines it with weight allocation, so that the system can give priority to certain key features, flexibly adapt to different business needs and data structure changes, and ensure that the document interface template that best meets the business logic of the target data is found, better meeting the needs of specific business scenarios, and effectively avoiding the limitations of single-dimensional evaluation, thereby improving the accuracy and reliability of document interface generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A diagram of the system composition for generating a document interface type based on target data of the present invention; Figure 2This is a flow chart of the method for generating a document interface type based on target data according to the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] In order to solve the technical problems of the existing technology, the traditional document interface generation method mainly relies on manual design and fixed templates. This method still has insufficient flexibility and uneven data quality. Figure 1-Figure 2 , this embodiment provides the following technical solutions: The method for generating a document interface type based on target data comprises the following steps: Obtain historical data and design various document interface templates based on different business scenarios and document types in the historical data.

[0020] Based on dynamic metadata technology, historical data is analyzed and business logic features are defined for each document interface template. Business logic features include: field type, number of fields, and business logic mode. The defined document interface template is stored in the template matching library. The following steps are then included: Establish a fixed review cycle and review plan, such as monthly or quarterly, and clearly define the specific time, tasks and responsible persons for each review; Review document interface templates, including: checking whether the templates contain all necessary fields and components; verifying whether the field types and display order meet business requirements; evaluating whether the interface layout is reasonable and conforms to user operating habits; and testing the display effects of the templates on different devices and browsers; Review business logic features, including: verifying that field types are consistent with business requirements; checking that the number of fields complies with business rules; reviewing the dependencies between fields, the calculation logic, and validation rules; and checking that data complies with business logic and validation rules. Establish a user feedback mechanism to regularly collect user feedback on document interface templates and business logic features. Provide multiple feedback channels, such as online feedback forms and user forums. Set a fixed feedback collection cycle, such as monthly or quarterly. Categorize and analyze collected feedback to identify key issues and improvement suggestions. Based on the review results and user feedback, we developed specific improvement measures and made timely adjustments and optimizations. For example, when users reported that the order entry template lacked a customer_name field, the technical team added a customer_name field with a string type to the order entry template and conducted tests. This ensured the integrity of the document interface template and the user experience, as well as the accuracy and consistency of the business logic features.

[0021] The collected target data is mapped, cleaned, and quality-assessed, and then stored to form a global metadata directory for the target data. The quality assessment process includes: counting the amount of erroneous data in the target data and its copies, and combining field type entropy analysis and dynamic threshold adjustment algorithm to perform data quality grading assessment, specifically including: Extracting error data and corresponding quantities from the target data as a first data amount; Acquire data related to target data from multiple channels as copies of the target data; Extracting the error data and the corresponding quantity in the target data copy as the second data volume; Presetting a parameter threshold, comparing the first data volume and the second data volume with the preset parameter threshold respectively, to obtain a comparison result; Based on the comparison results, the collection quality of the target data and the target data copy is evaluated to obtain a data quality evaluation result; specifically: If the comparison result shows that the first data volume exceeds the parameter threshold, and the second data volume does not exceed the parameter threshold, it indicates that the collection quality of the target data corresponding to the first data volume is low, and the target data copy corresponding to the second data volume will be used first as the target data for the current document generation interface; If the comparison result shows that the first data volume does not exceed the parameter threshold, but the second data volume exceeds the parameter threshold, it indicates that the collection quality of the target data copy corresponding to the second data volume is low. In this case, the target data corresponding to the first data volume will be used as the target data for the current document generation interface. If the comparison result shows that both the first and second data volumes exceed the parameter threshold, the acquisition channels of the target data and the target data copy are verified to determine whether there are any omissions in the transmission process of the target data and the target data copy. If the transmission is verified to be consistent, the data with the lower percentage exceeding the parameter threshold is preferentially used as the target data for the current document generation interface. If the comparison result shows that both the first data volume and the second data volume do not exceed the parameter threshold, the target data corresponding to the first data volume is preferentially used as the target data for the current document generation interface; Based on the evaluation results, select the target data for generating the document interface.

[0022] The beneficial effects achieved by the above content are: by extracting the amount of erroneous data in the target data and its copies and comparing it with the parameter threshold, the data collection quality can be effectively evaluated, and high-quality data sources can be screened out, thereby ensuring the accuracy and completeness of the document interface, reducing business errors and decision-making errors caused by data quality issues, and improving the user experience; when using the document interface, users can obtain accurate information and improve work efficiency.

[0023] Specifically, the parameter thresholds are preset, including: Extract the total quantity of fields in the document interface; Classify the fields in the document interface and obtain the corresponding field types in the document interface; For each field type, extract the ratio of the number of fields corresponding to the field type to the total number of fields; The Shannon entropy value corresponding to each field type is obtained by using the ratio of the number of fields corresponding to each field type to the total number of fields; The Shannon entropy value is obtained by the following formula: Among them, E represents the Shannon entropy value; n represents the number of field types; p i Indicates the ratio of the number of fields corresponding to the i-th field type to the total number of fields; Normalizing the Shannon entropy value to obtain the normalized Shannon entropy value; Comparing the first data volume with the second data volume to obtain a data volume difference between the first data volume and the second data volume; Performing summing and averaging processing on the first data volume and the second data volume to obtain an average data volume corresponding to the first data volume and the second data volume as a whole; The data volume difference is compared with the data volume average value to obtain a data volume ratio coefficient; The parameter threshold is set using the data volume ratio coefficient combined with the normalized Shannon entropy value corresponding to the field type.

[0024] The technical effect of the above technical solution is that, compared to existing technologies that rely on manual experience to set fixed thresholds or rely solely on single-dimensional data indicators, this solution significantly improves the scientific and adaptable nature of parameter threshold setting by integrating business characteristics with a dual-source comparison mechanism for data quality. Firstly, by utilizing Shannon entropy to quantify the diversity of field type distributions within document interfaces, business logic complexity is converted into a computable numerical metric, allowing thresholds to be deeply tied to the complexity of business scenarios. This avoids misjudgments caused by a "one-size-fits-all" fixed threshold (e.g., overly strict thresholds for simple scenarios, overly loose thresholds for complex scenarios). Normalization also achieves a unified measurement across scenarios, allowing thresholds to be dynamically adjusted based on business complexity. After normalizing the Shannon entropy value, the business complexity of different document interfaces can be uniformly mapped to the [0, 1] interval, facilitating mathematical integration with data quality indicators (such as the data volume ratio coefficient) to achieve dynamic adjustment of thresholds across scenarios. Secondly, by comparing the amount of erroneous data between the target data and its replica, calculating the data volume difference and average, and obtaining a data volume ratio coefficient reflecting the stability of data collection, this achieves dual-source verification of data quality. Finally, the business complexity index is combined with the data quality index to construct a composite threshold setting model, so that when the business is complex and the difference between the dual-source data is large, the threshold is automatically tightened, and when the business is simple and the data consistency is high, the threshold is automatically relaxed, thereby realizing the transformation from "manual experience preset" to "automatic derivation of data features", effectively solving the problems of rigidity and poor adaptability of traditional thresholds, and improving the accuracy of data quality assessment and the level of intelligent document generation.

[0025] Specifically, the parameter threshold is set using the data volume ratio coefficient combined with the normalized Shannon entropy value corresponding to the field type, including: Retrieve the normalized Shannon entropy value corresponding to the data volume ratio coefficient and field type; Comparing the data volume ratio coefficient with the normalized Shannon entropy value corresponding to the field type; When the data volume ratio coefficient does not exceed the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database as a parameter threshold for subsequent collection quality assessment; wherein the preset benchmark parameter threshold is obtained based on actual application requirements combined with experiments or experience; When the data volume ratio coefficient exceeds the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database and the benchmark parameter threshold is adjusted; The parameter threshold obtained after adjustment is used as the parameter threshold for subsequent acquisition quality assessment.

[0026] The parameter threshold obtained after adjustment is obtained by the following formula: Where Y represents the parameter threshold obtained after adjustment; Y0 represents the preset baseline parameter threshold; E g represents the normalized Shannon entropy value; D represents the data volume ratio coefficient; m represents the number of target data collected from historical data; C i represents the normalized data volume difference between the first data volume and the second data volume corresponding to the target data acquired for the i-th time. Specifically, The square root of the product of these two factors reflects the degree to which threshold adjustments are influenced by a comprehensive consideration of business complexity and data quality stability. The more complex the business and the greater the fluctuation in data quality, the larger this value, and the greater the impact on threshold adjustment. This is similar to how multiple influencing factors in physics work together to produce a single, comprehensive impact. The normalized data volume difference C of each target data i Take the negative exponent and sum it. Negative exponent operation makes C i The smaller the value (i.e., the smaller the difference between the two-source data and the better the data quality), the greater its contribution. The sum reflects the comprehensive quality of all previous data. The above sum is divided by the number of historical data collection times m, and the average is performed to obtain an indicator reflecting the overall quality of the historical data. This is similar to the statistical averaging of multiple sample data in statistical physics to obtain a representative statistic. It will adjust the threshold adjustment range based on the quality of the historical data. The better the overall data quality, the closer this value is to 1, and the more significant the amplification effect on the threshold adjustment range. This formula combines the business complexity (E g ), the current quality status of the data (D) and the quality of historical data (through C i The preset baseline parameter threshold, Y0, is adjusted from multiple perspectives, comprehensively considering factors influencing data quality assessment. This avoids the irrationality of setting thresholds based solely on a single dimension or partial information, ensuring that the resulting parameter thresholds more scientifically reflect actual data quality requirements. A dynamic adjustment mechanism adjusts the thresholds in real time based on different business scenarios (e.g., changes in Shannon entropy) and data quality fluctuations (e.g., changes in the data volume ratio and historical data differences). This improves the adaptability and effectiveness of the thresholds during data quality assessment.

[0027] The technical solution described above accurately quantifies business logic complexity by extracting the total number of fields and categorized fields from the document interface and calculating the Shannon entropy value. Parameter thresholds are adjusted based on this information, ensuring they are deeply aligned with the business scenario. This prevents data quality assessment bias caused by inaccurate judgment of business complexity, significantly improving the accuracy of data quality assessment results. This in turn makes the target data selected based on accurate assessment results more reliable, providing a high-quality data foundation for document interface generation. The data volume ratio coefficient is determined by comparing the target data with its copy, and the Shannon entropy value is used to set the threshold. The cross-verification of dual-source data effectively identifies fluctuations and errors in the data collection process. Compared to single-source data evaluation, this method can more accurately locate erroneous data, improve the accuracy of erroneous data identification, and ensure that parameter thresholds are more closely aligned with actual data quality conditions, enhancing the accuracy of data quality assessments. As business operations develop and data changes, the data volume ratio coefficient, Shannon entropy value, and other factors change in real time. By adjusting parameter thresholds based on these dynamically changing indicators, the solution can respond to data fluctuations in real time. Maintaining a stable assessment of data quality despite changes in the data collection environment and adjustments to business logic, the system avoids data quality assessment failures or misjudgments due to fixed thresholds, ensuring the stability of data processing during document interface generation. The formula incorporates the number of historical data collection times and the difference in data volume between each collection, leveraging historical data patterns to assist in threshold adjustment. When data fluctuates short-term, comprehensive consideration of historical data prevents over-adjustment of thresholds and maintains a relatively stable assessment system. When data exhibits long-term trends, thresholds can be adjusted promptly based on accumulated historical data, achieving a balance between stability and adaptability. This technical solution automates the calculation and adjustment of parameter thresholds, eliminating the need for frequent human intervention and manual threshold modification. From data feature extraction and indicator calculation to threshold determination, the entire process is automated, reducing manual operation time and costs, accelerating the data quality assessment process, and improving the overall efficiency of data processing before document interface generation. Appropriate threshold settings more efficiently filter out target data that meets requirements, reducing repeated processing and error correction operations caused by data quality issues during subsequent document interface generation. By accurately assessing data quality, high-quality data can be directly selected for field mapping and interface generation, optimizing the data processing process, shortening the document interface generation cycle, and improving the overall operating efficiency of the system.

[0028] On the other hand, existing technologies often use fixed or simply adjusted parameter thresholds, which are difficult to adapt to complex and changing data environments. This technical solution dynamically adjusts the threshold based on the data volume ratio coefficient and the normalized Shannon entropy value. When the comparison results of the data volume ratio coefficient and the Shannon entropy value are different, different strategies are adopted to better adapt to the fluctuations in data collection quality and the complexity of business logic, so that the threshold setting is more in line with the actual data situation, and the accuracy of data quality assessment is improved. It integrates business characteristics (Shannon entropy reflects the diversity of field types, that is, business complexity) and data quality characteristics (data volume ratio coefficient reflects the consistency of dual-source data). Compared with existing technologies that only consider a single factor, this solution can set thresholds based on multi-dimensional information, comprehensively consider business scenarios and data collection stability, make thresholds more reasonable, and avoid misjudgments or inaccurate assessments caused by single-factor judgments. The formula introduces the number of historical data collection times m and the normalized data volume difference C of each target data. i , using historical data to adjust and optimize thresholds. Existing technologies often lack effective use of historical data. This solution can better learn the patterns of data changes and further improve the accuracy and adaptability of threshold settings.

[0029] Analyze the metadata of target data based on dynamic metadata technology and extract the business logic features of target data; specifically: Determine the specific goals of metadata analysis and the source of the target data. Specific goals include: extracting the field type, number of fields, and business logic model of the target data; the source of the target data, for example: a relational database, and use metadata management tools to automatically collect metadata of the target data. Metadata includes: technical metadata and business metadata. Technical metadata includes: data table structure, field definition, storage location, etc. Business metadata includes: business meaning of fields, business rules, etc.

[0030] By analyzing the metadata of the target data, the field types in the target data, such as strings, dates, numbers, etc., are identified, and the total number of fields in the target data and the number of each field type are counted.

[0031] Analyze the dependencies between fields, for example, whether the value of a field depends on other fields; extract field validation rules, such as format validation, range validation, etc.; and identify the field calculation logic, for example, whether a field is calculated from other fields, to obtain the business logic characteristics of the target data.

[0032] Based on the feature matching algorithm, the best matching document interface template is found in the template matching library for the business logic features of the target data; specifically: Standardize the business logic features of each document interface template in the template matching library to facilitate quantitative comparison; for example, convert field types to unified encodings and normalize the number of fields.

[0033] Compare the field type similarity, field quantity similarity, and business logic pattern similarity between the target data and the template to obtain multi-dimensional similarity.

[0034] Assign different weights to field types, number of fields, and business logic patterns based on business needs; Multiply the similarity of each dimension by the corresponding weight and sum them up to get the comprehensive similarity; The template with the highest overall similarity is selected as the most matching document interface template; if there are multiple templates with the same similarity, the selection is made based on other factors, such as the frequency of template use and update time.

[0035] For example: Assume that the target data is an order system containing the following fields: string, date, string, table, and number.

[0036] In the template matching library, there are the following two templates: Template A: Field types: 2 string fields, 1 date field, 1 table field, 1 number field; Number of fields: 5 fields; Business logic mode: order_total depends on order_items.

[0037] Template B: Field type: 3 string fields, 1 date field, 1 number field; Number of fields: 5 fields; Business logic mode: No field dependency.

[0038] Through the feature matching algorithm: field type similarity: template A is 0.8, template B is 0.6; field quantity similarity: template A is 0.9, template B is 0.9; business logic model similarity: template A is 0.9, template B is 0.1.

[0039] Comprehensive similarity calculation, assuming weights are 0.4, 0.3, and 0.3 respectively: The comprehensive similarity of template A is: 0.8×0.4+0.9×0.3+0.9×0.3=0.82; The comprehensive similarity of template B is: 0.6×0.4+0.9×0.3+0.1×0.3=0.51; Finally, Template A is selected as the most matching document interface template.

[0040] The beneficial effects achieved by the above content are: by calculating multi-dimensional similarity and combining it with weight distribution, the system can flexibly adapt to different business needs and data structure changes; by adjusting the weight distribution, certain key features can be given priority to better meet the needs of specific business scenarios, enabling the system to more comprehensively and accurately evaluate the matching degree between target data and templates, effectively avoiding the limitations of single-dimensional evaluation, and ensuring that the document interface template that best conforms to the business logic of the target data is found, thereby improving the accuracy and reliability of document interface generation.

[0041] According to the matching document interface template, the fields of the target data are mapped to the form components defined in the document interface template; based on the field mapping relationship and the document interface template layout, the document interface is dynamically generated.

[0042] A document interface type generation system based on target data includes: a data acquisition unit and a data processing unit.

[0043] The data collection unit is configured to obtain target data from a relational database and obtain copies of the target data from multiple channels; wherein the target data serves as the optimal data source for generating a document interface, and the copies of the target data serve as backup data; the data collection unit includes: The data cleaning module is configured to map and clean the target data and its metadata and target data copies, remove duplicate or inconsistent information, and ensure the consistency of data between different sources.

[0044] The data storage module is configured to store the metadata of the target data to form a global metadata directory to facilitate subsequent query and management.

[0045] The data processing unit includes: a quality assessment module, a template construction module, a feature extraction module, a feature matching module, an interface generation module, a user feedback module and a review optimization module.

[0046] The quality assessment module is configured to obtain the data volume corresponding to the erroneous data in the target data and the target data copy, evaluate the data quality of the target data and the target data copy based on the preset parameter threshold, and select the optimal data source according to the evaluation results; selecting a high-quality data source to generate the document interface can ensure the accuracy and completeness of the document interface, reduce business errors and decision-making errors caused by data quality issues, and thus improve the user experience.

[0047] The template construction module is configured to design multiple document interface templates based on historical data and business scenarios, and store multiple document interface templates to form a template matching library; so that after obtaining the target data, the corresponding document interface template can be quickly matched in the template matching library according to its business logic characteristics.

[0048] The feature extraction module is configured to analyze the metadata of the target data based on dynamic metadata technology, extract the field type, number of fields and business logic pattern, identify the dependency relationship, verification rules and calculation logic between fields, and obtain the business logic characteristics of the target data.

[0049] The feature matching module is configured to analyze the business logic features of the target data based on dynamic metadata technology, and use the feature matching algorithm to find the most matching document interface template in the template matching library; by comparing the field type similarity, field quantity similarity and business logic pattern similarity between the target data and the template, different weights are assigned to the field type, field quantity and business logic pattern according to business needs, and the comprehensive similarity is calculated. The template with the highest comprehensive similarity is selected as the most matching document interface template.

[0050] The interface generation module is configured to associate the fields of the target data with the form components defined by the matching document interface template, and to construct the document interface based on the association results of the fields and the layout of the document interface template.

[0051] The user feedback module is configured to regularly collect user feedback on document interface templates and business logic features; the review and optimization module is configured to regularly review document interface templates and business logic features, and formulate improvement measures based on user feedback information.

[0052] Working principle: Build a template matching library, obtain target data and target data copies, extract the corresponding number of erroneous data, compare the erroneous data volume of the two with the parameter threshold, and select the data with the highest quality as the basic data for generating the document interface; extract features of the selected target data based on dynamic metadata technology, match the corresponding document interface template in the template matching library based on its business logic features, and map its fields to the form component defined in the document interface template to dynamically generate the document interface.

[0053] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0054] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating a document interface type based on target data, characterized in that: The method comprises the following steps: Obtain historical data and design various document interface templates based on different business scenarios and document types in the historical data; Analyze historical data based on dynamic metadata technology and define business logic features for each document interface template. Business logic features include: field type, number of fields, and business logic mode. The defined document interface templates are stored in the template matching library. The collected target data is mapped, cleaned, and quality-assessed, and then stored to form a global metadata directory of the target data. The quality assessment process includes: counting the amount of erroneous data in the target data and its copies, and performing a graded data quality assessment using field type entropy analysis and a dynamic threshold adjustment algorithm. Analyze the metadata of target data based on dynamic metadata technology and extract the business logic characteristics of target data; Based on the feature matching algorithm, the best matching document interface template is found in the template matching library for the business logic features of the target data; According to the matching document interface template, map the target data fields to the form components defined in the document interface template; The document interface is dynamically generated based on the field mapping relationship and the document interface template layout.

2. The method for generating a document interface type based on target data according to claim 1, characterized in that: When mapping and cleaning the collected target data, it also includes: Extracting error data and corresponding quantities from the target data as a first data amount; Acquire data related to target data from multiple channels as copies of the target data; Extracting the error data and the corresponding quantity in the target data copy as the second data volume; Presetting a parameter threshold, comparing the first data volume and the second data volume with the preset parameter threshold respectively, to obtain a comparison result; Based on the comparison result, the collection quality of the target data and the target data copy is evaluated to obtain a data quality evaluation result; Based on the evaluation results, select the target data for generating the document interface.

3. The method for generating a document interface type based on target data according to claim 2, characterized in that: Preset parameter thresholds, including: Extract the total quantity of fields in the document interface; Classify the fields in the document interface and obtain the corresponding field types in the document interface; For each field type, extract the ratio of the number of fields corresponding to the field type to the total number of fields; The Shannon entropy value corresponding to each field type is obtained by using the ratio of the number of fields corresponding to each field type to the total number of fields; Normalizing the Shannon entropy value to obtain the normalized Shannon entropy value; Comparing the first data volume with the second data volume to obtain a data volume difference between the first data volume and the second data volume; Performing summing and averaging processing on the first data volume and the second data volume to obtain an average data volume corresponding to the first data volume and the second data volume as a whole; The data volume difference is compared with the data volume average value to obtain a data volume ratio coefficient; The parameter threshold is set using the data volume ratio coefficient combined with the normalized Shannon entropy value corresponding to the field type.

4. The method for generating a document interface type based on target data according to claim 3, characterized in that: The parameter threshold is set using the data volume ratio coefficient combined with the normalized Shannon entropy value corresponding to the field type, including: Retrieve the normalized Shannon entropy value corresponding to the data volume ratio coefficient and field type; Comparing the data volume ratio coefficient with the normalized Shannon entropy value corresponding to the field type; When the data volume ratio coefficient does not exceed the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database as a parameter threshold for subsequent collection quality assessment; When the data volume ratio coefficient exceeds the normalized Shannon entropy value corresponding to the field type, a preset benchmark parameter threshold is retrieved from the database and the benchmark parameter threshold is adjusted; The parameter threshold obtained after adjustment is used as the parameter threshold for subsequent acquisition quality assessment.

5. The method for generating a document interface type based on target data according to claim 1, characterized in that: Based on the feature matching algorithm, the best matching document interface template is found in the template matching library for the business logic features of the target data. Specifically: Standardize the business logic features of each document interface template in the template matching library; Compare the field type similarity, field quantity similarity, and business logic model similarity between the target data and the template to obtain multi-dimensional similarity; Assign different weights to field types, number of fields, and business logic patterns based on business needs; Multiply the similarity of each dimension by the corresponding weight and sum them up to get the comprehensive similarity; Select the template with the highest overall similarity as the most matching document interface template; If there are multiple templates with the same similarity, the selection is made based on other factors.

6. The method for generating a document interface type based on target data according to claim 1, characterized in that: Analyze the metadata of the target data based on dynamic metadata technology and extract the business logic features of the target data, specifically: Determine the specific goals of metadata analysis and the source of target data, and use metadata management tools to automatically collect metadata of target data, including technical metadata and business metadata; Identify the field types in the target data by analyzing the metadata of the target data, and count the total number of fields in the target data and the number of each field type; Analyze the dependencies between fields, extract the validation rules of the fields, and identify the calculation logic of the fields to obtain the business logic characteristics of the target data.

7. The method for generating a document interface type based on target data according to claim 2, characterized in that: Based on the comparison results, the collection quality of the target data and the target data copy is evaluated, specifically: If the comparison result shows that the first data volume exceeds the parameter threshold, and the second data volume does not exceed the parameter threshold, it indicates that the collection quality of the target data corresponding to the first data volume is low, and the target data copy corresponding to the second data volume will be used first as the target data for the current document generation interface; If the comparison result shows that the first data volume does not exceed the parameter threshold, but the second data volume exceeds the parameter threshold, it indicates that the collection quality of the target data copy corresponding to the second data volume is low. In this case, the target data corresponding to the first data volume will be used as the target data for the current document generation interface. If the comparison result shows that both the first and second data volumes exceed the parameter threshold, the acquisition channels of the target data and the target data copy are verified to determine whether there are any omissions in the transmission process of the target data and the target data copy. If the transmission is verified to be consistent, the data with the lower percentage exceeding the parameter threshold is preferentially used as the target data for the current document generation interface. If the comparison result shows that both the first data volume and the second data volume do not exceed the parameter threshold, the target data corresponding to the first data volume is preferentially used as the target data for the current document generation interface.

8. The method for generating a document interface type based on target data according to claim 1, characterized in that: After the defined document interface template is stored in the template matching library, it includes: Review document interface templates and business logic features, regularly collect user feedback on document interface templates and business logic features, formulate improvement measures based on review results and user feedback, and make timely adjustments and optimizations.

9. A document interface type generation system based on target data, applied to a document interface type generation method based on target data according to any one of claims 1 to 8, characterized in that: The system includes: a data acquisition unit and a data processing unit, and the data processing unit includes: a quality assessment module, a feature extraction module and a feature matching module; a data acquisition unit configured to obtain target data from a relational database and obtain copies of the target data from multiple channels; a quality assessment module configured to obtain the amount of data corresponding to erroneous data in the target data and the target data copy, evaluate the data quality of the target data and the target data copy based on a preset parameter threshold, and select the optimal data source based on the evaluation result; A feature extraction module is configured to analyze the metadata of the target data based on dynamic metadata technology, identify the dependencies between fields, verification rules and calculation logic, and obtain the business logic features of the target data; A feature matching module is configured to use a feature matching algorithm to find the best matching document interface template in the template matching library; The template construction module is configured to design multiple document interface templates based on historical data and business scenarios, and store the multiple document interface templates to form a template matching library; An interface generation module is configured to associate the fields of the target data with the form components defined by the matching document interface template, and to construct the document interface based on the association results of the fields and the layout of the document interface template; A user feedback module is configured to regularly collect user feedback on document interface templates and business logic features; The review and optimization module is configured to regularly review document interface templates and business logic features, and formulate improvement measures based on user feedback information.

10. The document interface type generation system based on target data according to claim 9, characterized in that: The data acquisition unit includes: a data cleaning module configured to map and cleanse target data and its metadata and a copy of the target data; The data storage module is configured to store metadata of target data to form a global metadata directory.

Citation Information

Patent Citations

  • Receipt development method and device, computer equipment and storage medium

    CN108196921A

  • Multi-data-source data processing method and system

    CN119415503A

  • Big data processing-based document information input method and system

    CN119990719A

  • Form processing method, form processing device, and computer product

    US20080025618A1