Report data asset identification method, system, equipment and medium
Through the identification mechanism and knowledge graph technology of multi-technology integration, the problem of multi-source heterogeneous data recognition in the banking industry has been solved, the accurate identification and standardized conversion of report data has been achieved, and data processing efficiency and accuracy have been improved.
Patent Information
- Application Number
- CN202510863864.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the digital transformation of the banking industry, multi-source heterogeneous data have poor compatibility. The existing OCR technology lacks the accuracy of complex tables, fuzzy images and handwritten characters, making it difficult to extract the relationship between financial statements and the causal chain of customer behavior data, and cannot support risk control modeling or asset value mining.
Through the identification mechanism of multi-technology integration, including data preprocessing, multi-modal data recognition, field-level extraction, automatic mapping and verification storage, a data association network is built with knowledge graphs to achieve accurate identification and standardized conversion of report data.
Improve the accuracy and depth of data identification, ensure the integrity and standardization of data, reduce errors and redundancy, and support efficient data management and analysis.
Smart Images

Figure CN120371926A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of report data processing, and in particular relates to a report data asset identification method, system, device and medium. Background Art
[0002] The current digital transformation of the banking industry is accelerating, and the data generated in business scenarios presents multimodal and highly complex characteristics, including: structured reports, unstructured texts such as contracts, bills, scanned copies, signatures, and cross-system interaction logs. Traditional data recognition technology relies on manual entry or a single OCR tool, and has the following problems: Poor compatibility of multi-source heterogeneous data: Bank data sources are scattered, such as counter systems, mobile terminals, and third-party interfaces. The formats vary greatly, such as PDF, images, and Excel. The existing OCR technology has insufficient recognition accuracy for complex tables, blurred images, and handwriting, resulting in missed detection or misjudgment of amounts and account numbers.
[0003] Simple character recognition cannot extract the correlation between financial statements or the causal chain of customer behavior data, and it is difficult to support risk control modeling or asset value mining. Summary of the invention
[0004] The present invention provides a report data asset identification method, which manages the entire process from raw data acquisition to identification, conversion, verification, storage and evaluation. The accuracy and depth of data identification are improved through the identification mechanism of multi-technology integration; the standardized conversion and automatic mapping mechanism improves data processing efficiency.
[0005] Methods include: S101: Retrieve the original report and perform a preprocessing process on the data in the original report; S102: Identify the character information of the data in the original report, analyze the semantic content of the character information, and build a correlation network between the data in combination with the knowledge graph to achieve multi-modal data recognition; S103: Perform field-level extraction on the identified report data, and convert the identification results into standardized data assets through an automatic mapping mechanism to complete structured conversion and mapping; S104: Verify the converted report data and store the verified data assets in the database; S105: The verified data assets are classified and stored in the data asset directory according to the preset business attributes, technical attributes, management attributes, report information and application attributes; S106: Establish a data asset identification effect evaluation system based on the data asset identification results, calculate the score of the data asset attribute through the technical complexity and identification effect algorithm, compare the score with the preset score threshold, and determine whether to trigger the data asset early warning mechanism based on the comparison result.
[0006] Further, it should be noted that the methods for identifying data in the original report in S102 specifically include: S201: Compare the data in the original report with the asset catalog item by item; If the report data is consistent with the asset catalog data, proceed to step S202; If they are inconsistent or do not exist, proceed to step S203; S202: Mark the current data as unrecognized and terminate the process; S203: If the code or name of the report data exists in the asset catalog, proceed to step S204; If both the code and name do not exist, proceed to step S205; S204: Update the new code or name of the report data to be consistent with the asset catalog and return to step S201 for re-comparison; S205: Determine whether the report needs to be updated. The condition for judgment is: If the report needs to be updated, proceed to step S206; If no update is required, proceed to step S207; S206: Update the content of the report data and return to step S201 for re-comparison; S207: Start the identification operation on data assets; S208: Identify data assets based on the identification mechanism; S209: Verify, compare, and merge duplicate or similar data assets; S210: Extract and define the business attributes, technical attributes, management attributes, report information, and application attributes of data assets; S211: Add the identified data assets to the resource catalog and establish an index; S212: Supplement and improve the information of data assets, including: unique resource code, ownership, and resource classification label.
[0007] Further, it should be noted that the identification mechanism in step S208 specifically includes: S701: Extract the original data, including fields such as serial number, data date, customer number, affiliated institution, affiliated first-level branch, account manager number, and demand deposit; S702: Define the business dimension, management dimension, and technical dimension of the report; S703: Format the data according to the preset standards, unify the field naming rules, data types, and units, and eliminate abnormal values or missing values that do not meet the specifications; S704: Apply OCR technology to recognize character information in images, extract key data fields through the PRA algorithm, parse the text semantics using NLP, generate a structured description in combination with a large language model, and establish the association relationship between data through a knowledge graph; S705: Execute data verification policies, including comparison of the repeatability of new resources with existing resources, verification of version consistency, and merging of duplicate records, to ensure data uniqueness and timeliness; S706: Extract and classify resource attributes; S707: Establish a mapping relationship between the extracted resource attributes and the header type numbers.
[0008] Furthermore, it should be noted that step S103 specifically includes: Retrieve the data from the original report and perform field-level extraction on the data in the original report; Select an appropriate mapping template according to the report type and data characteristics; Use the mapping template and custom rules to perform mapping conversion on the extracted report data; After the conversion is completed, perform standardization verification and format correction on the mapping results to ensure the standardization and usability of data assets.
[0009] Furthermore, it should be noted that step S104 also includes: Conduct a data quality assessment on the converted report data and generate a data quality report; According to the data quality report, perform feedback processing on the report data that does not meet the quality requirements; Perform preprocessing on the report data that passes the verification, including operations such as data desensitization and data encryption; Store the preprocessed report data in the database according to the preset storage strategy, and at the same time establish a data index.
[0010] Furthermore, it should be noted that step S105 specifically includes: Establish a dynamic attribute association mapping library, and according to the changes in the business scenarios of data assets, update the association relationship between data assets and business attributes, technical attributes, management attributes, report information, and application attributes in real time; Adopt an attribute priority determination strategy, and based on factors such as the usage frequency and importance degree of data assets, determine the priority order of different attributes during classified storage; Define an intelligent classification storage scheduler, and according to the attribute priority and association relationship, efficiently store the data assets that pass the verification to the corresponding positions in the data asset catalog and generate a storage index; Construct a data asset catalog version management mechanism to record the versions of changes in the classified storage structure, attribute association relationship, etc. of data assets, and support historical version backtracking and difference comparison.
[0011] Further, it should be noted that in the technical complexity and recognition effect algorithm of step S106, the data asset attribute is defined as A, the weight of the attribute is represented by W, the recognition effect score is represented by S, and the technical parameter of a certain attribute Ai is set as W ij The scores of data assets identified by different technologies are S ij ; The total score S of Ai i The calculation method is: ; The weighted average score of all attributes in the report form is: ; is the attribute weight; The calculation method of the recognition result scores of all report form resources is: ; is the recognition result score of a certain report form, is the recognition result weight.
[0012] The present application also provides a report form data asset recognition system, which includes: A data preprocessing module, which is used to retrieve the original report form and execute a preprocessing process on the data in the original report form; A multi-modal data recognition module, which is used to recognize the character information of the data in the original report form, analyze the semantic content of the character information, and construct an association network between the data in combination with the knowledge graph to achieve multi-modal data recognition; A structured conversion and mapping module, which is used to perform field-level extraction on the recognized report form data, and convert the recognition result into a standardized data asset through an automatic mapping mechanism to complete structured conversion and mapping; A data verification and storage module, which is used to verify the converted report form data and store the data assets that pass the verification in the database; A data asset classification and storage module, which is used to classify and store the data assets that pass the verification into a data asset directory according to preset business attributes, technical attributes, management attributes, report form information, and application attributes; A recognition effect evaluation and early warning module, which establishes a data asset recognition effect evaluation system based on the data asset recognition result, calculates the scores of the data asset attributes through the technical complexity and recognition effect algorithm, compares the scores with a preset score threshold, and determines whether to trigger a data asset early warning mechanism based on the comparison result.
[0013] According to another embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the report form data asset recognition method are implemented.
[0014] According to another embodiment of the present application, a storage medium is further provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the report data asset identification method are implemented.
[0015] As can be seen from the above technical solutions, the present invention has the following advantages: The report data asset identification method provided by the present invention can effectively remove noise and redundant information in the data by retrieving the original report and executing the preprocessing process, improving the data processing efficiency and accuracy.
[0016] The identification mechanism realizes the accurate identification and structured conversion of complex report data through multi-technology integration and strict processing procedures. The complete process from raw data extraction to the establishment of the final mapping relationship ensures the accuracy, integrity, and standardization of the data. The data verification strategy guarantees the uniqueness and timeliness of the data, reducing data redundancy and errors.
[0017] The field-level extraction and automatic mapping mechanism convert the identified report data into standardized data assets. Verifying and storing the converted data ensures the accuracy and integrity of the data. The verification process can detect and correct errors in the data, preventing incorrect data from entering the database and affecting subsequent applications. Classifying and storing the data assets in the data asset directory according to preset attributes makes the management of data assets more orderly. Business personnel or data analysts can quickly locate and find the required data based on the attributes, improving the data retrieval efficiency. Establishing a data asset identification effect evaluation system and an early warning mechanism can monitor the data asset identification effect in real time. By calculating the attribute scores and comparing them with the thresholds, problems existing in the data identification process can be detected in a timely manner, thereby continuously improving the quality of data assets. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0019] Figure 1 is a system architecture diagram; Figure 2 is a flowchart of the report data asset identification method; Figure 3 is a flowchart of the data format identification in the original report; Figure 4 is a flowchart of the identification mechanism; Figure 5 is a schematic diagram of an electronic device. Specific Embodiments
[0020] The method for identifying report data assets involved in the application aims at the technical bottlenecks in the identification of multi-modal data in the banking industry, such as Figure 1 The system architecture diagram for the implementation method is given.
[0021] Specifically, the system has a data collection layer. It supports multiple data formats, including bank bills, financial reports, spreadsheets, image files, and structured database interfaces. The original data is cleaned, denoised, corrected for bias, and redundant data is removed. Image enhancement and text tokenization are selected according to the data type.
[0022] Intelligent character recognition is performed on non-standard formats, and the recognition accuracy is improved by combining an adaptive layout analysis algorithm. Based on NLP parsing, text semantic information is extracted, and business logic is parsed, such as the consistency check of transaction flows. Based on the dynamic knowledge graph technology, an association network between data items is established, such as the mapping relationship between customer information and transaction records.
[0023] Semi-structured / unstructured data is converted into standardized fields and automatically mapped to the bank's internal data asset catalog. Based on predefined business rules, such as statistical caliber and data format, data integrity verification is performed to ensure data quality. Based on the five-dimensional attribute set, a standardized classification system is constructed for business, technology, management, report information, and application attributes to support fast retrieval and invocation. The recognition effect is dynamically monitored, and early warnings are issued for abnormal data and optimization suggestions are pushed. Through the robotic process automation (RPA) technology, each technical module is connected in series to achieve an end-to-end automated recognition process, improving efficiency and reducing error rates. Through OCR enhancement, adaptive layout analysis, and deep learning models, the limitations of traditional OCR in complex formats are broken through, and field-level accurate extraction is achieved. By combining NLP and dynamic knowledge graph technology, a multi-dimensional association network of data, business, and risks is constructed to avoid information silos.
[0024] The method for identifying report data assets involved in the present application will be described in detail below. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details.
[0025] It should be understood that when used in the specification of the present application, the term "including" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "including", "comprising", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0026] It should be understood that "one or more" mentioned in this application refers to one, two or more than two, and "a plurality of" mentioned in this application refers to two or more than two. In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. "And / or" herein is merely a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0027] The statements such as "an embodiment" or "some embodiments" described in this application mean that the specific features, structures or characteristics described in the embodiment are included in one or more embodiments of this application. Thus, the statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different places in this application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0028] In embodiments of the present invention, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or a power server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (exemplarily, connected through the Internet using an Internet service provider).
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Please refer to Figure 2 The following is a flowchart of a method for identifying report data assets in a specific embodiment. The method includes: Step S101: Retrieve the original report and perform a preprocessing process on the data in the original report.
[0031] In some embodiments, when retrieving the original report, the system connects to the report storage location through the data interface, obtains data, and performs targeted preprocessing processes based on text, images, values, and formats. For text data, operations such as deduplication, removal of invalid characters, and unified encoding formats are performed; for image data, image enhancement processing such as grayscale, noise reduction, and tilt correction is performed; if it is numerical data, the data value range is checked, and abnormal values and duplicate records are removed, thereby effectively removing noise and redundant information. Step S102: Identify the character information of the data in the original report, parse the semantic content of the character information, build an association network between the data in combination with the knowledge graph, and realize multimodal data recognition.
[0032] In some embodiments, OCR technology is used to perform character recognition on the image data in the original report. Here, the image is analyzed for layout, and text areas, table areas, etc. are identified. Then, through character segmentation and feature extraction, the characters in the image are converted into computer-editable text information. For text data, NLP natural language processing technology is applied for semantic analysis, and the meaning expressed by the text is understood through operations such as word segmentation, part-of-speech tagging, and named entity recognition. Combined with knowledge graph technology, the identified data and the parsed semantic information are used to build an association network between data based on the relationship between entities. For example, after identifying information such as customers and transaction records in the report, the association relationship between customers and transaction records is clarified through the knowledge graph, so as to achieve comprehensive recognition and understanding of the multimodal information of the report data. Step S103: Perform field-level extraction on the identified report data, and convert the identification results into standardized data assets through an automatic mapping mechanism to complete structured conversion and mapping.
[0033] In some embodiments, the recognition results output by the multimodal data recognition module are accurately extracted at the field level. According to the business rules and data structure of the report, the extraction rules and scope of each field are determined, and the required field information is accurately extracted from the recognized data. Then, through the automatic mapping mechanism, the extracted field information is matched and converted with the pre-defined standardized data model, and the recognition results are converted into standardized data assets.
[0034] For example, the customer name, transaction amount and other fields in the report are converted according to the format and requirements of the standardized data model to ensure that the format and meaning of the data are consistent in different systems and applications, and complete the conversion from raw unstructured or semi-structured data to structured data assets. The extracted data is mapped with the standard format to achieve data structuring and standardization. Step S103 of the present application also involves the following steps: Step S301: Retrieve the data from the original report and perform field-level extraction on the data in the original report.
[0035] In this embodiment, the original report data is obtained from the data source, which can be a database, a file system, etc. Each field in the report is identified and extracted, such as customer number, affiliated branch, account manager number, etc.
[0036] Step S302: Select an appropriate mapping template according to the report type and data characteristics.
[0037] In this embodiment, the structure and content of the report are analyzed to determine its type. The format, content, and business significance of the data are evaluated. The most suitable mapping template for the current report is selected from the predefined template library.
[0038] Step S303: Use the mapping template and custom rules to perform mapping conversion on the extracted report data.
[0039] In this embodiment, the extracted data is converted according to the definition of the mapping template. Specific conversion rules are applied according to business requirements, such as data format conversion, data calculation, etc. The standardization and normalization of the data are realized, and the usability of the data is improved.
[0040] Step S304: After the conversion is completed, perform standardization verification and format correction on the mapping result to ensure the standardization and usability of the data assets.
[0041] Check whether the converted data meets the predefined standards and specifications. Adjust and correct the format of the data that does not meet the standards. Ensure the quality and consistency of the data assets and improve the credibility of the data.
[0042] Step S104: Verify the converted report data and store the data assets that pass the verification in the database.
[0043] In some embodiments, the standardized data assets output by the structured conversion and mapping module are verified. The verification process includes field-level verification to check whether the data format and value range of each field meet the requirements; record-level verification to verify the integrity and logical consistency of each data record; and cross-table-level verification to ensure the correct data association relationship between different data tables.
[0044] In this embodiment, through various verification methods, the accuracy and integrity of the data are comprehensively checked. For the data assets that pass the verification, they are stored in the database, and the database performs reasonable table structure design and index optimization according to the type and use of the data to improve the storage efficiency and query performance of the data. As an implementation manner of step S104 of this application, the following steps are also involved: Step S401: Perform data quality assessment on the converted report data and generate a data quality report.
[0045] Step S402: According to the data quality report, perform feedback processing on the report data that does not meet the quality requirements.
[0046] Step S403: Preprocess the report data that passes the verification, including operations such as data desensitization and data encryption.
[0047] Step S404: Store the preprocessed report data in the database according to the preset storage strategy, and establish a data index at the same time.
[0048] Specifically, when performing data quality assessment on the converted report data, an assessment system will be constructed from multiple dimensions. First, perform integrity assessment to check whether there are missing data records. For example, in the customer information field of the report, if a customer lacks key contact information or identity information, the record is considered incomplete; check whether there are null values in the field values. For numeric fields, if meaningless default values appear, they will also be marked. Then is the accuracy assessment, by comparing with authoritative data sources to verify the correctness of the data; This embodiment will also perform consistency assessment to check whether the expressions of the same data in different reports or between tables are consistent. For example, the writing of the customer name in different related reports must be unified. After the assessment is completed, summarize the assessment results of each item to generate a data quality report containing detailed information such as the type of data quality problems, specific locations, and severity levels. In this way, the usability and credibility of the data are improved.
[0049] This embodiment formulates corresponding feedback processing strategies based on the problems presented in the data quality report. For the problem of missing data, if the missing data can be supplemented through other data sources, the system will automatically initiate a data supplementation request to obtain the missing data from the connected relevant databases or files and fill it; if it cannot be automatically supplemented, the problem will be feedback to the data entry personnel or relevant business departments, along with detailed problem descriptions and data requirements for manual data supplementation.
[0050] When this embodiment preprocesses the report data that passes the verification, the data desensitization operation will adopt different desensitization methods according to the sensitivity of the data and the usage scenario. For highly sensitive information such as customer names and ID numbers, the replacement desensitization method is adopted to replace the real information with fictional but consistent formats. The data encryption operation uses mature encryption algorithms, such as symmetric encryption algorithms, to encrypt the data. According to the importance and usage frequency of the data, select an appropriate encryption strength to ensure the security of the data during storage and transmission.
[0051] The preset storage strategy in this embodiment is formulated according to factors such as the type of data, usage frequency, and importance. For the report data that is frequently queried and used, it is stored on high-performance storage devices and an appropriate data storage format is adopted. When storing the preprocessed report data into the database, according to the pre-designed database table structure, the data is accurately inserted into the corresponding tables and fields. Appropriate data indexes are established, and the indexes are regularly maintained and optimized to ensure the effectiveness of the indexes.
[0052] Step S105: Classify and store the data assets that pass the verification into the data asset catalog according to the preset business attributes, technical attributes, management attributes, report information, and application attributes.
[0053] In some embodiments, obtain the data assets that pass the verification from the data verification and storage module, and classify and store the data assets into the data asset catalog according to the preset business attributes, technical attributes, management attributes, report information, and application attributes. During the classification process, a clear directory hierarchy structure is established to facilitate the rapid retrieval and management of data assets. For example, classify the data assets related to customers under the customer business attributes, and further subdivide and store them according to technical attributes such as data type, and at the same time record the management and application related information of the data assets to form a complete data asset catalog system. Specifically, each attribute of the data assets identified based on the report and its sub-categories are an indispensable part of the data asset catalog. According to the attributes, classification, statistical caliber, and data characteristics of the report, referring to the structured definition of data assets in the data asset management platform, define that the report data asset attributes include five major categories of 30 detailed attribute contents: business attributes, technical attributes, management attributes, report information, and application attributes. Based on different attributes, select appropriate data asset identification methods and technologies.
[0054] Exemplarily, use OCR technology to identify technical attributes such as the table header, row and column structure, and fields of the report; apply the large model Q&A interface to identify business attributes such as the business theme and business purpose of the report. The data asset attribute items are shown in Table 1 below.
[0055] Table 1: Information Table of Data Resource Attribute Items
[0056] Step S106: Establish an evaluation system for the data asset identification effect based on the data asset identification results, calculate the scores of the data asset attributes through the technical complexity and identification effect algorithms, compare the scores with the preset score threshold, and judge whether to trigger the data asset warning mechanism based on the comparison results.
[0057] In some embodiments, an evaluation system for the identification effect of data assets is established based on the data asset identification results. Evaluation indicators are determined, including multiple dimensions such as the accuracy, integrity, timeliness, and identification efficiency of data. Through the technical complexity and identification effect algorithm, the difficulty level of the technology adopted in the data identification process and the actual identification effect are comprehensively considered to calculate the scores of data asset attributes.
[0058] In this embodiment, the calculated scores are compared with the preset score thresholds, and according to the comparison results, it is judged whether to trigger the data asset warning mechanism.
[0059] Exemplarily, if the accuracy score of the data is lower than the preset threshold, a warning is triggered to notify relevant personnel to conduct data checks and repairs to improve the quality of data assets. It can be seen that the evaluation system for the identification effect of data assets is based on the principles of quality assessment and threshold judgment. By setting reasonable evaluation indicators and weights, the data asset identification results are quantitatively evaluated; the warning mechanism is based on the principles of threshold comparison and event triggering. When the evaluation results do not meet the requirements, corresponding warning notifications are automatically triggered. In some specific embodiments, establishing an evaluation system for the identification effect of data assets and a warning mechanism can improve the quality and value of data assets.
[0060] This embodiment can evaluate the effect of a single attribute of a report, then evaluate the effect of a report, and finally evaluate the identification effect of all report data assets.
[0061] Combining factors such as the technical complexity and identification effect of identifying a certain attribute of data assets in the existing technology, technical parameters of book resources attributes are assigned. The data asset attribute is represented by A, the weight of the attribute is represented by W, and the identification effect score is represented by S.
[0062] Define the technical parameter of a certain attribute Ai as W ij , and the score of the data asset identified by this attribute using different technologies is S ij , as shown in Table 2 specifically. Table 2
[0063] Suppose the identification result score of a certain attribute Ai of a report is , and the weight is , then the weighted average score of all attributes of this report is .
[0064] Suppose the identification result score of a certain report is , and the weight is Qi , then the identification result score of all report resources is: .
[0065] It can be seen that the evaluation is carried out at three levels: from individual attributes of reports, single reports to all report data assets. At the level of individual attribute evaluation, for each attribute of the data asset, combined with the corresponding identification technology used to identify the attribute currently, a corresponding weight W is assigned to each attribute ij .
[0066] For example, for the attribute of report header recognition, if the OCR technology is used for recognition, the weight W is assigned according to its technical implementation difficulty, historical recognition accuracy, etc i1 . At the same time, record the scores (S ij ) of data assets with the same attribute recognized by different technologies, such as OCR, RPA, etc. For example, the score of the report header recognition attribute recognized by the OCR technology is S i1 . The total score S i of a single attribute is obtained through weighted calculation, that is, the sum of the products of the scores S ij of each technology and the corresponding weight W ij .
[0067] At the level of single report evaluation, the weighted average score of all attributes in the report is used as the recognition result score of the report. The calculation method is to sum the products of the total scores Si of each attribute and the weight Qi of the attribute in the report. At the level of all report data asset evaluation, the recognition result scores of each single report are summarized to comprehensively evaluate the overall data asset recognition effect. Finally, the scores calculated at each level are compared with the preset score threshold to determine whether to trigger the data asset warning mechanism. The preset score threshold here is used as the judgment criterion. When the calculated score is lower than the threshold, it indicates that the recognition effect does not meet the expectation, triggering the warning mechanism to remind relevant personnel to take measures. This makes the evaluation result more objective and operable. After triggering the warning mechanism, problems existing in the data assets can be discovered in time, avoiding the impact of data quality problems on subsequent data analysis and decision-making, and effectively improving the quality and value of data assets
[0068] For this embodiment, the recognition result score is positively correlated with the recognition effect. The higher the score, the better the data asset recognition effect. When the recognition result score reaches 90 points and above, the recognition effect is determined to be excellent and no additional processing is required; when the score is in the range of 80 points (inclusive) to 90 points, it indicates that there is room for improvement in the recognition process and the recognition process needs to be analyzed and optimized; if the score is lower than 60 points, it is determined that the recognition effect is poor, which means that there are serious deficiencies or errors in the data assets, and the existing technical solution should be improved in time, and the relevant data should be re-recognized and processed
[0069] In an embodiment of the present invention, based on the method of identifying data in the original report in step S102, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation scheme
[0070] Such asFigure 3 As shown in the figure, step S102 specifically includes: Step S201: Compare the data in the original report item by item with the asset catalog; If the report data is consistent with the asset catalog data, proceed to step S202; If they are inconsistent or do not exist, proceed to step S203; S202: Mark the current data as unrecognized and terminate the process; S203: If the code or name of the report data exists in the asset catalog, proceed to step S204; If both the code and name do not exist, proceed to step S205; S204: Update the new code or name of the report data to be consistent with the asset catalog and return to step S201 for re-comparison; S205: Determine whether the report needs to be updated. The condition judgment is as follows: If the report needs to be updated, proceed to step S206; If no update is required, proceed to step S207; S206: Update the content of the report data and return to step S201 for re-comparison; Step S207: Start the identification operation on the data assets; Step S208: Based on the identification mechanism, identify the data assets; Step S209: Verify, compare, and merge duplicate or similar data assets; Step S210: Extract and define the business attributes, technical attributes, management attributes, report information, and application attributes of the data assets; Step S211: Add the identified data assets to the resource catalog and establish an index; Step S212: Supplement and improve the information of the data assets, including: unique resource code, ownership, and resource classification label.
[0071] It can be seen that through the data comparison and processing process, it is ensured that the data assets entering the identification link are accurate and standardized. The processing of duplicate or similar data assets avoids data redundancy and improves the quality and storage efficiency of data assets. Adding the identified data assets to the resource catalog and establishing an index constructs an ordered data asset system and improves the accessibility and usability of data assets.
[0072] Such as Figure 4 As shown in the figure, the identification mechanism of step S208 in this embodiment specifically includes the following steps: Step S701: Extract the original data, including fields such as serial number, data date, customer number, affiliated institution, affiliated first-level branch, account manager number, and demand deposit; Step S702: Define the business dimension, management dimension, and technical dimension of the report; Step S703: Format the data according to preset standards, unify the field naming rules, data types, and units, and eliminate abnormal values or missing values that do not meet the specifications; Step S704: Apply OCR technology to recognize character information in images, extract key data fields through the PRA algorithm, use NLP to parse text semantics, generate structured descriptions in combination with large language models, and establish association relationships between data through knowledge graphs; Step S705: Execute data verification policies, including comparison of the repeatability of new resources and existing resources, verification of version consistency, and merging of duplicate records, to ensure data uniqueness and timeliness; Step S706: Extract and classify resource attributes; Step S707: Establish a mapping relationship between the extracted resource attributes and the header type numbers.
[0073] It should be noted that the large language model can adopt AIGC (Artificial Intelligence Generated Content), which is the process of automatically generating content using artificial intelligence technology. AIGC is based on large language models, generative adversarial networks, etc., and automatically generates text content that meets specific logical and format requirements by learning the patterns and rules of a large number of data samples.
[0074] In the scenario of report data processing, the AIGC system can understand the character information recognized by OCR, the keyword fields extracted by PRA, and the semantic content parsed by NLP, and integrate these fragmented information into a coherent and structured description.
[0075] The AIGC of this embodiment forms a closed loop with OCR, PRA, NLP, and knowledge graph technologies. OCR is used to provide original character information, PRA is used to locate key data fields. NLP parses text semantics. The knowledge graph provides data association rules, and AIGC is responsible for integrating the above information into a structured and understandable description.
[0076] It can be seen that the recognition mechanism realizes the accurate recognition and structured conversion of complex report data through multi-technology integration and strict processing procedures. The complete process from raw data extraction to the establishment of the final mapping relationship ensures the accuracy, integrity, and standardization of the data. The data verification policy guarantees data uniqueness and timeliness, reducing data redundancy and errors. The extraction and classification of resource attributes and the mapping relationship with the header type support the decision-making and operation of the enterprise.
[0077] Based on the above embodiments, in order to further improve the reliability of the report data asset identification method provided in the above embodiments, the following is an implementable manner of step S105, and step S105 specifically includes: Step S501: Establish a dynamic attribute association mapping library, and according to the changes in the business scenarios of data assets, update the association relationships between data assets and business attributes, technical attributes, management attributes, report information, and application attributes in real time; Step S502: Adopt an attribute priority determination strategy, and determine the priority order of different attributes during classified storage based on factors such as the usage frequency and importance of data assets; Step S503: Define an intelligent classified storage scheduler, and according to the attribute priority and association relationships, store the data assets that pass the verification efficiently to the corresponding positions in the data asset directory and generate storage indexes; Step S504: Construct a data asset directory version management mechanism to record the versions of changes such as the classified storage structure and attribute association relationships of data assets, and support historical version backtracking and difference comparison.
[0078] It can be seen that the established dynamic attribute association mapping library, like a real-time updated relationship graph, can automatically adjust the association between data assets and various attributes according to the changes in business scenarios, ensuring that the classification of data assets always conforms to business requirements. By evaluating factors such as the usage frequency and importance of data assets, priorities are assigned to different attributes, making the allocation of storage resources more reasonable. The intelligent classified storage scheduler accurately and efficiently stores data assets to the corresponding positions in the directory and generates indexes based on attribute priorities and association relationships, improving data retrieval efficiency. The data asset directory version management mechanism completely records the structure and relationship changes, supports historical version backtracking and difference comparison, and ensures that the evolution process of data assets is traceable.
[0079] It should be understood that the magnitudes of the sequence numbers of the above steps in the embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0080] The following are embodiments of the report data asset identification system provided by the embodiments of the present disclosure. This system and the report data asset identification method of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the report data asset identification system, reference can be made to the embodiments of the above report data asset identification method.
[0081] The system includes: a data preprocessing module, which is used to retrieve the original report and perform a preprocessing process on the data in the original report.
[0082] The multimodal data recognition module is used to recognize the character information of the data in the original report, parse the semantic content of the character information, construct an association network between the data in combination with the knowledge graph, and realize multimodal data recognition.
[0083] The structured conversion and mapping module is used to perform field-level extraction on the recognized report data, and convert the recognition result into a standardized data asset through an automatic mapping mechanism to complete structured conversion and mapping.
[0084] The data verification and storage module is used to verify the converted report data and store the data assets that pass the verification in the database.
[0085] The data asset classification and storage module is used to classify and store the data assets that pass the verification in the data asset directory according to preset business attributes, technical attributes, management attributes, report information, and application attributes.
[0086] The recognition effect evaluation and early warning module is used to establish a data asset recognition effect evaluation system based on the data asset recognition result, calculate the scores of the data asset attributes through the technical complexity and recognition effect algorithms, compare the scores with the preset score threshold, and judge whether to trigger the data asset early warning mechanism based on the comparison result.
[0087] As Figure 5 shown, the present application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored on the memory and executable on the processor 101. When the processor 101 executes the program, the steps of the report data asset recognition method are implemented.
[0088] In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0089] In the embodiments of the present application, the processor 101 may be implemented by using at least one of an application specific integrated circuit, a programmable logic device, a field programmable gate array, a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to execute the functions described herein. In some cases, such an implementation may be implemented in a controller. For a software implementation, an implementation of a process or function may be implemented with a separate software module that allows execution of at least one function or operation. The software code may be implemented by a software application (or program) written in any suitable programming language. The software code may be stored in a memory and executed by the controller.
[0090] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, and the display panel may be configured in the form of a liquid crystal display, an organic light emitting diode, etc.
[0091] The memory 102 may be used to store software programs and various data. The memory 102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid state storage devices.
[0092] The present application also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the report data asset identification method are implemented.
[0093] The storage medium may be any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0094] In the storage medium, the readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, and the readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0095] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying report data assets, characterized in that, The method includes: S101: Retrieve the original report and perform a preprocessing process on the data in the original report; S102: Identify the character information of the data in the original report, parse the semantic content of the character information, and construct an association network between the data in combination with the knowledge graph to achieve multi-modal data recognition; S103: Perform field-level extraction on the recognized report data, and convert the recognition results into standardized data assets through an automatic mapping mechanism to complete structured conversion and mapping; S104: Verify the converted report data, and store the data assets that pass the verification in the database; S105: Classify and store the data assets that pass the verification in the data asset catalog according to preset business attributes, technical attributes, management attributes, report information, and application attributes; S106: Establish a data asset recognition effect evaluation system based on the data asset recognition results, calculate the scores of the data asset attributes through the technical complexity and recognition effect algorithms, compare the scores with the preset score thresholds, and judge whether to trigger the data asset warning mechanism based on the comparison results.
2. The method for recognizing report data assets according to claim 1, wherein The method for identifying the data in the original report in S102 specifically includes: S201: Compare the data in the original report with the asset catalog item by item; If the report data is the same as the asset catalog data, go to step S202; If they are inconsistent or do not exist, go to step S203; S202: Mark the current data as not recognized, and the process terminates; S203: If the code or name of the report data exists in the asset catalog, go to step S204; If both the code and name do not exist, go to step S205; S204: Update the new code or name of the report data to be the same as the asset catalog, and return to step S201 for re-comparison; S205: Judge whether the report needs to be updated, and the condition judgment is: If the report needs to be updated, go to step S206; If no update is required, go to step S207; S206: Update the content of the report data, and return to step S201 for re-comparison; S207: Start the recognition operation on the data assets; S208: Based on the recognition mechanism, recognize the data assets; S209: Verify, compare, and merge duplicate or similar data assets; S210: Extract and define the business attributes, technical attributes, management attributes, report information, and application attributes of the data assets; S211: Add the recognized data assets to the resource catalog and establish an index; S212: Supplement and improve the information of the data assets, including: unique resource code, ownership attribution, resource classification label.
3. The method for recognizing report data assets according to claim 2, wherein The recognition mechanism in step S208 specifically includes: S701: Extract the original data, including fields such as serial number, data date, customer number, affiliated institution, affiliated first-level branch, account manager number, and demand deposit; S702: Define the business dimension, management dimension, and technical dimension of the report; S703: Format the data according to the preset standards, unify the field naming rules, data types, and units, and eliminate outliers or missing values that do not meet the specifications; S704: Apply OCR technology to identify character information in images, extract key data fields through the PRA algorithm, parse the text semantics using NLP, generate structured descriptions in combination with large language models, and establish association relationships between data through knowledge graphs; S705: Execute data verification strategies, including comparison of the repeatability of new resources with existing resources, verification of version consistency, and merging of duplicate records, to ensure the uniqueness and timeliness of data; S706: Extract and classify resource attributes; S707: Establish a mapping relationship between the extracted resource attributes and the header type numbers.
4. The method for identifying report data assets according to claim 1, characterized in that, Step S103 specifically includes: Retrieve the data in the original report and perform field-level extraction on the data in the original report; Select an appropriate mapping template according to the report type and data characteristics; Use the mapping template and custom rules to perform mapping conversion on the extracted report data; After the conversion is completed, perform standardization verification and format correction on the mapping results to ensure the standardization and usability of the data assets.
5. The method for identifying report data assets according to claim 1, characterized in that, Step S104 further includes: Conduct a data quality assessment on the converted report data to generate a data quality report; According to the data quality report, perform feedback processing on the report data that does not meet the quality requirements; Perform preprocessing on the report data that passes the verification, including data desensitization and data encryption operations; Store the preprocessed report data in the database according to the preset storage strategy, and establish a data index at the same time.
6. The method for identifying report data assets according to claim 1, characterized in that, Step S105 specifically includes: Establish a dynamic attribute association mapping library, and update the association relationships between data assets and business attributes, technical attributes, management attributes, report information, and application attributes in real time according to changes in the business scenarios of the data assets; Adopt an attribute priority determination strategy, and determine the priority order of different attributes during classification storage based on the usage frequency and importance of the data assets; Define an intelligent classification storage scheduler, and efficiently store the data assets that pass the verification in the corresponding positions of the data asset directory according to the attribute priority and association relationships, and generate a storage index; Construct a version management mechanism for the data asset directory to record the classification storage structure and attribute association relationships of the data assets, and support historical version backtracking and difference comparison.
7. The method for identifying report data assets according to claim 1, characterized in that, In the technical complexity and recognition effect algorithm of step S106, the data asset attribute is defined as A, the weight of the attribute is defined as W, the recognition effect score is defined as S, and the technical parameter of a certain attribute Ai is set as W ij The score of the data asset identified by different technologies is S ij ; The total score S of Ai i The calculation method is as follows: ; The weighted average score of all attributes of the report is: ; is the attribute weight; The calculation method of the recognition result score for all report resources is as follows: ; is the recognition result score of a certain report form, is the recognition result weight.
8. A report data asset identification system, characterized in that, The system is used to implement the method for identifying report data assets as described in any one of claims 1 to 7; The system includes: A data preprocessing module for retrieving the original report and performing a preprocessing process on the data in the original report; A multi-modal data recognition module for identifying the character information of the data in the original report, parsing the semantic content of the character information, and constructing an association network between data in combination with a knowledge graph to achieve multi-modal data recognition; A structured conversion and mapping module, which is used to perform field-level extraction on the identified report data, and convert the recognition results into standardized data assets through an automatic mapping mechanism to complete structured conversion and mapping; A data verification and storage module, which is used to verify the converted report data and store the verified data assets in a database; A data asset classification and storage module, which is used to classify and store the verified data assets in a data asset directory according to preset business attributes, technical attributes, management attributes, report information, and application attributes; An identification effect evaluation and early warning module, which establishes a data asset identification effect evaluation system based on the data asset identification results, calculates the scores of data asset attributes through a technical complexity and identification effect algorithm, compares the scores with a preset score threshold, and determines whether to trigger a data asset early warning mechanism based on the comparison results.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the report data asset identification method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the report data asset identification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Report generation method and device based on report template and computer equipment
CN111797605A
Data asset management system and method
CN118917565A
Character recognition and automatic auditing method and system
CN119091460A
Natural language rule table information extraction system based on large model
CN120031004A
Data processing method, device and system, and storage medium
WO2025039361A1
Cited By
Measurement and control equipment data mapping method and device, equipment and medium
CN121051141A
A bank data access semantic modeling method
CN122364232A