Intelligent data value mining method and system, storage medium and computer program product

By combining locally deployed intelligent agent parsing and dynamic lineage analyzers with asset prediction models, the problem of insufficient unstructured document processing capabilities and compliance requirements in existing technologies is solved, enabling efficient and accurate assessment of enterprise data asset value and compliance risks.

CN120873447APending Publication Date: 2025-10-31GUANGZHOU VISIONA DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510898576.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing modular data entry systems struggle to handle unstructured documents and lack compliance requirements and value prediction capabilities for enterprise data assets, resulting in inefficient and highly subjective data asset valuation.

Method used

The system employs a pre-built, locally deployed intelligent agent to parse enterprise internal data documents and extract key metadata sets; it calls a dynamic lineage analyzer to perform lineage analysis and generate a data value assessment matrix; and it loads an asset prediction model to perform hierarchical predictive analysis and generate an intelligent data value disclosure report.

Benefits of technology

It enables efficient parsing and vectorization of multi-source heterogeneous data documents, automatically tracks data sources and flows, assesses the value of data assets in the table and compliance risks from multiple dimensions, and generates intelligent data value disclosure reports that comply with regulations, thus solving the problem of ambiguity in the value of data assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873447A_ABST
    Figure CN120873447A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent data value mining method and system, a storage medium and a computer program product, and relates to the technical field of data asset management.The method comprises the steps that an enterprise internal source data document is obtained, a pre-constructed localized deployment agent is adopted to analyze the enterprise internal source data document, and a key metadata set is extracted; calling a pre-trained dynamic consanguinity analyzer to perform consanguinity analysis on the key metadata set to obtain a data value evaluation matrix; loading a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value evaluation matrix to obtain a table entry feasibility prediction result, the table entry feasibility prediction result including a table entry success rate and a recommended asset type; and generating an intelligent data value disclosure report based on the entry success rate and the recommended asset type. A localized deployment agent, a data asset prediction engine and a dynamic blood relationship analyzer are mutually combined, and the data asset value and compliance risk of an enterprise are efficiently and accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data asset management technology, and in particular to intelligent data value mining methods, systems, storage media and computer program products. Background Technology

[0002] With the acceleration of enterprise digital transformation, the importance of intelligent data assets is becoming increasingly prominent, and data asset entry into tables is an important means of data asset valuation. Traditional data asset valuation, which relies on manual review, is inefficient and highly subjective, making it difficult to automatically transform multi-source heterogeneous data into compliant assets. Although modular entry systems have been proposed, current modular entry systems for data asset valuation rely on rule engines, making it difficult to handle unstructured documents. Furthermore, the ARG technology used in modular entry systems is mainly geared towards general document retrieval and is not adapted to the compliance requirements and value prediction scenarios in the data asset field, lacking the ability to intelligently predict the feasibility of entry into tables.

[0003] Therefore, how to efficiently and accurately assess the value of a company's data assets and compliance risks has become an urgent problem to be solved in this application.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide an intelligent data value mining method, system, storage medium, and computer program product, which aims to solve the technical problem of how to efficiently and accurately assess the value of an enterprise's data assets and compliance risks.

[0006] To achieve the above objectives, this application proposes an intelligent data value mining method, the method comprising: Obtain enterprise internal data documents, and use a pre-built localized deployment intelligent agent to parse the enterprise internal data documents and extract key metadata sets; A pre-trained dynamic lineage analyzer is invoked to perform lineage analysis on the key metadata set to obtain a data value assessment matrix; A pre-trained asset prediction model is loaded to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table. The feasibility prediction results for inclusion in the table include the success rate of inclusion in the table and the recommended asset type. A smart data value disclosure report is generated based on the table entry success rate and the recommended asset types.

[0007] In one embodiment, before the step of obtaining enterprise internal data documents, parsing the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extracting a set of key metadata, the method further includes: Collect multi-source heterogeneous data from enterprises, perform data preprocessing and feature extraction on the multi-source heterogeneous data from enterprises, and construct a data asset fingerprint database; The sentence embedding model is trained using the data asset fingerprint database. The vectorized sentence embedding model is integrated with the retrieval enhancement framework to obtain a locally deployed intelligent agent; Based on the graph neural network model, the data asset fingerprint database is analyzed for lineage relationships to construct a full-link graph. The graph neural network model is trained using the full-link graph, and the trained graph neural network model is encapsulated to obtain a dynamic lineage analyzer. A structured training dataset is constructed based on the data asset fingerprint database. The parameters of the large language model are fine-tuned using the structured training dataset to obtain a domain-adapted large language model. The domain-adaptive large language model is trained hierarchically using the structured training dataset to obtain the asset prediction model.

[0008] In one embodiment, the step of acquiring enterprise internal data documents, parsing the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extracting a set of key metadata includes: Obtain internal data documents from the enterprise, perform data anonymization and recursive character segmentation on the internal data documents to obtain a set of short text fragments; The localized deployment agent is guided to vectorize the set of short text fragments using predefined domain-adaptive prompts, and the document vectors obtained after vectorization are searched and matched to obtain similar metadata patterns. Extract a set of key metadata based on the document vector and the similar metadata pattern.

[0009] In one embodiment, the step of invoking a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set to obtain a data value assessment matrix includes: A pre-trained dynamic lineage analyzer is invoked to identify entities and relationships in the key metadata set, and a list of entity relationships is constructed. Using entities in the entity relationship list as graph neural network nodes and relationships in the entity relationship list as graph neural network nodes, a full-link lineage graph of data flow is constructed based on the graph neural network nodes and graph neural network nodes. The value contribution coefficient of the full-link kinship map of data flow is calculated by the dynamic kinship analyzer to obtain a weighted full-link kinship map, and a data value assessment matrix is ​​generated based on the weighted full-link kinship map.

[0010] In one embodiment, the asset prediction model includes a preliminary screening layer, an actuarial layer, and a decision layer. The step of loading the pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table includes: The data value assessment matrix is ​​standardized to obtain a standardized feature vector; Load the pre-built asset prediction model, input the data value assessment matrix into the initial screening layer to retrieve similar cases of historical data assetization, and determine matching similar case information; The similar case information is input into the actuarial layer to calculate the cost-benefit ratio; Based on the decision-making level, and combined with the cost-benefit ratio and preset compliance verification rules, the success rate of table entry and asset type recommendation are predicted to obtain the table entry feasibility prediction result.

[0011] In one embodiment, the step of generating a smart data value disclosure report based on the table entry success rate and the recommended asset type includes: The built-in disclosure report generator is invoked to evaluate and analyze the success rate of data entry and the recommended asset type, and to generate data ownership certificates, evaluation method descriptions, and accounting treatment schemes. By integrating the data ownership verification, the assessment method description, and the accounting treatment scheme, an intelligent data value disclosure report is generated.

[0012] In one embodiment, the intelligent data value mining method further includes: Obtain the wireless signal characteristics of the terminal to be located, and input the wireless signal characteristics into the locally deployed intelligent agent to retrieve similar wireless signal characteristic data; The dynamic lineage analyzer is used to analyze the lineage relationship of the similar wireless signal feature data, and the coarse positioning area of ​​the terminal to be located is determined based on the lineage relationship. The wireless signal features are compared with the fingerprints of specific data assets in the coarse positioning area to calculate a similarity score; Based on the similarity score, the fingerprint of the selected data asset is input into the asset prediction model for prediction, thereby obtaining the positioning result of the terminal to be located. Intelligent data value mining is performed by combining the positioning results and the intelligent data value disclosure report.

[0013] Furthermore, to achieve the above objectives, this application also proposes an intelligent data value mining system, which includes: The localized intelligent engine module is used to acquire enterprise internal data documents, and uses a pre-built localized deployment intelligent agent to parse the enterprise internal data documents and extract key metadata sets; The dynamic lineage analysis module is used to call a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set and obtain a data value assessment matrix. The asset prediction module is used to load a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix and obtain the feasibility prediction results for inclusion in the table. The feasibility prediction results for inclusion in the table include the success rate of inclusion in the table and the recommended asset type. The disclosure report generation module is used to generate an intelligent data value disclosure report based on the table entry success rate and the recommended asset type.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the intelligent data value mining method described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the intelligent data value mining method described above.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: The process involves acquiring enterprise-internal data documents, parsing them using a pre-built, locally deployed intelligent agent, and extracting a set of key metadata. A pre-trained dynamic lineage analyzer is then invoked to perform lineage analysis on the key metadata set, yielding a data value assessment matrix. A pre-trained asset prediction model is loaded to perform hierarchical predictive analysis on the data value assessment matrix, resulting in a feasibility prediction result for data entry into the database, including the success rate and recommended asset types. Based on the success rate and recommended asset types, an intelligent data value disclosure report is generated. Firstly, the locally deployed intelligent agent enables efficient parsing and vectorization of multi-source heterogeneous data documents, including unstructured documents. It quickly retrieves and processes enterprise-internal data documents, extracting a set of key metadata. Furthermore, the dynamic lineage analyzer can rapidly generate a full-link graph of the data source processing chain, automatically tracking the source and flow of data. This paper utilizes a pre-trained dynamic lineage analyzer to perform lineage analysis on a key metadata set, quantifying the value contribution coefficient of the data and obtaining a data value assessment matrix. Further, a pre-trained asset prediction model is loaded to perform hierarchical predictive analysis on the data value assessment matrix, thereby assessing the data asset's value and compliance risks from multiple dimensions, obtaining a feasibility prediction result for inclusion in the table, and measuring the value of the data asset through the inclusion value. Finally, based on the feasibility prediction result, a compliant intelligent data value disclosure report is generated. This report serves as the core credential for data asset inclusion in the table and is also a key basis for internal value management, resolving the "value ambiguity" of data assets and achieving efficient and accurate assessment of enterprise data asset value and compliance risks. In summary, this application combines a locally deployed intelligent agent, a data asset prediction engine, and a dynamic lineage analyzer to achieve efficient and accurate assessment of enterprise data asset value and compliance risks. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating the first embodiment of the intelligent data value mining method of this application; Figure 2 A flowchart illustrating the second embodiment of the intelligent data value mining method of this application; Figure 3 A flowchart illustrating the third embodiment of the intelligent data value mining method of this application; Figure 4 A simplified flowchart illustrating the intelligent data value mining method provided in this application; Figure 5 This is a schematic diagram of the module structure of the intelligent data value mining device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent data value mining method in this application embodiment.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of this application embodiment is as follows: acquire enterprise internal data documents, parse the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extract key metadata sets; call a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata sets to obtain a data value assessment matrix; load a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction result for inclusion in the table; and generate an intelligent data value disclosure report based on the feasibility prediction result for inclusion in the table.

[0024] This application's embodiments take into account that: as enterprises accelerate their digital transformation, the importance of intelligent data assets is becoming increasingly prominent, and data asset entry into tables is a crucial means of data asset value assessment. Traditional data asset assessment, which relies on manual review, is inefficient and highly subjective, making it difficult to automatically transform multi-source heterogeneous data into compliant assets. Although modular entry systems have been proposed, current modular entry systems for data asset value assessment rely on rule engines, making it difficult to handle unstructured documents. Furthermore, the retrieval enhancement ARG technology used in modular entry systems is mainly geared towards general document retrieval, failing to adapt to the compliance requirements and value prediction scenarios in the data asset field, and lacking the ability to intelligently predict the feasibility of entry into tables.

[0025] Therefore, this application provides a solution to acquire enterprise internal data documents, parse the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extract a key metadata set; call a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set to obtain a data value assessment matrix; load a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain a table entry feasibility prediction result, which includes the table entry success rate and recommended asset type; and generate an intelligent data value disclosure report based on the table entry success rate and the recommended asset type. First, the locally deployed localized deployment intelligent agent can achieve efficient parsing and vectorization of multi-source heterogeneous data documents, including unstructured documents, and quickly retrieves and processes enterprise internal data documents to extract a key metadata set; furthermore, the dynamic lineage analyzer can quickly generate a full-link map of the data source processing chain, automatically tracking the source and flow of data. This paper utilizes a pre-trained dynamic lineage analyzer to perform lineage analysis on a key metadata set, quantifying the value contribution coefficient of the data and obtaining a data value assessment matrix. Further, a pre-trained asset prediction model is loaded to perform hierarchical predictive analysis on the data value assessment matrix, thereby assessing the data asset's value and compliance risks from multiple dimensions, obtaining a feasibility prediction result for inclusion in the table, and measuring the value of the data asset through the inclusion value. Finally, based on the feasibility prediction result, a compliant intelligent data value disclosure report is generated. This report serves as the core credential for data asset inclusion in the table and is also a key basis for internal value management, resolving the "value ambiguity" of data assets and achieving efficient and accurate assessment of enterprise data asset value and compliance risks. In summary, this application combines a locally deployed intelligent agent, a data asset prediction engine, and a dynamic lineage analyzer to achieve efficient and accurate assessment of enterprise data asset value and compliance risks.

[0026] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or intelligent data value mining system capable of performing the above functions. The following description uses an intelligent data value mining system as an example to illustrate this embodiment and the subsequent embodiments.

[0027] Based on this, the embodiments of this application provide an intelligent data value mining method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the intelligent data value mining method of this application.

[0028] In this embodiment, the intelligent data value mining method includes steps S10 to S40: Step S10: Obtain enterprise internal data documents, and use a pre-built localized deployment intelligent agent to parse the enterprise internal data documents and extract key metadata sets; It should be noted that the localized deployment agent involved in this application embodiment is a RAG retrieval enhancement agent based on the Retriever-Reader-Generator architecture. This agent utilizes a Retriever to retrieve relevant information from large-scale documents, a Reader to perform in-depth reading and analysis of the retrieved information, and a Generator to generate the final result, enabling efficient parsing of enterprise internal data documents. Enterprise internal data documents refer to documents generated and used internally by the enterprise, such as data dictionaries and API logs, that record enterprise business activities and data processes. The key metadata set contains a collection of important characteristics used to identify and manage data assets, such as field descriptions, update frequency, and privacy levels.

[0029] Understandably, due to data security protections within enterprise management, and to prevent the leakage of internal data, internal data documents can also be rough data generated and recorded within the enterprise to document business activities. Furthermore, for internal data documents that are not rough data, to prevent information leakage during intelligent data value mining, the internal data documents can be anonymized after acquisition, and then RAG intelligent agents can be used to perform intelligent data mining on the anonymized internal data documents.

[0030] Additionally, it's important to note that internal enterprise data documents are crucial carriers of business activities and data assets, encompassing various key information throughout the enterprise's operations. In one possible implementation, these documents can be structured database files containing detailed business data and process information; semi-structured log files recording the time, subject, and data changes of business operations; or unstructured text files, such as data dictionary documents detailing the definitions and uses of each data item. The system connects to the enterprise's ERP and CRM systems via a multi-source access module to acquire these internal data documents. The RAG intelligent agent flexibly adjusts its parsing strategy based on different document types and formats, using a retrieval tool to quickly locate sections of the document that may contain key metadata. A reader then performs in-depth analysis to accurately extract key metadata sets, such as field descriptions, update frequency, and privacy levels, ensuring the accuracy and efficiency of subsequent data processing.

[0031] Step S20: Call the pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set to obtain a data value assessment matrix; The Dynamic Lineage Analyzer is an intelligent analysis tool based on data lineage tracing technology. It identifies and records the source, flow, and interrelationships of data throughout its entire lifecycle, from generation to use, processing, and storage. Pre-trained with extensive enterprise data and process case studies, the Dynamic Lineage Analyzer can perform in-depth analysis of key metadata sets to construct a data value assessment matrix. This matrix includes quantitative indicators for dimensions such as data scarcity, monetization potential, and compliance costs, used to evaluate the value of data assets.

[0032] Additionally, it's important to note that data lineage analysis essentially traces the origin and flow path of data, assessing its role and value contribution within the overall business processes. In one possible implementation, a dynamic lineage analyzer utilizes graph database technology to construct a data lineage graph, treating data sources, data processing procedures, and data consumption stages as nodes, and connecting the data flow relationships between these nodes as edges. By analyzing the attributes and weights of nodes and edges, the value of data assets is dynamically evaluated. For example, data scarcity can be understood as the uniqueness and substitutability of the data in the market or within the enterprise; the scarcer the data, the higher its value assessment. Monetization potential refers to the ability of data to be converted into actual economic benefits, quantified by analyzing the correlation between data and business revenue. Compliance costs involve the enterprise's investment in using and managing data in compliance with regulatory requirements, including costs related to privacy protection and data security; higher costs may affect the net value of data assets.

[0033] For example, in one specific implementation, the dynamic lineage analyzer first uses each data item in the key metadata set as an initial node to identify its source in the enterprise's business systems. It then traces the processing steps, tools, and algorithms used, and ultimately, which business departments or application systems consume this data. During this process, the dynamic lineage analyzer dynamically calculates the values ​​of three dimensions: data scarcity, monetization potential, and compliance cost, based on information such as the frequency of data use and importance scores at different stages. For instance, a company's customer transaction data might have a scarcity score of 0.8 (out of 1) because this type of data is unique within the industry; a monetization potential score of 0.75 because it can be directly used for targeted marketing activities; and a compliance cost score of 0.6 because it must meet strict customer privacy regulations. These indicators together constitute a data value assessment matrix, providing a quantitative basis for subsequent predictive models.

[0034] Step S30: Load the pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table. The feasibility prediction results for inclusion in the table include the success rate of inclusion in the table and the recommended asset type. The asset prediction model is a machine learning-based model that undergoes multi-layered training. Its pre-training process encompasses a large amount of historical data asset cases, corporate financial indicators, and market data. The model employs a hierarchical predictive analysis approach, dividing the prediction process into multiple layers. Each layer focuses on prediction objectives of different dimensions and depths, processes the data value assessment matrix, and ultimately outputs the feasibility prediction results for asset inclusion. These results include indicators such as inclusion success rate, recommended asset types, expected returns, and risk levels, guiding companies on whether to include their data assets in their financial statements and, if so, how.

[0035] Understandably, enterprises can only enter valuable data assets into their tables. Therefore, the feasibility prediction results of data entry can quantify the value of enterprise data, thereby enabling better intelligent data value mining.

[0036] Furthermore, it should be noted that the hierarchical predictive analysis architecture of the asset prediction model draws on multi-stage decision theory. By decomposing complex prediction tasks into multiple relatively simple sub-tasks, it delves deeper layer by layer, gradually narrowing the prediction scope and improving prediction accuracy. In one possible implementation, the initial screening layer of the asset prediction model uses a RAG agent to retrieve historical similar cases, such as successful cases of public transportation data assetization, to make a preliminary judgment on the basic feasibility of data assets being included in the balance sheet, and to screen out data assets with basic potential for inclusion. The actuarial layer introduces corporate financial indicators, such as development costs, maintenance costs, and expected financing and credit enhancement returns, to perform detailed cost-benefit ratio calculations and assess the economic feasibility and expected benefits of including data assets in the balance sheet. The decision layer finally integrates the results from each layer and outputs a comprehensive feasibility prediction result for inclusion in the balance sheet, including the percentage of success rate, recommended asset types (such as current assets, intangible assets), expected return range, and risk level assessment, providing a strong basis for accurate corporate decision-making.

[0037] Step S40: Generate an intelligent data value disclosure report based on the table entry success rate and the recommended asset type.

[0038] The Intelligent Data Value Disclosure Report is a formal report document prepared in accordance with relevant regulations. The report covers asset ownership verification (proof of ownership and usage rights of data assets), valuation methods (detailed explanations of the cost approach, income approach, or market approach used), and accounting treatment (including the selection of accounting items for data asset entry, determination of amortization period, and impairment testing methods). Additionally, the Intelligent Data Value Disclosure Report includes suggestions for data asset implementation scenarios, data asset value estimates, and data asset implementation paths. Through systematic and automated report generation, it ensures that the enterprise's data asset entry process complies with regulatory requirements and improves the enterprise's data asset management level.

[0039] Additionally, it should be noted that the report generation process integrates natural language generation and template matching technologies. Template matching refers to selecting appropriate text templates based on regulatory requirements and report type, and then filling the generated content into the corresponding positions. In one possible implementation, the system automatically generates an asset ownership certificate section based on various indicators in the feasibility prediction results, such as the success rate of inclusion, recommended asset type, expected return, and risk level. This certificate details key information such as the source, development process, and ownership of the data asset. The valuation method section generates detailed descriptions based on the valuation models and methods used in the prediction process (such as the specific cost structure in the cost approach and the revenue prediction model in the income approach). The accounting treatment scheme section automatically matches the relevant accounting standards requirements based on the recommended asset type, determining the accounting items for inclusion (such as specific sub-items under the intangible assets category), amortization period (determined based on the expected useful life of the data asset and industry practices), and impairment testing methods (such as the impairment indicator assessment and testing process conducted at the end of each year), ensuring the completeness and accuracy of the report.

[0040] This embodiment provides an intelligent data value mining method. It acquires enterprise internal data documents, uses a pre-built localized deployment intelligent agent to parse these documents, and extracts a set of key metadata. A pre-trained dynamic lineage analyzer is then invoked to perform lineage analysis on the key metadata set, yielding a data value assessment matrix. A pre-trained asset prediction model is loaded to perform hierarchical predictive analysis on the data value assessment matrix, obtaining a feasibility prediction result for data entry into the database. This feasibility prediction result includes the success rate of data entry and recommended asset types. An intelligent data value disclosure report is generated based on the success rate and recommended asset types. Firstly, the locally deployed intelligent agent enables efficient parsing and vectorization of multi-source heterogeneous data documents, including unstructured documents. It quickly retrieves and processes enterprise internal data documents, extracting a set of key metadata. Furthermore, the dynamic lineage analyzer can quickly generate a full-link graph of the data source processing chain, automatically tracking the source and flow of data. This paper utilizes a pre-trained dynamic lineage analyzer to perform lineage analysis on a key metadata set, quantifying the value contribution coefficient of the data and obtaining a data value assessment matrix. Furthermore, a pre-trained asset prediction model is loaded to perform hierarchical predictive analysis on the data value assessment matrix, thereby assessing the on-balance-sheet value and compliance risks of data assets from multiple dimensions and obtaining on-balance-sheet feasibility prediction results. Finally, based on the on-balance-sheet feasibility prediction results, a compliant intelligent data value disclosure report is generated. This report serves as the core credential for data asset on-balance-sheet entry and is also a key basis for internal enterprise value management, resolving the "value ambiguity" of data assets and thus achieving efficient and accurate assessment of the on-balance-sheet value and compliance risks of enterprise data assets. In summary, this application combines locally deployed intelligent agents, a data asset prediction engine, and a dynamic lineage analyzer to achieve efficient and accurate assessment of the on-balance-sheet value and compliance risks of enterprise data assets.

[0041] In one feasible implementation, step S10 may include steps S11 to S13: Step S11: Obtain the enterprise internal data document, perform data desensitization and recursive character segmentation on the enterprise internal data document to obtain a set of short text fragments; To improve processing efficiency and accuracy, the acquired internal enterprise data documents are preprocessed, transforming them into a set of short text fragments suitable for subsequent processing. Character segmentation refers to recursively dividing long text into shorter fragments layer by layer. This method can handle text of varying lengths, ensuring that the length of each short text fragment is within a processable range. The set of short text fragments refers to the shorter text fragments obtained after recursive character segmentation.

[0042] Specifically, the system first obtains documents related to various data assets stored internally by the enterprise through a multi-source access module. After de-identifying the documents, a recursive segmentation algorithm is used to perform multiple rounds of character cutting according to preset rules for long text documents. For example, sentences are segmented by periods / semicolons. If a single sentence is too long, phrases are further separated by commas until the length of all segments does not exceed the preset processing limit.

[0043] Step S12: Use predefined domain adaptation prompts to guide the localized deployment agent to perform vectorization processing on the short text fragment set, and perform retrieval and matching on the document vectors obtained after vectorization processing to obtain similar metadata patterns; Domain-specific prompts are a set of prompts designed for a specific data asset domain, guiding locally deployed RAG agents to focus on metadata information related to the data asset. Vectorization refers to converting short text fragments into a computer-processable numerical form, i.e., vectors, for subsequent retrieval and matching operations. Similar metadata patterns refer to patterns that share similar characteristics with metadata extracted from internal enterprise data documents.

[0044] Using predefined domain-adaptation prompts, such as "extract ownership information of data assets" or "identify compliance nodes in the data processing chain," which are professional instructions tailored to the data asset entry scenario, the locally deployed RAG agent is guided to perform vectorization processing, transforming text fragments into high-dimensional semantic document vectors. Based on these document vectors, similarity searches are performed on validated metadata structure templates from historical projects to select similar metadata patterns that best match the semantics of the current document vector, providing a structural reference for subsequent metadata extraction.

[0045] Step S13: Extract a set of key metadata based on the document vector and the similar metadata pattern.

[0046] Document vectors are numerical representations of short text fragments obtained through vectorization. Similar metadata patterns are patterns found through retrieval and matching that are similar to document vectors. A key metadata set refers to a collection of key metadata extracted from internal enterprise data documents; this metadata includes, but is not limited to, field descriptions, update frequency, and privacy levels.

[0047] By combining document vectors carrying semantic information of short texts and similar metadata patterns, the RAG agent is invoked to perform extraction operations: on the one hand, the key content in the text is located through the semantic information of the document vectors, and on the other hand, the located content is structured and mapped according to the structure of the similar metadata patterns, and finally outputs a set of key metadata containing core information such as data source, collection time, ownership, value level, and compliance status.

[0048] In one feasible implementation, step S40 may include steps S41-S42: Step S41: Call the built-in disclosure report generator to evaluate and analyze the table entry success rate and the recommended asset type, and generate data ownership certificate, evaluation method description and accounting treatment plan; The system's built-in disclosure report generator is invoked. This generator is pre-configured with relevant policy databases and performs in-depth evaluation and analysis of the success rate of data entry and recommended asset types, specifically including: First, based on the success rate of data entry, extract the ownership information of recommended asset types (such as intellectual property registration number and source system authorization document), and refer to typical data ownership confirmation cases to generate a data ownership confirmation certificate that includes data collection time, ownership party, and scope of use authorization. Furthermore, by combining the data value assessment matrix (scarcity, monetization potential, compliance costs) and the cost-benefit ratio calculation results, the assessment method used is clearly defined (such as the cost approach to calculate development costs and the income approach to predict financing credit enhancement returns), and the specific calculation formula and parameter source of the cost-benefit ratio are attached; The specific formula for calculating the cost-benefit ratio is: ROI = (financing valuation increase - compliance costs) / development costs, where ROI represents the cost-benefit ratio.

[0049] Finally, based on relevant accounting requirements, the amortization period and impairment test standards for data assets are determined, and an accounting treatment plan is generated.

[0050] Step S42: Integrate the data ownership certificate, the evaluation method description, and the accounting treatment scheme to generate an intelligent data value disclosure report.

[0051] The system integrates data ownership verification, assessment methodology, and accounting treatment in a structured manner: it uses a built-in template engine for unified formatting to ensure consistent terminology and logical coherence across all parts; secondly, it calls the compliance verification module to verify whether the report content covers all elements required by policy; finally, it outputs an "Intelligent Data Value Disclosure Report" that includes triple verification (data authenticity, assessment scientificity, and accounting compliance), serving as the core basis for enterprises to declare data assets to the finance department and quantify the value of enterprise data assets.

[0052] In addition, the Intelligent Data Value Disclosure Report also includes suggestions for data asset implementation scenarios, data asset value estimates, and data asset implementation paths. After integrating data ownership certificates, assessment method descriptions, and accounting treatment schemes to obtain the Intelligent Data Value Disclosure Report, it proposes application scenarios for data assets both internally and externally, based on the company's business needs and the characteristics of the data assets. Using the results of the asset prediction module, it estimates the current and potential value of the data assets. Employing multiple assessment models and methods, and comprehensively considering factors such as the market environment and industry development trends, it provides companies with a value reference for their data assets. It establishes an implementation path for data assets from generation to application, clarifying the key steps and timelines at each stage to ensure that data assets can be effectively transformed into corporate value. The aforementioned data asset implementation scenario suggestions, data asset value estimates, and data asset implementation paths are added to the Intelligent Data Value Disclosure Report to form a complete Intelligent Data Value Disclosure Report, thereby realizing the mining of the company's intelligent data value.

[0053] Based on the first embodiment of this application, a second embodiment of this application is proposed. In the second embodiment of this application, content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter.

[0054] Based on this, please refer to Figure 2 , Figure 2 This is a schematic flowchart of the second embodiment of the intelligent data value mining method provided in this application.

[0055] like Figure 2 As shown, before step S10, the intelligent data value mining method further includes steps S01 to S06: Step S01: Collect multi-source heterogeneous data from enterprises, perform data preprocessing and feature extraction on the multi-source heterogeneous data from enterprises, and construct a data asset fingerprint database; Enterprise multi-source heterogeneous data refers to data of different types and structures generated from multiple business systems and data sources within an enterprise. These data sources are wide-ranging, including but not limited to ERP system data, CRM system data, OA system data, PDF files, database logs, etc. Data preprocessing refers to cleaning, transforming, and standardizing the collected raw data to eliminate noise, missing values, duplicate data, and other issues, ensuring data quality and consistency. Feature extraction refers to selecting representative and discriminative features from the preprocessed data for subsequent analysis and processing. The data asset fingerprint database is a database that stores the feature information of data assets after preprocessing and feature extraction. It records the key features of data assets in the form of feature vectors to facilitate rapid retrieval and matching.

[0056] In one possible implementation, the system can employ a multi-threaded acquisition mechanism to simultaneously collect data from multiple business systems and data storage media (such as relational databases, non-relational databases, and file systems). Multi-source heterogeneous data is collected from different systems within the enterprise via data interfaces or manual import. The collected data is then preprocessed and feature extracted. Preprocessing includes cleaning redundant and erroneous values, removing duplicate records, and standardizing the format. Feature extraction includes extracting unique identifier hash values, core semantic summaries, and metadata tags such as "customer information" and "equipment operation logs." Finally, the extracted features are stored according to a unified rule of data type-source system-timestamp, constructing a "data asset fingerprint database" covering all of the enterprise's data assets, which serves as the basic input for subsequent model training and lineage analysis.

[0057] Step S02: Use the data asset fingerprint database to perform vectorized training on the sentence embedding model; The sentence embedding model used in this embodiment refers to the Ollam deployment nomicembedtext model, which is a sentence embedding model based on the SentenceTransformers library and is widely used in text similarity analysis, information retrieval, semantic retrieval and other fields.

[0058] Data features such as hash values, semantic summaries, and metadata tags from the data asset fingerprint database are used as training samples input into the model. By adjusting the initial parameters of the nomicembedtext model, the model learns the semantic associations and structural features of the data assets, and finally outputs a high-dimensional vector that can accurately represent the semantics of the data assets.

[0059] Step S03: Integrate the vectorized trained sentence embedding model with the retrieval enhancement framework to obtain a locally deployed intelligent agent; The trained sentence embedding model serves as the "retrieval module encoder" of the RAG retrieval enhancement framework, responsible for converting user queries or input text into vectors; at the same time, the data asset fingerprint database serves as the "knowledge base" of RAG, used to store and retrieve relevant data asset information; through interface integration and setting parameters such as retrieval thresholds and relevance ranking, the RAG agent is finally integrated and deployed locally.

[0060] Step S04: Perform lineage analysis on the data asset fingerprint database based on the graph neural network model, construct a full-link graph, train the graph neural network model using the full-link graph, and encapsulate the trained graph neural network model to obtain a dynamic lineage analyzer. Graph neural network models are neural network models based on graph structures, capable of processing and analyzing the relationships between nodes and edges in graph data, capturing complex dependencies and interaction patterns within the graph. A full-link graph is a visual representation of the complete lineage of data assets, encompassing all stages and interrelationships from data source to final use. A dynamic lineage analyzer refers to an encapsulated graph neural network model that can analyze changes in the lineage of data assets in real-time or near real-time, promptly identifying anomalies and problems in the data lineage chain.

[0061] Based on a data asset fingerprint database, each data asset is treated as a "node" in a graph neural network, and the relationships between data assets are treated as "edges," constructing a "full-link graph" that reflects the entire lifecycle of data. This graph is used to train a GNN model, and the loss function is optimized to capture complex dependencies, enabling the model to identify the source, processing path, and potential impact of data assets. After training, the model parameters and inference logic are encapsulated to obtain a "dynamic lineage analyzer" that can analyze the lineage of data assets in real time, used for dependency path tracing in subsequent data value assessment.

[0062] Step S05: Construct a structured training dataset based on the data asset fingerprint database, and use the structured training dataset to fine-tune the parameters of the large language model to obtain a domain-adapted large language model. Structured training datasets refer to data extracted from data asset fingerprint databases that has been organized and labeled. They exist in the form of tables, records, etc., and contain various characteristics and label information of data assets, such as data type, data volume, and data quality indicators. Parameter fine-tuning refers to further adjusting and optimizing the parameters of a large language model using a domain-specific structured training dataset, making the model adaptable to the language style, terminology, and task requirements of that specific domain. Domain-adapted large language models, after parameter fine-tuning, are large language models that can adapt to the specific needs of the data asset domain, possessing a deeper understanding and more accurate generation capabilities for data asset-related text.

[0063] Standardized features and corresponding labels are extracted from the data asset fingerprint database, and a structured training dataset is constructed in an "input-output" format. The dataset is then used to fine-tune the parameters of a general-purpose large language model, enabling it to learn the professional terminology and task logic of the data asset table entry domain. The professional terminology includes "table entry feasibility" and "cost method measurement," while the task logic includes metadata extraction rules. Ultimately, a domain-adaptive large language model that can accurately handle data asset-related tasks is obtained.

[0064] Step S06: Use the structured training dataset to train the domain-adapted large language model hierarchically to obtain the asset prediction model.

[0065] Layered training refers to dividing the training process of a domain-adaptive large language model into multiple layers or stages, with each layer or stage focusing on learning different features and tasks. Based on a structured training dataset, the domain-adaptive large language model is progressively trained according to task complexity: first, simple tasks, such as "identifying data sources," are used to optimize the model's basic classification capabilities; then, medium-complexity tasks, such as "calculating costs based on processing links," are used to improve feature fusion capabilities; finally, complex tasks, such as "combining compliance and value levels to determine whether to include data in the table," are used to strengthen multi-factor reasoning capabilities. After training, through parameter solidification and interface encapsulation, an "asset prediction model" is obtained that can output the probability of data inclusion feasibility based on data asset characteristics (source, value, compliance), providing core predictive basis for the final generation of the data inclusion disclosure report.

[0066] In this embodiment, multi-source heterogeneous data from enterprises is collected to construct a data asset fingerprint database, laying a solid foundation for subsequent operations. Through vectorized training and integration with the RAG framework, the resulting locally deployed RAG intelligent agent can accurately retrieve and generate text information related to data assets. Based on graph neural network model-based lineage analysis and the encapsulation of a dynamic lineage analyzer, the entire lineage of data assets can be clearly traced, ensuring data quality traceability and compliance. Fine-tuning and hierarchical training of the large language model parameters ultimately yields an asset prediction model that accurately predicts the value of data assets, feasibility of data entry, and compliance risks. Overall, this significantly improves the efficiency, accuracy, and reliability of enterprise data asset entry, meeting the complex needs of data asset management and entry during enterprise digital transformation.

[0067] Based on the first and / or second embodiments of this application, a third embodiment of this application is proposed. In the third embodiment of this application, content that is the same as or similar to the first and / or second embodiments described above can be referred to the above description and will not be repeated hereafter.

[0068] Based on this, please refer to Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the intelligent data value mining method provided in this application.

[0069] like Figure 3 As shown, step S20, which involves calling a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set to obtain a data value assessment matrix, may include steps S21 to S23: Step S21: Invoke the pre-trained dynamic lineage analyzer to identify entities and relationships in the key metadata set and construct an entity-relationship list; The pre-trained dynamic lineage analyzer is invoked to perform semantic parsing and relationship mining on the extracted key metadata set: on the one hand, core entities in the metadata are identified, such as data sources, data processing steps, and data asset entities; on the other hand, the relationships between entities are extracted, and finally, a list of entity relationships containing entity names and relationship types is output, providing structured input for subsequent graph construction.

[0070] Step S22: Using the entities in the entity relationship list as graph neural network nodes and the relationships in the entity relationship list as graph neural network nodes, construct a full-link lineage graph of data flow based on the graph neural network nodes and graph neural network nodes; Using entities in the entity relationship list as "nodes" and the relationships between entities as "edges," the graph neural network leverages its topological modeling capabilities to integrate discrete entity relationships into a "full-link lineage graph" that reflects the entire lifecycle of data flow. This full-link lineage graph visually presents the complete data flow from initial acquisition (e.g., in an ERP system) to processing (e.g., data cleaning and desensitization) and finally to the formation of assets, enabling visualization and traceability of the data flow path.

[0071] Step S23: Calculate the value contribution coefficient of the full-link kinship map of data flow using the dynamic kinship analyzer to obtain a weighted full-link kinship map, and generate a data value assessment matrix based on the weighted full-link kinship map.

[0072] By utilizing the graph neural network inference module built into the dynamic lineage analyzer, the "value contribution coefficient" (ranging from 0 to 1, e.g., the edge weight for "merchant data → tourist profile" is 0.78) is calculated for each edge based on features such as the connection strength and influence range of nodes in the full-link lineage graph, forming a weighted full-link lineage graph. Subsequently, by combining the attributes of each node in the graph (data scarcity, compliance cost) and the value contribution coefficient of the edge weights, key indicators (scarcity score, expected monetization revenue, and compliance rectification cost) are integrated from three dimensions: "data scarcity, monetization potential, and compliance cost," ultimately generating a "data value assessment matrix" for data asset valuation, providing a quantitative basis for subsequent feasibility prediction and value measurement.

[0073] In this embodiment, a clear list of entity relationships is constructed using a dynamic lineage analyzer, and based on this, a full-link lineage graph of data flow is built, visually presenting the data flow path and relationships in the form of graph neural network nodes. By calculating the value contribution coefficient and generating a weighted, scored full-link lineage graph, the contribution of each node and relationship to the data value is further quantified. The resulting data value assessment matrix comprehensively evaluates the value of data assets from multiple dimensions. This provides enterprises with a complete, quantitative, and visualized data asset value assessment system, helping them to deeply understand the value of their data assets, thereby optimizing data asset management strategies and improving the utilization efficiency and value creation capabilities of data assets.

[0074] Based on the above embodiments of this application, a fourth embodiment of this application is proposed. In this fourth embodiment, content that is the same as or similar to that in the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0075] In this embodiment, the asset prediction model includes a preliminary screening layer, an actuarial layer, and a decision layer. Step S30, which loads the pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table, may include steps S31 to S34: Step S31: Standardize the data value assessment matrix to obtain a standardized feature vector; To eliminate the dimensional differences between indicators of different dimensions such as data scarcity, monetization potential, and compliance costs, a normalization algorithm or Z-score standardization algorithm is used to convert each indicator in the data value assessment matrix into a value with a unified dimension. This results in a standardized feature vector that can be directly input into the model for processing, ensuring that the subsequent prediction model can fairly and accurately analyze the impact of each dimension on the data asset entry into the table.

[0076] Step S32: Load the pre-built asset prediction model, input the standardized feature vector into the initial screening layer to retrieve similar cases of assetization in historical data, and determine the matching similar case information; A pre-built asset prediction model is loaded, and standardized feature vectors are input into the initial screening layer. Based on the retrieval capabilities of a localized RAG agent, the initial screening layer matches case information of similar scenarios from a historical data assetization case library. By comparing the similarity between the features of the current data asset and the feature vectors of historical cases, historical cases that highly match the current data asset are selected, providing a reference for subsequent cost-benefit analysis.

[0077] Step S33: Input the similar case information into the actuarial layer to calculate the cost-benefit ratio; After receiving similar case information from the initial screening layer, the actuarial layer uses the pre-set cost-benefit ratio calculation formula ROI = (financing valuation increase - compliance cost) / development cost to quantitatively calculate and assess the economic feasibility of adding data assets to the table, based on the development costs recorded in the cases: such as labor / equipment costs for data collection, cleaning, and storage; maintenance costs: such as long-term investment in data updates and security protection; and financing credit enhancement benefits: such as increased financing amount and lower interest rate due to improved corporate credit rating resulting from adding data assets to the table.

[0078] Step S34: Based on the decision-making layer, and combining the cost-benefit ratio and preset compliance verification rules, predict the success rate of table entry and recommend asset types to obtain the table entry feasibility prediction result.

[0079] The decision-making level combines the cost-benefit ratio output by the actuarial level with the built-in compliance verification rules to comprehensively judge the success rate of data asset entry into the table, and recommends suitable asset types based on historical successful cases. Finally, it forms a feasibility result for entry into the table that includes success rate prediction and asset type suggestions, providing direct support for enterprise decision-making.

[0080] In this embodiment, through hierarchical analysis of the asset prediction model, similar historical cases are accurately retrieved, the cost-benefit ratio is quantitatively calculated, and the success rate of data asset entry into the table and the recommended asset types are predicted in combination with compliance rules. Finally, reliable data asset entry feasibility prediction results are obtained, which effectively improves the scientificity, accuracy and compliance of data asset entry into the table, reduces decision-making risks, and provides strong support for the reasonable entry of enterprise data assets into the table.

[0081] Based on the above embodiments of this application, a fifth embodiment of this application is proposed. In this fifth embodiment, content that is the same as or similar to that in the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0082] In this embodiment, the intelligent data value mining method further includes steps A10 to A50: Step A10: Obtain the wireless signal characteristics of the terminal to be located, and input the wireless signal characteristics into the localized deployment agent to retrieve similar wireless signal characteristic data; Data asset mining requires processing multi-source heterogeneous data from enterprises, such as equipment logs, sensor signals, and business system files. Some of this data (wireless signals, IoT device data) is strongly correlated with physical location (e.g., sensor data for a certain area only corresponds to the production scenario in that area). To efficiently and accurately assess the value and compliance risks of enterprises' data assets on the balance sheet, and to assist enterprises in data asset valuation, the RAG agent and dynamic lineage analyzer pre-trained in this application can be reused.

[0083] Specifically, the wireless signal characteristics generated by the terminal to be located in the network environment are first obtained, such as Wi-Fi strength, Bluetooth beacon, and cellular network parameters. These characteristics are then input into a locally deployed RAG agent (based on the Ollama framework's nomicembedtext model). The RAG agent transforms the wireless signal characteristics into high-dimensional vectors using vectorization techniques and retrieves semantically similar historical wireless signal characteristic data from a data asset fingerprint database, providing a foundation of related data for subsequent lineage analysis. Step A20: The dynamic lineage analyzer is used to analyze the lineage relationship of the similar wireless signal feature data, and the coarse positioning area of ​​the terminal to be located is determined based on the lineage relationship. A dynamic lineage analyzer is invoked to perform lineage analysis on similar wireless signal feature data, identifying the data source, processing link, and associated asset entities. By analyzing the flow relationships between data, a coarse positioning area is determined, quickly narrowing down the positioning range.

[0084] Step A30: Compare the wireless signal features with the fingerprints of specific data assets in the coarse positioning area and calculate a similarity score; Within the coarse positioning area, the wireless signal characteristics of the terminal to be located are compared feature by feature with the fingerprints of specific data assets pre-stored in the area. The similarity scores between the two are calculated, and the top N data asset fingerprints with the highest similarity are selected to provide candidate basis for precise positioning.

[0085] Step A40: Based on the similarity score, select the data asset fingerprint and input it into the asset prediction model for prediction to obtain the positioning result of the terminal to be located; The selected N highly similar data asset fingerprints are input into the asset prediction model (including the initial screening layer, the actuarial layer, and the decision layer): the initial screening layer quickly eliminates obviously mismatched fingerprints; the actuarial layer optimizes the matching weights by combining factors such as the time stability of signal features and regional coverage; the decision layer makes a comprehensive judgment through the fine-tuned large language model and outputs the specific location result of the terminal to be located.

[0086] Step A50: Combine the positioning results and the intelligent data value disclosure report to perform intelligent data value mining.

[0087] The system associates the location results of the terminal to be located with the intelligent data value disclosure report automatically generated by the system: if the location results indicate that the terminal is associated with data assets, then the data assets are included in the enterprise data asset list, and the entry process is completed; if there are risks, such as insufficient compliance, then rectification suggestions are triggered, ultimately achieving efficient and compliant entry of data assets into the list.

[0088] In this embodiment, RAG intelligent agents are used to accurately locate similar wireless signal feature data; a dynamic lineage analyzer is used to determine a coarse location area, narrowing the location range; similarity scores are calculated through comparison, further improving the accuracy of location; data asset fingerprints are selected for prediction based on the similarity scores, ensuring the accuracy of the location results; and the location results are combined with an intelligent data value mining disclosure report to realize intelligent data value mining. This improves the efficiency and accuracy of data asset entry into the table, while simultaneously enhancing the accuracy and reliability of location, providing strong support for data asset management.

[0089] For example, to help understand the implementation process of the intelligent data value mining method obtained by combining this embodiment with the above embodiments, please refer to... Figure 4 , Figure 4 A simplified flowchart of an intelligent data value mining method is provided, specifically: like Figure 4 As shown, the data access layer module is responsible for accessing multi-source heterogeneous data from enterprises, providing raw data for subsequent processing, connecting with enterprise business systems such as ERP and CRM, and collecting enterprise data documents such as data dictionaries and API logs. It uses an unstructured parsing engine combined with RAG agents to parse unstructured data such as PDFs and database logs, extracting key metadata such as field descriptions, update frequency, and privacy levels. Specifically, the RAG agent layer performs data vectorization processing to achieve efficient retrieval and preliminary analysis, segmenting long documents or large datasets and vectorizing documents based on the Ollam nomicembedtext model.

[0090] Furthermore, the core processing layer includes a lineage analysis engine, namely a dynamic lineage analyzer, and a compliance verifier. The lineage analysis engine utilizes graph neural networks to construct a data lineage graph, identify entities and relationships, quantify value contribution, and generate a weighted, scored end-to-end lineage graph. A matrix is ​​constructed from three dimensions—data scarcity, monetization potential, and compliance costs—to assess the value of data assets, resulting in an asset value matrix.

[0091] Furthermore, the predictive decision-making module is responsible for predicting the feasibility of data asset entry into the table, providing support for decision-making: the initial screening case library searches for historical data assetization cases to find cases similar to the current data; the cost-benefit model calculates the development and maintenance costs and financing and credit enhancement benefits of the data asset to obtain the cost-benefit ratio; the decision-making layer combines the cost-benefit ratio and compliance verification rules to output the success rate of entry into the table and the recommended asset type.

[0092] Furthermore, the output layer automatically generates an intelligent data value disclosure report based on the prediction results of the prediction decision layer and relevant regulations, using the disclosure report generator. The report includes asset ownership certificates, valuation method descriptions, and accounting treatment schemes. It also provides an asset registration interface to connect with the asset registration system, facilitating the registration of assessed data assets by enterprises.

[0093] In addition, this application provides a feedback optimization layer module for continuous optimization and improvement of the system. The manual annotation interface allows auditors and other professionals to manually annotate and correct the prediction results. Incremental model training updates the model parameters by incrementally training the prediction model based on manually annotated data and new data asset cases.

[0094] In a company's data asset entry project, this system was used for data asset management and entry. This company has multiple business systems with complex data sources, including data from ERP, CRM, OA systems, and data in formats such as PDF and database logs. Through the system's multi-source access module, it successfully connected to the company's various business systems, achieving unified data collection and management. In the offline phase, a data asset fingerprint database was built, and a localized RAG engine, dynamic lineage analyzer, and asset prediction module were trained. In the online phase, coarse and fine positioning of the wireless signal characteristics of the target terminal were performed, obtaining accurate positioning results. Ultimately, by combining the locally deployed RAG agent with the data asset prediction engine, the data leakage risk of public cloud RAG services was overcome. The company's data asset entry efficiency improved by 300% (15 days manually → 4 hours system), the entry success rate and prediction accuracy increased from 62% to 93%, and the policy compliance verification completeness increased from 80% to 98%.

[0095] It should be noted that the percentages of 300%, 15 days of manual labor to 4 hours of system operation, 93%, and 98% are only used to illustrate that this application has achieved efficient and accurate assessment of the value of an enterprise's data assets on the balance sheet and compliance risks. The above examples and specific data are only for understanding this application and do not constitute a limitation on this application. Any simple modifications based on this technical concept are within the scope of protection of this application.

[0096] This application also provides an intelligent data value mining device, please refer to... Figure 5 The intelligent data value mining device includes: The localized intelligent engine module 10 is used to acquire enterprise internal data documents, and to parse the enterprise internal data documents using a pre-built localized deployment intelligent agent to extract key metadata sets; The dynamic lineage analysis module 20 is used to call a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set and obtain a data value assessment matrix. The asset prediction module 30 is used to load a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix and obtain the feasibility prediction result for inclusion in the table. The feasibility prediction result for inclusion in the table includes the success rate of inclusion in the table and the recommended asset type. The disclosure report generation module 40 is used to generate an intelligent data value disclosure report based on the table entry success rate and the recommended asset type.

[0097] The intelligent data value mining device provided in this application, employing the intelligent data value mining method in the above embodiments, can solve the technical problems of intelligent data value mining. Compared with the prior art, the beneficial effects of the intelligent data value mining device provided in this application are the same as those of the intelligent data value mining method provided in the above embodiments, and other technical features in the intelligent data value mining device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0098] This application provides an intelligent data value mining device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the intelligent data value mining method in the first embodiment described above.

[0099] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an intelligent data value mining device suitable for implementing embodiments of this application. The intelligent data value mining device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The intelligent data value mining device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0100] like Figure 6As shown, the intelligent data value mining device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the intelligent data value mining device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the intelligent data value mining device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows intelligent data value mining devices with various systems, it should be understood that implementing or possessing all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0101] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0102] The intelligent data value mining device provided in this application, employing the intelligent data value mining method described in the above embodiments, can solve the technical problems of intelligent data value mining. Compared with the prior art, the beneficial effects of the intelligent data value mining device provided in this application are the same as those of the intelligent data value mining method provided in the above embodiments, and other technical features of the intelligent data value mining device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0103] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0105] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the intelligent data value mining method in the above embodiments.

[0106] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0107] The aforementioned computer-readable storage medium may be included in the intelligent data value mining device; or it may exist independently and not be assembled into the intelligent data value mining device.

[0108] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the intelligent data value mining device, the intelligent data value mining device performs the following actions: acquires internal enterprise data documents; parses the internal enterprise data documents using a pre-built, locally deployed intelligent agent; extracts a set of key metadata; calls a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set to obtain a data value assessment matrix; loads a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain a feasibility prediction result for inclusion in the asset table, the feasibility prediction result including the inclusion success rate and recommended asset type; and generates an intelligent data value disclosure report based on the inclusion success rate and the recommended asset type.

[0109] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0111] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0112] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described intelligent data value mining method, and is capable of solving the technical problems of intelligent data value mining. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the intelligent data value mining method provided in the above embodiments, and will not be repeated here.

[0113] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent data value mining method described above.

[0114] The computer program product provided in this application can solve the technical problem of intelligent data value mining. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the intelligent data value mining method provided in the above embodiments, and will not be repeated here.

[0115] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for intelligent data value mining, characterized in that, The intelligent data value mining method includes: Obtain enterprise internal data documents, and use a pre-built localized deployment intelligent agent to parse the enterprise internal data documents and extract key metadata sets; A pre-trained dynamic lineage analyzer is invoked to perform lineage analysis on the key metadata set to obtain a data value assessment matrix; A pre-trained asset prediction model is loaded to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table. The feasibility prediction results for inclusion in the table include the success rate of inclusion in the table and the recommended asset type. A smart data value disclosure report is generated based on the table entry success rate and the recommended asset types.

2. The intelligent data value mining method as described in claim 1, characterized in that, Before the step of acquiring enterprise internal data documents, parsing the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extracting key metadata sets, the following steps are also included: Collect multi-source heterogeneous data from enterprises, perform data preprocessing and feature extraction on the multi-source heterogeneous data from enterprises, and construct a data asset fingerprint database; The sentence embedding model is trained using the data asset fingerprint database. The vectorized sentence embedding model is integrated with the retrieval enhancement framework to obtain a locally deployed intelligent agent; Based on the graph neural network model, the data asset fingerprint database is analyzed for lineage relationships to construct a full-link graph. The graph neural network model is trained using the full-link graph, and the trained graph neural network model is encapsulated to obtain a dynamic lineage analyzer. A structured training dataset is constructed based on the data asset fingerprint database. The parameters of the large language model are fine-tuned using the structured training dataset to obtain a domain-adapted large language model. The domain-adaptive large language model is trained hierarchically using the structured training dataset to obtain the asset prediction model.

3. The intelligent data value mining method as described in claim 1, characterized in that, The steps of acquiring enterprise internal data documents, parsing the enterprise internal data documents using a pre-built localized deployment intelligent agent, and extracting key metadata sets include: Obtain internal data documents from the enterprise, perform data anonymization and recursive character segmentation on the internal data documents to obtain a set of short text fragments; The localized deployment agent is guided to vectorize the set of short text fragments using predefined domain-adaptive prompts, and the document vectors obtained after vectorization are searched and matched to obtain similar metadata patterns. Extract a set of key metadata based on the document vector and the similar metadata pattern.

4. The intelligent data value mining method as described in claim 1, characterized in that, The step of calling a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set and obtaining a data value assessment matrix includes: A pre-trained dynamic lineage analyzer is invoked to identify entities and relationships in the key metadata set, and a list of entity relationships is constructed. Using entities in the entity relationship list as graph neural network nodes and relationships in the entity relationship list as graph neural network nodes, a full-link lineage graph of data flow is constructed based on the graph neural network nodes and graph neural network nodes. The value contribution coefficient of the full-link kinship map of data flow is calculated by the dynamic kinship analyzer to obtain a weighted full-link kinship map, and a data value assessment matrix is ​​generated based on the weighted full-link kinship map.

5. The intelligent data value mining method as described in claim 1, characterized in that, The asset prediction model includes a preliminary screening layer, an actuarial layer, and a decision-making layer. The steps of loading the pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix to obtain the feasibility prediction results for inclusion in the table include: The data value assessment matrix is ​​standardized to obtain a standardized feature vector; Load the pre-built asset prediction model, input the standardized feature vector into the initial screening layer to retrieve similar cases of assetization in historical data, and determine the matching similar case information; The similar case information is input into the actuarial layer to calculate the cost-benefit ratio; Based on the decision-making level, and combined with the cost-benefit ratio and preset compliance verification rules, the success rate of table entry and asset type recommendation are predicted to obtain the table entry feasibility prediction result.

6. The intelligent data value mining method as described in claim 1, characterized in that, The steps for generating a smart data value disclosure report based on the table entry success rate and the recommended asset type include: The built-in disclosure report generator is invoked to evaluate and analyze the success rate of data entry and the recommended asset type, and to generate data ownership certificates, evaluation method descriptions, and accounting treatment schemes. By integrating the data ownership verification, the assessment method description, and the accounting treatment scheme, an intelligent data value disclosure report is generated.

7. The intelligent data value mining method as described in any one of claims 1 to 6, characterized in that, The intelligent data value mining method also includes: Obtain the wireless signal characteristics of the terminal to be located, and input the wireless signal characteristics into the locally deployed intelligent agent to retrieve similar wireless signal characteristic data; The dynamic lineage analyzer is used to analyze the lineage relationship of the similar wireless signal feature data, and the coarse positioning area of ​​the terminal to be located is determined based on the lineage relationship. The wireless signal features are compared with the fingerprints of specific data assets in the coarse positioning area to calculate a similarity score; Based on the similarity score, the fingerprint of the selected data asset is input into the asset prediction model for prediction, thereby obtaining the positioning result of the terminal to be located. Intelligent data value mining is performed by combining the positioning results and the intelligent data value disclosure report.

8. An intelligent data value mining system, characterized in that, The intelligent data value mining system includes: The localized intelligent engine module is used to acquire enterprise internal data documents, and uses a pre-built localized deployment intelligent agent to parse the enterprise internal data documents and extract key metadata sets; The dynamic lineage analysis module is used to call a pre-trained dynamic lineage analyzer to perform lineage analysis on the key metadata set and obtain a data value assessment matrix. The asset prediction module is used to load a pre-trained asset prediction model to perform hierarchical prediction analysis on the data value assessment matrix and obtain the feasibility prediction results for inclusion in the table. The feasibility prediction results for inclusion in the table include the success rate of inclusion in the table and the recommended asset type. The disclosure report generation module is used to generate an intelligent data value disclosure report based on the table entry success rate and the recommended asset type.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the intelligent data value mining method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the intelligent data value mining method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Data value evaluation method and system based on knowledge mining large model and analogue simulation agent

    CN121526720A

  • Data asset assessment method

    CN122243554A