A method for incorporating data assets into the balance sheet based on large language models

Through the data asset table entry method based on the large language model, the problem of efficient analysis and semantic labeling of large-scale text data and real-time streaming data is solved, efficient data processing and dynamic optimization are achieved, automated evaluation and visualization tools are provided, and the commercial value of data assets is enhanced.

CN119623442BActive Publication Date: 2025-06-17KAIXINDAI FINANCING SERVICES JIANGSU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510168222.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-17
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

It is difficult for the existing technology to efficiently analyze and process large-scale, high-dimensional complex text data and real-time streaming data, and in enterprise-level data assetization scenarios, there are difficulties in automated semantic annotation, aggregation processing and dynamic optimization.

Method used

The data asset table entry method based on a large language model is adopted, and the data analysis module is used to analyze and identify text data and real-time stream data, and semantic annotation is used using zero-sample learning and few-sample learning methods. The data processing and asset evaluation are carried out in combination with DBSCAN clustering algorithm and causal inference method, and the data assetization process is continuously optimized through the dynamic optimization module.

Benefits of technology

It realizes efficient analysis and semantic annotation of large-scale text data and real-time streaming data, reduces manual intervention, improves data processing efficiency and accuracy, can dynamically optimize the data assetization process, meets the needs of low latency and high throughput, and provides automated evaluation and visualization tools for data assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623442B_ABST
    Figure CN119623442B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for incorporating data assets into the balance sheet based on large language models, belonging to the technical field of data processing. It includes data parsing, semantic annotation, data processing, asset evaluation, and dynamic optimization, realizing an automated data assetization process from data extraction, semantic annotation, structured processing to asset evaluation and optimization. It solves the technical problems of efficient parsing, semantic annotation, data processing, asset evaluation, disclosure of data assets incorporated into the balance sheet, and dynamic optimization for text data and real-time stream data. The present invention reduces the need for manual annotation, realizes automated annotation and processing of data, can efficiently process large-scale text data and real-time stream data, meets the low-latency requirements, provides a comprehensive scoring and visualization tool to help users intuitively understand the value of data assets, and endows data assets with higher commercial value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method for incorporating data assets into the balance sheet based on large language models. Background Art

[0002] With the advent of the big data era, enterprises and organizations have accumulated a vast amount of structured and unstructured data, which contains rich commercial value. How to effectively convert this data into usable assets is a key issue in the current field of data management and processing. Traditional data assetization methods mainly rely on manual annotation and rule-driven algorithms, usually requiring a large amount of manual intervention and professional knowledge, and having low processing efficiency and accuracy for massive and complex data sets.

[0003] Most existing technologies rely on standardized data cleaning and formatting, but have limited processing capabilities for unstructured text data and real-time stream data. In the processing of text data, entity recognition and relation extraction based on machine learning have made certain progress, such as models like BiLSTM-CRF and BERT, but there are still problems of low efficiency and low accuracy for large-scale and high-dimensional complex data. For real-time stream data, traditional data stream parsing technologies often cannot meet the requirements of low latency and high throughput.

[0004] In addition, although current natural language processing (NLP) technologies have achieved success in tasks such as sentiment analysis and text classification, in the enterprise-level data assetization scenario, how to effectively perform automated semantic annotation, aggregation processing, and dynamic optimization remains a difficult point. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for incorporating data assets into the balance sheet based on large language models, which solves the technical problems of efficient parsing, semantic annotation, data processing, asset evaluation, and dynamic optimization for text data and real-time stream data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for incorporating data assets into the balance sheet based on large language models, comprising the following steps:

[0008] Step 1: Establish a data parsing module. The data parsing module obtains text data and real-time stream data, splits the text data, and uses BiLSTM - CRF a model to identify entities, and extract entities, attributes, and their relationships;

[0009] Step 2: Establish a semantic annotation module. The semantic annotation module automatically adds semantic labels to the data through zero-shot learning methods and few-shot learning methods;

[0010] Step 3: Establish a data processing module, which semantically aggregates entities with semantic tags and processes them into structured data assets;

[0011] Step 4: Establish an asset evaluation module, which quantitatively evaluates the value of data resources, generates a data asset statement disclosure report and provides a visual display;

[0012] Step 5: Establish a dynamic optimization module, which continuously optimizes the data assetization process from Step 1 to Step 4.

[0013] Preferably, when performing Step 1, it specifically includes the following steps:

[0014] Step 1-1: Data input, obtaining text data and real-time stream data; the text data includes enterprise documents, logs and contracts; the real-time stream data includes transaction records and statements;

[0015] Step 1-2: Semantic segmentation, splitting long text data by paragraphs and sentences, and the splitting function is as follows:

[0016] Seg ( T )={ S 1, S 2,..., S n}

[0017] where, T represents the text data, S n represents the split sentence, n represents the number of split sentences;

[0018] Step 1-3: Entity recognition, combining BiLSTM - CRF to extract key entities:

[0019] ;

[0020] where, S represents the input sentence, h i represents the semantic representation of the BiLSTM -th word extracted from i , e i represents the entity label of the i -th word, CRF represents the conditional random field, which is used to calculate the label sequence probability; E represents all entity label sequences extracted from the sentence S ; P(E | S)Indicates the probability of generating an entity tag sequence S under the condition of a given input sentence E ;

[0021] Extract the semantic relationship between entities through the relationship classification model:

[0022] ;

[0023] Among them, h i , h j is the context feature representation of entities e i , e j ; r is the relationship category, W represents the weight matrix, b is the bias term;

[0024] Steps 1-4: Real-time stream data parsing, and low-latency data parsing is achieved through the Streaming model. The model is shown in the following formula: Streaming

[0025] Event (( D ) = {transaction ID , amount, timestamp,...};

[0026] Among them D is the transaction data stream;

[0027] Steps 1-5: Relationship extraction, using the Attention - based model to extract the semantic relationship between entities. The specific formula is as follows:

[0028] ;

[0029] Among them, Q , K , V respectively represent the query, key, and value vectors of the entity, d represents the vector dimension, R represents the final attention value, i and j both represent the index position, T is the transpose operation.

[0030] Preferably, when performing step 2, the specific steps are as follows:

[0031] Step 2-1: Extract entities and relationships from the data parsing module;

[0032] Step 2-2: Predefined Tag Generation. Based on LLM the model, perform inference on the business scenario to generate a set of possible semantic tags as follows:

[0033] l = LLM ( Context );

[0034] Among them, l represents the generated semantic tag, Context represents the business scenario context;

[0035] The obtained set of semantic tags is L = { l 1, l 2,... l n};

[0036] Step 2-3: Use the zero-shot learning model to match entities and relationships to the optimal tags. The specific formula is as follows:

[0037] P ( l ∣ e ) = Softmax ( Sim ( h e , h l ));

[0038] Among them, Sim is the cosine similarity function, h e and h l are the semantic embeddings of the entity and the tag respectively;

[0039] Step 2-4: Few-shot Learning Correction. Use a small amount of labeled data to fine-tune the results obtained by the zero-shot learning model.

[0040] Preferably, when performing Step 3, it specifically includes the following steps:

[0041] Step 3-1: Retrieve the labeled entity and relationship data;

[0042] Step 3-2: Semantic Aggregation. Adopt an improved DBSCAN clustering algorithm to classify entities with similar semantics. The improved DBSCAN clustering algorithm is as follows:

[0043] Sim ( x , y ) > ϵ and∣ x − y ∣ < Δ ;

[0044] Among them, Sim represents a similarity function, specifically the cosine similarity function, x and y are both input entities, ϵ is the semantic similarity threshold, Δ is the numerical difference threshold; if the value obtained by the improved DBSCAN clustering algorithm is true, then x and y are of the same class;

[0045] Step 3-3: Pipeline generation, design an automated data processing pipeline, which is divided into three sub-processes: cleaning, integration, and standardization;

[0046] Step 3-4: Asset conversion, introduce a causal inference method, and convert it into high-value data assets based on the causal relationship between data. The formula of the causal inference method is as follows:

[0047] Y = αX + δ ;

[0048] Among them, δ represents the error term, α represents the causal effect, X represents the capital flow, Y represents the performance risk.

[0049] Preferably, when executing Step 4, it specifically includes the following steps:

[0050] Step 4-1: Obtain the processed structured data assets;

[0051] Step 4-2: Weight assignment, based on the LLM scoring system of the

[0052] ;

[0053] Among them, w i is the business indicator weight, v i ( d ) is the scoring value of the i th indicator; d represents the structured data asset, n represents the number of indicators, S ( d) Represents the comprehensive score of data assets;

[0054] Step 4-3: Asset sorting, prioritize data assets according to the comprehensive score, and screen out high-value assets;

[0055] Step 4-4: User visualization, generate a data asset statement disclosure report and provide visual display, the specific content includes the basic information of the data assets and the value assessment results of the data assets.

[0056] Preferably, when performing Step 5, the specific steps are as follows:

[0057] Step 5-1: Obtain user feedback and system evaluation indicators;

[0058] Step 5-2: Semantic feedback, collect feedback data through user interaction, and LLM fine-tune and update the model, specifically including:

[0059] Step 5-2-1: Collect a dataset with feedback annotations;

[0060] Step 5-2-2: Adjust the model objective and add a feedback loss function;

[0061] Step 5-2-3: Use mini-batch stochastic gradient descent SGD to optimize the model parameters, the specific formula is as follows:

[0062] L fine-tune = L original + βL feedback ;

[0063] Where, L feedback represents the loss based on user feedback, β is the weight coefficient, L fine-tune represents the loss function after fine-tuning, L original represents the original loss function;

[0064] Step 5-3: Model optimization, use the knowledge distillation method to LLM lightweight optimize the model;

[0065] Step 5-4: Asset update, design an automatic update mechanism, update the data asset library based on real-time feedback, and eliminate low-value or obsolete assets;

[0066] Step 5-5: Adapt to multi-task scenarios, through multi-task learning Multi - task LearningEnhance the adaptability of the model to multi-business scenarios.

[0067] A method for incorporating data assets into the balance sheet based on large language models according to the present invention solves the technical problems of efficient parsing, semantic annotation, data processing, asset evaluation, and dynamic optimization for text data and real-time stream data. By adopting zero-shot learning and few-shot learning methods, the present invention reduces the need for manual annotation, realizes automated annotation and processing of data, can efficiently process large-scale text data and real-time stream data, meets the low-latency requirements, combines the LLM model and business rules for automated evaluation of data assets, provides a comprehensive score and visualization tool to help users intuitively understand the value of data assets, the system can continuously optimize according to user feedback and real-time data streams, continuously improve the accuracy and adaptability of the model through multi-task learning and knowledge distillation methods, and uses an improved DBSCAN clustering algorithm and causal inference method to perform semantic aggregation and value transformation on data, endowing higher commercial value to data assets. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is the main flow chart of the present invention;

[0069] Figure 2 is the flow chart of step 1 of the present invention;

[0070] Figure 3 is the flow chart of step 2 of the present invention;

[0071] Figure 4 is the flow chart of step 3 of the present invention;

[0072] Figure 5 is the flow chart of step 4 of the present invention;

[0073] Figure 6 is the flow chart of step 5 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] By Figures 1 - 6 A method for incorporating data assets into the balance sheet based on large language models shown includes the following steps:

[0075] Step 1: Establish a data parsing module. The data parsing module obtains text data and real-time stream data, splits the text data, and uses BiLSTM - CRF the model to identify entities, extract entities, attributes, and their relationships;

[0076] When performing step 1, it specifically includes the following steps:

[0077] Step 1-1: Data input, obtain text data and real-time stream data; the text data includes enterprise documents, logs, and contracts; the real-time stream data includes transaction records and reports;

[0078] Step 1-2: Semantic segmentation. The long text data is segmented into paragraphs and sentences. The segmentation function is as follows:

[0079] Seg ( T ) = { S 1, S 2,..., S n}

[0080] Among them, T represents the text data, S n represents the segmented sentence, n represents the number of segmented sentences;

[0081] Step 1-3: Entity recognition. Combine BiLSTM - CRF to extract key entities:

[0082] ;

[0083] Among them, S represents the input sentence, h i represents the semantic representation of the BiLSTM th word extracted by i , e i represents the entity label of the i th word, CRF represents the conditional random field used to calculate the label sequence probability; E represents all entity label sequences extracted from the sentence S ; P(E | S) represents the probability of generating the entity label sequence S E under the condition of the given input sentence;

[0084] Extract the semantic relationship between entities through the relationship classification model:

[0085] ;

[0086] Among them, h i , h j is the context feature representation of the entities e i , e j , r is the relationship category, W represents the weight matrix, b is the bias term;

[0087] Steps 1-4: Real-time stream data parsing. Through Streaming the model, low-latency data parsing is achieved. Streaming The model is shown by the following formula:

[0088] Event ( D ) = {transaction ID , amount, timestamp, …};

[0089] where D is the transaction data stream;

[0090] Steps 1-5: Relationship extraction. Use Attention - based the model to extract the semantic relationships between entities. The specific formula is as follows:

[0091] ;

[0092] where, Q , K , V respectively represent the query, key, and value vectors of the entity. d represents the vector dimension. R represents the final attention value. i and j both represent the index positions. T is the transpose operation.

[0093] In this embodiment, an Attention-based model is used. First, the similarity between the query and the key is calculated for each pair of entities to generate the attention weight (RRR). Then, the relationship representation between entities is obtained through weighted summation. This representation reflects the semantic association between entities and is further used for the relationship extraction task to identify the specific relationships between entities.

[0094] In the relationship extraction task, the query and the key are usually representations of different entities. For example, assuming that the relationship between two entities in the text needs to be extracted, the word vectors or context vectors of these two entities can be used as the query and the key.

[0095] By calculating the similarity between the query vector and the key vector, the attention weight matrix is obtained. This similarity reflects the correlation between entities in the semantic space.

[0096] Based on the attention weight (i.e., R), the value vectors are weighted and summed, and finally the relationship representation between entities is extracted. This representation contains the semantic relationship between entities.

[0097] Step 2: Establish a semantic annotation module. The semantic annotation module automatically adds semantic labels to the data through zero-shot learning and few-shot learning methods;

[0098] When performing Step 2, the specific steps are as follows:

[0099] Step 2-1: Extract entities and relationships from the data parsing module;

[0100] Step 2-2: Generate predefined labels. Based on LLM the model infers the business scenario to generate a set of possible semantic labels, specifically as follows:

[0101] l = LLM ( Context );

[0102] Among them, l represents the generated semantic label, Context represents the business scenario context;

[0103] The obtained set of semantic labels is L = { l 1, l 2,..., l n};

[0104] Step 2-3: Use the zero-shot learning model to match entities and relationships to the optimal labels. The specific formula is as follows:

[0105] P ( l ∣ e ) = Softmax ( Sim ( h e , h l ));

[0106] Among them, Sim is the cosine similarity function, h e 、 h l are the semantic embeddings of the entity and the label respectively;

[0107] Zero-shot learning refers to the ability of the model to make predictions or inferences without labeled data for a specific task. Usually, it performs inferences by converting the task into a general text representation (such as through a large language model LLM).

[0108] In the zero-shot learning model, P ( l ∣ e ) represents that under the zero-shot condition, the entity e is assigned the labell probability

[0109] Step 2-4: Few-shot learning correction, using a small amount of labeled data to fine-tune the results obtained by the zero-shot learning model.

[0110] In this embodiment, in the case of having a small amount of labeled data, the prediction results are corrected by fine-tuning the zero-shot learning model. At this time, the few-shot data (usually labeled entity or relation pairs) are used to adjust the parameters of the model so that it can make more accurate predictions based on these few samples.

[0111] The fine-tuning process is to compare the output of the model with the labels of a small amount of labeled data by adjusting the loss function. For example, using cross-entropy loss to update the parameters of the model:

[0112] ;

[0113] where y i is the distribution of the true labels (few-shot labels),

[0114] is the label distribution predicted by the model.

[0115] The specific steps of fine-tuning are as follows:

[0116] Step 2-4-1: Train the model with a small amount of labeled data, calculate the gap between the prediction result and the true label, and optimize it using the cross-entropy loss function;

[0117] Step 2-4-2: Use the mini-batch stochastic gradient descent (SGD) optimizer to update the model parameters;

[0118] Step 2-4-3: Use the representation generated by the pre-trained large language model (LLM) as the initial state, and adjust the model parameters so that the model can better adapt to specific tasks.

[0119] Step 3: Establish a data processing module, which semantically aggregates the entities with semantic labels and processes them into structured data assets;

[0120] When performing Step 3, it specifically includes the following steps:

[0121] Step 3-1: Retrieve the labeled entity and relation data;

[0122] Step 3-2: Semantic aggregation, using an improved DBSCAN clustering algorithm to classify entities with similar semantics. The improved DBSCAN clustering algorithm is as follows:

[0123] Sim (x , y ) > ϵ and ∣ x − y ∣ < Δ ;

[0124] Among them, Sim represents a similarity function, specifically the cosine similarity function, x and y are both input entities, ϵ is the semantic similarity threshold, Δ is the numerical difference threshold; if the value obtained by the improved DBSCAN clustering algorithm is true, then x and y are of the same class;

[0125] Step 3-3: Pipeline generation, design an automated data processing pipeline, which is divided into three sub-processes: cleaning, integration, and standardization;

[0126] In this embodiment, the goal of pipeline generation is to design an automated data processing process to process data through three sub-processes: cleaning, integration, and standardization. These sub-processes can ensure that the original data can be converted into structured data suitable for data assetization. The following are the specific steps and functions of each sub-process:

[0127] Data cleaning is the first step in data preprocessing, aiming to delete or correct errors, duplicates, missing values, and inconsistencies in the data. It ensures that the input data meets the quality standards and provides a clean basis for subsequent processing. The specific tools used are the Pandas library in Python (such as functions like.drop_duplicates(),.fillna(), etc.), and regular expressions are used to clean text data. The specific operations are as follows:

[0128] Duplicate removal: Delete duplicate data records to ensure that each piece of data is unique. For example, remove duplicate records of the same transaction or duplicate paragraphs of the same text;

[0129] Missing value filling: For missing fields, interpolation, mean, median filling, or other appropriate strategies can be used for supplementation. For text data, it may be necessary to fill in missing entities or relationships through context;

[0130] Outlier detection: Detect and correct outliers in the data. For example, for transaction data, abnormal fluctuations in amounts may need to be corrected or deleted;

[0131] Format Standardization: Uniformly process fields such as dates, times, currencies, etc., to ensure data consistency throughout the processing. For example, convert all dates to the "YYYY-MM-DD" format.

[0132] The purpose of data integration is to combine data from different sources into a unified dataset, ensuring that all relevant data can be processed within the same data model. This step typically involves merging multiple data sources (such as databases, log files, APIs, etc.). The specific tools used are database query languages (SQL) for cross-table and cross-database integration, ETL tools (such as Apache Nifi, Talend, etc.) for data extraction, transformation, and loading, and graph databases (such as Neo4j) for relationship modeling and merging. The specific operations are as follows:

[0133] Data Source Merging: Merge data sources from different systems or formats. For example, merge entities extracted from enterprise documents with data from real-time transaction flows to form a complete database;

[0134] Entity Mapping: Map the same entities in different data sources. For example, extract customer IDs and names from different databases and ensure they refer to the same customer;

[0135] Relationship Fusion: Process relationships in different data sources to ensure they remain consistent after merging. For instance, if the transaction relationships of "Customer A" are mentioned in both data sources, these pieces of information need to be merged into a unified representation;

[0136] Time Alignment: For multi-source time series data, it is necessary to align their timestamps to ensure the consistency of data from different sources in the time dimension.

[0137] The purpose of data standardization is to convert data into a unified format and structure suitable for analysis and modeling. At this stage, we convert data of different dimensions and formats into a unified standard. The tools used are StandardScaler or MinMaxScaler in scikit-learn for processing. For preprocessing text data, NLP toolkits such as spaCy or NLTK can be used for lexical normalization. For field mapping and encoding of structured data, Pandas and SQL can achieve data cleaning and transformation. The specific operation steps are as follows:

[0138] Field Standardization: Unify all field names, units, data types, etc. For example, map the "transaction_amount" and "amount" fields to the same standard field "amount".

[0139] Data Scaling and Normalization: Standardize numerical data (e.g., standardize the amount field to ensure it has the same scale) to ensure that numerical values from different data sources can be compared. For example, use the Z-score standardization or min-max normalization method.

[0140] Lexical Standardization: For text data, perform lexical standardization. For example, normalize synonyms (e.g., consider "purchase" and "buy" as the same concept) and remove stop words.

[0141] Coding Standardization: Standardize the encoding of categorical variables to ensure that each category in the data has a consistent representation. For example, map "male" and "female" in the "gender" field to numbers (e.g., 0 and 1), or use one-hot encoding to convert them into multiple binary features.

[0142] Steps 3 - 4: Asset Transformation, introduce causal inference methods, and transform based on the causal relationships between data into high-value data assets. The formula for the causal inference method is as follows:

[0143] Y = αX + δ ;

[0144] where, δ represents the error term, α represents the causal effect, X represents the fund flow, Y represents the performance risk.

[0145] In this embodiment, it is necessary to determine which variables are participants in possible causal relationships. For example, assume we are evaluating the impact of fund flow (X) on performance risk (Y), and select key variables: fund flow, performance risk, etc.

[0146] Based on data analysis and domain knowledge, construct a causal inference model. A simple linear regression model (such as the above formula) can be used to represent the causal relationship, or more variables and control terms can be introduced according to the complexity of the data. For example, other factors (such as market conditions, enterprise scale, etc.) may need to be introduced as control variables into the model.

[0147] Use appropriate causal inference methods to estimate the causal effect. For example, estimate the value of α through a regression model to represent the impact of fund flow on performance risk.

[0148] After estimating the causal effect, the contributions of these effects to the business objectives can be evaluated. For example, if α represents the impact of fund flow on performance risk, then this information can be used to evaluate the level of performance risk for an enterprise or system and rank the data assets based on this evaluation.

[0149] By combining causal effects with other business metrics (such as cash flow, profit, etc.), high-value data assets can be generated. These data assets can help decision-makers formulate risk management strategies, optimize cash flow, etc.

[0150] Step 4: Establish an asset evaluation module. The asset evaluation module quantitatively evaluates the value of data resources, generates a data asset statement disclosure report, and provides visual display;

[0151] When implementing Step 4, it specifically includes the following steps:

[0152] Step 4-1: Obtain the processed structured data assets;

[0153] Step 4-2: Weight assignment. Based on LLM the scoring system of the model, combine business rules to assign weights to data assets and calculate the comprehensive score of data assets:

[0154] ;

[0155] Among them, w i is the business metric weight, v i ( d ) is the score value of the i th indicator; d represents the structured data asset, n represents the number of indicators, S ( d ) represents the comprehensive score of the data asset;

[0156] Step 4-3: Asset ranking. Rank the data assets according to the comprehensive score to screen out high-value assets;

[0157] Step 4-4: User visualization. Generate a data asset statement disclosure report and provide visual display. The specific content includes the basic information of the data asset and the value evaluation results of the data asset.

[0158] Step 5: Establish a dynamic optimization module. The dynamic optimization module continuously optimizes the data assetization process from Step 1 to Step 4.

[0159] When implementing Step 5, the specific steps are as follows:

[0160] Step 5-1: Obtain user feedback and system evaluation metrics;

[0161] In this embodiment, the collection of user feedback is the basis for optimizing the model. There are mainly the following ways:

[0162] Direct user feedback: Through the interaction between users and the system, directly collect users' feedback on the data assetization process.

[0163] Direct user feedback can be achieved in the following ways:

[0164] Feedback button / form: Provide a feedback button or form in the system, allowing users to submit their opinions, questions, or evaluations of the data asset assessment results at any time (e.g., scoring, suggestions, improvement requirements, etc.);

[0165] Scoring system: According to users' usage experiences, allow users to score each data asset generated by the system (e.g., 0 - 5 points) and add additional comments;

[0166] Questionnaire: Regularly send questionnaires to users to collect their feedback on data asset management, data assessment, data processing, and other aspects.

[0167] Indirect user feedback: User feedback can be obtained through the following indirect methods:

[0168] User behavior analysis: By analyzing users' behaviors in the system (such as click-through rate, access frequency, dwell time, etc.), infer users' interests and attentions to certain data assets. Frequent viewing and operation of specific assets by users can be regarded as positive feedback, while neglect of certain data assets can be regarded as negative feedback.

[0169] Decision result feedback: When users make decisions in the system (such as selecting specific data assets), record the decision results and use them to evaluate the quality of the current system output. If the decision results do not match users' expectations, it can reflect that the system needs further optimization.

[0170] System evaluation metrics are used to quantify the current performance of the system, evaluate the accuracy, reliability, efficiency, etc. of the evaluation model. Key metrics include but are not limited to:

[0171] Prediction accuracy: Measure the accuracy of the model's data asset assessment results, that is, the matching degree between the prediction results and the actual results;

[0172] Recall rate and precision rate: In the process of data processing, semantic annotation, and relation extraction, the recall rate and precision rate of the system are also very critical evaluation metrics;

[0173] Response time: Evaluate the response speed of the system to users' requests, especially the real-time nature of data processing and assessment;

[0174] Computing efficiency: Evaluate the efficiency of the system in processing a large amount of data, such as processing speed, resource consumption (CPU, memory, etc.);

[0175] Satisfaction score: Collect the satisfaction of users with the data asset management system through user evaluations, rating systems, or questionnaires;

[0176] User activity: Measure the user activity of the system and judge the popularity of the system by monitoring information such as user logins and usage frequencies.

[0177] Step 5-2: Semantic feedback. Collect feedback data through user interaction and fine-tune and update the LLM model, specifically including:

[0178] Step 5-2-1: Collect a dataset with feedback annotations;

[0179] In this embodiment, a dataset with clear feedback is collected through user interaction. User feedback can be ratings based on data assets, evaluations of model outputs by users, error identifications, requirement descriptions, etc.; Generally, user feedback can include positive (e.g., correct predictions or useful data annotations) and negative feedback (e.g., incorrect predictions or labels that do not meet business requirements). These feedbacks will be annotated on the dataset and used as the input for model fine-tuning.

[0180] For example, users rate "high" or "low" to mark the quality of the model output; users annotate errors or irrelevant information for certain data items.

[0181] Collecting these annotated datasets with feedback will be used as the training data for subsequent fine-tuning.

[0182] Step 5-2-2: Adjust the model objective and add a feedback loss function;

[0183] In this embodiment, the objective of the model is adjusted by introducing a feedback loss function so that it not only optimizes the original task (such as text classification, relation extraction, etc.) but also considers user feedback information. In this embodiment, a loss function needs to be designed to combine the original objective of the model with the feedback-based optimization objective.

[0184] Step 5-2-3: Use mini-batch stochastic gradient descent (SGD) to optimize the model parameters. The specific formula is as follows:

[0185] L fine-tune = L original + βL feedback ;

[0186] Where, L feedback represents the loss based on user feedback, β is the weight coefficient, L fine-tune represents the loss function after fine-tuning,L original represents the original loss function;

[0187] L fine-tune The fine-tuned loss function is used to find a balance between the feedback data and the loss of the original task, and is minimized by SGD L fine-tune , thereby updating the parameters of the model.

[0188] Step 5-3: Model optimization, using the knowledge distillation method to LLM lightweight optimize the model;

[0189] In this embodiment, the knowledge distillation method is a prior art, so it will not be described in detail.

[0190] Step 5-4: Asset update, design an automatic update mechanism to update the data asset library based on real-time feedback, and eliminate low-value or outdated assets;

[0191] In this embodiment, the specific steps of Step 5-4 are as follows:

[0192] Step 5-4-1: First, the system needs to receive real-time feedback information from different sources. Usually, these feedbacks can come from user behavior, system logs, changes in business metrics, etc.;

[0193] For example:

[0194] User behavior feedback: The query, download, modification or usage frequency of certain data assets by users, etc.;

[0195] System performance feedback: The usage rate, access frequency, query duration of data assets, etc.;

[0196] Business feedback: The change in the demand for data assets by the business department or the emergence of new business scenarios.

[0197] Step 5-4-2: Based on the real-time feedback data, design a dynamic asset evaluation mechanism to evaluate the current value of data assets. It includes the following contents:

[0198] Value scoring: Assign a value score to each data asset according to the preset evaluation criteria (for example, data timeliness, quality, relevance, etc.);

[0199] Timeliness evaluation: Judge whether it is outdated according to the update time and usage frequency of the data. Data with poor timeliness should be considered of low value;

[0200] User demand matching degree: Evaluate the fit between the data asset and the actual demand according to the current business demand and user feedback;

[0201] Step 5-4-3: Based on real-time feedback and evaluation results, set some rules to determine the update strategy of the data asset library. The rules include:

[0202] Elimination of assets with value below the threshold: When the comprehensive score or value of a certain data asset is lower than the set threshold, the asset can be automatically eliminated or marked as obsolete. For example, if the query frequency of a certain data drops to a certain proportion or the relevance is insufficient, its value can be considered to have decreased;

[0203] Elimination of assets with expired timeliness: Regularly check the update time of each data asset. If some data has not been updated for a long time and does not reflect new business requirements, these data can be considered obsolete;

[0204] Priority promotion: For data with higher value or meeting current business requirements, its priority can be automatically promoted to ensure that it is processed or used first.

[0205] Step 5-4-4: Once the evaluation and elimination or promotion rules of the data assets are determined, automatic updates can be achieved in the following ways:

[0206] Scheduled tasks: Set data update tasks for regular inspection and evaluation, and regularly re-evaluate and update the data asset library;

[0207] Event trigger: When there is new user feedback or business requirement changes, trigger real-time update tasks to automatically adjust the asset value in the data asset library;

[0208] Integrated data quality check: Each time data is updated, combine data quality checks (such as data integrity, accuracy, etc.) to ensure that the updated data meets the expected quality requirements;

[0209] Step 5-4-5: To avoid losing historical data due to excessive automated updates, a version control mechanism can be implemented:

[0210] Version management: Each time a data asset is updated, save the data version before the update to facilitate tracing and comparing data changes at different time points;

[0211] Data history record: Retain the historical record of each asset update, including the historical data of deletion, modification, and re-annotation, to ensure that decisions on obsolete or low-value data are well-founded.

[0212] Step 5-5: Adapt to multi-task scenarios and enhance the adaptability of the model to multi-business scenarios through Multi-task Learning.

[0213] When performing Step 5-5, the Multi-task Learning enhanced model specifically includes:

[0214] Step 5-5-1: Identify the task types of multiple tasks to be processed and define the goals of each task. For example:

[0215] Data classification task: Identify data types, label classification, etc.;

[0216] Data regression task: Predict the numerical results of data, such as risk score, transaction amount prediction, etc.;

[0217] Relationship extraction task: Extract the relationships between entities in the data;

[0218] Data evaluation task: Evaluate the value, accuracy, etc. of data;

[0219] Step 5-5-2: In multi-task learning, the Multi-task Learning enhanced model includes a shared layer and task-specific layers:

[0220] Shared layer: By sharing part of the network structure (e.g., the convolutional layer, fully connected layer, etc. in the first few layers), the model can learn the common features between multiple tasks. This shared knowledge is the basis for different tasks to help each other;

[0221] Task-specific layer: Each task has its own dedicated layer for specific output prediction. For example, for a classification task, the last output layer is a classifier; for a regression task, the last output layer is a regression model;

[0222] Step 5-5-3: Define an appropriate loss function and combine the losses of multiple tasks with weights. The loss functions for each task can be different, and the specific design can refer to the following methods:

[0223] Task-independent loss: The loss function for each task is calculated separately. For example:

[0224] The cross-entropy loss function (Cross-Entropy Loss) is used for classification tasks;

[0225] The mean squared error loss function (Mean Squared Error) is used for regression tasks;

[0226] Weighted loss: The losses of different tasks can be weighted by weight coefficients and combined into an overall loss function. The formula for the overall loss is as follows:

[0227] ;

[0228] where L i is the loss function of task i and λi is the weight of each task. By adjusting the weight coefficient, the model can balance the learning objectives among different tasks.

[0229] Step 5-5-4: In multi-task learning, there may be certain correlations between different tasks. The higher the correlation between tasks, the more useful the shared knowledge is. The shared features or network structure levels can be determined by analyzing the correlations between tasks. For example:

[0230] Strong task association: If tasks are highly correlated, more model layers can be shared, reducing the design of independent models;

[0231] Weak task association: If the correlation between tasks is low, fewer shared layers can be used, retaining more independent model structures.

[0232] Step 5-5-5: Adopt a joint training strategy to optimize the shared layers and task-specific layers. Specifically, by optimizing the loss functions of all tasks in each training step, a common network model is trained. The model will benefit from learning multiple tasks simultaneously.

[0233] Step 5-5-6: Prepare the dataset for each task, ensuring that the data for each task can be merged or share the same input features (for example, different labels of the same dataset can be used as the targets for different tasks);

[0234] Design a neural network model with shared layers and task-specific layers. The shared layers can be convolutional layers, fully connected layers, etc., and the task-specific layers are usually output layers.

[0235] Training process:

[0236] Sum the weighted loss functions of each task to obtain the total loss function;

[0237] Use the total loss for backpropagation and optimization, gradually adjusting the parameters of the network;

[0238] Regularly evaluate the performance of each task to ensure that the objectives of each task can be balanced and optimized.

[0239] In this embodiment, through multi-task learning, the model can effectively adapt to multiple task scenarios. For example: in financial data analysis, tasks such as trading prediction, credit scoring, and risk assessment can be carried out simultaneously; in enterprise data management, multiple tasks can include data classification, data cleaning, and data value evaluation.

[0240] A method for incorporating data assets into the balance sheet based on large language models according to the present invention solves the technical problems of efficient parsing, semantic annotation, data processing, asset evaluation, and dynamic optimization of text data and real-time stream data. By adopting zero-shot learning and few-shot learning methods, the present invention reduces the need for manual annotation, realizes automated annotation and processing of data, can efficiently process large-scale text data and real-time stream data, meets the low-latency requirements, combines the LLM model and business rules for automated evaluation of data assets, provides comprehensive scoring and visualization tools to help users intuitively understand the value of data assets, the system can continuously optimize according to user feedback and real-time data streams, continuously improve the accuracy and adaptability of the model through multi-task learning and knowledge distillation methods, and uses an improved DBSCAN clustering algorithm and causal inference method to perform semantic aggregation and value transformation on data, endowing higher commercial value to data assets.

Claims

1. A method for entering data assets into a table based on a large language model, characterized in that: The steps include: Step 1: Establish a data parsing module. The data parsing module obtains text data and real-time streaming data, segments the text data, uses the BiLSTM-CRF model to identify entities, and extracts entities, attributes, and their relationships. Step 2: Establish a semantic annotation module, which automatically adds semantic labels to the data through zero-shot learning method and few-shot learning method; Step 3: Establish a data processing module, which aggregates entities with semantic tags and processes them into structured data assets; Step 4: Establish an asset evaluation module, which quantitatively evaluates the value of data resources, generates a data asset disclosure report, and provides a visual display; Step 5: Establish a dynamic optimization module, which continuously optimizes the data assetization process from step 1 to step 4; When executing step 5, the specific steps are as follows: Step 5-1: Obtain user feedback and system evaluation indicators; Step 5-2: Semantic feedback: Collect feedback data through user interaction and fine-tune the LLM model, including: Step 5-2-1: Collect a dataset with feedback annotations; Step 5-2-2: Adjust the model objective and add feedback loss function; Step 5-2-3: Use mini-batch stochastic gradient descent SGD to optimize model parameters. The specific formula is as follows: L fine-tune =L original +βL feedback ; Among them, L feedback represents the loss based on user feedback, β is the weight coefficient, and L fine-tune represents the loss function after fine-tuning, L original represents the original loss function; Step 5-3: Model optimization, use the knowledge distillation method to perform lightweight optimization on the LLM model; Step 5-4: Asset update: design an automatic update mechanism to update the data asset library based on real-time feedback and remove low-value or outdated assets; Step 5-5: Adapt to multi-task scenarios and learn multi-task Learning enhances the model's adaptability to multiple business scenarios.

2. The method for entering data assets into a table based on a large language model according to claim 1, characterized in that: When executing step 1, the specific steps include: Step 1-1: Data input, obtaining text data and real-time streaming data; text data includes corporate documents, logs and contracts; real-time streaming data includes transaction records and reports; Step 1-2: Semantic segmentation: segment the long text data into paragraphs and sentences. The segmentation function is as follows: Seg(T)={S1,S2,…,S n } Among them, T represents text data, S n represents the segmented sentences, and n represents the number of segmented sentences; Step 1-3: Entity recognition, combined with BiLSTM-CRF for key entity extraction: Among them, S represents the input sentence, h i represents the semantic representation of the i-th word extracted by BiLSTM, e i represents the entity label of the i-th word, and CRF represents conditional random field, which is used to calculate the probability of label sequence; Extracting semantic relations between entities through relational classification models: Among them, h i ,h j It is entity e i , e j The context feature representation of r is the relationship category, W represents the weight matrix, b is the bias term; Step 1-4: Real-time streaming data analysis, low-latency data analysis is achieved through the Streaming model. The Streaming model is shown in the following formula: Event(D) = {transaction ID, amount, timestamp, ...}; Where D is the transaction data stream; Step 1-5: Relationship extraction, using the Attention-based model to extract the semantic relationship between entities. The specific formula is as follows: Among them, Q, K, V represent the query, key and value vectors of the entity respectively, d represents the vector dimension, R represents the final attention value, i and j represent the index position, and T is the transposition operation.

3. The method for entering data assets into a table based on a large language model according to claim 1, characterized in that: When executing step 2, the specific steps are as follows: Step 2-1: Entities and relations extracted from the data parsing module; Step 2-2: Generate predefined tags. Reason about the business scenario based on the LLM model and generate a possible set of semantic tags. The details are as follows: l = LLM(Context); Among them, l represents the generated semantic label, and Context represents the business scenario context; The resulting semantic label set is L = {l1, l2, ..., l n }; Step 2-3: Use the zero-shot learning model to match entities and relations to the optimal labels. The specific formula is as follows: P(l∣e)=Softmax(Sim(h e ,h l )); Among them, Sim is the cosine similarity function, h e 、h l They are the semantic embeddings of entities and tags respectively; Step 2-4: Few-shot learning correction, use a small amount of labeled data to fine-tune the results obtained by the zero-shot learning model.

4. The method for entering data assets into a table based on a large language model according to claim 1, characterized in that: When executing step 3, the specific steps include: Step 3-1: Retrieve the labeled entity and relationship data; Step 3-2: Semantic aggregation: Use the improved DBSCAN clustering algorithm to classify entities with similar semantics. The improved DBSCAN clustering algorithm is as follows: Sim(x, y)>∈and∣xy∣<Δ; Wherein, Sim represents a similarity function, specifically a cosine similarity function, x and y are both input entities, ∈ is a semantic similarity threshold, and Δ is a numerical difference threshold; if the value obtained by the improved DBSCAN clustering algorithm is true, then x and y are of the same class; Step 3-3: Pipeline generation, designing an automated data processing pipeline, which is divided into three sub-processes: cleaning, integration, and standardization; Step 3-4: Asset conversion, introduce causal inference method, and convert it into high-value data assets based on the causal relationship between data. The formula of causal inference method is as follows: Y = αX + δ; Among them, δ represents the error term, α represents the causal effect, X represents the capital flow, and Y represents the performance risk.

5. The method for entering data assets into a table based on a large language model according to claim 1, characterized in that: When executing step 4, the specific steps include: Step 4-1: Obtain processed structured data assets; Step 4-2: Weight allocation: Based on the LLM model scoring system, weights are allocated to data assets in combination with business rules to calculate the comprehensive score of data assets: Among them, w i is the business indicator weight, v i (d) is the score value of the i-th indicator; d represents the structured data asset, n represents the number of indicators, and S(d) represents the comprehensive score of the data asset; Step 4-3: Asset sorting: prioritize data assets according to comprehensive scores and select high-value assets; Step 4-4: User visualization, generate a data asset disclosure report and provide a visual display, including the basic information of the data assets and the value assessment results of the data assets.

Citation Information

Patent Citations

  • Automatic tagging method for enterprise data assets based on artificial intelligence large language model

    CN117235292A

  • System and method for inputting data assets into table based on multi-level knowledge graph

    CN118710084A