Financial data category prediction method and device, equipment, storage medium and product
By generating knowledge graph query statements and performing text graph conversion methods, the problem of inaccessibility and lack of transparency of model parameters in financial data classification is solved, the comprehensiveness and accuracy of financial data categories is achieved, and a transparent classification decision-making process is provided.
Patent Information
- Application Number
- CN202510015556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems in the classification of financial data, such as inaccessibility and lack of transparency in model parameters, making it difficult to adapt to data classification tasks in specific fields, and in high-risk data classification, classification decisions are crucial.
By obtaining the financial data to be classified, a knowledge graph query statement is generated, and inputting it into the pre-trained financial data question and answer model, the knowledge graph entity relationship is obtained. Then, the text graph conversion of the entity relationship is performed, the node set, edge set, feature matrix and adjacency matrix are generated, and finally input into the financial data classification model to obtain the financial data category.
It realizes comprehensiveness and accuracy of financial data categories, provides a transparent and clear classification decision-making process, and enhances the interpretability of the decision-making process.
Smart Images

Figure CN119939424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial technology, and in particular to financial data category prediction methods, devices, equipment, storage media and products. Background Art
[0002] At present, artificial intelligence is developing rapidly and has made great achievements in many fields, such as natural language processing, image processing, data classification, etc. Among them, data classification is one of the most basic tasks in natural language processing, and is the basis of many tasks such as recommendation tasks and question-answering systems. As the amount of financial data increases at a geometric rate, the research on data classification tasks, which is the basis of many applications, becomes increasingly important.
[0003] Most existing data classification methods use generative artificial intelligence to identify and classify data. However, as a black box model, generative artificial intelligence has two major limitations: first, the inaccessibility of model parameters makes it impossible for generative artificial intelligence models to be more flexibly refined on specific data sets, making it difficult to adapt to data classification tasks in specific fields; second, generative artificial intelligence lacks transparency in the decision-making process of data classification, making it impossible to explain the classification results of data. At the same time, generative artificial intelligence is not open source, and existing decision explanation methods cannot be applied. In high-risk data classification, classification decisions are crucial.
[0004] Therefore, how to predict the categories of financial data, realize the classification of financial data, ensure the comprehensiveness and accuracy of data classification, and at the same time make a more transparent representation of the decision-making process of data classification has become an urgent problem that technical personnel need to solve. Summary of the invention
[0005] The present invention provides a financial data category prediction method, apparatus, device, storage medium and product to ensure the comprehensiveness and accuracy of the prediction of financial data categories and to present the decision-making process of financial data classification in a more transparent and clear manner.
[0006] According to one aspect of the present invention, a method for predicting financial data categories is provided, comprising:
[0007] Obtaining financial data to be classified, and generating a knowledge graph query statement based on the financial data to be classified;
[0008] Inputting the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model;
[0009] Performing text-graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix;
[0010] The node set, the edge set, the feature matrix and the adjacency matrix are input into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
[0011] According to another aspect of the present invention, there is provided a financial data category prediction device, comprising:
[0012] A query statement generation module, used to obtain the financial data to be classified, and generate a knowledge graph query statement based on the financial data to be classified;
[0013] An entity relationship acquisition module, used to input the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model;
[0014] A text-to-graph conversion module, used to perform text-to-graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix;
[0015] The financial data category acquisition module is used to input the node set, the edge set, the feature matrix and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the financial data category prediction method described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the financial data category prediction method described in any embodiment of the present invention when executed.
[0021] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the financial data category prediction method described in any embodiment of the present invention is implemented.
[0022] The technical solution of the embodiment of the present invention obtains the financial data to be classified, and generates a knowledge graph query statement based on the financial data to be classified; inputs the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model; performs text graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix; inputs the node set, the edge set, the feature matrix, and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model. This technical solution can extract detailed and structured data content from the financial data to be classified, and generate a knowledge graph entity relationship. The generated knowledge graph entity relationship can also be converted into a corresponding text graph to obtain the data category of the financial data, thereby ensuring the comprehensiveness and accuracy of the financial data classification results and providing a transparent and clear classification decision process.
[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 is a flowchart of a financial data category prediction method provided according to Embodiment 1 of the present invention;
[0026] Figure 2 is a flowchart of a financial data category prediction method provided according to Embodiment 2 of the present invention;
[0027] Figure 3 is a schematic diagram of the structure of a financial data category prediction device provided according to Embodiment 3 of the present invention;
[0028] Figure 4 It is a schematic diagram of the structure of an electronic device for implementing the financial data category prediction method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment 1
[0032] Figure 1 A flowchart of a method for predicting financial data categories is provided for the first embodiment of the present invention. This embodiment can be applied to classify financial data under different financial business classification demand scenarios. The method can be executed by a financial data category prediction device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110. Obtain the financial data to be classified, and generate a knowledge graph query statement based on the financial data to be classified.
[0034] Financial data can be data generated by an enterprise or individual during financial transactions, or it can be the information used by an enterprise or individual during financial transactions, including company information, personal information, and transaction information. Company information can include company name, industry, stock code, etc.; personal information can include name, position, company, and transaction behavior; transaction information can include transaction amount, transaction time, participants, and transaction type (buy, sell, transfer, etc.).
[0035] Specifically, the generated knowledge graph query statement can be entity recognition and relationship extraction for the financial data to be classified, where entities can represent objects or concepts with clear meanings, such as companies, tasks, places, transactions, etc.; relationships can be used to represent the association between different entities, such as the "work at" relationship between people and companies, the "production" relationship between products and companies, the "occurrence at" relationship between time and events, etc. Exemplarily, based on the financial data to be classified, the generated knowledge graph query statement can be "query whether user A works at company a".
[0036] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with relevant laws, regulations and standards in relevant regions.
[0037] S120. Input the knowledge graph query statement into the pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model.
[0038] Among them, the financial data question-answering model can be a digital model used to identify and extract named entities and corresponding relationships in classified financial data based on knowledge graph query statements.
[0039] Specifically, when the knowledge graph query statement is input into the financial data question-answering model, the entities in the classified financial data can be identified according to the knowledge graph query statement, and the associations between different entities can be extracted, and then triples can be generated according to the identified entities and the relationships between entities. The triples can be data sets used to describe entities and relationships in the knowledge graph. Exemplarily, continuing the above example, the knowledge graph query statement is "query whether user A works for company a", then the triple "user, work at, company a" can be obtained through the output of the financial data question-answering model "user A works at company a", where entities can include user A and company a, and the relationship between entities can be work at.
[0040] S130. Perform text graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix and an adjacency matrix.
[0041] Specifically, the text graph conversion of the knowledge graph entity relationship can be to express the entities and relationships in the knowledge graph in the form of a graph structure. The text graph representation can be G = {V, E, X, A}, where G can be represented as a text graph, V can be represented as a node set, X can be represented as a feature matrix, and A can be represented as an adjacency matrix.
[0042] Among them, the node set can be represented as a set for recording all entities; the edge set can be represented as a set for recording the relationship between entities; the feature matrix can be used to represent the node features of each node, and the specific node features can be the attribute information corresponding to the entity; the adjacency matrix can represent the connection relationship between each node in the text graph.
[0043] Optionally, the knowledge graph entity relationships include entity sets and relationship sets; accordingly, the knowledge graph entity relationships are converted from text to graph to obtain node sets, feature matrices and adjacency matrices, including: converting the knowledge graph entity relationships from text to graph to generate node sets based on the entity sets; generating edge sets based on the relationship sets; generating feature matrices based on node feature information of each node in the node set; determining the association relationships between each node based on the node sets and edge sets, and generating an adjacency matrix based on the association relationships.
[0044] Among them, the entity set can be a set used to record the entities in the data to be classified; the relationship set can be a set that records the relationships between different entities.
[0045] Among them, generating a node set including an entity set and a relationship set may specifically be to extract entities and relationships in the acquired triples, generate an entity set and a relationship set according to the extracted entities and relationships, respectively, and then generate a node set according to the entity set and the relationship set.
[0046] For example, if the first triple (person M, founded, XX technology company) and the second triple (person M, birthplace, XX country) are obtained, the entity set (person M, XX technology company, XX country) and the relationship set (founded, birthplace) can be obtained, and the node set can be obtained as (person M, XX technology company, XX country). Further, an edge set is generated according to the relationship set, and the edge set can be determined as (founded, birthplace).
[0047] Specifically, generating the feature matrix can be to define the attribute information of the entity corresponding to each node in the node set, and determine the feature matrix X according to the definition result of the entity attribute information. For example, continuing the above example, the character M can be defined as: person (1), founder (1), male (1); XX technology company: company (1), technology company (1); XX country: country (1). Then, the feature matrix X can be generated according to the node set (character M, XX technology company, XX country), specifically:
[0048]
[0049] Among them, generating the adjacency matrix based on the association relationship can specifically be based on whether there is an association relationship between each node in the node set, and then based on the association relationship between entities, generating the adjacency matrix A. Exemplarily, continuing the above example, in the first triple (person M, founded, XX technology company), it can be determined that there is an association relationship "founded" between person M and XX technology company, and then it can be determined that there is an edge between person M and XX technology company; it can also be determined in the second triple (person M, birthplace, XX country) that there is an association relationship "birthplace" between person M and XX country, and then it can be determined that there is an edge between person M and XX country, and then the adjacency matrix A can be generated. Specifically,
[0050]
[0051] This technical solution can convert the knowledge graph into a text graph and obtain the node set, edge set, feature matrix and adjacency matrix in the text graph. It can make the data structure of the classified financial data more intuitive, simplify the presentation form of the financial data to be processed, and improve the efficiency of data classification and the display effect of the data classification decision process.
[0052] S140, inputting the node set, edge set, feature matrix and adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
[0053] Specifically, after the knowledge graph is converted into a text graph, the node set, edge set, feature matrix and adjacency matrix extracted from the financial data to be classified can be used as the model input of the financial data classification model. Furthermore, the data category of the financial data to be classified can be output by the financial data classification model. The financial data classification model can be a graph convolutional neural network model or a graph attention network model.
[0054] Optionally, before inputting the knowledge graph query statement into the pre-trained financial data question and answer model to obtain the knowledge graph entity relationship output by the financial data question and answer model, it also includes: obtaining the financial sample data to be trained, and determining the open source question and answer model; the open source question and answer model is pre-trained based on the open source sample data; using the financial sample data to be trained and the open source sample data to train the open source question and answer model to obtain a benchmark question and answer model; using the financial sample data to be trained to train the benchmark question and answer model to obtain the financial data question and answer model.
[0055] The financial sample data to be trained may be sample data obtained from the financial data to be classified and used for model training. The open source sample data may be a publicly available benchmark text classification data set mainly used for model training of a financial data classification model.
[0056] Specifically, the pre-built network model can be open source trained with open source sample data, so that the network model can be applied to the application scenario of data classification of financial data, and an open source question-answering model is obtained. Then, the financial sample data to be trained and the open source sample data are integrated, the model training sample data is expanded, and the open source question-answering model is trained with the expanded training sample data to obtain a benchmark question-answering model. Then, the benchmark question-answering model is trained with the sample data to be trained, so as to obtain a financial data classification model.
[0057] Exemplarily, the pre-built neural network model can be open source trained by the obtained open source sample data, such as the neural network model can be trained by the obtained 11,314 financial sample data, and the model can be tested by 7,532 financial sample data to obtain an open source question-answering model. Furthermore, financial data training samples can be selected from the financial literature disclosed in the historical period, and a data set with single-label classification can be formed, such as 3,357 financial sample data to be trained and 4,043 model test data. Furthermore, the 11,314 financial sample data obtained and the 3,357 financial sample data to be trained selected from the disclosed financial literature can be sample fused to train the open source question-answering model and obtain a benchmark question-answering model. Then, the benchmark question-answering model can be trained by selecting 3,357 financial sample data to be trained from the disclosed financial literature, thereby obtaining a financial data classification model.
[0058] After obtaining the open source question-answering model trained with the open source sample data, this technical solution uses the financial sample data to be trained and the open source sample data to train the open source question-answering model to obtain a benchmark question-answering model; and then trains the benchmark question-answering model with the sample data to be trained to obtain a financial data classification model. The open source question-answering model can be trained twice with the financial sample data to be trained and the open source sample data, further improving the comprehensiveness and accuracy of financial data classification using the generative model.
[0059] The technical solution of the embodiment of the present invention obtains the financial data to be classified, and generates a knowledge graph query statement based on the financial data to be classified; inputs the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model; performs text graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix; inputs the node set, the edge set, the feature matrix, and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model. This technical solution can extract refined and structured data content from the financial data to be classified, and generate a knowledge graph entity relationship. The generated knowledge graph entity relationship can also be converted into a corresponding text graph to obtain the data category of the financial data, thereby ensuring the comprehensiveness and accuracy of the financial data classification results, providing a transparent and clear classification decision process, and enhancing the explainability of the decision process.
[0060] Embodiment 2
[0061] Figure 2 This is a flow chart of a financial data category prediction method provided in the second embodiment of the present invention. Based on the above embodiment, this embodiment further optimizes the above financial data category prediction method.
[0062] Furthermore, before the step of "obtaining the financial data to be classified, and generating a knowledge graph query statement based on the financial data to be classified", add the step of "obtaining the financial sample data to be trained, and determining the true classification label of the financial sample data to be trained; performing data processing on the financial sample data to be trained, and generating a text graph including a sample node set, a sample edge set, a sample feature matrix, and a sample adjacency matrix; inputting the text graph into a pre-built graph convolutional network model to obtain the predicted classification result output by the graph convolutional network model; training the graph convolutional network model based on the true classification label and the predicted classification result, and updating the model weight matrix in the graph convolutional network model until the model training end condition is met to obtain a financial data classification model." to improve the method of predicting the category of financial data. Figure 2 As shown, the method includes:
[0063] S210: Obtain financial sample data to be trained, and determine true classification labels of the financial sample data to be trained.
[0064] The real classification label may be the financial data classification category corresponding to the financial sample data to be trained, and the specific real classification label may be the classification label used by the technician to classify the financial sample data to be trained based on experience.
[0065] Specifically, when extracting the financial sample data to be trained from the financial data to be classified, the data category of the financial data of the financial sample data to be trained may be labeled.
[0066] S220, performing data processing on the financial sample data to be trained to generate a text graph including a sample node set, a sample edge set, a sample feature matrix, and a sample adjacency matrix.
[0067] Specifically, data processing can be to process the extracted financial sample data to be trained using a natural language model. Specifically, the acquired financial sample data to be trained can be subjected to grammatical correction, spelling error checking, synonym replacement, and text sentence structure clarification to ensure the accuracy and coherence of the text of the financial sample data to be trained.
[0068] Furthermore, the entity relationships of the knowledge graph of the financial sample data to be trained can be obtained by identifying and extracting named entities and extracting relationships of the financial sample data to be trained. Then, the knowledge graph entity relationships are converted into text graphs to generate sample node sets, sample edge sets, sample feature matrices, and sample adjacency matrices. The specific generation of the knowledge graph entity relationships of the financial sample data to be trained and the conversion and generation process of the text graph can refer to the relevant description in the aforementioned embodiment 1, and will not be repeated here.
[0069] S230: Input the text graph into a pre-built graph convolutional network model to obtain a predicted classification result output by the graph convolutional network model.
[0070] Among them, the predicted classification result can be inputting the text graph of the financial sample data to be trained into the graph convolutional network model, and outputting the data classification result of the financial sample data to be trained.
[0071] Specifically, the graph convolutional network model can be used as the execution body for executing data classification of the financial sample data to be trained, and the data classification results of the financial sample data to be trained can be obtained by using the sample node set and sample edge set extracted from the financial sample data to be trained.
[0072] Optionally, the number of graph convolutional layers in the graph convolutional network model is and only is one.
[0073] Specifically, in this technical solution, sufficient data classification information can be obtained by extracting text graphs from the training financial sample data. Therefore, the number of graph convolution layers can be set to one and only one in the pre-built graph convolution network model, thereby avoiding overfitting of the graph convolution network model, improving the computational efficiency of the graph convolution network model, and enhancing the interpretability of the decision-making process of data classification through the graph convolution network model.
[0074] S240. Perform model training on the graph convolutional network model according to the true classification labels and the predicted classification results, and update the model weight matrix in the graph convolutional network model until the model training end conditions are met to obtain a financial data classification model.
[0075] Specifically, the deviation value of the model prediction result of the graph convolution network model can be determined according to the financial data classification category corresponding to the financial sample data to be trained and the data classification result of the financial sample data to be trained output by the graph convolution network model, and the model weight matrix in the graph convolution network model is updated by using the deviation value, and the model parameters of the graph convolution network model are updated until the model training end condition is met to obtain the financial data classification model. Among them, the model training end condition can be that the deviation value is less than a preset threshold, or it can be that the number of model training times is met, which can be set by technical personnel based on experience.
[0076] In an optional scheme of the present embodiment, it can be combined with one or more optional schemes of the present embodiment. Optionally, the sample node set includes a sample entity set; the sample entity set consists of at least one sample vocabulary; the initial model weight matrix in the graph convolutional network model is determined as follows: obtain reference financial sample data that is associated with the financial sample to be trained; perform text processing on the reference financial sample data to obtain at least one text vocabulary; determine the target vocabulary in each text vocabulary that is associated with each sample vocabulary; determine the word frequency score value corresponding to each target vocabulary; determine the initial model weight matrix in the graph convolutional network model according to the word frequency score value of each target vocabulary.
[0077] The reference financial sample data may be sample data extracted from financial data other than the financial sample data to be trained.
[0078] The target vocabulary may be a vocabulary used to describe financial data, such as a job title, a company name, or an industry name, etc. The word frequency score value may be a score value of the number of times the target vocabulary appears in the text.
[0079] Specifically, the named entity recognition can be performed on the reference financial sample data, thereby extracting different text entity words. Then, it is determined whether there is a correlation between the text entity words and the sample words to determine the target word. For example, the sample word can be XX Company, and the obtained text entity word can be XX Technology Co., Ltd., then it can be determined that there is a correlation between the text entity word and the sample word, and both are expressed as company names, and the text entity word can be determined as the target word.
[0080] Among them, the word frequency score values corresponding to each target vocabulary are determined, and the specific calculation method can be as follows:
[0081] tf*idf(w, W)=tf(w, W)×idf(w);
[0082] Among them, tf*idf(w, W) can be expressed as the word frequency data value of entity word w in text W; tf(w, W) can be expressed as the number of times entity word w appears in text W; idf(w) can be expressed as the inverse document frequency of entity word w in the corpus. Specifically,
[0083]
[0084] Where count(w, W) can be expressed as the number of times the entity word w appears in the text segment W; |W| can represent the total number of entity occurrences in the text segment W. Specifically,
[0085]
[0086] Among them, |D| can represent the total number of text segments in the corpus, and df(d) can represent the number of text segments in the corpus that contain the entity vocabulary w.
[0087] Furthermore, a model weight matrix in the graph convolutional network model can be generated according to the calculated word frequency score values of each target word in the text.
[0088] This technical solution obtains each named entity vocabulary from the reference financial sample data, and then determines the target vocabulary associated with the sample vocabulary from each named entity vocabulary, and calculates the word frequency score value of each target vocabulary to obtain the initial model weight matrix. The model weight matrix can be initialized by the word frequency score value, so that high-frequency entity vocabulary has a higher weight, thereby improving the performance of the model in classifying financial data and enhancing the interpretability of the classification decision process of financial data to a certain extent.
[0089] S250: Obtain the financial data to be classified, and generate a knowledge graph query statement based on the financial data to be classified.
[0090] S260. Input the knowledge graph query statement into the pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model.
[0091] S270. Perform text graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix and an adjacency matrix.
[0092] S280, inputting the node set, edge set, feature matrix and adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
[0093] The technical solution of the embodiment of the present invention obtains the financial sample data to be trained and determines the real classification labels of the financial sample data to be trained, thereby generating a text graph including a sample node set, a sample edge set, a sample feature matrix and a sample adjacency matrix. The text graph is input into a pre-built graph convolutional network model to obtain the predicted classification results output by the graph convolutional network model; the graph convolutional network model is trained according to the real classification labels and the predicted classification results, and the model weight matrix in the graph convolutional network model is updated until the model training end condition is met to obtain a financial data classification model. Then, the obtained financial data to be classified is processed and a knowledge graph entity relationship is generated, and then the knowledge graph entity relationship is converted into a text graph to obtain a node set, an edge set, a feature matrix and an adjacency matrix; the node set, the edge set, the feature matrix and the adjacency matrix are input into the pre-trained financial data classification model to obtain the financial data category output by the financial data classification model. This technical solution can train the pre-built graph network model by referencing the model weight matrix, thereby obtaining a financial data classification model and outputting data categories for the acquired financial data to be classified, thereby improving the comprehensiveness and accuracy of classifying financial data using a generative model, and further enhancing the interpretability of the financial data classification decision process.
[0094] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with relevant laws, regulations and standards in relevant regions.
[0095] Embodiment 3
[0096] Figure 3 A schematic diagram of the structure of a financial data category prediction device provided in the third embodiment of the present invention. A financial data category prediction device provided in the embodiment of the present invention can be applied to classify financial data under different financial business classification demand scenarios. The financial data category prediction device can be implemented in the form of hardware and / or software, such as Figure 3 As shown, it specifically includes: a query statement generation module 310, an entity relationship acquisition module 320, a text graph conversion module 330 and a financial data category acquisition module 340. Among them,
[0097] The query statement generation module 310 is used to obtain the financial data to be classified and generate a knowledge graph query statement based on the financial data to be classified.
[0098] An entity relationship acquisition module 320 is used to input the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model;
[0099] The text-to-graph conversion module 330 is used to perform text-to-graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix and an adjacency matrix.
[0100] The financial data category acquisition module 340 is used to input the node set, the edge set, the feature matrix and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
[0101] This technical solution can extract detailed and structured data content from the financial data to be classified and generate knowledge graph entity relationships. The generated knowledge graph entity relationships can also be converted into corresponding text graphs to obtain the data categories of financial data, ensuring the comprehensiveness and accuracy of the financial data classification results and providing a transparent and clear classification decision process.
[0102] Optionally, the knowledge graph entity relationship includes an entity set and a relationship set; accordingly, the text graph conversion module 330 is specifically used to:
[0103] Performing text graph conversion on the knowledge graph entity relationship to generate a node set including the entity set and the relationship set;
[0104] Generate an edge set according to the relationship set;
[0105] Generate a feature matrix according to node feature information of each node in the node set;
[0106] The association relationship between the nodes is determined according to the node set and the edge set, and an adjacency matrix is generated based on the association relationship.
[0107] Optionally, the device further includes a financial data question-answering model training module, which is used, before inputting the knowledge graph query statement into the pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model,
[0108] Obtaining financial sample data to be trained, and determining an open source question-answering model; the open source question-answering model is pre-trained based on the open source sample data;
[0109] Using the financial sample data to be trained and the open source sample data, the open source question-answering model is trained to obtain a benchmark question-answering model;
[0110] The benchmark question-answering model is trained using the financial sample data to obtain a financial data question-answering model.
[0111] Optionally, the device further includes a financial data classification model training module, which is used for, before inputting the node set, the edge set, the feature matrix and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model,
[0112] Acquire financial sample data to be trained, and determine the true classification labels of the financial sample data to be trained;
[0113] Performing data processing on the financial sample data to be trained to generate a text graph including a sample node set, a sample edge set, a sample feature matrix, and a sample adjacency matrix;
[0114] Inputting the text graph into a pre-built graph convolutional network model to obtain a predicted classification result output by the graph convolutional network model;
[0115] According to the true classification label and the predicted classification result, the graph convolutional network model is trained, and the model weight matrix in the graph convolutional network model is updated until the model training end condition is met to obtain a financial data classification model.
[0116] Optionally, the number of graph convolutional layers in the graph convolutional network model is and is only one layer.
[0117] Optionally, the sample node set includes a sample entity set; the sample entity set consists of at least one sample vocabulary; the device further includes: a model weight matrix determination module, used to:
[0118] Acquire reference financial sample data associated with the financial sample to be trained;
[0119] Performing text processing on the reference financial sample data to obtain at least one text word;
[0120] Determining a target word in each of the text words that has an associated relationship with each of the sample words;
[0121] Determine the word frequency score value corresponding to each of the target words;
[0122] According to the word frequency score value of each of the target words, an initial model weight matrix in the graph convolutional network model is determined.
[0123] The financial data category prediction device provided in the embodiment of the present invention can execute the financial data category prediction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0124] Embodiment 4
[0125] Figure 4 A schematic diagram of the structure of an electronic device 40 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0126] like Figure 4 As shown, the electronic device 40 includes at least one processor 41, and a memory connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 to the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0127] A number of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0128] The processor 41 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 performs the various methods and processes described above, such as the financial data category prediction method.
[0129] In some embodiments, the financial data category prediction method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the financial data category prediction method described above may be performed. Alternatively, in other embodiments, the processor 41 may be configured to execute the financial data category prediction method in any other appropriate manner (e.g., by means of firmware).
[0130] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0132] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0134] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0135] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0136] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0137] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for predicting financial data categories, characterized in that: include: Obtaining financial data to be classified, and generating a knowledge graph query statement based on the financial data to be classified; Inputting the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model; Performing text-graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix; The node set, the edge set, the feature matrix and the adjacency matrix are input into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
2. The method according to claim 1, characterized in that The knowledge graph entity relationship includes an entity set and a relationship set; accordingly, the knowledge graph entity relationship is converted into a text graph to obtain a node set, a feature matrix and an adjacency matrix, including: Performing text-graph conversion on the knowledge graph entity relationship, and generating a node set according to the entity set; Generate an edge set according to the relationship set; Generate a feature matrix according to node feature information of each node in the node set; The association relationship between the nodes is determined according to the node set and the edge set, and an adjacency matrix is generated based on the association relationship.
3. The method according to claim 1, characterized in that Before inputting the knowledge graph query statement into the pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model, the method further includes: Obtaining financial sample data to be trained, and determining an open source question-answering model; the open source question-answering model is pre-trained based on the open source sample data; Using the financial sample data to be trained and the open source sample data, the open source question-answering model is trained to obtain a benchmark question-answering model; The benchmark question-answering model is trained using the financial sample data to obtain a financial data question-answering model.
4. The method according to claim 1, characterized in that Before inputting the node set, the edge set, the feature matrix and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model, the method further includes: Acquire financial sample data to be trained, and determine the true classification labels of the financial sample data to be trained; Performing data processing on the financial sample data to be trained to generate a text graph including a sample node set, a sample edge set, a sample feature matrix, and a sample adjacency matrix; Inputting the text graph into a pre-built graph convolutional network model to obtain a predicted classification result output by the graph convolutional network model; According to the true classification label and the predicted classification result, the graph convolutional network model is trained, and the model weight matrix in the graph convolutional network model is updated until the model training end condition is met to obtain a financial data classification model.
5. The method according to claim 4, characterized in that The number of graph convolutional layers in the graph convolutional network model is and is only one.
6. The method according to claim 4, characterized in that The sample node set includes a sample entity set; the sample entity set consists of at least one sample vocabulary; the initial model weight matrix in the graph convolutional network model is determined as follows: Acquire reference financial sample data associated with the financial sample to be trained; Performing text processing on the reference financial sample data to obtain at least one text word; Determining a target word in each of the text words that has an associated relationship with each of the sample words; Determine the word frequency score value corresponding to each of the target words; According to the word frequency score value of each of the target words, an initial model weight matrix in the graph convolutional network model is determined.
7. A financial data category prediction device, characterized in that: include: A query statement generation module, used to obtain the financial data to be classified, and generate a knowledge graph query statement based on the financial data to be classified; An entity relationship acquisition module, used to input the knowledge graph query statement into a pre-trained financial data question-answering model to obtain the knowledge graph entity relationship output by the financial data question-answering model; A text-to-graph conversion module, used to perform text-to-graph conversion on the knowledge graph entity relationship to obtain a node set, an edge set, a feature matrix, and an adjacency matrix; The financial data category acquisition module is used to input the node set, the edge set, the feature matrix and the adjacency matrix into a pre-trained financial data classification model to obtain the financial data category output by the financial data classification model.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the financial data category prediction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the financial data category prediction method according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the financial data category prediction method according to any one of claims 1 to 6.
Citation Information
Cited By
Knowledge graph data multi-level classification method and system and medium
CN122020323A
A knowledge graph data multi-level classification method and system, and a medium
CN122020323B