A financial big data processing method, device, equipment and medium
By constructing an asset tree and performing semantic decomposition and rule-based transformation on mismatched information, the impact of financial asset trees on financial big data processing was resolved, achieving efficient and accurate data processing and user queries.
Patent Information
- Application Number
- CN202311043366.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-08-17
AI Technical Summary
The inconsistencies in how investors construct financial asset trees lead to a decrease in the accuracy of financial indicator results in financial big data platforms, affecting data processing efficiency.
By acquiring users' financial asset information, an asset tree is constructed, and mismatched information is semantically split and rule-based transformed. Matching and computational information is then obtained and fused for computation, thus isolating the asset tree from the financial big data platform.
It achieves flexible decoupling between the asset tree and the financial big data platform, improves data processing efficiency and accuracy, and facilitates direct querying and modification by users.
Smart Images

Figure CN116975112B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the financial field, and more particularly to a financial big data processing method, apparatus, equipment, and medium. Background Technology
[0002] Currently, to help investors understand and manage their financial assets, financial asset trees are constructed to divide financial assets into different categories and subcategories, and to classify and organize them according to the relationships and characteristics between them. A financial asset tree is a tool that can be used to describe financial asset categories, calculate financial asset value, and analyze financial asset risk. By performing operations between nodes in the financial asset tree, indicators such as the total value, return, and risk of different combinations of financial assets can be obtained.
[0003] Currently, financial big data platforms obtain various financial indicator results through the financial asset trees of individual investors. However, due to differences in the descriptions investors use when constructing their financial asset trees, there are discrepancies with the data in the financial big data platform, which affects the accuracy of the obtained financial indicator results and reduces the data processing efficiency of the platform. Therefore, how to avoid the impact of financial asset trees on financial big data processing and improve the data processing efficiency of the financial big data platform has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a financial big data processing method, apparatus, device, and medium to address the impact of financial asset trees on financial big data processing.
[0005] In a first aspect, embodiments of the present invention provide a financial big data processing method, the financial big data processing method comprising:
[0006] Obtain the user's financial asset information, and construct an asset tree corresponding to the user based on the financial asset information and preset asset hierarchy rules, wherein the node information of any node in the asset tree is the asset information of the corresponding asset class;
[0007] For any node in the asset tree, obtain the description rules of the asset class corresponding to the node in the asset big data, match the node information of the node with the description rules, and obtain matching information and non-matching information;
[0008] The mismatched information is semantically split to obtain N sub-information. Based on the description rules corresponding to the nodes, the sub-information is transformed into regularized information.
[0009] The information to be calculated in the asset big data is obtained, and the information to be calculated, the matching information, and the transformed information are fused and calculated to obtain the fusion calculation result.
[0010] Secondly, embodiments of the present invention provide a financial big data processing device, the financial big data processing device comprising:
[0011] The asset tree configuration module is used to obtain the user's financial asset information and construct the corresponding asset tree for the user based on the financial asset information and preset asset hierarchy rules. The node information of any node in the asset tree is the asset information of the corresponding asset class.
[0012] The information matching module is used to obtain the description rules of the asset class corresponding to any node in the asset tree in the asset big data, match the node information of the node with the description rules, and obtain matching information and non-matching information.
[0013] The information conversion module is used to semantically split the mismatched information to obtain N sub-information, and to perform regularization conversion on the sub-information according to the description rules corresponding to the node to obtain the converted information;
[0014] The information fusion module is used to acquire information to be calculated from the asset big data, and to perform fusion calculation on the information to be calculated, the matching information, and the transformed information to obtain the fusion calculation result.
[0015] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the financial big data processing method as described in the first aspect.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the financial big data processing method as described in the first aspect.
[0017] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0018] This invention acquires a user's financial asset information and constructs an asset tree for the corresponding user based on the financial asset information and preset asset hierarchy rules. The node information of any node in the asset tree represents the asset information of the corresponding asset class. For any node in the asset tree, the description rules of the asset class corresponding to the node in the asset big data are obtained. The node information is matched with the description rules to obtain matching and non-matching information. The non-matching information is semantically split to obtain N sub-information. Based on the description rules corresponding to the node, the sub-information is transformed into rule-based information. Information to be calculated in the asset big data is acquired, and the information to be calculated, the matching information, and the transformed information are fused and calculated to obtain the fused calculation result. By classifying and combining the user's financial assets into an asset tree, it facilitates direct querying and modification by the user. Simultaneously, based on the description rules of the asset classes corresponding to each node in the asset tree in the asset big data, the information of each node in the asset tree is transformed into information that the financial platform can recognize. This isolates the user's asset tree configuration from the financial platform's big data processing, ensuring that they do not interfere with each other and can flexibly handle their respective tasks, thereby achieving flexible decoupling and efficient processing. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an application environment for a financial big data processing method provided in Embodiment 1 of the present invention;
[0021] Figure 2 This is a flowchart illustrating a financial big data processing method provided in Embodiment 1 of the present invention;
[0022] Figure 3 This is a schematic diagram of the structure of a financial big data processing device provided in Embodiment 2 of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0025] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0028] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0031] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0032] The financial big data processing method provided in Embodiment 1 of this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. The client includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0033] This financial big data processing method is applied in financial scenarios. The client can be a mobile phone, which connects to the server via a wireless network to send the user's financial asset information. Then, the server constructs a corresponding asset tree based on the user's financial asset information and converts the asset information in the asset tree into a description in the asset big data, thereby realizing the management of asset big data.
[0034] See Figure 2 This is a flowchart illustrating a financial big data processing method provided in Embodiment 1 of the present invention. The aforementioned financial big data processing method can be applied to... Figure 1 The server-side component connects to a corresponding database to retrieve backend data, such as the user's financial asset information, which is used to construct an asset tree composed of various asset information. Additionally, the server connects to the client-side component to obtain financial asset information sent by the user from the client. Figure 2 As shown, this financial big data processing method may include the following steps:
[0035] Step S201: Obtain the user's financial asset information, and construct the corresponding user's asset tree based on the financial asset information and preset asset hierarchy rules.
[0036] In this asset tree, the node information of any node represents the asset information of the corresponding asset class. A user's financial asset information can include the holdings of various financial products, such as wealth management products, government bonds, demand deposits, agency referrals, salary deposits, insurance, time deposits, precious metals, funds, and loans. Asset types include domestic equity assets, domestic fixed-income assets, overseas equity assets, overseas fixed-income assets, real estate, cash, and cash equivalents. Domestic equity assets refer to investments made within China to acquire the equity or net assets of other companies. Domestic fixed-income assets refer to investments in fixed-income assets within China, such as bank time deposits, negotiated deposits, government bonds, financial bonds, corporate bonds, convertible bonds, and bond funds. Conversely, overseas equity assets refer to investments made outside China to acquire the equity or net assets of other companies. Overseas fixed-income assets refer to investments outside China, such as bank time deposits, negotiated deposits, government bonds, financial bonds, corporate bonds, convertible bonds, and bond funds. Real estate refers to properties purchased domestically or internationally. This embodiment of the invention does not impose specific limitations on these categories. The preset asset hierarchy rules are asset classification dimensions pre-set based on the hierarchical relationships between asset types. Users can customize these rules according to their needs, and this invention does not impose any restrictions. Therefore, by obtaining a user's financial asset information and constructing the corresponding user's asset tree based on this information and the preset asset hierarchy rules, the invention achieves this.
[0037] In this invention, the method for constructing a user's asset tree is as follows: Based on the user's financial asset information, determine the asset information corresponding to each asset type. Treat each asset type as an asset node, and the asset information of each asset node as node information. Determine the subordinate types of each asset type from preset asset hierarchy rules, and treat each subordinate type as a subordinate node. Based on the parent-child relationship between subordinate nodes and asset nodes, form corresponding subordinate nodes on all asset nodes to form an asset tree. Treat subordinate types as asset types, and subordinate nodes as asset nodes, and return to the steps of determining the subordinate types of each asset type from preset asset hierarchy rules and treating each subordinate type as a subordinate node, until all asset types have no subordinate types in the preset rules, thus obtaining the final asset tree. For any subordinate node corresponding to a subordinate type, aggregate the node information of the asset nodes corresponding to the asset types under the subordinate type to obtain the aggregation result, and use the aggregation result as the node information of the subordinate node. Update the final asset tree based on the node information of all subordinate nodes to obtain the updated asset tree.
[0038] Step S202: For any node in the asset tree, obtain the description rules of the asset class corresponding to the node in the asset big data, match the node information with the description rules, and obtain matching information and non-matching information.
[0039] The description rules define the specific attributes of each data item in the target document to be generated, including but not limited to area information and conversion mechanisms. Area information defines the cell range occupied by the data field, supporting data field arrangement across rows and columns. The conversion mechanism defines the data format corresponding to various data types, such as the precision of decimals and the display format of dates. Description rules describe the information corresponding to each asset class. Asset big data refers to data provided by a financial big data platform. Matching information refers to asset information that conforms to the description in the asset big data, i.e., asset information that the asset big data platform can identify through its rule engine. Conversely, non-matching information refers to asset information that does not conform to the description in the asset big data, i.e., asset information that the asset big data platform cannot identify through its rule engine. Therefore, in this invention, for any node in the asset tree, the description rules of the asset class corresponding to the node in the asset big data are first obtained. Then, the node information is matched with the description rules to obtain matching and non-matching information, thereby filtering out asset information in the asset tree that does not conform to the description in the asset big data, improving the efficiency of subsequent information conversion.
[0040] Optionally, the node information is matched against the description rules to obtain matching and non-matching information, including:
[0041] The description rules are parsed to obtain the corresponding standard information word segmentation results;
[0042] The node information is segmented into words to obtain the segmentation results;
[0043] The word segmentation results are matched with the standard information word segmentation results to obtain matching and non-matching information.
[0044] Word segmentation is a fundamental task in lexical analysis. Word segmentation algorithms are mainly divided into two categories based on their core ideas: one is dictionary-based segmentation, which first segments the text data into words according to a dictionary and then finds the optimal word combination; the other is character-based segmentation, which constructs words from characters, first dividing the sentence into individual characters and then combining them into words to find the optimal segmentation strategy. Therefore, firstly, the description rules are parsed to obtain the standard information segmentation results corresponding to the description rules. Using the standard information segmentation results as templates, the node information is then segmented to obtain the segmentation results. Based on template matching, the segmentation results are matched with the standard information segmentation results to obtain matching and non-matching information. Using the standard information segmentation results as templates, template matching ensures the accuracy of obtaining matching and non-matching information.
[0045] In this invention, an understanding-based word segmentation method is used to segment node information to obtain segmentation results. The understanding-based method achieves word recognition by having the computer simulate human understanding of sentences. The basic idea of this method is to perform syntactic and semantic analysis simultaneously with word segmentation, using syntactic and semantic information to handle ambiguity.
[0046] In one implementation, rule-based word segmentation methods (such as string matching-based segmentation) can also be used. Rule-based segmentation methods match the Chinese character string to be analyzed against entries in a sufficiently large dictionary according to a certain strategy. If a string is found in the dictionary, the match is successful (a word is identified). Commonly used rule-based word segmentation methods include: forward maximum matching (from left to right); backward maximum matching (from right to left); and minimum segmentation (minimizing the number of words segmented in each sentence). Forward maximum matching involves dividing a string into segments with a limited length, then matching the substrings against words in the dictionary. If a match is successful, the next round of matching is performed until all strings have been processed; otherwise, a character is removed from the end of the substring, and the matching process is repeated. Backward maximum matching is similar to forward maximum matching.
[0047] Step S203: Semantically split the mismatched information to obtain N sub-information. Based on the description rules corresponding to the nodes, perform rule-based transformation on the sub-information to obtain transformed information.
[0048] Semantic decomposition is used to separate different sub-information from the information, where different semantics may represent descriptions of different information. The transformed information refers to information that conforms to the representation in the asset big data. Therefore, this invention performs semantic decomposition on mismatched information to obtain N sub-information. Based on the description rules corresponding to the nodes, the sub-information is transformed according to rules to obtain transformed information. By transforming mismatched information into information that conforms to the representation in the asset big data, the rule representations of the asset tree and the asset big data are consistent, which can improve the efficiency of subsequent asset big data management.
[0049] Optionally, the mismatched information can be semantically split to obtain N sub-information, including:
[0050] Retrieve key information fields from mismatched information;
[0051] The key information fields in the mismatched information are independently semantically divided to obtain at least one first sub-information.
[0052] The information other than the first sub-information is divided sequentially according to independent semantics to obtain at least one second sub-information.
[0053] The process begins by parsing the mismatched information to identify key information fields. Then, the starting position of each key information field within the mismatched information is determined. Based on these starting positions, the key information fields in the mismatched information are independently semantically segmented to obtain at least one first sub-information. Next, for information other than the first sub-information, a word segmentation method is used to sequentially segment the information according to independent semantics, obtaining at least one second sub-information.
[0054] In this invention, the method for obtaining key information fields from mismatched information is as follows: First, stop words in the matched information are filtered out. Then, a keyword extraction algorithm is used to extract keywords from the filtered matched information to obtain at least one key information field. Stop words refer to certain characters or words that are automatically filtered out before or after processing natural language data (or text) in information retrieval to save storage space and improve search efficiency. Filtering stop words can improve search efficiency and save computation time.
[0055] It should be noted that keyword extraction algorithms are generally divided into supervised and unsupervised categories. Unsupervised keyword extraction algorithms do not require manual generation and maintenance of vocabulary lists, nor do they require manually labeled corpora for training. Currently, commonly used unsupervised keyword extraction algorithms include TF-IDF (Term Frequency-Inverse Text Frequency), TextRank (Text Ranking Algorithm), and LDA (Linear Discriminant Analysis).
[0056] TF-IDF is a commonly used weighting technique for information retrieval and data mining. TF stands for Term Frequency, and IDF stands for Inverse Document Frequency. TF-IDF is the product of the two. The TextRank algorithm is based on the PageRank algorithm, a webpage ranking algorithm with two basic principles: link quantity and link quality. Link quantity refers to the number of other webpages linking to a webpage, indicating its importance; link quality refers to the number of higher-weighted webpages linking to a webpage, also indicating its importance. TextRank is a text ranking algorithm that treats words as nodes on the internet and calculates the importance of each word based on their co-occurrence relationships. The TextRank algorithm can extract keywords and keyword phrases from a given text.
[0057] In one embodiment, a pre-trained keyword extraction model can be used to extract keywords from mismatched information to obtain at least one key information field. The pre-trained keyword extraction model can be a machine learning-based model, whereby labels are generated by manually annotating keywords in the information, and sample information from the training set is input into the pre-built keyword extraction model to obtain keyword output results. The loss calculated based on the keyword output results and labels is then used to train the keyword extraction model, resulting in a trained keyword extraction model.
[0058] Optionally, the mismatched information can be semantically split to obtain N sub-information, including:
[0059] Calculate the string length of each sub-information, filter the N sub-information pieces based on the string length, and obtain the target sub-information;
[0060] Based on the description rules corresponding to the nodes, the sub-information is transformed into rule-based information, including:
[0061] Based on the description rules corresponding to the nodes, the target sub-information is transformed into regularized information.
[0062] In this invention, the string length of each sub-information is calculated and compared with a preset length threshold. If the string length is greater than or equal to the preset length threshold, the corresponding sub-information is considered valid and retained; if the string length is less than the preset length threshold, the corresponding sub-information is considered invalid and filtered out. Therefore, by filtering N sub-information based on string length, target sub-information with string lengths greater than or equal to the preset length threshold is obtained. Then, according to the description rules corresponding to the nodes, the target sub-information is transformed into regularized information, avoiding interference from invalid information and improving information transformation efficiency. For example, assuming that the sub-information with a string length of 1 is useless, by setting the length threshold to 2, sub-information with a string length less than 2 is filtered out, while sub-information with a string length greater than or equal to 2 is retained.
[0063] Optionally, based on the description rules corresponding to the nodes, the target sub-information is transformed into rule-based information, including:
[0064] Obtain the information template of the description rule corresponding to the node, and fill the information template with the target sub-information to obtain the transformed information.
[0065] In this invention, an information template refers to a fixed and standardized information structure that clearly defines the information requirements, i.e., selecting useful data. Therefore, this invention first obtains the information template corresponding to the description rules of the nodes, fills the target sub-information into the information template, and thus generates the transformed information. This achieves customized information transformation, reduces the resources and time consumed in processing information transformation, and thereby improves information processing efficiency.
[0066] Optionally, based on the description rules corresponding to the nodes, the target sub-information is transformed into rule-based information, including:
[0067] Obtain a trained information reconstruction model corresponding to the description rules of the nodes. The training set includes sample information and labels obtained based on the description rules. The information reconstruction model includes an encoder and a decoder. The encoder encodes the sample information to obtain the encoding result. The decoder decodes the encoding result to obtain the decoding result. The information reconstruction model is trained based on the loss calculated from the decoding result and the labels to obtain the trained information reconstruction model.
[0068] The target sub-information is input into the trained information reconstruction model, and the transformed information is output.
[0069] In this context, a well-trained information reconstruction model refers to an information reconstruction model that has been fully trained and possesses strong information reconstruction capabilities, accurately reconstructing corresponding transformation information even for new information. The well-trained information reconstruction model is obtained through training on large sample data. Therefore, before obtaining the well-trained information reconstruction model corresponding to the node description rules, the information reconstruction model needs to be trained. The training process is as follows: acquire sample information and labels obtained based on the description rules; encode the sample information using an encoder to obtain the corresponding encoding result; decode the encoding result using a decoder to obtain the corresponding decoding result; calculate the reconstruction loss using the cross-entropy loss function based on the decoding result of the sample information and the corresponding labels; and correct the parameters in the information reconstruction model using the gradient descent method until the reconstruction loss converges, thus obtaining the well-trained information reconstruction model. Finally, the target sub-information is input into the well-trained information reconstruction model, and the transformed information is output.
[0070] Step S204: Obtain the information to be calculated from the asset big data, and perform fusion calculation on the information to be calculated, the matching information, and the transformed information to obtain the fusion calculation result.
[0071] The information to be calculated refers to the original information in the asset big data that needs to participate in the fusion calculation, that is, asset information in the asset big data that belongs to the same category as the asset class corresponding to the node. Fusion calculation is used to merge multiple objectives into one objective. Fusion calculation can include, but is not limited to, summation calculation, maximum value calculation, average value calculation, and minimum value calculation, which can be set according to actual needs. Therefore, this invention can obtain the information to be calculated in the asset big data from the financial big data platform, perform fusion calculation on the information to be calculated, the matching information, and the transformed information, and obtain the fusion calculation result.
[0072] It should be noted that, for each node in the user's asset tree, using steps S202 to S204, matching information and transformation information are obtained according to the node information corresponding to each node. Then, the information to be calculated corresponding to the asset class of each node is obtained in the asset big data. The information to be calculated, the matching information and the transformation information are fused and calculated to obtain the fusion calculation result of the corresponding node.
[0073] Optionally, the information to be calculated, the matching information, and the transformed information are fused together to obtain the fused calculation result, which includes:
[0074] Receive the user's query request for fusion computing results, confirm the user's preset table structure based on the query request, generate a fusion computing result table based on the preset table structure and all fusion computing results, and return the fusion computing result table to the user.
[0075] The query request for fused computation results carries specified query rules. These rules define how the fused computation results should be queried. For example, the rules can specify that the results are queried in ascending order or according to user-selected results. Each preset table corresponds one-to-one with the query request. Therefore, after obtaining all fused computation results, the user can send a query request to the financial big data platform. Upon receiving this request, the platform confirms the user's preset table structure, fills the preset table structure with the fused computation results, and then returns the retrieved results to the user.
[0076] It's worth noting that after receiving a user's query request for fusion computing results, the system can verify the request, providing query services only to those that meet the query criteria. This effectively controls information access behavior, reduces the risk of misuse of external information, and thus effectively protects user privacy. For example, by setting user identification information, the system can uniquely identify the corresponding user. Matching the identification information verifies the user's identity. If the verification passes, the user meets the query criteria; if the verification fails, the user does not meet the query criteria, and the corresponding query request can be ignored.
[0077] This invention acquires a user's financial asset information and constructs an asset tree for the corresponding user based on the financial asset information and preset asset hierarchy rules. The node information of any node in the asset tree represents the asset information of the corresponding asset class. For any node in the asset tree, the description rules of the asset class corresponding to the node in the asset big data are acquired. The node information is matched with the description rules to obtain matching and non-matching information. The non-matching information is semantically split to obtain N sub-information. Based on the description rules corresponding to the node, the sub-information is transformed into rule-based information. Information to be calculated in the asset big data is acquired, and the information to be calculated, the matching information, and the transformed information are fused and calculated to obtain the fused calculation result. By classifying and combining the user's financial assets into an asset tree, it facilitates direct querying and modification by the user. Simultaneously, based on the description rules of the asset classes corresponding to each node in the asset tree in the asset big data, the information of each node in the asset tree is transformed into information recognizable by the financial platform. This isolates the user's asset tree configuration from the financial platform's big data processing, ensuring that they do not interfere with each other and can flexibly handle their respective tasks, thereby achieving flexible decoupling and efficient processing.
[0078] Corresponding to the financial big data processing method in the above embodiments, Figure 3 This diagram illustrates a structural block diagram of a financial big data processing device according to Embodiment 2 of the present invention. (See also...) Figure 3 The financial big data processing device includes:
[0079] The asset tree configuration module 31 is used to obtain the user's financial asset information and construct an asset tree corresponding to the user based on the financial asset information and preset asset hierarchy rules, wherein the node information of any node in the asset tree is the asset information of the corresponding asset class.
[0080] The information matching module 32 is used to obtain the description rules of the asset class corresponding to any node in the asset tree in the asset big data, match the node information of the node with the description rules, and obtain matching information and non-matching information.
[0081] The information conversion module 33 is used to semantically split the mismatched information to obtain N sub-information, and to perform rule-based conversion on the sub-information according to the description rules corresponding to the node to obtain the converted information;
[0082] The information fusion module 34 is used to acquire information to be calculated in the asset big data, and to perform fusion calculation on the information to be calculated, the matching information and the transformed information to obtain the fusion calculation result.
[0083] Optionally, the information matching module 32 includes:
[0084] The rule parsing unit is used to parse the description rules to obtain the corresponding standard information word segmentation results;
[0085] The word segmentation processing unit is used to perform word segmentation processing on the node information of the node to obtain the word segmentation result;
[0086] The word segmentation matching unit is used to match the word segmentation result with the standard information word segmentation result to obtain matching information and non-matching information.
[0087] Optionally, the information conversion module 33 includes:
[0088] A key information extraction unit is used to obtain key information fields from the mismatched information;
[0089] The first partitioning unit is used to independently semantically partition the key information fields in the mismatched information to obtain at least one first sub-information.
[0090] The second partitioning unit is used to partition the information other than the first sub-information according to independent semantics in sequence to obtain at least one second sub-information.
[0091] Optionally, the information conversion module 33 includes:
[0092] An information filtering unit is used to perform semantic segmentation on the mismatched information to obtain N sub-information, then count the string length of each sub-information, and filter the N sub-information according to the string length to obtain the target sub-information;
[0093] The target transformation unit is used to perform rule-based transformation on the target sub-information according to the description rules corresponding to the node, so as to obtain transformed information.
[0094] Optionally, the target transformation unit includes:
[0095] The template filling subunit is used to obtain the information template of the description rule corresponding to the node, and fill the information template with the target sub-information to obtain the transformed information.
[0096] Optionally, the target transformation unit includes:
[0097] The model acquisition subunit is used to acquire the trained information reconstruction model corresponding to the description rule of the node. The training set includes sample information and labels obtained based on the description rule. The information reconstruction model includes an encoder and a decoder. The encoder encodes the sample information to obtain an encoding result. The decoder decodes the encoding result to obtain a decoding result. The information reconstruction model is trained based on the loss calculated from the decoding result and the label to obtain the trained information reconstruction model.
[0098] The information reconstruction subunit is used to input the target sub-information into the trained information reconstruction model and output the transformed information.
[0099] Optionally, the financial big data processing device includes:
[0100] The result query module is used to perform fusion calculation on the information to be calculated, the matching information, and the transformed information, and after obtaining the fusion calculation result, receive the user's fusion calculation result query request, confirm the user's preset table structure according to the fusion calculation result query request, generate a fusion calculation result table according to the preset table structure and all fusion calculation results, and feed back the fusion calculation result table to the user.
[0101] It should be noted that the information interaction and execution process between the above modules, units, and sub-units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0102] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the diagram), a memory, and a computer program stored in the memory and capable of running on at least one processor, wherein the processor executes the computer program to implement the steps in any of the above embodiments of the financial big data processing methods.
[0103] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0104] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0105] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of a computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0106] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0107] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0108] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0109] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0110] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0111] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0112] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for processing financial big data, characterized in that, The financial big data processing method includes: Obtain the user's financial asset information, and construct an asset tree corresponding to the user based on the financial asset information and preset asset hierarchy rules, wherein the node information of any node in the asset tree is the asset information of the corresponding asset class; For any node in the asset tree, obtain the description rules of the asset class corresponding to the node in the asset big data, match the node information of the node with the description rules, and obtain matching information and non-matching information; The mismatched information is semantically split to obtain N sub-information. Based on the description rules corresponding to the nodes, the sub-information is transformed into regularized information. Obtain the information to be calculated from the asset big data, and perform fusion calculation on the information to be calculated, the matching information, and the transformed information to obtain the fusion calculation result; The step of matching the node information of the node with the description rule to obtain matching information and non-matching information includes: The description rules are parsed to obtain the corresponding standard information word segmentation results; The node information of the node is processed by word segmentation to obtain the word segmentation result; The word segmentation results are matched with the standard information word segmentation results to obtain matching information and non-matching information; The semantic segmentation of the mismatched information yields N sub-information items, including: Retrieve key information fields from the mismatched information; The key information fields in the mismatched information are independently semantically divided to obtain at least one first sub-information; The information other than the first sub-information is divided sequentially according to independent semantics to obtain at least one second sub-information; After semantically splitting the mismatched information to obtain N sub-information, the following is included: The string length of each sub-information is counted, and the N sub-information are filtered according to the string length to obtain the target sub-information; Based on the description rules corresponding to the nodes, the sub-information is transformed into rule-based information to obtain the transformed information, including: Based on the description rules corresponding to the nodes, the target sub-information is transformed into regularized information.
2. The financial big data processing method according to claim 1, characterized in that, The step of performing rule-based transformation on the target sub-information according to the description rules corresponding to the node to obtain transformed information includes: Obtain the information template of the description rule corresponding to the node, and fill the information template with the target sub-information to obtain the transformed information.
3. The financial big data processing method according to claim 1, characterized in that, The step of performing rule-based transformation on the target sub-information according to the description rules corresponding to the node to obtain transformed information includes: Obtain a trained information reconstruction model corresponding to the description rule of the node, wherein the training set includes sample information and labels obtained based on the description rule, the information reconstruction model includes an encoder and a decoder, the encoder is used to encode the sample information to obtain an encoding result, the decoder is used to decode the encoding result to obtain a decoding result, and the information reconstruction model is trained based on the loss calculated according to the decoding result and the label to obtain a trained information reconstruction model; The target sub-information is input into the trained information reconstruction model, and the transformed information is output.
4. The financial big data processing method according to claim 1, characterized in that, After fusing the information to be calculated, the matching information, and the transformed information to obtain the fusion calculation result, the process includes: The system receives a user's query request for fusion calculation results, confirms the user's preset table structure based on the query request, generates a fusion calculation result table based on the preset table structure and all fusion calculation results, and sends the fusion calculation result table back to the user.
5. A financial big data processing apparatus, used to implement the financial big data processing method as described in any one of claims 1 to 4, characterized in that, The financial big data processing device includes: The asset tree configuration module is used to obtain the user's financial asset information and construct the corresponding asset tree for the user based on the financial asset information and preset asset hierarchy rules. The node information of any node in the asset tree is the asset information of the corresponding asset class. The information matching module is used to obtain the description rules of the asset class corresponding to any node in the asset tree in the asset big data, match the node information of the node with the description rules, and obtain matching information and non-matching information. The information conversion module is used to semantically split the mismatched information to obtain N sub-information, and to perform regularization conversion on the sub-information according to the description rules corresponding to the node to obtain the converted information; The information fusion module is used to acquire information to be calculated from the asset big data, and to perform fusion calculation on the information to be calculated, the matching information, and the transformed information to obtain the fusion calculation result.
6. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the financial big data processing method as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the financial big data processing method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Enterprise data acquisition and governance method
CN108769255A
Intelligent data asset identification method
CN113673889A