Index-Based Scalable Counterparty Identification Method, System, and Medium

The method leverages pre-trained models and decision trees to enhance the scalability and accuracy of counterparty identification in bond trading by structuring transaction texts, addressing the limitations of hard-coded rules in existing systems.

CN114398879BActive Publication Date: 2025-07-15BEIJING KUAQUO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210039442.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-07-15
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

The prior art has high maintenance costs and poor scalability in securities trading, making it difficult to accurately judge counterparty and adapt to various expression forms.

Method used

The index-based method is adopted to extract text features through pre-training models, and text structured processing is performed using structured rule indexes and decision tree models. It combines logical indexes to identify counterparties, reducing the dependence of hard-coded.

Benefits of technology

It improves the generalization ability and scalability of counterparty recognition, reduces training and maintenance costs, and achieves efficient and accurate counterparty recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398879B_ABST
    Figure CN114398879B_ABST
Patent Text Reader

Abstract

The present invention discloses an index-based scalable counterparty identification method, system and medium. By obtaining the text to be identified, feature extraction is performed on the text to be identified to obtain the entity elements in the text to be identified; based on a preset structured rule index, the text to be identified is structurally processed according to the entity elements to obtain structured order information; according to a pre-constructed and trained first decision tree model and a logical index, counterparty identification is performed on the trading institutions in the structured order information to confirm the counterparty in the text to be identified. By accurately extracting the entity element features in the text to be identified and realizing efficient structured processing of order texts based on a structured rule index, and further using a decision tree model in the form of soft coding to achieve counterparty identification with strong generalization ability, the scalability of counterparty identification is effectively improved and the training and maintenance costs are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an index-based scalable counterparty identification method, system and medium. Background Art

[0002] Spot trading is one of the most active trading forms in the financial trading market. In the text understanding of spot trading, the identification of spot trading counterparties is a very important issue. For the text understanding of spot trading, a series of intelligent algorithms and business rule indexes generally need to be integrated to complete the comprehensive understanding of spot trading texts. For example, for a common spot trading corpus "Bond A, 0000001.IB 1Y 3.901 XX institution sells to YY institution\n Bond C 0000002.IB 2+1Y 3.5 YY institution sells to ZZ institution requests HH institution", to understand this text and provide convenient and intelligent operations for downstream tasks, this text needs to be structured into the structure shown in Table 1:

[0003] Table 1 Structure after text structuring

[0004]

[0005] This includes several subtask processes. First is intention understanding, to judge whether the user is inquiring or making a deal; second is element extraction, to extract core elements provided by the user such as "bond name, interest rate, trading volume, counterparty", etc. For intention recognition and element extraction, such problems can be better solved based on deep neural networks.

[0006] However, for the identification of counterparties, based on "dialogue text", it is impossible to accurately judge the counterparty and the counterparty direction, and often "financial professional knowledge" outside the field is required for comprehensive judgment. Since usually a hard-coded calculation method is adopted and the rules are fixed and difficult to change, the maintenance cost of the counterparty identification model is very high and the scalability is poor.

[0007] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0008] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide an index-based scalable counterparty identification method, system and medium, aiming to improve the generalization ability and scalability of counterparty identification.

[0009] The technical solution of the present invention is as follows:

[0010] An index-based scalable counterparty identification method, comprising:

[0011] Obtain the text to be recognized, extract features from the text to be recognized, and obtain the entity elements in the text to be recognized;

[0012] Based on a preset structured rule index, perform structured processing on the text to be recognized according to the entity elements to obtain structured order information;

[0013] According to the pre-constructed and trained first decision tree model and logical index, identify the counterparty of the trading institution in the structured order information to confirm the counterparty in the text to be recognized.

[0014] In one embodiment, the obtaining the text to be recognized, extracting features from the text to be recognized, and obtaining the entity elements in the text to be recognized includes:

[0015] Perform character-by-character segmentation on the text to be recognized through a pre-trained model to obtain corresponding character vector encodings;

[0016] After extracting features from the character vector encodings, output the semantic feature vectors of each character;

[0017] Perform label prediction on the entity elements in the semantic feature vectors to label and obtain the entity elements in the text to be recognized.

[0018] In one embodiment, the based on a preset structured rule index, performing structured processing on the text to be recognized according to the entity elements to obtain structured order information includes:

[0019] Perform line break splitting on the text to be recognized to obtain several sub-texts per line and the entity elements included in each sub-text per line;

[0020] Perform category judgment on each sub-text per line according to the entity elements in each sub-text per line to obtain the information category of each sub-text per line;

[0021] According to the preset structured rule index and the information category of each sub-text per line, perform element aggregation and complement on each sub-text per line to obtain at least one structured order information.

[0022] In one embodiment, the performing category judgment on each sub-text per line according to the entity elements in each sub-text per line to obtain the information category of each sub-text per line includes:

[0023] Pre-construct a second decision tree model and entity feature training samples for training the second decision tree model;

[0024] Adopt the minimum Gini index as the feature selection criterion, and perform learning and training on the second decision tree model through the entity feature training samples;

[0025] Based on the entity elements in each line of sub-text, the trained second decision tree model judges the category of each line of sub-text to obtain the information category of each line of sub-text.

[0026] In one embodiment, the information categories include order information, shared information, supplementary information, and other information.

[0027] In one embodiment, before identifying the counterparty of the trading institution in the structured order information according to the pre-constructed and trained first decision tree model and the logical index and confirming the counterparty in the text to be identified, the method further includes:

[0028] Pre-construct a first decision tree model and institution feature training samples for training the first decision tree model, where the institution feature training samples include the location and direction of the trading institution;

[0029] Adopt the information gain strategy of the ID3 algorithm to train the first decision tree model through the institution feature training samples until the information gain of the first decision tree model reaches the maximum and the training is completed.

[0030] In one embodiment, identifying the counterparty of the trading institution in the structured order information according to the pre-constructed and trained first decision tree model and the logical index and confirming the counterparty in the text to be identified includes:

[0031] Obtain the location and trading direction of each trading institution in the structured order information;

[0032] According to the location and trading direction of each trading institution, judge the institution category of each trading institution through the trained first decision tree model, and the institution categories include buyer institution, seller institution, and bridging institution;

[0033] Based on the current visual institution, confirm the counterparty in the text to be identified relative to the visual institution according to the logical index and the institution category of each trading institution.

[0034] In one embodiment, based on the current visual institution, confirming the counterparty in the text to be identified relative to the visual institution according to the logical index and the institution category of each trading institution specifically includes:

[0035] If the current visual institution is a buyer institution, then the counterparty is the seller institution in the text to be identified;

[0036] If the current visual institution is a seller institution, then the counterparty is the buyer institution in the text to be identified;

[0037] If the current visual mechanism is a bridge mechanism, the counterparty is the buyer institution and the seller institution in the text to be recognized.

[0038] An index-based scalable counterparty recognition system, the system includes at least one processor; and,

[0039] A memory communicatively connected to the at least one processor; wherein,

[0040] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned index-based scalable counterparty recognition method.

[0041] A non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned index-based scalable counterparty recognition method.

[0042] Beneficial effects: The present invention discloses an index-based scalable counterparty recognition method, system and medium. Compared with the prior art, the embodiments of the present invention accurately extract the entity element features in the text to be recognized, and perform efficient order text structuring processing based on structured rule indexing. Further, a decision tree model in the form of soft coding is used to achieve strong generalization ability for counterparty recognition, effectively improving the scalability of counterparty recognition and reducing the training and maintenance costs. Description of the Drawings

[0043] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0044] Figure 1 is a flowchart of an index-based scalable counterparty recognition method provided by an embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of a second decision tree model in the index-based scalable counterparty recognition method provided by an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of functional modules of an index-based scalable counterparty recognition device provided by an embodiment of the present invention;

[0047] Figure 4 is a schematic diagram of the hardware structure of an index-based scalable counterparty recognition system provided by an embodiment of the present invention. Detailed Embodiments

[0048] To make the objectives, technical solutions and effects of the present invention more clear and definite, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments of the present invention will be introduced below with reference to the accompanying drawings.

[0049] Please refer to Figure 1 , Figure 1 which is a flowchart of an embodiment of the index-based scalable counterparty identification method provided by the present invention. The index-based scalable counterparty identification method provided in this embodiment is applicable to the situation of automatically identifying counterparties in the trading process. As Figure 1 shown, the method specifically includes the following steps:

[0050] S100. Obtain the text to be identified, extract features from the text to be identified, and obtain the entity elements in the text to be identified.

[0051] In this embodiment, the text to be identified may be the trading dialogue text information in the spot trading process, such as order information, consultation information, etc. sent between different trading institutions. By obtaining the text information in the trading dialogue as the text to be identified and automatically extracting the features therein, the entity elements in the text to be identified can be obtained. The specific entity elements may include institution names, trading directions, trading terms, trading volumes, trading methods, trading product names, etc., so as to achieve efficient text feature extraction as the data basis for subsequent counterparty identification. Of course, in other embodiments, the text to be identified is not limited to the text information in spot trading, but may also be the text information in other transactions, or the text information obtained by identifying and converting trading voice information, etc. This embodiment does not make any limitations in this regard.

[0052] In one embodiment, step S100 includes:

[0053] Segment the text to be identified character by character through a pre-trained model to obtain the corresponding character vector encoding;

[0054] Extract features from the character vector encoding and output the semantic feature vector of each character;

[0055] Perform label prediction on the entity elements in the semantic feature vector to label and obtain the entity elements in the text to be identified.

[0056] In this embodiment, when extracting text features, the to-be-recognized text is segmented character by character through a pre-trained model to obtain corresponding character vector encodings. Specifically, in this embodiment, a pre-trained model based on Bert is used for character encoding. Bert is a pre-trained language representation model, which emphasizes that instead of using the traditional unidirectional language model or the method of shallow splicing of two unidirectional language models for pre-training as in the past, a new MLM (masked language model) is adopted to generate deep bidirectional language representations. That is, for the input text, some words in the text are randomly masked with a certain probability, and then the Bert model is used to predict these masked words for pre-training to obtain the vector encoding of each character. Of course, in other embodiments, pre-trained models such as RoBerta and Ernie can also be used for character encoding, and this embodiment does not limit this.

[0057] After obtaining the character vector encodings, feature learning and extraction are performed to obtain the semantic feature vectors of each character. Specifically, the features of the sequence text after character encoding can be learned through an LSTM (Long Short-Term Memory) network structure to extract the semantic feature vectors of each character. Specifically, the character vector encodings obtained by the BERT pre-trained model are used as the input of the LSTM network, and an N-dimensional LSTM network structure is used for feature learning and extraction, and an N-dimensional semantic feature vector is output.

[0058] After obtaining the semantic feature vectors of the characters, the trained label prediction model can be used to decode and label the semantic feature vectors to obtain information such as the positions and categories of entity elements in the to-be-recognized text. Specifically, the label prediction training model can adopt the mature CRF (conditional random fields) algorithm to achieve efficient and accurate sequence labeling. That is, in this embodiment, through a deep learning model based on Bert+LSTM+CRF and targeted optimization for existing bonds, entity element information such as "institutional name, trading direction, trading term, trading volume, trading method, trading product name" in the sequence text is extracted to obtain various core element information in the sequence text. Of course, in other embodiments, other deep learning models with the same extraction effect can also be used, and this embodiment does not limit this.

[0059] S200. Based on a preset structured rule index, the to-be-recognized text is structured according to the entity elements to obtain structured order information.

[0060] In this embodiment, for unstructured text, after obtaining the entity elements in the text to be recognized, automatic structuring processing is performed based on a preset structured rule index. For the "multiple bonds" situation in a piece of information, multiple orders are cut to form structured order information with corresponding structures and complete information, so as to regularize the messy and possibly information - missing text to be recognized, so that subsequent efficient counter - party recognition processing can be based on accurate structured order information, improving the efficiency and accuracy of counter - party recognition.

[0061] In one embodiment, step S200 includes:

[0062] Perform line - break splitting on the text to be recognized to obtain several lines of sub - texts and the entity elements included in each line of sub - text;

[0063] Perform category judgment on each line of sub - text according to the entity elements in each line of sub - text to obtain the information category of each line of sub - text;

[0064] Perform element aggregation and complement on each line of sub - text according to the preset structured rule index and the information category of each line of sub - text to obtain at least one piece of structured order information.

[0065] In this embodiment, during the structuring process, first, based on the line - break character "\n" existing in the text to be recognized, the text to be recognized is split by the "\n" character to obtain several lines of sub - texts, thereby obtaining the entity elements included in each line of sub - text. Based on the entity elements in each line of sub - text, category judgment is performed on each line of information, and each line of information is judged as order information, shared information, supplementary information, or other information, that is, the information category of each line of sub - text is obtained. Among them, order information is defined as the text information with the core spot - bond trading theme such as "bond name or bond code" in this line of information; supplementary information is defined as the supplementary information without the core information of spot - bond trading in this line of information; shared information is the public shared information shared to each order information in this line of information; other information is some invalid text information, such as information without any valid entities.

[0066] For the judgment of each line of information, the usual method is to use hard - coded rule logic for judgment, which has poor flexibility and high maintenance cost. In the embodiment of the present invention, a decision - tree machine - learning model is adopted. Different from deep learning, a decision tree does not require a large number of training samples and has relatively good generalization ability.

[0067] Specifically, the performing category judgment on each line of sub - text according to the entity elements in each line of sub - text to obtain the information category of each line of sub - text includes:

[0068] Pre - construct a second decision - tree model and entity feature training samples for training the second decision - tree model;

[0069] Using the minimum Gini index as the feature selection criterion, the second decision tree model is trained and learned through the entity feature training samples;

[0070] According to the entity elements in each line of sub-text, the second decision tree model after training is used to judge the category of each line of sub-text, and the information category of each line of sub-text is obtained.

[0071] In this embodiment, the second decision tree model is pre-constructed and trained through entity feature training samples. The definition of the entity feature training samples is shown in Table 2. Based on the relevant experience and knowledge in the financial field, the entities in Table 2 are extracted as key features in this embodiment. At the same time, whether each entity exists in each line is used as a feature, quantified as "yes / no" (yes for existence, no for non-existence), and finally the entity feature training samples in the following table form are formed.

[0072] Table 2 Entity Feature Training Samples

[0073]

[0074]

[0075] Among them, category 1 represents order information, category 2 represents shared information, category 3 represents supplementary information, and category 4 represents other information. Considering the calculation efficiency and effect, the CART method is used to generate the decision tree in this embodiment. Since CART can only generate binary trees, the generated decision tree diagram is as Figure 2 shown. The criterion for generating the decision tree is to use the minimum Gini index as the feature selection. The second decision tree model is trained and learned through the entity feature training samples, and its formula is as follows:

[0076]

[0077] The Gini index of the set D under the condition of feature A is:

[0078]

[0079] Among them, D is the sample set, L k is the subset of samples belonging to the k-th category in D, K is the number of categories, D1 is the sample set when the feature exists, D2 is the sample set when the feature does not exist. Based on the second decision tree model after training, the category of each line of sub-text can be judged, and the information category of each line of sub-text can be obtained.

[0080] After obtaining the information categories of each line of sub-text, the sub-texts of each line are aggregated and supplemented with elements according to the preset structured rule index and the information categories of each line of sub-text, so as to combine them into one or more complete orders. The logic of the specific structured rule index is: with the "order as the core", the "supplementary information" is spliced into the "order information" according to the "upward supplement" logic. For example, "Buy: \n Bond A 000001 30 million XX institution sells to YY institution\n Sell:\n Bond B 000002 40 million XX institution sells to KK institution", where "\n" is the line break character. In this corpus, both "Buy" and "Sell" are "supplementary information". "Buy" will first be supplemented to the "order information" closest to it. At this time, the transaction direction of this order information is no longer missing, so the "Sell" transaction direction is no longer supplemented upward. Moreover, the next "order information" is also missing the transaction direction. Therefore, "Sell" is supplemented to the next "order information". The "shared information" is directly copied and supplemented into each "order information", thus forming one or more structured order information with "order" as the dimension.

[0081] S300. According to the pre-constructed and trained first decision tree model and logical index, identify the trading counterpart of the trading institution in the structured order information, and confirm the trading counterpart in the text to be identified.

[0082] In this embodiment, after obtaining one or more structured order information through structured processing, the order information contains core order elements, supplementary elements, and shared elements. In these element sequences, there are one or more "institution names". Subsequently, through the pre-constructed and trained first decision tree model and logical index, the trading counterpart of the trading institution in the structured order information is identified to achieve the identification of the trading counterpart.

[0083] For example, based on the established financial rules, "XX Bank sends to YY Fund" can express this type of structure as [Institution A to Institution B]; in the hard-coded logic, the solution sorts out various similar expression structures, such as [Institution A Institution B to Institution C], [To Institution A To Institution B], etc. According to various expressions, judge the directions of "Institution A, Institution B", etc. Since there are infinite expression forms in natural language, a large number of judgment logic rules are required for hard coding. In the embodiment of the present invention, a soft-coding form based on the "decision tree algorithm" is adopted, and the institution is judged through the pre-trained first decision tree model.

[0084] Specifically, before step S100, the method further includes:

[0085] Pre-construct a first decision tree model and an institution feature training sample for training the first decision tree model. The institution feature training sample includes the position and direction of the trading institution.

[0086] Adopt the information gain strategy of the ID3 algorithm, and train the first decision tree model with the institutional feature training samples until the training is completed when the information gain of the first decision tree model reaches the maximum.

[0087] In this embodiment, by pre - constructing the first decision tree model and training it with institutional feature training samples, since the judgment of the opponent's direction mainly relies on features such as "the position of the institution, direction, position of the direction, and the whole sentence pattern", therefore, in this embodiment, these features are innovatively expressed as the institutional feature training samples shown in Table 3, as the input features required by the decision tree algorithm.

[0088] Table 3 Institutional Feature Training Samples

[0089]

[0090] Among them, 1 represents existence, 0 represents non - existence, and the institutional feature training samples in Table 3 represent the positions of institutions and trading directions in the corpus. If the corpus has two institutions and one direction, such as "A bond 000001 3000W XX institution sells to YY institution", it conforms to the first item, the left institution 1 is "XX institution", the trading direction 1 is "sells to", and the institution 1 is "YY institution"; if the corpus is "A bond 000001 3000W Buyer: XX institution, Seller: YY institution", it conforms to the fifth item. At this time, "buys" is the trading direction 1, "XX institution" is the institution 1, "sells" is the trading direction 2, and "YY institution" is the right institution 1. Therefore, the first decision tree model is optimized for decision training by representing the positions of institutions and directions in the corpus.

[0091] In terms of the optimal decision, in this embodiment, the information gain strategy of ID3 is adopted. When the information gain of the entire decision tree increases after splitting according to a feature, it means that splitting according to this feature is correct. When the information gain of the decision tree can no longer increase, it means that the decision tree at this time is optimal. The specific formula is as follows:

[0092] g(E,B) = H(E)-H(E|B)

[0093]

[0094]

[0095] Among them, H(E) is the entropy of the data set E, H(E i ) is the entropy of the data set E i , H(E|B) is the conditional entropy of the data set E with respect to the feature B. E i is the sample subset of the data set E where the feature B takes the i - th value, Ck It is a sample subset of class k in E. n is the number of values that feature B can take, and K is the number of classes. Based on the trained first decision tree model, the trading institutions in the structured order information can be judged, and the institution category of the trading institution in each piece of structured order information can be obtained.

[0096] Specifically, step S300 includes:

[0097] Obtain the positions and trading directions of each trading institution in the structured order information;

[0098] According to the positions and trading directions of each trading institution, use the trained first decision tree model to judge the institution category of each trading institution. The institution categories include buyer institutions, seller institutions, and bridge institutions;

[0099] Based on the current visual institution, confirm the trading counterparty in the text to be recognized relative to the visual institution according to the logical index and the institution category of each trading institution.

[0100] In this embodiment, for each trading institution, first obtain its position and trading direction in the structured order information. For example, according to the position where each institution name is located, obtain the institution name and trading direction before its position, the institution name and trading direction after it, etc. as the features for judging the current institution's direction. Then, based on the positions and trading directions of each trading institution, use the trained first decision tree model to judge whether each trading institution is a buyer institution, a seller institution, or a bridge institution.

[0101] For each trading institution, based on the determination of the three categories of "buyer institution, seller institution, bridge institution", the identification of the trading counterparty can be obtained by combining the constructed logical index with the information of the trading direction. The identification of the trading counterparty needs to be determined according to the visual institution, that is, the construction of the logical index needs to be determined according to the visual party. The visual institution is the institution that sees this text to be recognized. The specific logical index is that if the current visual institution is a buyer institution, then the trading counterparty is the seller institution in the text to be recognized; if the current visual institution is a seller institution, then the trading counterparty is the buyer institution in the text to be recognized; if the current visual institution is a bridge institution, then the trading counterparty is the buyer institution and the seller institution in the text to be recognized. Therefore, in this embodiment, by combining financial logic to construct feature information and then using the decision tree algorithm in the form of soft coding for the identification and judgment of the counterparty, on the one hand, the requirements for training samples are reduced, and the cost is reduced; on the other hand, it has stronger generalization ability than the existing hard coding scheme, can solve more different types of problems, and solves the problem of high rule maintenance cost.

[0102] Another embodiment of the present invention provides an index-based scalable counterparty identification device, as Figure 3 shown. The device includes:

[0103] A feature extraction module 11, configured to obtain the text to be identified, extract features from the text to be identified, and obtain entity elements in the text to be identified;

[0104] A structuring module 12, configured to perform structuring processing on the text to be identified based on the entity elements according to a preset structuring rule index, and obtain structured order information;

[0105] A counterparty identification module 13, configured to identify the counterparty of the trading institution in the structured order information according to a pre-constructed and trained first decision tree model and a logical index, and confirm the counterparty in the text to be identified.

[0106] The feature extraction module 11, the structuring module 12, and the counterparty identification module 13 are connected in sequence. The modules referred to in the present invention refer to a series of computer program instruction segments that can complete specific functions, which are more suitable for describing the execution process of index-based scalable counterparty identification than programs. For the specific implementation manners of each module, please refer to the corresponding method embodiments above, and will not be elaborated here.

[0107] Another embodiment of the present invention provides an index-based scalable counterparty identification system, as Figure 4 shown. The system 10 includes:

[0108] One or more processors 110 and a memory 120. Figure 4 Here, taking one processor 110 as an example for introduction, the processor 110 and the memory 120 can be connected through a bus or other means. Figure 4 Here, taking the connection through a bus as an example.

[0109] The processor 110 is configured to complete various control logics of the system 10. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISCMachine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Further, the processor 110 can also be any conventional processor, microprocessor, or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.

[0110] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the index-based scalable counterparty identification method in the embodiments of the present invention. The processor 110 executes various functional applications and data processing of the system 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, that is, implements the index-based scalable counterparty identification method in the above method embodiments.

[0111] The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the system 10, etc. In addition, the memory 120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 120 optionally includes memories remotely provided relative to the processor 110, and these remote memories can be connected to the system 10 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0112] One or more units are stored in the memory 120 and, when executed by one or more processors 110, execute the index-based scalable counterparty identification method in any of the above method embodiments. For example, execute the Figure 1 method steps S100 to step S300 described above.

[0113] The embodiments of the present invention provide a non-volatile computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and these computer-executable instructions are executed by one or more processors. For example, execute the Figure 1 method steps S100 to step S300 described above.

[0114] By way of example, the non-volatile storage medium can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. The volatile memory can include random access memory (RAM) as an external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM, (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The disclosed memory components or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable types of memories.

[0115] In summary, in the index-based scalable counterparty identification method, system and medium disclosed by the present invention, the method obtains a text to be identified, extracts features of the text to be identified to obtain entity elements in the text to be identified; based on a preset structured rule index, the text to be identified is structured according to the entity elements to obtain structured order information; according to a first decision tree model and a logical index that are pre-constructed and trained, the trading institutions in the structured order information are identified as counterparties to confirm the counterparties in the text to be identified. By accurately extracting the entity element features in the text to be identified and implementing efficient order text structuring through order cutting based on the structured rule index, a decision tree model in the form of soft coding is further used to achieve counterparty identification with strong generalization ability, effectively improving the scalability of counterparty identification and reducing the training and maintenance costs.

[0116] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical memory, etc.

[0117] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. An index-based scalable counterparty identification method, characterized in that, Including: Obtain the text to be recognized, extract features from the text to be recognized, and obtain entity elements in the text to be recognized; Based on a preset structured rule index, perform structured processing on the text to be recognized according to the entity elements to obtain structured order information; According to a pre-constructed and trained first decision tree model and a logical index, identify the counterparty of the trading institution in the structured order information, and confirm the counterparty in the text to be recognized; The performing structured processing on the text to be recognized according to the entity elements based on a preset structured rule index to obtain structured order information includes: Perform line break splitting on the text to be recognized to obtain several lines of sub-texts and the entity elements included in each line of sub-text; Perform category judgment on each line of sub-text according to the entity elements in each line of sub-text to obtain the information category of each line of sub-text; According to the preset structured rule index and the information category of each line of sub-text, perform element aggregation and supplementation on each line of sub-text to obtain at least one piece of structured order information; The performing category judgment on each line of sub-text according to the entity elements in each line of sub-text to obtain the information category of each line of sub-text includes: Pre-construct a second decision tree model and entity feature training samples for training the second decision tree model; Adopt the minimum Gini index as the feature selection criterion, and perform learning and training on the second decision tree model through the entity feature training samples; According to the entity elements in each line of sub-text, perform category judgment on each line of sub-text through the trained second decision tree model to obtain the information category of each line of sub-text; Perform element aggregation and supplementation on each line of sub-text according to the preset structured rule index and the information category of each line of sub-text, and combine them into one or more orders.

2. The index-based scalable counterparty identification method according to claim 1, wherein The obtaining the text to be recognized, extracting features from the text to be recognized, and obtaining entity elements in the text to be recognized includes: Perform character-by-character segmentation on the text to be recognized through a pre-trained model to obtain corresponding character vector encodings; Extract features from the character vector encodings and output the semantic feature vectors of each character; Perform label prediction on the entity elements in the semantic feature vectors, and mark and obtain the entity elements in the text to be recognized.

3. The index-based scalable counterparty identification method according to claim 1, wherein The information categories include order information, shared information, supplementary information, and other information.

4. The index-based scalable counterparty identification method according to claim 1, wherein Before the identifying the counterparty of the trading institution in the structured order information according to a pre-constructed and trained first decision tree model and a logical index, and confirming the counterparty in the text to be recognized, the method further includes: Pre-construct a first decision tree model and institution feature training samples for training the first decision tree model, and the institution feature training samples include the location and direction of the trading institution; Adopt the information gain strategy of the ID3 algorithm, and perform training on the first decision tree model through the institution feature training samples until the information gain of the first decision tree model reaches the maximum and the training is completed.

5. The index-based scalable counterparty identification method according to claim 1, wherein Performing counterparty identification on the trading institutions in the structured order information according to the pre-constructed and trained first decision tree model and logical index, and confirming the counterparty in the text to be identified, including: Obtaining the positions and trading directions of each trading institution in the structured order information; Judging the institution category of each trading institution through the trained first decision tree model according to the positions and trading directions of each trading institution, where the institution category includes a buyer institution, a seller institution, and a bridging institution; Based on the current visual institution, confirming the counterparty in the text to be identified relative to the visual institution according to the logical index and the institution category of each trading institution.

6. The index-based scalable counterparty identification method according to claim 5, wherein The step of confirming the counterparty in the text to be identified relative to the visual institution based on the current visual institution according to the logical index and the institution category of each trading institution specifically includes: If the current visual institution is a buyer institution, the counterparty is the seller institution in the text to be identified; If the current visual institution is a seller institution, the counterparty is the buyer institution in the text to be identified; If the current visual institution is a bridging institution, the counterparty is the buyer institution and the seller institution in the text to be identified.

7. An index-based scalable counterparty identification system, characterized in that, The system includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the index-based extensible counterparty identification method according to any one of claims 1-6.

8. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can be enabled to execute the index-based extensible counterparty identification method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Block chain security event public opinion monitoring method and system

    CN112883734A

  • Management method and system for improving security of big data transaction platform

    CN113806350A