Method and device for realizing investor associated account detection processing based on large-scale time sequence graph, processor and readable storage medium thereof

By building a time sequence knowledge graph of large-scale investor trading behaviors and a multi-task learning framework, identifying and monitoring related transactions in the securities market, the problem of difficulty in identifying related transactions in the existing technology is solved, and the level of supervision intelligence and transaction fairness are improved.

CN120011418APending Publication Date: 2025-05-16GUOTAI JUNAN SECURITIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411850426.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively identify and monitor related transactions in the securities market, resulting in the difficulty of abnormal behaviors such as market manipulation and insider trading to be discovered in a timely manner.

Method used

By constructing a time sequence knowledge graph for large-scale investors' trading behavior, we calculate the similarity of related accounts based on basic information and trading terminals, and use a multi-task learning framework to learn investor trading behavior patterns, and finally divide the related account groups based on joint similarity.

Benefits of technology

It improves the reliability and interpretability of investor relationships, allowing regulators to promptly discover abnormal relationships among investors in the market and maintain fairness in financial market transactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011418A_ABST
    Figure CN120011418A_ABST
Patent Text Reader

Abstract

The invention relates to a method for realizing investor associated account detection processing based on a large-scale time sequence graph. The method comprises the following steps: constructing a large-scale investor transaction behavior time sequence knowledge graph; calculating an associated account similarity based on the basic information and the transaction terminal; calculating the similarity of associated accounts embedded based on the time sequence interaction diagram; and dividing associated account groups based on the joint similarity. The invention also relates to a device for realizing investor associated account detection processing based on the large-scale time sequence graph, a processor and a readable storage medium thereof. According to the method and device for realizing investor associated account detection processing based on the large-scale time sequence graph, the processor and the computer readable storage medium thereof, multi-view investor information similarity calculation is utilized, and an investor transaction behavior mode is learned based on a multi-task learning framework; according to the scheme, the reliability and interpretability of investor association relationship mining are enhanced, the reliability of investor association relationship mining is improved, and financial market transaction fairness is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of securities, and in particular to the field of knowledge graphs, and specifically refers to a method, device, processor and computer-readable storage medium thereof for detecting and processing investor-associated accounts based on a large-scale time series graph. Background Art

[0002] With the widespread application of artificial intelligence and big data technologies in the securities field, trading product systems and technological innovations have developed rapidly, investors' trading behaviors have become increasingly complex, and the concealment of abnormal trading behaviors such as market manipulation and insider trading has continued to increase, and new abnormal trading behaviors have emerged in an endless stream. In particular, related-party transactions (also known as coordinated transactions) pose a serious threat to the integrity of the financial market. Regulators regard related-party transactions as highly risky behaviors because orders that appear to come from independent traders may actually be manipulated by the same entity or organization. Such coordinated transactions often cause abnormal fluctuations in prices and trading volumes, further triggering financial fraud such as insider trading, price manipulation and deception. Related-party transactions are mainly manifested as coordinated behaviors between multiple traders or organizations, aimed at misleading regulators and other market participants. Traditional methods that rely on manual experience and rules have exposed deficiencies in comprehensiveness, accuracy and timeliness when facing trading big data. They can no longer meet the needs of identifying abnormal trading behaviors under the new situation, and it is urgent to improve the level of intelligence in securities market supervision.

[0003] As a structured form of human knowledge expression, knowledge graph plays an important role in achieving semantic interoperability of multi-source heterogeneous data and provides effective support for tasks such as data analysis. In recent years, it has become a research hotspot in academia and industry. At present, most knowledge graphs are built based on static, non-real-time data, without fully considering the temporal attributes of entities and relationships. However, data in application scenarios such as social network communications, financial transactions, and epidemic spread present real-time dynamics and complex temporal characteristics. How to use time series data to build and effectively model knowledge graphs is an important challenge. At present, many research works have enriched the features of knowledge graphs by introducing temporal information, expanding them into a four-tuple form containing temporal attributes (head entity, relationship, tail entity, time), forming a temporal knowledge graph to better support the knowledge representation and application of dynamic data.

[0004] In the construction of the financial time-series transaction knowledge graph, on the one hand, based on the real-time and complexity of abnormal transaction identification in the securities market, the graph needs to cover key business entities and features that are closely related to anomaly identification; on the other hand, the graph database needs to support distributed storage of data at the scale of hundreds of millions of nodes, and have the ability to expand and shrink online horizontally, ensuring low query latency when processing ultra-large-scale data to meet the needs of efficient real-time queries.

[0005] Self-supervised learning of graph data is a method that automatically generates supervision signals through the structure or attribute information of the graph itself without manual labeling for model training. In graph data, self-supervised learning usually uses graph structural characteristics (such as node neighbor relationships, connection patterns, paths, etc.) and node or edge attributes to generate pseudo labels to construct learning tasks. For example, by constructing self-supervised objectives such as node similarity tasks, edge prediction tasks, and node reconstruction tasks, the model can be guided to capture semantic relationships and topological information in the graph during the learning process. In practical applications, self-supervised learning of graph data can generate high-quality node vectors or graph embedding representations that not only retain the attribute information of the nodes, but also contain complex structural relationships. Compared with traditional supervised learning, self-supervised learning methods show stronger generalization ability and data utilization efficiency in the absence of labels or when labels are scarce, and have been widely used in social network analysis, recommendation systems, financial risk control and other fields. Summary of the invention

[0006] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method, device, processor and computer-readable storage medium thereof for detecting and processing investor-associated accounts based on large-scale time series graphs, which have good reliability, high accuracy and a wide range of applications.

[0007] In order to achieve the above-mentioned purpose, the method, device, processor and computer-readable storage medium thereof for detecting and processing investor-related accounts based on a large-scale time series graph of the present invention are as follows:

[0008] The method for detecting and processing investor-related accounts based on a large-scale time series graph is mainly characterized in that the method comprises the following steps:

[0009] (1) Construct a time-series knowledge graph of large-scale investor trading behaviors;

[0010] (2) Calculate the similarity of associated accounts based on basic information and transaction terminals;

[0011] (3) Calculate the similarity of associated accounts based on the temporal interaction graph embedding;

[0012] (4) Divide the associated account groups based on joint similarity.

[0013] Preferably, the large-scale investor transaction behavior time-series knowledge graph records and updates the relationship edges between investors and their transaction terminals, the fund transfer relationship edges between fund accounts and bank accounts, and the buy and sell transaction relationship data between securities accounts and stocks at an appropriate time granularity.

[0014] Preferably, the step (2) is specifically as follows:

[0015] Use Jaccard similarity to calculate the similarity scores between investment account i and investment account j;

[0016] The similarity score between investment account i and investment account j is calculated according to the following formula:

[0017]

[0018] in, Indicates that investment account i in attribute set C within a certain period of time T k The cross feature set on Indicates that investment account j in attribute set C in a certain period of time T k The cross feature set on .

[0019] Preferably, the step (3) specifically comprises the following steps:

[0020] (3.1) Group investors’ trading behaviors according to different trading frequencies;

[0021] (3.2) The model adopts a multi-task learning strategy to jointly learn and reconstruct investor trading behavior data from two different perspectives and conduct comparative learning.

[0022] Preferably, the step (3.2) specifically includes the following steps:

[0023] The reconstruction-based module randomly masks the information of investors’ trading objects, i.e. the target nodes, and reconstructs them through an autoencoder based on a graph neural network to learn the global information of the graph structure.

[0024] The module based on contrastive learning learns the temporal correlation of individual investors and the differences between different investors from two aspects: temporal difference contrastive learning and investor difference contrastive learning. From the perspective of temporal difference contrast, two pairs of behavioral data with close transaction time intervals are regarded as positive pairs, and those with far transaction time intervals are regarded as negative pairs. From the perspective of investor difference contrast, the transaction behavior data of two investors who buy and sell related targets are regarded as positive pairs, and those of two investors who buy and sell unrelated targets are regarded as negative pairs.

[0025] Preferably, the step (4) specifically comprises the following steps:

[0026] (4.1) Using investor accounts as nodes and similarity coefficients between investor accounts as edges, construct an investor account similarity graph;

[0027] (4.2) Based on the weighted account similarity graph, multiple community discovery methods are used to group the account groups of related transactions.

[0028] Preferably, the step (4.1) specifically comprises the following steps:

[0029] (4.1.1) If either the similarity of the trading terminal associated accounts or the trading model similarity reaches a threshold, an edge is connected between the two investor nodes;

[0030] (4.1.2) Set high weights for edges related to basic information and transaction terminals, and set low weights for edges based on transaction behavior patterns.

[0031] Preferably, the step (4.2) is specifically as follows:

[0032] The Louvain algorithm is used to divide closely related accounts into a group, and several groups of suspected related accounts are obtained.

[0033] Preferably, the Louvain algorithm is specifically:

[0034] Detect the community structure of the network by maximizing modularity;

[0035] Modularity is calculated according to the following formula:

[0036]

[0037] Among them, A ij Indicates whether there is an edge between node i and node j, where 1 indicates existence and 0 indicates non-existence; k i and k j are the weighted degrees of nodes i and j respectively; m is the total number of edges in the network; C i and C j Respectively represent the communities where nodes i and j are located; δ(C i ,C j ) is the Kronecker function, when C i =C j The value is 1 when , otherwise it is 0.

[0038] The main feature of the device for detecting and processing investor-related accounts based on a large-scale time series graph is that the device comprises:

[0039] a processor configured to execute computer executable instructions;

[0040] The memory stores one or more computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned method for detecting and processing investor-associated accounts based on a large-scale time series graph are implemented.

[0041] The processor for implementing investor-associated account detection and processing based on a large-scale time series graph has the main feature that the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the above-mentioned method for implementing investor-associated account detection and processing based on a large-scale time series graph are implemented.

[0042] The main feature of the computer-readable storage medium is that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for detecting and processing investor-associated accounts based on a large-scale time series graph.

[0043] The method, device, processor and computer-readable storage medium for detecting and processing investor-related accounts based on large-scale time series graphs of the present invention have been adopted to solve the pain point that it is difficult to capture hidden investor relationships using traditional manual or static rules under the current massive transaction data. The scheme enhances the reliability and interpretability of investor relationship mining by calculating the similarity of investor information from multiple perspectives and learning investor trading behavior patterns based on a multi-task learning framework. The present invention improves the reliability of investor relationship mining, enabling regulators or financial service companies to promptly discover abnormal investor relationships in the market and maintain fairness in financial market transactions. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 The present invention is a flowchart of a method for detecting and processing investor-associated accounts based on a large-scale time series graph.

[0045] Figure 2 1) A schematic diagram of a large-scale investor transaction behavior time series knowledge graph of the method for realizing investor associated account detection and processing based on a large-scale time series graph of the present invention. DETAILED DESCRIPTION

[0046] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.

[0047] The method for detecting and processing investor-associated accounts based on a large-scale time series graph of the present invention comprises the following steps:

[0048] (1) Construct a time-series knowledge graph of large-scale investor trading behaviors;

[0049] (2) Calculate the similarity of associated accounts based on basic information and transaction terminals;

[0050] (3) Calculate the similarity of associated accounts based on the temporal interaction graph embedding;

[0051] (4) Divide the associated account groups based on joint similarity.

[0052] As a preferred embodiment of the present invention, the large-scale investor trading behavior time-series knowledge graph records and updates the relationship edges between investors and their trading terminals, the fund transfer relationship edges between fund accounts and bank accounts, and the buying and selling transaction relationship data between securities accounts and stocks at an appropriate time granularity.

[0053] As a preferred embodiment of the present invention, the step (2) is specifically as follows:

[0054] Use Jaccard similarity to calculate the similarity score between investment account i and investment account j;

[0055] The similarity score between investment account i and investment account j is calculated according to the following formula:

[0056]

[0057] in, Indicates that investment account i in attribute set C in a certain period of time T k The cross feature set on Indicates that investment account j in attribute set C in a certain period of time T k The cross feature set on .

[0058] As a preferred embodiment of the present invention, the step (3) specifically comprises the following steps:

[0059] (3.1) Group investors’ trading behaviors according to different trading frequencies;

[0060] (3.2) The model adopts a multi-task learning strategy to jointly learn and reconstruct investor trading behavior data from two different perspectives and conduct comparative learning.

[0061] As a preferred embodiment of the present invention, the step (3.2) specifically includes the following steps:

[0062] The reconstruction-based module randomly masks the information of investors’ trading objects, i.e. the target nodes, and reconstructs them through an autoencoder based on a graph neural network to learn the global information of the graph structure.

[0063] The module based on contrastive learning learns the temporal correlation of individual investors and the differences between different investors from two aspects: temporal difference contrastive learning and investor difference contrastive learning. From the perspective of temporal difference contrast, two pairs of behavioral data with close transaction time intervals are regarded as positive pairs, and those with far transaction time intervals are regarded as negative pairs. From the perspective of investor difference contrast, the transaction behavior data of two investors who buy and sell related targets are regarded as positive pairs, and those of two investors who buy and sell unrelated targets are regarded as negative pairs.

[0064] As a preferred embodiment of the present invention, the step (4) specifically comprises the following steps:

[0065] (4.1) Using investor accounts as nodes and similarity coefficients between investor accounts as edges, construct an investor account similarity graph;

[0066] (4.2) Based on the weighted account similarity graph, multiple community discovery methods are used to group the account groups of related transactions.

[0067] As a preferred embodiment of the present invention, the step (4.1) specifically includes the following steps:

[0068] (4.1.1) If either the similarity of the trading terminal associated accounts or the trading model similarity reaches a threshold, an edge is connected between the two investor nodes;

[0069] (4.1.2) Set high weights for edges related to basic information and transaction terminals, and set low weights for edges based on transaction behavior patterns.

[0070] As a preferred embodiment of the present invention, the step (4.2) is specifically as follows:

[0071] The Louvain algorithm is used to divide closely related accounts into a group, and several groups of suspected related accounts are obtained.

[0072] As a preferred embodiment of the present invention, the Louvain algorithm is specifically:

[0073] Detect the community structure of the network by maximizing modularity;

[0074] Modularity is calculated according to the following formula:

[0075]

[0076] Among them, A ij Indicates whether there is an edge between node i and node j, where 1 indicates existence and 0 indicates non-existence; k i and k j are the weighted degrees of nodes i and j respectively; m is the total number of edges in the network; C i and C j Respectively represent the communities where nodes i and j are located; δ(C i ,C j ) is the Kronecker function, when C i =C j The value is 1 when , otherwise it is 0.

[0077] The device for detecting and processing investor-associated accounts based on a large-scale time series graph of the present invention comprises:

[0078] a processor configured to execute computer executable instructions;

[0079] The memory stores one or more computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned method for detecting and processing investor-associated accounts based on a large-scale time series graph are implemented.

[0080] The processor of the present invention is used to implement investor-associated account detection and processing based on a large-scale time series graph, wherein the processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, the various steps of the above-mentioned method for implementing investor-associated account detection and processing based on a large-scale time series graph are implemented.

[0081] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for detecting and processing investor-associated accounts based on a large-scale time series graph.

[0082] In a specific implementation of the present invention, by constructing a large-scale financial time-series transaction knowledge graph and introducing graph structure information, time-series information and graph information are effectively integrated, and investor characteristics are modeled from multiple angles such as investor basic information, transaction terminal information, and transaction behavior information, so as to more accurately detect abnormal related accounts in the securities market.

[0083] The present invention provides a complete solution for mining abnormal related transactions from the massive investor transaction behavior records in the securities and stock markets. The solution mainly includes three modules, and the overall flow chart of the solution is shown in the figure below:

[0084] (1) Constructing a time-series knowledge graph of large-scale investor trading behaviors:

[0085] In order to mine potential abnormal related transaction behaviors in the stock trading market, the present invention designs a large-scale investor transaction behavior time series knowledge graph, covering multiple entities such as individual investors, institutional investors, product investors, and multiple relationships such as contact addresses and operating IP addresses. In the graph structure, the relationship edges between investors and their trading terminals (such as mobile phones, IP, MAC, hard disks, etc.), the fund transfer relationship edges between fund accounts and bank accounts, and the buy and sell transaction relationship data between securities accounts and stocks will be recorded and updated at an appropriate time granularity to support efficient abnormal behavior identification.

[0086] (2) Calculate the similarity of associated accounts based on basic information and transaction terminals:

[0087] The basic information of investors and the information of trading terminals include the investor's contact number, contact address, IP address of the operation when the transaction occurs, MAC address of the device, etc. This type of information is usually easy to obtain, does not require complex expert rules or model processing, can directly reflect the static attributes of investors, and is the "first line of defense" for detecting investor association relationships. In actual business scenarios, there are often situations where one person or a group of actual controllers control multiple accounts to conduct transactions through the same device or under the same IP network segment. By querying the association path of basic information and trading terminals, such groups of associated trading accounts can be effectively identified.

[0088] There are many similarity calculations for this type of discrete data, such as the Jaccard Similarity method to mine the basic information and transaction terminal associations between accounts. Jaccard Similarity is a statistical method to measure the similarity and difference between two sets, and is widely used in text mining, recommendation systems, cluster analysis and other fields. Its core idea is to evaluate the similarity of two sets by calculating the ratio of their intersection to their union. Because it is based on the ratio of intersection and union, Jaccard Similarity performs well when processing sparse data (such as Boolean vectors of documents or features), and is more suitable for association detection in this scenario.

[0089] Specifically, the attribute set C is formed by using the personalized attributes of the account such as the operation IP, MAC address, account address, and mobile phone number. Indicates that investment account i in attribute set C within a certain period of time T k Using the Jaccard similarity calculation method, the similarity scores of investment account i and investment account j are calculated, which can be formally expressed as:

[0090]

[0091] (3) Calculate the similarity of associated accounts based on the temporal interaction graph embedding:

[0092] Based on the time series heterogeneous graph composed of investors' historical entrustment and transaction data, and combined with investors' attribute information, modeling is carried out. By learning the topological structure of the graph data, the machine learning model is trained in a self-supervised manner to obtain user node vector representations containing rich information. This vector combines the user's personal attributes and trading behavior information, so that the similarity of the investor's associated accounts in trading behavior patterns can be effectively calculated. Therefore, this part includes two aspects: investor trading behavior embedding representation learning and similarity calculation.

[0093] The core innovation of this invention is to embed and represent investors’ trading behaviors. In order to effectively learn investors’ trading behavior information, this solution combines two common proxy tasks of self-supervised learning: reconstruction and constructive learning, and is first implemented on a large-scale investor trading behavior knowledge graph.

[0094] Specifically, first, investors' trading behaviors are grouped according to different trading frequencies to ensure a unified scale during model training. Secondly, the model adopts a multi-task learning strategy to jointly learn the investor trading behavior data information from two different perspectives: reconstruction and contrastive learning. The reconstruction-based module will randomly mask the investor's trading object information, that is, the target node, and reconstruct it through an autoencoder based on a graph neural network to learn the global information of the graph structure; on the other hand, the contrastive learning-based module will learn the temporal correlation of a single investor and the differences between different investors from two aspects: temporal difference contrastive learning and investor difference contrastive learning. Specifically, the difference between the two types of contrastive learning lies in the screening of positive and negative samples. The current transaction of an investor is used as a positive sample. The temporal difference contrastive learning aims to learn the historical trading changes of the investor, using his long-term trading behavior as a negative example pair and recent trading behavior data as a positive example pair; investor difference contrastive learning uses other investors' behaviors as negative examples for comparison. Such joint learning can effectively learn the attribute information of investor trading behavior, so that the model focuses on the local information of the trading pattern itself. From the perspective of time series difference comparison, two pairs of behavioral data with close transaction time intervals are regarded as positive pairs, and those with far transaction time intervals are regarded as negative pairs; from the perspective of investor difference comparison, the transaction behavior data of two investors who buy and sell related targets are regarded as positive pairs, and those of two investors who buy and sell unrelated targets are regarded as negative pairs.

[0095] Based on the above multi-task learning framework pre-training, an embedding representation learning model of investor trading behavior patterns can be obtained. Based on the embedding representation of investor accounts learned by this model, a feasible solution is to use the Pearson correlation coefficient similarity calculation method to calculate the similarity between two accounts. The Pearson correlation coefficient is widely used in statistics, machine learning, and economic analysis to evaluate the correlation between variables. At the same time, the Pearson correlation coefficient similarity calculation method can efficiently operate on a large number of data sets, is independent of the dimension of the data during calculation, and can discover the bidirectional relationship of correlation between investors.

[0096] (4) Divide the associated account groups based on joint similarity:

[0097] Based on the above-mentioned basic information, similarity of associated accounts of trading terminals, and similarity of trading models, two parts of data, the present invention uses investor accounts as nodes and similarity coefficients between investor accounts as edges to construct an investor account similarity map. Specifically, if a certain type of similarity reaches the threshold, an edge is connected between two investor nodes. Since basic information and trading terminals can often reflect associated accounts with certainty, this scheme first sets a higher weight for such edges, and conversely sets a lower weight for edges based on trading behavior patterns. Therefore, there are four possibilities for the investor node relationship in the similarity map: there is no associated edge, that is, the edge weight is 0; there is only a suspected association in the trading model, that is, the edge weight is low; there is only a suspected association relationship between the basic information and the trading terminal, that is, the edge weight is high; there is an association relationship in the trading model, basic information and trading terminal, and the weight is the highest.

[0098] Based on the above-mentioned weighted account similarity graph, this solution can use a variety of community discovery methods to group account groups of related transactions. For example, the community discovery algorithm Louvain algorithm (Louvain Method for Community Detection) can be used to divide closely related accounts into a group, thereby obtaining several suspected related account groups. The Louvain algorithm is a community discovery algorithm based on modularity optimization, which is used to identify community structures with high-density node clusters in the network. It is widely used in social network analysis, biological networks, information dissemination, and recommendation systems due to its high efficiency and good performance in large and complex networks. The core idea of ​​the Louvain algorithm is to detect the community structure of the network by maximizing modularity. Modularity measures the difference between the network partition result and the random partition. The larger the value, the more likely the nodes in the network are to be closely connected to the nodes in the same community. The calculation formula for modularity is as follows:

[0099]

[0100] Among them, A ij Indicates whether the edge between node i and node j exists (1 means it exists, 0 means it does not exist); k i and k j are the weighted degrees of nodes i and j respectively; m is the total number of edges in the network; C i and C j Respectively represent the communities where nodes i and j are located; δ(C i ,C j ) is the Kronecker function, when C i =C j The value is 1 when , otherwise it is 0.

[0101] Construction of a large-scale investor trading behavior knowledge graph: In order to explore potential abnormal related trading behaviors in the stock trading market, the present invention designs a large-scale investor trading behavior time-series knowledge graph, covering multiple entities such as individual investors, institutional investors, product investors, and multiple relationships such as contact addresses and operating IP addresses.

[0102] Similarity calculation method for investor trading behavior from two different perspectives: This scheme designs a set of similarity calculation methods from two different perspectives. On the one hand, this scheme uses the similarity calculation method for discrete data (such as JaccardSimilarity) to mine the basic information between accounts and the relationship between trading terminals. On the other hand, this scheme uses a self-supervised learning framework to learn the dynamic trading patterns of investors. Finally, the similarity map is constructed using two joint similarities.

[0103] Learning the representation of investors’ dynamic trading patterns based on a self-supervised learning framework: In order to learn investors’ dynamic trading patterns, this scheme designs a multi-task self-supervised training framework based on reconstruction and contrastive learning, which jointly learns the global topological structure and local trading attribute information of the knowledge graph of investors’ trading behavior.

[0104] The specific implementation scheme of this embodiment can refer to the relevant description in the above embodiment, which will not be repeated here.

[0105] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0106] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.

[0107] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0108] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0109] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the corresponding program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.

[0110] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0111] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0112] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0113] The method, device, processor and computer-readable storage medium for detecting and processing investor-related accounts based on large-scale time series graphs of the present invention have been adopted to solve the pain point that it is difficult to capture hidden investor relationships using traditional manual or static rules under the current massive transaction data. The scheme enhances the reliability and interpretability of investor relationship mining by calculating the similarity of investor information from multiple perspectives and learning investor trading behavior patterns based on a multi-task learning framework. The present invention improves the reliability of investor relationship mining, enabling regulators or financial service companies to promptly discover abnormal investor relationships in the market and maintain fairness in financial market transactions.

[0114] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it is apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.

Claims

1. A method for detecting and processing investor-related accounts based on a large-scale time series graph, characterized in that: The method comprises the following steps: (1) Construct a time-series knowledge graph of large-scale investor trading behaviors; (2) Calculate the similarity of associated accounts based on basic information and transaction terminals; (3) Calculate the similarity of associated accounts based on the temporal interaction graph embedding; (4) Divide the associated account groups based on joint similarity.

2. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 1 is characterized in that: The large-scale investor transaction behavior time series knowledge graph records and updates the relationship edges between investors and their transaction terminals, the fund transfer relationship edges between fund accounts and bank accounts, and the buy and sell transaction relationship data between securities accounts and stocks at an appropriate time granularity.

3. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 1 is characterized in that: The step (2) is specifically as follows: Use Jaccard similarity to calculate the similarity scores between investment account i and investment account j; The similarity score between investment account i and investment account j is calculated according to the following formula: in, Indicates that investment account i in attribute set C within a certain period of time T k The cross feature set on Indicates that investment account j in attribute set C in a certain period of time T k The cross feature set on .

4. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 1, characterized in that: The step (3) specifically comprises the following steps: (3.1) Group investors’ trading behaviors according to different trading frequencies; (3.2) The model adopts a multi-task learning strategy to jointly learn and reconstruct investor trading behavior data from two different perspectives and conduct comparative learning.

5. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 4 is characterized in that: The step (3.2) specifically comprises the following steps: The reconstruction-based module randomly masks the information of investors’ trading objects, i.e. the target nodes, and reconstructs them through an autoencoder based on a graph neural network to learn the global information of the graph structure. The module based on contrastive learning learns the temporal correlation of individual investors and the differences between different investors from two aspects: temporal difference contrastive learning and investor difference contrastive learning. From the perspective of temporal difference contrast, two pairs of behavioral data with close transaction time intervals are regarded as positive pairs, and those with far transaction time intervals are regarded as negative pairs. From the perspective of investor difference contrast, the transaction behavior data of two investors who buy and sell related targets are regarded as positive pairs, and those of two investors who buy and sell unrelated targets are regarded as negative pairs.

6. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 1, characterized in that: The step (4) specifically comprises the following steps: (4.1) Using investor accounts as nodes and similarity coefficients between investor accounts as edges, construct an investor account similarity graph; (4.2) Based on the weighted account similarity graph, multiple community discovery methods are used to group the account groups of related transactions.

7. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 6 is characterized in that: The step (4.1) specifically comprises the following steps: (4.1.1) If either the similarity of the trading terminal associated accounts or the trading model similarity reaches a threshold, an edge is connected between the two investor nodes; (4.1.2) Set high weights for edges related to basic information and transaction terminals, and set low weights for edges based on transaction behavior patterns.

8. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 6 is characterized in that: The step (4.2) is specifically: The Louvain algorithm is used to divide closely related accounts into a group, and several groups of suspected related accounts are obtained.

9. The method for detecting and processing investor-related accounts based on a large-scale time series graph according to claim 8, characterized in that: The Louvain algorithm is specifically: Detect the community structure of the network by maximizing modularity; Modularity is calculated according to the following formula: Among them, A ij Indicates whether there is an edge between node i and node j, where 1 indicates existence and 0 indicates non-existence; k i and k j are the weighted degrees of nodes i and j respectively; m is the total number of edges in the network; C i and C j Respectively represent the communities where nodes i and j are located; δ(C i , C j ) is the Kronecker function, when C i =C j The value is 1 when , otherwise it is 0.

10. A device for detecting and processing investor-related accounts based on a large-scale time series graph, characterized in that: The device comprises: a processor configured to execute computer executable instructions; A memory storing one or more computer executable instructions, wherein when the computer executable instructions are executed by the processor, the steps of the method for detecting and processing investor associated accounts based on a large-scale time series graph as described in any one of claims 1 to 9 are implemented.

11. A processor for detecting and processing investor-related accounts based on a large-scale time series graph, characterized in that: The processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for detecting and processing investor-associated accounts based on a large-scale time series graph as described in any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the method for detecting and processing investor-associated accounts based on a large-scale time series graph as described in any one of claims 1 to 9.