Method, apparatus, device and medium for identifying shell companies
By building a directed knowledge graph and analyzing the number of suspicious entities, the problem of low shell company identification efficiency is solved, and fast and convenient shell company identification is achieved.
Patent Information
- Application Number
- CN202111577520.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The identification efficiency of the prior art hollow shell companies is low, the identification cost is high and time-consuming, making it difficult to quickly and accurately identify multiple shell companies.
By obtaining the entity information of the enterprise to be identified, the enterprise investment relationship information and the legal representative relationship information, a directed knowledge graph is constructed, and whether there are suspicious entities in the directed knowledge graph is determined, and whether it is a shell company is judged based on the number of suspicious entities.
It realizes the identification of shell companies quickly, conveniently and accurately, and improves the identification efficiency of shell companies.
Smart Images

Figure CN114240634B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of financial technology (Fintech), and particularly to a method, device, equipment and medium for identifying shell companies. Background Art
[0002] Shell companies are one of the important risks that financial institutions need to be vigilant about. However, due to the limitations of identification methods and due diligence costs, the chaos of shell companies has trapped financial institutions in multiple scenarios. Currently, financial institutions often need to investigate multiple dimensions such as the financial statements, assets, taxes, and liabilities of a single enterprise to determine whether a single enterprise is a shell company. In the existing solutions, when identifying multiple shell companies, the identification cost is high and the time consumption is long, resulting in a low current identification efficiency of shell companies. Summary of the Invention
[0003] The main purpose of this application is to provide a method, device, equipment and medium for identifying shell companies, aiming to solve the technical problem of low current identification efficiency of shell companies.
[0004] To achieve the above purpose, an embodiment of this application provides a method for identifying shell companies, and the method for identifying shell companies includes:
[0005] Obtain the entity information, enterprise investment relationship information and legal representative relationship information of the enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information and the legal representative relationship information;
[0006] Determine whether there are suspicious entities in the directed knowledge graph;
[0007] If there are the suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities.
[0008] Preferably, the step of determining whether there are suspicious entities in the directed knowledge graph includes:
[0009] Determine whether there is a central node in the directed knowledge graph;
[0010] If there is the central node, perform entity screening on all entities in the directed knowledge graph according to the central node to obtain an entity set, where the entity set includes each connected enterprise corresponding to the entity of the central node;
[0011] Determine the similarity between every two connected enterprises in the entity set respectively;
[0012] Determine whether each connected enterprise is a suspicious entity according to each similarity respectively;
[0013] If any of the outgoing enterprises among the said outgoing enterprises is a suspicious entity, it is determined that there is a suspicious entity in the said directed knowledge graph.
[0014] Preferably, the step of respectively determining whether each of the outgoing enterprises is a suspicious entity according to each of the similarities includes:
[0015] For each of the outgoing enterprises, the following steps are respectively executed:
[0016] Compare the name similarity between the current outgoing enterprise and the outgoing enterprises in the entity set with a first preset similarity threshold;
[0017] If the similarity is greater than or equal to the first preset similarity threshold, determine whether the registered addresses of the two outgoing enterprises corresponding to the similarity match;
[0018] If the registered addresses of the two outgoing enterprises corresponding to the similarity match, determine that the current outgoing enterprise and the corresponding other outgoing enterprises are suspicious entities.
[0019] Preferably, the step of determining whether the suspicious entity is a shell company according to the number of the suspicious entities includes:
[0020] Determine the suspicious entities with the same registered address in the entity set as the target suspicious entities, and determine the number of the target suspicious entities;
[0021] Compare the number of the target suspicious entities with a preset number threshold;
[0022] If the number of the target suspicious entities is greater than or equal to the preset number threshold, determine that each of the target suspicious entities is a shell company.
[0023] Preferably, after the step of determining that each of the target suspicious entities is a shell company, it further includes:
[0024] Obtain a first sub-graph corresponding to the shell company from the directed knowledge graph;
[0025] Determine a second sub-graph of the central node in the directed knowledge graph;
[0026] Based on the first sub-graph, determine whether the entities in the second sub-graph are the gangs of the shell company.
[0027] Preferably, the step of determining whether the entities in the second sub-graph are the gangs of the shell company based on the first sub-graph includes:
[0028] Extract a first feature to be compared of the first sub-graph and a second feature to be compared of the second sub-graph;
[0029] Determine the structural similarity, attribute similarity, and central node importance similarity of the first sub-graph and the second sub-graph based on the first feature to be compared and the second feature to be compared, respectively;
[0030] Determine whether the entities in the second sub-graph are the gangs of shell companies according to the structural similarity, the attribute similarity, and the central node importance similarity.
[0031] Preferably, the step of determining whether the entities in the second sub-graph are the gangs of shell companies according to the structural similarity, the attribute similarity, and the central node importance similarity includes:
[0032] Perform a weighted operation on the structural similarity, the attribute similarity, and the central node importance similarity to obtain a weighted similarity;
[0033] Compare the weighted similarity with a second preset similarity threshold;
[0034] If the weighted similarity is greater than or equal to the second preset similarity threshold, determine that the entities in the second sub-graph are the gangs of shell companies.
[0035] To achieve the above object, the present application further provides a shell company identification device, and the shell company identification device includes:
[0036] A construction module, configured to obtain entity information, enterprise investment relationship information, and legal representative relationship information of an enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information, and the legal representative relationship information;
[0037] A first determination module, configured to determine whether there are suspicious entities in the directed knowledge graph;
[0038] A second determination module, configured to, if there are the suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities.
[0039] Furthermore, to achieve the above object, the present application further provides a shell company identification device, and the shell company identification device includes a memory, a processor, and a shell company identification program stored on the memory and executable on the processor. When the shell company identification program is executed by the processor, the steps of the above shell company identification method are implemented.
[0040] Further, to achieve the above object, the present application also provides a medium, which is a computer-readable storage medium. An empty shell company identification program is stored on the computer-readable storage medium. When the empty shell company identification program is executed by a processor, the steps of the above-mentioned empty shell company identification method are implemented.
[0041] Further, to achieve the above object, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned empty shell company identification method are implemented.
[0042] The embodiments of the present application provide an empty shell company identification method, device, equipment and medium. Obtain the entity information, enterprise investment relationship information and legal representative relationship information of the enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information and the legal representative relationship information; determine whether there are suspicious entities in the directed knowledge graph; if there are the suspicious entities in the directed knowledge graph, determine whether the suspicious entities are empty shell companies according to the number of the suspicious entities. The present application can construct a directed knowledge graph according to the entity information, enterprise investment relationship information and legal representative relationship information of the enterprise to be identified, and when detecting that there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are empty shell companies according to the number of the suspicious entities, and can quickly, conveniently and accurately identify empty shell companies, effectively improving the identification efficiency of empty shell companies. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic structural diagram of the hardware operating environment involved in the embodiment scheme of the empty shell company identification method of the present application;
[0044] Figure 2 It is a schematic flowchart of the first embodiment of the empty shell company identification method of the present application;
[0045] Figure 3 It is a schematic flowchart of the second embodiment of the empty shell company identification method of the present application;
[0046] Figure 4 It is a schematic flowchart of the third embodiment of the empty shell company identification method of the present application;
[0047] Figure 5 It is a schematic diagram of the functional modules of the preferred embodiment of the empty shell company identification device of the present application.
[0048] The realization, functional features and advantages of the object of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] The embodiments of the present application provide a method, apparatus, device, and medium for identifying shell companies, which obtain the entity information, enterprise investment relationship information, and legal representative relationship information of an enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information, and the legal representative relationship information; determine whether there are suspicious entities in the directed knowledge graph; if there are suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities. The present application can construct a directed knowledge graph according to the entity information, enterprise investment relationship information, and legal representative relationship information of an enterprise to be identified, and when detecting that there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of the suspicious entities, and can quickly, conveniently, and accurately identify shell companies, effectively improving the identification efficiency of shell companies.
[0051] Technical terms involved in the embodiments of the present application:
[0052] Entity: A typed node in a knowledge graph;
[0053] Relationship: A typed edge connecting entities in a knowledge graph;
[0054] Out-degree: The number of outward connections of an entity in a knowledge graph;
[0055] Central node: An entity in a knowledge graph with an out-degree exceeding m;
[0056] n-hop: Starting from the current node, perform a breadth-first traversal outward at most n times;
[0057] Directed knowledge graph: The relationships in the knowledge graph have directions.
[0058] As Figure 1 shown, Figure 1 is a schematic structural diagram of a shell company identification device in the hardware operating environment involved in the embodiments of the present application.
[0059] In subsequent descriptions, suffixes such as "module", "component", or "unit" used to represent components are only for the convenience of description of the present application, and have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0060] The shell company identification device in the embodiments of the present application can be a PC, or a mobile terminal device such as a tablet computer or a portable computer.
[0061] As Figure 1As shown in the figure, the shell company identification device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0062] Those skilled in the art can understand that Figure 1 the structure of the shell company identification device shown in does not constitute a limitation on the shell company identification device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0063] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a shell company identification program.
[0064] In Figure 1 the device shown in the figure, the network interface 1004 is mainly used to connect to the background server and perform data communication with the background server; the user interface 1003 is mainly used to connect to the client (user side) and perform data communication with the client; and the processor 1001 may be used to call the shell company identification program stored in the memory 1005 and perform the following operations:
[0065] Obtain the entity information, enterprise investment relationship information, and legal representative relationship information of the enterprise to be identified, and construct a directed knowledge graph based on the entity information, the enterprise investment relationship information, and the legal representative relationship information;
[0066] Determine whether there are suspicious entities in the directed knowledge graph;
[0067] If there are such suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities.
[0068] Further, the step of determining whether there are suspicious entities in the directed knowledge graph includes:
[0069] Determine whether there is a central node in the directed knowledge graph;
[0070] If there is such a central node, entity screening is performed on all entities in the directed knowledge graph according to the central node to obtain an entity set, where the entity set includes each outgoing enterprise corresponding to the entity of the central node;
[0071] The similarity between every two outgoing enterprises in the entity set is determined respectively;
[0072] Whether each of the outgoing enterprises is a suspicious entity is determined respectively according to each of the similarities;
[0073] If any of the outgoing enterprises is a suspicious entity among all the outgoing enterprises, it is determined that there is a suspicious entity in the directed knowledge graph.
[0074] Further, the step of determining whether each of the outgoing enterprises is a suspicious entity according to each of the similarities includes:
[0075] For each of the outgoing enterprises, the following steps are performed respectively:
[0076] The name similarity between the current outgoing enterprise and the outgoing enterprises in the entity set is compared with a first preset similarity threshold;
[0077] If the similarity is greater than or equal to the first preset similarity threshold, it is determined whether the registered addresses of the two outgoing enterprises corresponding to the similarity match;
[0078] If the registered addresses of the two outgoing enterprises corresponding to the similarity match, it is determined that the current outgoing enterprise and the corresponding other outgoing enterprises are suspicious entities.
[0079] Further, the step of determining whether the suspicious entity is a shell company according to the number of the suspicious entities includes:
[0080] The suspicious entities with the same registered address in the entity set are determined as target suspicious entities, and the number of the target suspicious entities is determined;
[0081] The number of the target suspicious entities is compared with a preset number threshold;
[0082] If the number of the target suspicious entities is greater than or equal to the preset number threshold, it is determined that each of the target suspicious entities is a shell company.
[0083] Further, after the step of determining that each of the target suspicious entities is a shell company, the processor 1001 can be used to call the shell company identification program stored in the memory 1005 and perform the following operations:
[0084] Obtain a first sub-graph corresponding to the shell company from the directed knowledge graph;
[0085] Determine the second sub-graph of the central node in the directed knowledge graph;
[0086] Based on the first sub-graph, determine whether the entities in the second sub-graph are groups of shell companies.
[0087] Further, the step of determining whether the entities in the second sub-graph are groups of shell companies based on the first sub-graph includes:
[0088] Extract the first feature to be compared of the first sub-graph and the second feature to be compared of the second sub-graph;
[0089] According to the first feature to be compared and the second feature to be compared, respectively determine the structural similarity, attribute similarity, and central node importance similarity between the first sub-graph and the second sub-graph;
[0090] According to the structural similarity, the attribute similarity, and the central node importance similarity, determine whether the entities in the second sub-graph are groups of shell companies.
[0091] Further, the step of determining whether the entities in the second sub-graph are groups of shell companies according to the structural similarity, the attribute similarity, and the central node importance similarity includes:
[0092] Perform a weighted operation on the structural similarity, the attribute similarity, and the central node importance similarity to obtain a weighted similarity;
[0093] Compare the weighted similarity with a second preset similarity threshold;
[0094] If the weighted similarity is greater than or equal to the second preset similarity threshold, determine that the entities in the second sub-graph are groups of shell companies.
[0095] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0096] To better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.
[0097] Refer to Figure 2 , Figure 2The flowchart of a method for identifying shell companies provided by the first embodiment of the present application. In this embodiment, the method for identifying shell companies includes the following steps:
[0098] Step S10: Obtain the entity information, enterprise investment relationship information, and legal representative relationship information of the enterprise to be identified, and construct a directed knowledge graph based on the entity information, the enterprise investment relationship information, and the legal representative relationship information.
[0099] In this embodiment, the method for identifying shell companies is applied to a shell company identification system, which can be deployed in a smart terminal or a server to execute the method for identifying shell companies. By executing the method for identifying shell companies through the shell company identification system in this embodiment, a directed knowledge graph can be constructed based on the entity information and enterprise investment relationship information of the enterprise to be identified. When detecting suspicious entities in the directed knowledge graph, it can accurately determine whether the suspicious entities are shell companies according to the number of suspicious entities, and can quickly, conveniently, and accurately identify shell companies, effectively improving the identification efficiency of shell companies.
[0100] Specifically, when an enterprise needs to conduct financial business such as loan, insurance, etc. with a financial institution, it needs to provide its enterprise information to the financial institution. The enterprise information may include the enterprise business license, so that the financial institution can review the enterprise based on the enterprise information to determine whether to conduct business with the enterprise. However, current shell companies are one of the important risks that financial institutions need to be vigilant about. Therefore, the financial institution needs to judge whether the enterprise is associated with a shell company based on the enterprise information provided by the enterprise. Therefore, the financial institution can use the shell company identification system to regard the enterprise as the enterprise to be identified and input the enterprise information of the enterprise to be identified into the shell company identification system.
[0101] Further, the shell company identification system obtains the business license of the enterprise to be identified, constructs the enterprise entity of the enterprise to be identified based on the business license. Specifically, it can extract the enterprise information in the business license as various attribute information of the enterprise entity, and construct the personal entity. Specifically, it can extract the legal representative in the business license as the personal entity information, and obtain the entity information of the enterprise to be identified composed of the enterprise entity and the personal entity. At the same time, it constructs the legal representative relationship between the enterprise entity and the personal entity, and extracts the enterprise investment relationship information. Specifically, it can construct the investment relationship between enterprise entities. Further, it constructs a directed knowledge graph based on the entity information, enterprise investment relationship information and legal representative relationship information. Specifically, it can use the enterprise entity and personal entity in the entity information as entities, and the investment relationship between enterprises in the enterprise investment relationship information and the legal representative relationship between the enterprise and the individual in the legal representative relationship information as edges to construct a directed knowledge graph. This is to facilitate subsequent determination of whether there are suspicious entities in the directed knowledge graph, and when there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of suspicious entities, so as to quickly, conveniently and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0102] Step S20, determine whether there are suspicious entities in the directed knowledge graph;
[0103] After constructing the directed knowledge graph based on the entity information, enterprise investment relationship information and legal representative relationship information, in this embodiment, it can identify whether there are shell companies by combining the rule model with the directed knowledge graph. Among them, the rule model mainly identifies the scenarios of one legal representative / investing in multiple shell companies and one company investing in multiple shell companies. Specifically, it can first determine whether there are central nodes in the directed knowledge graph; if there are no central nodes, the current shell company identification process can be ended. If there are central nodes, entity screening is performed on all entities in the directed knowledge graph according to the central nodes to obtain an entity set. It should be noted that the entity set includes each outgoing enterprise corresponding to the entity of the central node, and the central node may be one or more; the similarity between every two outgoing enterprises in the entity set is determined respectively; according to each similarity, it is determined whether each outgoing enterprise is a suspicious entity, and one similarity can determine whether the corresponding two outgoing enterprises are suspicious entities; if none of the outgoing enterprises is a suspicious entity, it is determined that there are no suspicious entities in the directed knowledge graph, and the current shell company identification process can be ended. If any of the outgoing enterprises is a suspicious entity, it is determined that there are suspicious entities in the directed knowledge graph. When there are suspicious entities in the directed knowledge graph, it is determined whether the suspicious entities are shell companies according to the number of suspicious entities, so as to quickly, conveniently and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0104] Step S30. If the suspicious entity exists in the directed knowledge graph, determine whether the suspicious entity is a shell company according to the number of the suspicious entities.
[0105] After determining whether a suspicious entity exists in the directed knowledge graph, if it is determined that a suspicious entity exists in the directed knowledge graph, further determine the number of suspicious entities in the entity set, specifically, the number of suspicious entities with the same registered address in the entity set. Determine whether the suspicious entity is a shell company by determining whether the number of suspicious entities reaches the shell company standard. Among them, if the number of suspicious entities with the same registered address reaches the shell company standard, determine that the suspicious entity is a shell company; if the number of suspicious entities does not reach the shell company standard, determine that the suspicious entity is not a shell company, so as to quickly, conveniently and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0106] Further, the step of determining whether the suspicious entity is a shell company according to the number of the suspicious entities includes:
[0107] Step S31. Determine the suspicious entities with the same registered address in the entity set as the target suspicious entities, and determine the number of the target suspicious entities;
[0108] Step S32. Compare the number of the target suspicious entities with a preset quantity threshold;
[0109] Step S33. If the number of the target suspicious entities is greater than or equal to the preset quantity threshold, determine that each of the target suspicious entities is a shell company.
[0110] After determining that a suspicious entity exists in the directed knowledge graph, among all the suspicious entities in the entity set, determine the suspicious entities with the same registered address as the target suspicious entities, and count the number of suspicious entities with the same registered address among all the suspicious entities, that is, the number of the target suspicious entities. Specifically, count the number of outgoing enterprises with a similarity greater than or equal to the first preset similarity threshold and a registered address matching the registered address of other outgoing enterprises, that is, determine the number of outgoing enterprises with a higher similarity and the same registered address. The first preset similarity threshold is a value set according to actual needs. After determining the number of the target suspicious entities, compare the number of the target suspicious entities with the preset quantity threshold to determine the size relationship between the number of the target suspicious entities and the preset quantity threshold. More specifically, a difference operation can be performed on the number of the target suspicious entities and the preset quantity threshold, and the size relationship between the number of the target suspicious entities and the preset quantity threshold can be determined according to the result of the difference operation. If it is determined through comparison that the number of the target suspicious entities is greater than or equal to the preset quantity threshold, it means that the suspicious entity reaches the shell company determination standard, and then determine that each target suspicious entity is a shell company. Among them, the preset quantity threshold is a value set according to actual needs.
[0111] This embodiment provides a method for identifying shell companies, which obtains the entity information, enterprise investment relationship information, and legal representative relationship information of the enterprise to be identified, and constructs a directed knowledge graph based on the entity information, the enterprise investment relationship information, and the legal representative relationship information; determines whether there are suspicious entities in the directed knowledge graph; if there are suspicious entities in the directed knowledge graph, determines whether the suspicious entities are shell companies according to the number of the suspicious entities. This application can construct a directed knowledge graph based on the entity information, enterprise investment relationship information, and legal representative relationship information of the enterprise to be identified, and when detecting that there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of the suspicious entities, and can quickly, conveniently, and accurately identify shell companies, effectively improving the identification efficiency of shell companies.
[0112] Further, referring to Figure 3 , based on the first embodiment of the shell company identification method of this application, a second embodiment of the shell company identification method of this application is proposed. In the second embodiment, the step of determining whether there are suspicious entities in the directed knowledge graph includes:
[0113] Step S21, determining whether there is a central node in the directed knowledge graph;
[0114] Step S22, if there is the central node, performing entity screening on all entities in the directed knowledge graph according to the central node to obtain an entity set, where the entity set includes each connected-out enterprise corresponding to the entity of the central node;
[0115] Step S23, respectively determining the similarity between two connected-out enterprises in the entity set;
[0116] Step S24, respectively determining whether each connected-out enterprise is a suspicious entity according to each similarity;
[0117] Step S25, if any of the connected-out enterprises is a suspicious entity, determining that there are suspicious entities in the directed knowledge graph.
[0118] After constructing a directed knowledge graph based on entity information, enterprise investment relationship information, and legal representative relationship information, first determine whether there is a central node in the directed knowledge graph. Specifically, determine whether there is an entity in the directed knowledge graph whose out-degree exceeds a certain range, such as an entity with an out-degree exceeding 3, an out-degree exceeding 4, an out-degree exceeding 5, etc. If not, it is determined that there is no central node in the directed knowledge graph, and the current shell company identification process can be terminated. If so, that entity is determined as the central node. If there are multiple such entities, it means there are multiple central nodes. Further, for each central node, entity screening can be performed on all entities in the directed knowledge graph through the central node. Specifically, determine the connected enterprises corresponding to all entities connected out by the entity of the central node, and form an entity set consisting of the connected enterprises corresponding to all entities connected out by the entity of the central node. If there are multiple central nodes, the entity set can also be directly formed by the connected enterprises corresponding to all entities connected out by the entities of multiple central nodes.
[0119] Further, respectively determine the similarity between every two connected enterprises of each central node in the entity set. Specifically, respectively compare the pairwise name similarity between each connected enterprise corresponding to each central node in the entity set and other connected enterprises, and determine the similarity between each two connected enterprises according to the pairwise name similarity between each two connected enterprises. For example, if the entity set includes the connected enterprises corresponding to 10 entities connected out by central node 1, then calculate the similarity between each connected enterprise and the other 9 connected enterprises respectively to determine whether each connected enterprise is a suspicious entity. Further, respectively determine whether the similarity between every two connected enterprises is greater than or equal to the first preset similarity threshold. When the similarity is greater than or equal to the first preset similarity threshold, further determine whether the registered addresses of the two connected enterprises corresponding to this similarity match; and when the registered addresses of the two connected enterprises corresponding to this similarity match, determine that this connected enterprise and the other connected enterprise corresponding to this similarity are suspicious entities. If the similarity is less than the first preset similarity threshold, or the similarity is greater than or equal to the first preset similarity threshold, but the registered addresses of the two connected enterprises corresponding to this similarity do not match, then determine that this connected enterprise is not a suspicious entity. And so on, until it is respectively determined whether all connected enterprises are suspicious entities. If none of the connected enterprises are suspicious entities, it is determined that there are no suspicious entities in the directed knowledge graph, and the current shell company identification process can be terminated. If any of the connected enterprises is a suspicious entity, it is determined that there are suspicious entities in the directed knowledge graph. When there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of suspicious entities, so as to quickly, conveniently, and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0120] Further, the step of respectively determining whether each of the outgoing enterprises is a suspicious entity according to each of the similarities includes:
[0121] Step S241, for each of the outgoing enterprises, respectively execute steps S242 - S244:
[0122] Step S242, compare the name similarity between the current outgoing enterprise and the outgoing enterprises in the entity set with a first preset similarity threshold;
[0123] Step S243, if the similarity is greater than or equal to the first preset similarity threshold, determine whether the registered addresses of the two outgoing enterprises corresponding to the similarity match;
[0124] Step S244, if the registered addresses of the two outgoing enterprises corresponding to the similarity match, determine that the current outgoing enterprise and the corresponding other outgoing enterprises are suspicious entities.
[0125] When determining whether each outgoing enterprise is a suspicious entity according to each similarity, determine the outgoing enterprise that needs to be determined currently, obtain the similarity between the current outgoing enterprise and one of the other outgoing enterprises, compare the similarity with the first preset similarity threshold, and determine the size relationship between the similarity and the first preset similarity threshold. More specifically, the similarity can be subjected to a difference operation with the first preset similarity threshold, and the size relationship between the similarity and the first preset similarity threshold can be determined according to the result of the difference operation. Further, if it is determined through comparison that the similarity is greater than or equal to the first preset similarity threshold, obtain the registered addresses of the two outgoing enterprises corresponding to the similarity, and further match the registered addresses of the two outgoing enterprises to determine whether the registered addresses of the two outgoing enterprises match. Specifically, the registered address can specifically be the registered city, that is, match the registered cities of the two outgoing enterprises to determine whether the registered cities of the two outgoing enterprises are the same. If they are the same, it is determined that the registered addresses of the two outgoing enterprises match; if they are not the same, it is determined that the registered addresses of the two outgoing enterprises do not match. Further, if it is determined that the registered addresses of the two outgoing enterprises match, that is, the similarity between the two outgoing enterprises is greater than or equal to the first preset similarity threshold and the registered cities are the same, then determine that the two outgoing enterprises corresponding to the similarity, that is, the current outgoing enterprise and a corresponding other outgoing enterprise, are suspicious entities. Then obtain the similarity between the next outgoing enterprise and other outgoing enterprises and compare it with the first preset similarity threshold, or obtain the similarity between the current outgoing enterprise and another other outgoing enterprise and compare it with the first preset similarity threshold. When the similarity is greater than or equal to the first preset similarity threshold and the registered addresses of the two outgoing enterprises corresponding to the similarity are the same, determine that the two outgoing enterprises are suspicious entities. And so on, until it is determined whether all outgoing enterprises are suspicious entities respectively. When there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of suspicious entities, so as to quickly, conveniently and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0126] In this embodiment, it can be first determined whether there is a central node in the directed knowledge graph; if there is a central node, entity screening is performed on all entities in the directed knowledge graph according to the central node to obtain an entity set; the similarities between pairwise outgoing enterprises in the entity set are determined respectively; whether each outgoing enterprise is a suspicious entity is determined according to each similarity; if any of the outgoing enterprises is a suspicious entity, it is determined that there are suspicious entities in the directed knowledge graph. When there are suspicious entities in the directed knowledge graph, accurately determine whether the suspicious entities are shell companies according to the number of suspicious entities, so as to quickly, conveniently and accurately identify shell companies and effectively improve the identification efficiency of shell companies.
[0127] Further, referring to Figure 4, based on the first embodiment of the shell company identification method of the present application, the third embodiment of the shell company identification method of the present application is proposed. In the third embodiment, after the step of determining that each of the target suspicious entities is a shell company, the following steps are further included:
[0128] Step S40: Obtain a first sub-graph corresponding to the shell company from the directed knowledge graph;
[0129] Step S50: Determine a second sub-graph of the central node in the directed knowledge graph;
[0130] Step S60: Based on the first sub-graph, determine whether the entities in the second sub-graph are gangs of the shell company.
[0131] It can be understood that after determining the shell company in the directed knowledge graph, in this embodiment, it is also possible to accurately determine whether there is a gang of the shell company in the directed knowledge graph. Specifically, obtain a first sub-graph composed of entities and their edges (i.e., relationships) that are n hops away from the shell company in the directed knowledge graph, where n is 0, 1, 2,..., and the corresponding value can be selected according to actual needs; at the same time, determine the central node from the directed knowledge graph, specifically, determine the individuals or enterprises in the directed knowledge graph whose out-degree exceeds m, where m is 0, 1, 2,..., and the corresponding value can be selected according to actual needs, and the central node may be one or more; and further obtain a second sub-graph composed of entities and their edges that are n hops away from the central node in the directed knowledge graph. Further, extract the first feature to be compared of the first sub-graph and the second feature to be compared of the second sub-graph; determine the structural similarity, attribute similarity, and central node importance similarity of the first sub-graph and the second sub-graph according to the first feature to be compared and the second feature to be compared respectively; determine whether the entities in the second sub-graph are gangs of the shell company according to the structural similarity, attribute similarity, and central node importance similarity. It can accurately discover the potential gang relationships behind shell companies, which is conducive to large-scale popularization and application, and helps financial institutions establish a unified and standardized shell company risk prevention ability and system.
[0132] Further, the step of determining whether the entities in the second sub-graph are gangs of the shell company based on the first sub-graph includes:
[0133] Step S61: Extract the first feature to be compared of the first sub-graph and the second feature to be compared of the second sub-graph;
[0134] Step S62: Determine the structural similarity, attribute similarity, and central node importance similarity of the first sub-graph and the second sub-graph according to the first feature to be compared and the second feature to be compared respectively;
[0135] Step S63: Determine whether the entities in the second sub-graph are the gang of the shell company according to the structural similarity, the attribute similarity, and the central node importance similarity.
[0136] After obtaining the first sub-graph corresponding to the shell company and the second sub-graph corresponding to the central node, extract the first core structure, the first graph attributes, and the first central node features of the first sub-graph to obtain the first features to be compared of the first sub-graph, and extract the second core structure, the second graph attributes, and the second central node features of the second sub-graph to obtain the second features to be compared of the second sub-graph. Among them, the core structure refers to the largest connected sub-graph with non-zero out-degree in the directed knowledge graph. The graph attributes cover the proportion of enterprises with similar names and the proportion of enterprises with the same registered address. The extraction of central node features covers the central node importance. Further, by using the Weisfeiler-Lehman kernel algorithm to combine the first core structure and the second core structure to obtain a unique feature set on the graph as the discriminant basis for whether the graphs are similar, compare the structural similarity between the first sub-graph and the second sub-graph; use the Levenshtein Distance edit distance algorithm to combine the first graph attributes and the second graph attributes to judge the specific attribute similarity of the enterprises connected by the central node, and determine the attribute similarity between the first sub-graph and the second sub-graph; use the Pagerank algorithm to combine the first central node features and the second central node features to calculate the node importance on different graphs, and obtain the central node importance similarity between the first sub-graph and the second sub-graph. Further, determine whether the entities in the second sub-graph are the gang of the shell company according to the structural similarity, the attribute similarity, and the central node importance similarity. It can accurately discover the potential gang relationships behind the shell companies, which is conducive to large-scale popularization and application, and helps financial institutions establish a unified and standardized risk prevention ability and system for shell companies.
[0137] Specifically, the step of determining whether the entities in the second sub-graph are the gang of the shell company according to the structural similarity, the attribute similarity, and the central node importance similarity includes:
[0138] Step S651: Perform a weighted operation on the structural similarity, the attribute similarity, and the central node importance similarity to obtain a weighted similarity.
[0139] Step S652: Compare the weighted similarity with a second preset similarity threshold.
[0140] Step S653: If the weighted similarity is greater than or equal to the second preset similarity threshold, determine that the entities in the second sub-graph are the gang of the shell company.
[0141] After determining the structural similarity, attribute similarity, and central node importance similarity between the first sub-graph and the second sub-graph, perform a weighted operation on the structural similarity, attribute similarity, and central node importance similarity to obtain a weighted similarity; compare the obtained weighted similarity with a second preset similarity threshold to determine the magnitude relationship between the weighted similarity and the second preset similarity threshold, where the second preset similarity threshold is a value set according to actual recognition requirements. Specifically, a difference operation can be performed between the weighted similarity and the second preset similarity threshold, and the magnitude relationship between the weighted similarity and the second preset similarity threshold can be determined according to the result of the difference operation. Further, if it is determined through comparison that the weighted similarity is greater than or equal to the second preset similarity threshold, then determine that the entity in the second sub-graph is a gang of the shell company. If there are multiple such central nodes, then determine that each of the multiple central nodes is a gang of the shell company.
[0142] After determining the shell company in the directed knowledge graph in this embodiment, it is also possible to accurately determine whether there is a gang of the shell company in the directed knowledge graph. By accurately excavating the potential gang relationship behind the shell company, it is beneficial to large-scale popularization and application, and helps financial institutions establish a unified and standardized risk prevention ability and system for shell companies.
[0143] Further, the present application also provides a shell company identification device.
[0144] Refer to Figure 5 , Figure 5 which is a schematic diagram of the functional modules of the first embodiment of the shell company identification device of the present application.
[0145] The shell company identification device includes:
[0146] A construction module 10, configured to obtain entity information, enterprise investment relationship information, and legal representative relationship information of an enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information, and the legal representative relationship information;
[0147] A first determination module 20, configured to determine whether there is a suspicious entity in the directed knowledge graph;
[0148] A second determination module 30, configured to, if there is the suspicious entity in the directed knowledge graph, determine whether the suspicious entity is a shell company according to the number of the suspicious entities.
[0149] In addition, the present application also provides a medium, which is preferably a computer-readable storage medium, on which a shell company identification program is stored. When the shell company identification program is executed by a processor, the steps of the above-mentioned shell company identification method in each embodiment are implemented.
[0150] In addition, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the steps of the above-described embodiments of the shell company identification method.
[0151] In the embodiments of the shell company identification device, computer-readable storage medium, and computer program product of the present application, all the technical features of the above-described embodiments of the shell company identification method are included, and the description and explanation content are basically the same as those of the above-described embodiments of the shell company identification method, and will not be repeated here.
[0152] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0153] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a fixed terminal, such as an Internet of Things intelligent device, including smart home appliances such as smart air conditioners, smart lights, smart power supplies, smart routers, etc.; or a mobile terminal, including smartphones, wearable networked AR / VR devices, smart speakers, autonomous driving cars, and many other networked devices) to execute the methods described in the various embodiments of the present application.
[0155] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for identifying shell companies, characterized in that, The shell company identification method includes: Obtain the entity information, enterprise investment relationship information, and legal representative relationship information of the enterprise to be identified, and construct a directed knowledge graph based on the entity information, the enterprise investment relationship information, and the legal representative relationship information; Determine whether there are suspicious entities in the directed knowledge graph; If there are such suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities; The step of determining whether there are suspicious entities in the directed knowledge graph includes: Determine whether there is a central node in the directed knowledge graph, where the central node is an entity in the knowledge graph with an out-degree exceeding m, and m is a corresponding value selected according to actual needs; If there is such a central node, perform entity screening on all entities in the directed knowledge graph according to the central node to obtain an entity set, where the entity set includes each connected-out enterprise corresponding to the entity of the central node; Determine the similarity between every two connected-out enterprises in the entity set respectively; Determine whether each of the connected-out enterprises is a suspicious entity according to each of the similarities; If any of the connected-out enterprises is a suspicious entity, determine that there are suspicious entities in the directed knowledge graph; If it is determined that the suspicious entity is a shell company, obtain the first sub-graph corresponding to the shell company from the directed knowledge graph; Determine the second sub-graph of the central node in the directed knowledge graph; Based on the structural similarity, attribute similarity, and central node importance similarity between the first sub-graph and the second sub-graph, determine whether the entities in the second sub-graph are the gangs of the shell company.
2. The shell company identification method according to claim 1, wherein The step of determining whether each of the connected-out enterprises is a suspicious entity according to each of the similarities includes: For each of the connected-out enterprises, perform the following steps respectively: Compare the name similarity between the current connected-out enterprise and the connected-out enterprises in the entity set with a first preset similarity threshold; If the similarity is greater than or equal to the first preset similarity threshold, determine whether the registered addresses of the two connected-out enterprises corresponding to the similarity match; If the registered addresses of the two connected-out enterprises corresponding to the similarity match, determine that the current connected-out enterprise and the corresponding other connected-out enterprises are suspicious entities.
3. The method for identifying a shell company according to claim 1, wherein, The step of determining whether the suspicious entity is a shell company according to the number of the suspicious entities includes: Determine the target suspicious entities as the suspicious entities with the same registered address in the entity set, and determine the number of the target suspicious entities; Compare the number of the target suspicious entities with a preset number threshold; If the number of the target suspicious entities is greater than or equal to the preset number threshold, determine that each of the target suspicious entities is a shell company.
4. The shell company identification method according to claim 1, wherein The step of determining whether the entities in the second sub-graph are the gangs of the shell company based on the first sub-graph includes: Extract the first features to be compared of the first sub-graph and the second features to be compared of the second sub-graph; Determine the structural similarity, attribute similarity, and central node importance similarity of the first sub-graph and the second sub-graph based on the first feature to be compared and the second feature to be compared respectively; Determine whether the entities in the second sub-graph are the gangs of the shell companies according to the structural similarity, the attribute similarity, and the central node importance similarity; 5. The method for identifying a shell company according to claim 4, wherein The step of determining whether the entities in the second sub-graph are the gangs of the shell companies according to the structural similarity, the attribute similarity, and the central node importance similarity includes: Perform a weighted operation on the structural similarity, the attribute similarity, and the central node importance similarity to obtain a weighted similarity; Compare the weighted similarity with a second preset similarity threshold; If the weighted similarity is greater than or equal to the second preset similarity threshold, determine that the entities in the second sub-graph are the gangs of the shell companies.
6. An empty shell company identification device, characterized in that, The shell company identification device includes: A construction module, configured to obtain entity information, enterprise investment relationship information, and legal representative relationship information of an enterprise to be identified, and construct a directed knowledge graph according to the entity information, the enterprise investment relationship information, and the legal representative relationship information; A first determination module, configured to determine whether there are suspicious entities in the directed knowledge graph; A second determination module, configured to, if there are the suspicious entities in the directed knowledge graph, determine whether the suspicious entities are shell companies according to the number of the suspicious entities; The determination of whether there are suspicious entities in the directed knowledge graph includes: Determine whether there is a central node in the directed knowledge graph, where the central node is an entity in the knowledge graph whose out-degree exceeds m, and m is a corresponding value selected according to actual needs; If there is the central node, perform entity screening on all entities in the directed knowledge graph according to the central node to obtain an entity set, where the entity set includes each connected enterprise corresponding to the entity of the central node; Determine the similarity between any two connected enterprises in the entity set respectively; Determine whether each connected enterprise is a suspicious entity according to each similarity respectively; If any one of the connected enterprises is a suspicious entity, determine that there are suspicious entities in the directed knowledge graph; If it is determined that the suspicious entity is a shell company, obtain the first sub-graph corresponding to the shell company from the directed knowledge graph; Determine the second sub-graph of the central node in the directed knowledge graph; Based on the structural similarity, attribute similarity, and central node importance similarity between the first sub-graph and the second sub-graph, determine whether the entities in the second sub-graph are the gangs of the shell companies.
7. An empty shell company identification device, characterized in that, The shell company identification device includes a memory, a processor, and a shell company identification program stored on the memory and executable on the processor. When the shell company identification program is executed by the processor, the steps of the shell company identification method according to any one of claims 1-5 are implemented.
8. A medium, the medium being a computer-readable storage medium, characterized in that, A shell company identification program is stored on the computer-readable storage medium. When the shell company identification program is executed by a processor, the steps of the shell company identification method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Method and system for identifying enterprise risks based on enterprise external features
CN112541698A
Knowledge graph-based hidden relationship collection method and device, equipment and medium
CN112732937A
Method for effectively identifying company with fake-licensed behavior
CN112989067A