Enterprise equity graph construction method and storage medium based on multi-dimensional data verification
Through a multi-dimensional data verification method, combined with OCR and large language model technology, the corporate equity map is constructed and optimized, and the difficulties in handling complex equity relationships in the existing technology are solved, and automation, accuracy and scalability are improved.
Patent Information
- Application Number
- CN202510128541.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-27
AI Technical Summary
The existing technology is difficult to effectively deal with complex corporate equity relations, especially in multi-level and multi-dimensional structures, loops and contradictions are prone to occur, and there is a lack of automated processing methods, resulting in data inconsistency and inefficiency.
Using a multi-dimensional data verification method, text items are extracted from PDF documents through OCR algorithm, and a large language model is used to extract corporate entity information and equity relationships from Markdown files to build an enterprise equity map, and rigorous algorithms to resolve contradictions to ensure the logical rigor of the map and the accuracy of the data.
It realizes the full process of enterprise equity relations, improves the accuracy and scalability of data extraction, reduces manual intervention, ensures data consistency and logical self-consistentness of the graph, and can quickly respond to large-scale enterprise data updates.
Smart Images

Figure CN119577196B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph technology, and specifically to a method and storage medium for constructing an enterprise equity graph based on multi-dimensional data verification. Background Art
[0002] In recent years, with the expansion of corporate scale and the global extension of the industrial chain, the equity relationship and control relationship between enterprises have become increasingly complex. Traditional technical means (such as manual review and semi-automatic text retrieval) usually rely on manual or simple rule-based tools to extract and associate information from corporate annual reports, equity disclosure documents, etc. Its main practices include: analysts manually read corporate annual reports or announcement documents in PDF format to extract information such as equity ratios and voting rights ratios between companies. Preliminary records and comparisons are made through Excel tables or relational databases, and the equity relationships between different companies are marked; for scanned PDFs, optical character recognition (OCR) technology is used to convert the files into text, and then keywords or regular expressions are used to locate equity-related information; some predefined rules are set to identify obviously erroneous data.
[0003] The above-mentioned existing technologies can, to a certain extent, complete the extraction and preliminary analysis of corporate equity relations, but lack large-scale, automated, and intelligent solutions. They are often stretched when dealing with multi-level, multi-dimensional corporate relationship structures, and are prone to loops and contradictions. They lack automated processing methods and are difficult to ensure the consistency, integrity, and practical value of data. Summary of the invention
[0004] The purpose of the present invention is to provide a method and storage medium for constructing an enterprise equity map based on multi-dimensional data verification to solve the problems raised in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for constructing an enterprise equity map based on multi-dimensional data verification, the method comprising the following steps:
[0007] S100, acquiring document data, extracting text from the acquired document data, performing layout analysis and table structure reconstruction on the extracted text, and then converting the format of the text;
[0008] Furthermore, the specific steps of performing layout analysis on the extracted text and reconstructing the table structure are as follows:
[0009] S101, extracting a PDF document from an enterprise annual report database, storing the extracted PDF document in a file management server, and dividing the PDF document into a scanned version and a native version with a text layer;
[0010] S102, calling the OCR engine to perform page-level text recognition on the scanned PDF document, and extracting text items in the scanned PDF; reading OCR data in the process of extracting text items, obtaining the confidence conf of each text item, and extracting the top position top, left position left, height height, and width width of the characters in the text item;
[0011] The confidence of each text item is used to screen the text items, retaining the text items with conf>60 and removing the text items with conf≤60;
[0012] Record the position information of each text item retained after filtering, and calculate the top, bottom, left, and right coordinates of the bounding box of the characters in the text item. The formula is: up=top, down=top+height, right=left+width, where up represents the top coordinate of the bounding box, down represents the bottom coordinate of the bounding box, right represents the right coordinate of the bounding box, and left is used as the left coordinate of the bounding box; set the standard value of the page size to 1;
[0013] For native PDF documents, directly extract low-level text items;
[0014] S103: After extracting text items from all stored PDF documents, calculate the rectangular area Area of each character in the text item. i ;
[0015] Normalize the rectangular area of each character to the relative value of the page size Standard_Area i , then calculate the font size of each character Font_Size i ;
[0016] S104. After obtaining the font size of each character, the average font size of all characters is calculated using the average formula, and then the characters are divided into two clusters according to the font size using the discrete clustering algorithm, where cluster 1 is the main text font and cluster 2 is the title font; the font size of cluster 2 is larger than that of cluster 1; after the characters are divided into two clusters, the characters in the two clusters are structurally marked, specifically: according to the clustering results, the text items that are identified as titles are marked with corresponding title tags in the Markdown file, and the text items that are identified as the main text are marked with ordinary paragraph tags; the marked text items are converted to generate a Markdown file.
[0017] In step S100, the characters are classified by calculating the character sizes in the extracted text items to obtain main text characters and title characters, which effectively overcomes the problem of row and column misalignment that is prone to occur in traditional OCR when recognizing complex tables, and greatly reduces manual intervention; the classified characters are marked in the markup language Markdown file, and automatic typesetting can be achieved according to the marks when the characters are extracted, which greatly improves the automation level of obtaining equity data from PDF annual reports and can quickly respond to the update and analysis of large-scale enterprise data.
[0018] The generated markup language Markdown file contains various information of the enterprise. To obtain various information of the enterprise in the Markdown file, it is necessary to use a large language model for extraction; step S200 illustrates the specific extraction method as follows:
[0019] S200, input the converted format into a large language model to extract the enterprise entity information and equity relationship in the text;
[0020] Furthermore, the specific steps for extracting the corporate entity information and equity relationship in the text are as follows:
[0021] S201, dividing the markup language Markdown file into blocks, converting each block into a vector using a text vector model, and storing all the block vectors in a retrieval-augmented generation (RAG) vector database to form an index that can be quickly queried;
[0022] S202, custom-build a prompt message containing the equity relationship of the enterprise, use the formed index to retrieve the Markdown block reflecting the equity relationship of the enterprise in the RAG vector database, splice the retrieved Markdown block with the prompt message, and then initiate an information extraction request to the LLM (LargeLanguageModel) large language model;
[0023] S203. After receiving the request, the LLM large language model parses the text, extracts the enterprise entity information and equity relationship, and converts different information formats into numerical forms; uses the extracted enterprise entity information and equity relationship to construct a triple structure, specifically: enterprise name-equity relationship-shareholder name.
[0024] When extracting various information about a company, the LLM large language model is selected. With the help of the LLM large language model's ability to understand the semantics of the text context, it can accurately identify company information such as company names, shareholder information, and equity relationships from unstructured or semi-structured documents; it avoids the low efficiency and high error rate of traditional reliance on manual reading or keyword retrieval, and significantly improves the accuracy and scalability of data extraction.
[0025] Through the combination of OCR algorithm and LLM algorithm, the OCR algorithm extracts text items from the PDF document and finally generates a Markdown file. The LLM algorithm extracts various types of information of the enterprise from the Markdown file, completing the automation of the entire process from data source to relationship extraction, greatly reducing manual workload and ensuring efficiency. Even in the scenario of a large number of enterprises and frequent updates of annual reports, it can still respond and update the map quickly.
[0026] Based on the enterprise entity information and equity relationship successfully extracted from the large language model by S200, in S300, these accurately extracted data are used as key matching items and deeply matched with the enterprise information in the external database, so as to gradually build a preliminary equity map to ensure the consistency of data and the accuracy of map construction. The specific matching method is as follows:
[0027] S300, matching the extracted enterprise entity information with the enterprise information in the external database to construct a preliminary equity map;
[0028] Furthermore, the specific steps to construct a preliminary equity map are:
[0029] S301, extracting enterprise entity information and enterprise information from an external database; the external database represents the sorted and cleaned enterprise information, including elements such as corporate legal person, corporate shareholders, corporate registered capital, and corporate business scope;
[0030] S302, determining whether the enterprise entity information contains an ID. If the ID is contained, accurately matching the ID with the ENTID in the enterprise information of the external database. If the ID = the ENTID in the enterprise information of the external database, the match is successful, and the information alignment is performed directly.
[0031] When the ID is not included, the similarity between the enterprise entity information and each enterprise information in the external database is calculated using the formula:
[0032] ;
[0033] In the formula, Jaccard (A, B) represents the similarity between the extracted enterprise entity information and each enterprise information in the external database, A represents the character set of the extracted enterprise entity information, and B represents the character set of the enterprise name of the enterprise information in the external database; when Jaccard (A, B) = 1, it means that they are exactly the same, and when Jaccard (A, B) = 0, it means that there are no identical characters; the user-defined similarity threshold β is used to judge the calculated similarity, and when Jaccard (A, B) max> β, the match is successful, the information is aligned, and the ENTID in the enterprise information of the external database is recorded;
[0034] Through the element information of the enterprise itself, establish the enterprise's unique id mapping entid, and then associate the extracted information with the database.
[0035] S303. After the extracted enterprise entity information is aligned, a preliminary equity map of the enterprise is constructed in combination with the extracted equity relationship.
[0036] After completing the construction of the preliminary equity map in S300, the equity relationship network in the map has taken shape. Then, in S400, a comprehensive loop detection is carried out on this preliminary equity map. With rigorous algorithms and logic, it is accurately determined whether there are contradictions, and the contradictions are resolved in a scientific and reasonable way, thereby ensuring the logical rigor of the equity map and the accuracy of the data. The specific contradiction resolution method is as follows:
[0037] S400, detecting loops in the preliminary equity map, determining whether there are contradictions and resolving the contradictions;
[0038] Furthermore, the specific steps to determine whether there is a contradiction and resolve it are:
[0039] S401. After generating a preliminary equity graph of an enterprise, take the names of shareholders in the preliminary equity graph as nodes and the equity relationships as edges, extract all relationship paths in the equity graph, and judge the relationship paths. When there is only one node and one edge connecting the enterprise in the relationship path, the corresponding path is judged to be a direct path. When there are more than one node and one edge connecting the enterprise in the relationship path, it is judged to be a multi-hop path.
[0040] S402, extracting the first nodes in the direct path and the multi-hop path, and using the first nodes to classify the relationship paths, and classifying the same first nodes as the same shareholder equity relationship;
[0041] In the same shareholder equity relationship, the equity map is used to calculate the shareholding ratio of each shareholder. For the direct path, the formula is: , X represents shareholder G r The shareholding ratio of enterprise Q through the direct path, w represents the equity relationship of the edge in the equity graph; for multi-hop paths, the formula is:
[0042] ;
[0043] In the formula, D (G r , Q) represents G r Based on the shareholding ratio of enterprise Q in the multi-hop path, P represents the total number of multi-hop paths in the equity relationship of the same shareholder, u represents the total number of edges in a multi-hop path, and G y represents the yth node in a multi-hop path, G y+1 represents the y+1th node in a multi-hop path; x belongs to 1 to P, y belongs to 1 to u;
[0044] S403. In the same shareholder equity relationship, the shareholder's shareholding ratio of enterprise Q through the direct path and the shareholding ratio of enterprise Q through the multi-hop path are judged. When X=D(G r , Q), determine whether there is any contradiction in the shareholder's shareholding ratio of enterprise Q, and eliminate the direct path; calculate all shareholder equity relations, eliminate all contradictory direct paths, and output the updated equity map.
[0045] First, the direct path and multi-hop path are determined to build the same shareholder relationship. Then, the equity ratios of different paths are calculated, and the equity ratios of different paths are compared to determine the contradictions. Finally, for conflicting equity ratios or multiple sources of shareholding in the same enterprise, a "deletion" strategy is adopted, and its confidence information is retained. This operation not only ensures data quality but also facilitates traceability. It effectively avoids the common equity "circulation" or "false high" problems in traditional systems, and can automatically handle shareholding ratio conflicts, ensuring that the graph structure is clear and logically self-consistent, providing a more reliable data basis for subsequent analysis and supervision.
[0046] After S400's rigorous loop detection and contradiction resolution, the internal logic of the equity map has been optimized and improved. At this point, S500, based on this, focuses on in-depth mining of the equity map after the contradictions are resolved, identifies the interest community through professional analysis methods, and accurately determines whether there is family control. Based on the judgment results, virtual nodes are constructed to associate family control relationships, and finally a complete and healthy equity map is output to provide solid and reliable data support for subsequent related research and decision-making. The specific method of identifying the interest community is as follows:
[0047] S500, identify the community of interests in the equity map after the contradiction is resolved, determine whether it is family control, build a virtual node to associate family control, and output a healthy equity map;
[0048] Furthermore, the specific steps to output a healthy equity map are:
[0049] S501. Use the community division algorithm to divide sub-communities in the preliminary equity graph after contradictions are resolved, extract the enterprises and all nodes in the sub-communities, and calculate the total number of edges when all nodes in the sub-communities are fully interconnected: , where L represents the number of all nodes and enterprises in the extracted sub-community. The equity density in the sub-community is calculated as follows: , where ED represents the equity density in the subcommunity, and w e represents the equity relationship of the e-th edge;
[0050] S502: For each node in the subcommunity, calculate the sum of the edge equity relationships between the node and all other nodes, take the sum of the edge equity relationships as the control contribution CC of the corresponding node, and extract the maximum control contribution CC among all nodes max , calculate the maximum control concentration in the subcommunity, the formula is:
[0051] ;
[0052] In the formula, represents the maximum control concentration of the subcommunity;
[0053] S503, extracting the identity information of all nodes in the sub-community, extracting the nodes belonging to the same family, calculating the sum of the equity relationships of the nodes in the same family as FC, and using FC as the family control degree in the sub-community;
[0054] S504, customizing the control density threshold to α D , the maximum control contribution threshold is α CC , the family control threshold is α FC ; When ED>α D and >α CC When FC>α, the subcommunity is judged as a community of interest and marked; FC When the subcommunity is judged to be family-controlled, a virtual node is constructed for the subcommunity in the database, and the virtual node is associated with the equity relationship of the node controlled by the family. The virtual node is highlighted in the preliminary equity map after the contradiction is resolved to obtain a healthy equity map.
[0055] The graph is divided into communities, and interest communities are identified based on indicators such as equity density and maximum control concentration; the subgraphs of enterprises controlled by families or groups are automatically annotated, which significantly improves the insight into complex equity networks. The originally dispersed equity connections are aggregated in the graph, and the actual control of "family control" over enterprise groups is discovered and annotated, which facilitates risk prevention or business decision-making.
[0056] After successfully completing the construction of a healthy equity map by S500, S600 will immediately start the real-time data update process to ensure that the map can always reflect the latest equity situation of the enterprise. By extracting enterprise entity information in real time and verifying the credibility of these real-time information with the help of the healthy equity map generated by S500, the real-time enterprise entity information that has passed the verification is accurately integrated into the healthy equity map, so that the equity map always maintains timeliness and accuracy, providing a continuous and reliable basis for dynamic monitoring and analysis of corporate equity. The specific method of updating the healthy equity map is as follows:
[0057] S600, extracting enterprise entity information in real time, using the healthy equity map to verify the credibility of the real-time extracted enterprise entity information, and updating the healthy equity map using the verified real-time enterprise entity information;
[0058] Furthermore, the specific steps for updating the healthy equity map using the verified real-time corporate entity information are as follows:
[0059] Extract enterprise entity information in real time, align the extracted enterprise entity information with the existing ID in the healthy equity map, search for all equity relationships in the extracted enterprise entity information in the healthy equity map, and if not found, judge the corresponding equity relationship as suspicious information;
[0060] When found, determine whether all equity relationships in the real-time extracted enterprise entity information are the same as the equity relationships in the healthy equity map. If they are not the same, mark them as suspicious information in the healthy equity map.
[0061] All corporate entity information extracted in real time is verified, suspicious information is marked in the healthy equity map, and the healthy equity map is updated.
[0062] In S600, after completing the credibility check of the real-time extracted enterprise entity information and updating the healthy equity map, in order to achieve persistent storage and convenient call of data, S700 immediately stores the updated healthy equity map;
[0063] S700. Output the updated healthy equity map and store it in the database.
[0064] Furthermore, the specific steps to output the updated healthy equity graph are:
[0065] S701. Save the healthy equity map after contradiction resolution, interest community annotation, and healthy equity map verification in the graph database; record the confidence and data source information of all edges and nodes in the equity map, visualize and output the equity map through the front-end visualization tool, and store it in the Neo4j database.
[0066] A storage medium stores a computer program, which, when executed by a processor, implements the execution steps of a construction and optimization method.
[0067] Combined with algorithmic cycle detection and credibility layering processing, it can not only retain traceable information, but also ensure that the final map is concise and consistent; it can present a penetrating equity structure to the outside world and conduct audit backtracking internally.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] 1. The overall steps of the present invention from data acquisition, OCR table reconstruction to LLM extraction, map construction, contradiction resolution, interest community identification and health map comparison are as follows: Figure 1 As shown, a complete set of automated technology chains has been formed, greatly reducing manual operation costs.
[0070] 2. The present invention combines the OCR algorithm and the LLM algorithm. The OCR algorithm extracts text items from the PDF document and finally generates a Markdown file. The LLM algorithm extracts various information of the enterprise from the Markdown file, and completes the whole process automation from the data source to the relationship extraction, which greatly reduces the manual workload and ensures high efficiency. It can still respond and update the map quickly even in the scenario of a large number of enterprises and frequent updates of annual reports.
[0071] 3. This invention adopts a "deletion" strategy for conflicting equity ratios or multiple sources of shareholdings in the same enterprise, and retains its confidence information, which not only ensures data quality but also facilitates traceability. It effectively avoids the common equity "circulation" or "false high" problems in traditional systems, and can automatically handle shareholding ratio conflicts, ensuring a clear graph structure and self-consistent logic, providing a more reliable data basis for subsequent analysis and supervision. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 A schematic diagram of the steps of the method for constructing an enterprise equity map based on multi-dimensional data verification according to the present invention;
[0073] Figure 2 It is a preliminary equity map of the enterprise equity map construction method based on multi-dimensional data verification of the present invention;
[0074] Figure 3 This is the equity map after the contradictions of the enterprise equity map construction method based on multi-dimensional data verification of the present invention are resolved. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0076] Example: Figure 1-Figure 3 As shown, the present invention provides a technical solution.
[0077] A method for constructing an enterprise equity map based on multi-dimensional data verification, the method comprising the following steps:
[0078] S100, acquiring document data, extracting text from the acquired document data, performing layout analysis and table structure reconstruction on the extracted text, and then converting the format of the text;
[0079] The specific steps for analyzing the layout of the extracted text and reconstructing the table structure are as follows:
[0080] S101, extracting a PDF document from an enterprise annual report database, storing the extracted PDF document in a file management server, and dividing the PDF document into a scanned version and a native version with a text layer;
[0081] S102, calling the OCR engine to perform page-level text recognition on the scanned PDF document, and extracting text items in the scanned PDF; reading OCR data in the process of extracting text items, obtaining the confidence conf of each text item, and extracting the top position top, left position left, height height, and width width of the characters in the text item;
[0082] The confidence of each text item is used to screen the text items, retaining the text items with conf>60 and removing the text items with conf≤60;
[0083] Record the position information of each text item retained after filtering, and calculate the top, bottom, left, and right coordinates of the bounding box of the characters in the text item. The formula is: up=top, down=top+height, right=left+width, where up represents the top coordinate of the bounding box, down represents the bottom coordinate of the bounding box, right represents the right coordinate of the bounding box, and left is used as the left coordinate of the bounding box; set the standard value of the page size to 1;
[0084] For native PDF documents, directly extract low-level text items;
[0085] S103, after extracting text items from all stored PDF documents, calculate the rectangular area of each character in the text item, using the formula: , in the formula, Area i Indicates the calculated rectangular area of the i-th character, up i 、down i 、left i and right i Represents the up, down, left, and right coordinates of the i-th character;
[0086] Normalize the rectangular area of each character to the relative value of the page size using the formula:
[0087] ;
[0088] In the formula, Standard_Areai Indicates the relative value normalized to the page size, and Page_Size indicates the standard value of the page size; then the font size of each character is calculated using the formula:
[0089] ;
[0090] In the formula, Font_Size i represents the font size of each character, and k represents the proportional constant;
[0091] S104. After obtaining the font size of each character, the average font size of all characters is calculated using the average formula, and then the characters are divided into two clusters according to the font size using the discrete clustering algorithm, where cluster 1 is the main text font and cluster 2 is the title font; the font size of cluster 2 is larger than that of cluster 1; after the characters are divided into two clusters, the characters in the two clusters are structurally marked, specifically: according to the clustering results, the text items that are identified as titles are marked with corresponding title tags in the Markdown file, and the text items that are identified as the main text are marked with ordinary paragraph tags; the marked text items are converted to generate a Markdown file.
[0092] Assume that the top, bottom, left, and right coordinates of a character are 0.78, 0.25, 1.5, and 2.1 respectively, and calculate the rectangular area of the character to be 0.318; normalize the rectangular area of the character to the relative value of the page size to be 0.318; set the scale factor to 1, and finally calculate the font size of the character to be 0.56;
[0093] In step S100, the characters are classified by calculating the character sizes in the extracted text items to obtain text characters and title characters, which effectively overcomes the problem of row and column misalignment that is prone to occur in traditional OCR when recognizing complex tables and greatly reduces manual intervention; the classified characters are marked in the Markdown file, and automatic typesetting can be achieved according to the marks when the characters are extracted, which greatly improves the automation level of obtaining equity data from PDF annual reports and can quickly respond to the update and analysis of large-scale enterprise data.
[0094] The generated Markdown file contains various information of the enterprise. To obtain various information of the enterprise in the Markdown file, it is necessary to use a large language model for extraction. Step S200 illustrates the specific extraction method as follows:
[0095] S200, inputting the format-converted text into a large language model to extract the corporate entity information and equity relationship in the text;
[0096] The specific steps for extracting corporate entity information and equity relationships in the text are:
[0097] S201, dividing the Markdown file into blocks, converting each block into a vector using a text vector model, and storing all block vectors in a RAG vector database to form an index that can be quickly queried;
[0098] S202, custom-build a prompt containing the equity relationship of the enterprise, use the formed index to retrieve the Markdown block reflecting the equity relationship of the enterprise in the RAG vector database, splice the retrieved Markdown block with the prompt, and then initiate an information extraction request to the LLM model;
[0099] S203. After receiving the request, the LLM model parses the text, extracts the enterprise entity information and equity relationship, and converts different information formats into numerical forms, converting 30% and 0.5 into 0.30 and 0.50; and uses the extracted enterprise entity information and equity relationship to construct a triple structure, specifically: enterprise name-equity relationship-shareholder name.
[0100] The example triple structure is:
[0101] (Company A) - [:SHAREHOLDING{percentage:0.30}] -> (Shareholder B)
[0102] (Company A) - [:VOTING_RIGHT{percentage: 0.20}] -> (Shareholder C)
[0103] When extracting various information about a company, the LLM large language model is selected. With the help of the LLM large language model's ability to understand the semantics of the text context, it can accurately identify company information such as company names, shareholder information, and equity relationships from unstructured or semi-structured documents; it avoids the low efficiency and high error rate of traditional reliance on manual reading or keyword retrieval, and significantly improves the accuracy and scalability of data extraction.
[0104] Through the combination of OCR algorithm and LLM algorithm, the OCR algorithm extracts text items from the PDF document and finally generates a Markdown file. The LLM algorithm extracts various types of information of the enterprise from the Markdown file, completing the automation of the entire process from data source to relationship extraction, greatly reducing manual workload and ensuring efficiency. Even in the scenario of a large number of enterprises and frequent updates of annual reports, it can still respond and update the map quickly.
[0105] Based on the enterprise entity information and equity relationship successfully extracted from the large language model by S200, in S300, these accurately extracted data are used as key matching items and deeply matched with the enterprise information in the external database, so as to gradually build a preliminary equity map to ensure the consistency of data and the accuracy of map construction. The specific matching method is as follows:
[0106] S300, matching the extracted enterprise entity information with the enterprise information in the external database to construct a preliminary equity map;
[0107] The specific steps to construct a preliminary equity map are:
[0108] S301, extracting enterprise entity information and enterprise information from an external database; the external database represents the sorted and cleaned enterprise information, including elements such as corporate legal person, corporate shareholders, corporate registered capital, and corporate business scope;
[0109] S302, determining whether the enterprise entity information contains an ID. If the ID is contained, accurately matching the ID with the ENTID in the enterprise information of the external database. If the ID = the ENTID in the enterprise information of the external database, the match is successful, and the information alignment is performed directly.
[0110] When the ID is not included, the similarity between the enterprise entity information and each enterprise information in the external database is calculated using the formula:
[0111] ;
[0112] In the formula, Jaccard (A, B) represents the similarity between the extracted enterprise entity information and each enterprise information in the external database, A represents the character set of the extracted enterprise entity information, and B represents the character set of the enterprise name of the enterprise information in the external database; when Jaccard (A, B) = 1, it means that they are exactly the same, and when Jaccard (A, B) = 0, it means that there are no identical characters; the user-defined similarity threshold β is used to judge the calculated similarity, and when Jaccard (A, B) max> β, the match is successful, the information is aligned, and the ENTID in the enterprise information of the external database is recorded;
[0113] Through the element information of the enterprise itself, establish the enterprise's unique id mapping entid, and then associate the extracted information with the database.
[0114] S303. After the extracted enterprise entity information is aligned, a preliminary equity map of the enterprise is constructed in combination with the extracted equity relationship.
[0115] Extract the equity relationship of a certain enterprise Y.
[0116] Assume that the equity and voting rights structure of Enterprise Y is as follows:
[0117] Shareholder 1 holds 70.5% of the voting rights in Company Y
[0118] Shareholder 2 holds 66.7% of the equity of Company Y
[0119] Shareholder 3 holds 3.1% of the equity of Company Y
[0120] Shareholder 4 holds 3.8% of the equity of Company Y
[0121] In addition, according to publicly disclosed information, Shareholder 1 holds 100% of Shareholder 2 and Shareholder 3, which results in a certain circular or double marking in the preliminary equity map: it shows that Shareholder 1 directly holds the voting rights of Enterprise Y, and at the same time holds the equity or voting rights of Enterprise Y through Shareholder 2 and Shareholder 3, which are wholly owned by Shareholder 1. The preliminary equity map is as follows: Figure 2 shown.
[0122] After completing the construction of the preliminary equity map in S300, the equity relationship network in the map has taken shape. Then, in S400, a comprehensive loop detection is carried out on this preliminary equity map. With rigorous algorithms and logic, it is accurately determined whether there are contradictions, and the contradictions are resolved in a scientific and reasonable way, thereby ensuring the logical rigor of the equity map and the accuracy of the data. The specific contradiction resolution method is as follows:
[0123] S400, detecting loops in the preliminary equity map, determining whether there are contradictions and resolving the contradictions;
[0124] The specific steps to determine whether there is a contradiction and resolve it are:
[0125] S401. After generating a preliminary equity graph of an enterprise, take the names of shareholders in the preliminary equity graph as nodes and the equity relationships as edges, extract all relationship paths in the equity graph, and judge the relationship paths. When there is only one node and one edge connecting the enterprise in the relationship path, the corresponding path is judged to be a direct path. When there are more than one node and one edge connecting the enterprise in the relationship path, it is judged to be a multi-hop path.
[0126] S402, extracting the first nodes in the direct path and the multi-hop path, and using the first nodes to classify the relationship paths, and classifying the same first nodes as the same shareholder equity relationship;
[0127] In the same shareholder equity relationship, the equity map is used to calculate the shareholding ratio of each shareholder. For the direct path, the formula is: , X represents shareholder G r The shareholding ratio of enterprise Q through the direct path, w represents the equity relationship of the edge in the equity graph; for multi-hop paths, the formula is:
[0128] ;
[0129] In the formula, D (G r , Q) represents G r Based on the shareholding ratio of enterprise Q in the multi-hop path, P represents the total number of multi-hop paths in the equity relationship of the same shareholder, u represents the total number of edges in a multi-hop path, and Gy represents the yth node in a multi-hop path, G y+1 represents the y+1th node in a multi-hop path; x belongs to 1 to P, y belongs to 1 to u;
[0130] S403. In the same shareholder equity relationship, the shareholder's shareholding ratio of enterprise Q through the direct path and the shareholding ratio of enterprise Q through the multi-hop path are judged. When X=D(G r , Q), determine whether there is any contradiction in the shareholder's shareholding ratio of enterprise Q, and eliminate the direct path; calculate all shareholder equity relations, eliminate all contradictory direct paths, and output the updated equity map.
[0131] In the initial stage of graph construction, if the relationships such as "Shareholder 1 → Enterprise Y (70.5% of voting rights)" and "Shareholder 2 → Enterprise Y (66.7%)" and "Shareholder 1 → Shareholder 2 (100%)" are directly placed in the graph in parallel, the system will generate a circular structure of "Shareholder 1 holds Enterprise Y both directly and indirectly".
[0132] Similarly, "Shareholder 1 → Shareholder 3 (100%)" and "Shareholder 3 → Company Y (3.8%)" will also lead to overlapping or conflicting descriptions of "Shareholder 1's control over Company Y". In the corporate equity map, this cycle will make people mistakenly believe that Shareholder 1 has "direct + indirect" dual holdings, and the total voting rights value will be ambiguous. If it is not handled, it will be mistakenly determined that Shareholder 1's voting rights ratio for Company Y should be added to the indirect holdings of Shareholders 2 and 3, resulting in "falsely high" or unreasonable data records.
[0133] Through the strongly connected component (SCC) algorithm, several loops with "shareholder 1-shareholder 2-enterprise Y" and "shareholder 1-shareholder 3-enterprise Y" as the core were scanned.
[0134] In the loop, the credibility and logical consistency of each edge are analyzed. According to the source of equity and actual disclosure, it is confirmed that shareholder 1 does not directly hold the equity of enterprise Y, but indirectly controls enterprise Y through shareholders 2 and 3.
[0135] Although public information often discloses that shareholder 1 has 70.5% of the voting rights in company Y, this does not mean that he directly holds shares in company Y;
[0136] After the graph optimization process, the edge "Shareholder 1 → Enterprise Y (direct shareholding)" will be downgraded or marked as "virtualized". The real structure after the contradiction is resolved should be:
[0137] “Shareholder 1 → Shareholder 2 (100% shareholding)”
[0138] “Shareholder 2 → Company Y (66.7% equity)”
[0139] “Shareholder 1 → Shareholder 3 (100% shareholding)”
[0140] “Shareholder 3 → Company Y (3.8% equity)”
[0141] Correct equity penetration diagram after replanning:
[0142] Shareholder 1 (through wholly owned shareholders 2 and 3) → Company Y
[0143] Shareholder 4 → Company Y (3.1% equity)
[0144] When displaying internally or externally, shareholder 1's 70.5% voting rights in company Y can be marked as "indirect control" and a note can be kept stating that "shareholder 1 does not have direct equity, but controls through two special purpose companies under its umbrella." Figure 3 shown.
[0145] First, the direct path and multi-hop path are determined to build the same shareholder relationship. Then, the equity ratios of different paths are calculated, and the equity ratios of different paths are compared to determine the contradictions. Finally, for conflicting equity ratios or multiple sources of shareholding in the same enterprise, a "deletion" strategy is adopted, and its confidence information is retained. This operation not only ensures data quality but also facilitates traceability. It effectively avoids the common equity "circulation" or "false high" problems in traditional systems, and can automatically handle shareholding ratio conflicts, ensuring that the graph structure is clear and logically self-consistent, providing a more reliable data basis for subsequent analysis and supervision.
[0146] After S400's rigorous loop detection and contradiction resolution, the internal logic of the equity map has been optimized and improved. At this point, S500, based on this, focuses on in-depth mining of the equity map after the contradictions are resolved, identifies the interest community through professional analysis methods, and accurately determines whether there is family control. Based on the judgment results, virtual nodes are constructed to associate family control relationships, and finally a complete and healthy equity map is output to provide solid and reliable data support for subsequent related research and decision-making. The specific method of identifying the interest community is as follows:
[0147] S500, identify the community of interests in the equity map after the contradiction is resolved, determine whether it is family control, build a virtual node to associate family control, and output a healthy equity map;
[0148] The specific steps to output a healthy equity map are:
[0149] S501. Use the community division algorithm to divide sub-communities in the preliminary equity graph after contradictions are resolved, extract the enterprises and all nodes in the sub-communities, and calculate the total number of edges when all nodes in the sub-communities are fully interconnected: , where L represents the number of all nodes and enterprises in the extracted sub-community. The equity density in the sub-community is calculated as follows: , where ED represents the equity density in the subcommunity, and w e represents the equity relationship of the e-th edge;
[0150] S502: For each node in the subcommunity, calculate the sum of the edge equity relationships between the node and all other nodes, take the sum of the edge equity relationships as the control contribution CC of the corresponding node, and extract the maximum control contribution CC among all nodes max , calculate the maximum control concentration in the subcommunity, the formula is:
[0151] ;
[0152] In the formula, represents the maximum control concentration of the subcommunity;
[0153] S503, extracting the identity information of all nodes in the sub-community, extracting the nodes belonging to the same family, calculating the sum of the equity relationships of the nodes in the same family as FC, and using FC as the family control degree in the sub-community;
[0154] S504, customizing the control density threshold to α D , the maximum control contribution threshold is α CC , the family control threshold is α FC ; When ED>α D and >α CC When FC>α, the subcommunity is judged as a community of interest and marked; FC When the subcommunity is judged to be family-controlled, a virtual node is constructed for the subcommunity in the database, and the virtual node is associated with the equity relationship of the node controlled by the family. The virtual node is highlighted in the preliminary equity map after the contradiction is resolved to obtain a healthy equity map.
[0155] To determine whether there is a community of interests for a certain enterprise Y, based on the example, the simplified equity relationship (nodes and shareholding ratios) is as follows:
[0156] Enterprise Y: target enterprise;
[0157] There are five natural person shareholders 1-5, and all five shareholders have the same surname. Each of them claims to hold 74.81% of the shares. It is necessary to determine whether the shareholdings are consolidated or there are multiple records based on the actual information. This is only used as an example.
[0158] Investment company nodes: Enterprise 1 (8.26%), Enterprise 2 (66.55%), Enterprise 3 (66.55%), Enterprise 4 (66.55%), Enterprise 5 (66.55%); many high-proportion edges of these nodes point to "Enterprise Y" or there is a certain degree of cross-holding between them.
[0159] The community partitioning algorithm is applied to the entire equity network. The results show that "Enterprise Y" is divided into the same subcommunity as the above-mentioned natural person shareholders and investment company nodes, indicating that they are highly interconnected in the graph. The number of nodes in the same subcommunity is 11, and the number of edges in the case of full interconnection = 11×10 / 2=55; calculate the equity density, assuming that the final sum of edge weights is 80.0, then ED=80.0 / 55≈1.45;
[0160] With α D =0.3, 1.45>>0.3, indicating that the internal connections of the community are extremely dense.
[0161] To calculate the maximum control concentration, the sum of the equity relationship between the node and all other nodes is designed to be 80, and the total control contribution of a "family merger" node or a single natural person reaches the maximum control contribution CC max is 50; the maximum control concentration is calculated to be 0.625;
[0162] Let α CC =0.6, 0.625>0.6, judging that this node (or merger) contributes 62.5% of the equity edge of the entire community, and the concentration is extremely high.
[0163] When the calculated equity density and maximum control concentration are both greater than the threshold, the subcommunity is judged to be a community of interests;
[0164] Since the five natural person shareholders all have the same surname, the family control degree FC is designed to be 0.7, and the family control degree threshold α FC =0.5; the subcommunity is judged to be family-controlled.
[0165] After successfully completing the construction of a healthy equity map by S500, S600 will immediately start the real-time data update process to ensure that the map can always reflect the latest equity situation of the enterprise. By extracting enterprise entity information in real time and verifying the credibility of these real-time information with the help of the healthy equity map generated by S500, the real-time enterprise entity information that has passed the verification is accurately integrated into the healthy equity map, so that the equity map always maintains timeliness and accuracy, providing a continuous and reliable basis for dynamic monitoring and analysis of corporate equity. The specific method of updating the healthy equity map is as follows:
[0166] S600, extracting enterprise entity information in real time, using the healthy equity map to verify the credibility of the real-time extracted enterprise entity information, and updating the healthy equity map using the verified real-time enterprise entity information;
[0167] The specific steps to update the healthy equity map using verified real-time corporate entity information are:
[0168] Extract enterprise entity information in real time, align the extracted enterprise entity information with the existing ID in the healthy equity map, search for all equity relationships in the extracted enterprise entity information in the healthy equity map, and if not found, judge the corresponding equity relationship as suspicious information;
[0169] When found, determine whether all equity relationships in the real-time extracted enterprise entity information are the same as the equity relationships in the healthy equity map. If they are not the same, mark them as suspicious information in the healthy equity map.
[0170] All corporate entity information extracted in real time is verified, suspicious information is marked in the healthy equity map, and the healthy equity map is updated.
[0171] In the real-time extracted corporate entity information, shareholder 1 holds 50% of shareholder 2, while in the healthy equity map, shareholder 1 holds 100% of shareholder 2, which is considered suspicious information. The equity relationship is marked in the healthy equity map.
[0172] In S600, after completing the credibility check of the real-time extracted enterprise entity information and updating the healthy equity map, in order to achieve persistent storage and convenient call of data, S700 immediately stores the updated healthy equity map;
[0173] S700. Output the updated healthy equity map and store it in the database.
[0174] The specific steps to output the updated healthy equity graph are:
[0175] S701. Save the healthy equity map after contradiction resolution, interest community annotation, and healthy equity map verification in the graph database; record the confidence and data source information of all edges and nodes in the equity map, visualize and output the equity map through the front-end visualization tool, and store it in the Neo4j database.
[0176] A storage medium stores a computer program, which, when executed by a processor, implements the execution steps of a construction and optimization method.
[0177] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A method for constructing an enterprise equity map based on multi-dimensional data verification, characterized in that: The method comprises the following steps: S100, acquiring document data, extracting text from the acquired document data, performing layout analysis and table structure reconstruction on the extracted text, and then converting the format of the text; S200, inputting the format-converted text into a large language model to extract the corporate entity information and equity relationship in the text; The specific steps for extracting corporate entity information and equity relationships in the text are: S201, dividing the markup file into blocks, converting each block into a vector using a text vector model, and storing all the block vectors in a vector database to form an index that can be quickly queried; S202, customizing and constructing prompt information including corporate equity relations, using the formed index to retrieve a marker block reflecting the corporate equity relations in the vector database, concatenating the retrieved marker block with the prompt information, and then initiating an information extraction request to the large language model; S203, after receiving the request, the large language model parses the text, extracts the enterprise entity information and equity relationship, and converts different information formats into numerical forms; uses the extracted enterprise entity information and equity relationship to construct a triple structure, specifically: enterprise name-equity relationship-shareholder name; S300, matching the extracted enterprise entity information with the enterprise information in the external database to construct a preliminary equity map; The specific steps to construct a preliminary equity map are: S301, extracting enterprise entity information and enterprise information from external databases; S302, determining whether the enterprise entity information contains an ID. If the ID is contained, accurately matching the ID with the enterprise entity code ENTID in the enterprise information of the external database. If the ID is equal to the enterprise entity code ENTID in the enterprise information of the external database, the match is successful, and information alignment is performed directly. When the ID is not included, the similarity Jaccard (A, B) between the enterprise entity information and each enterprise information in the external database is calculated using the formula: ; In the formula, Jaccard (A, B) represents the similarity between the extracted enterprise entity information and each enterprise information in the external database, A represents the character set of the extracted enterprise entity information, and B represents the character set of the enterprise name of the enterprise information in the external database; When Jaccard(A, B) = 1, it means they are exactly the same, and when Jaccard(A, B) = 0, it means there are no identical characters. The user-defined similarity threshold β is used to judge the calculated similarity. Jaccard(A, B) max When >β, the match is successful, information alignment is performed and the ENTID in the external database enterprise information is recorded; S303, after the extracted enterprise entity information is aligned, a preliminary equity map of the enterprise is constructed in combination with the extracted equity relationship; S400, detecting loops in the preliminary equity map, determining whether there are contradictions and resolving the contradictions; S500, identify the community of interests in the equity map after the contradiction is resolved, determine whether it is family control, build a virtual node to associate family control, and output a healthy equity map; S600, extracting enterprise entity information in real time, using the healthy equity map to verify the credibility of the real-time extracted enterprise entity information, and updating the healthy equity map using the verified real-time enterprise entity information; S700. Output the updated healthy equity map and store it in the database.
2. The method for constructing an enterprise equity map based on multi-dimensional data verification according to claim 1 is characterized in that: The specific steps of performing layout analysis and table structure reconstruction on the extracted text in S100 are: S101, extracting a PDF document from an enterprise annual report database, storing the extracted PDF document in a file management server, and dividing the PDF document into a scanned version and a native version with a text layer; S102, calling an algorithm engine to perform page-level text recognition on a scanned PDF document, and extracting text items in the scanned PDF document; reading algorithm data in the process of extracting text items, obtaining the confidence of each text item, and extracting the top position, left position, height, and width of characters in the text item; The confidence of each text item is used to screen the text items, the text items whose confidence exceeds the threshold are retained, and the text items whose confidence is less than the threshold are removed; Record the position information of each text item retained after filtering, calculate the upper, lower, left, and right coordinates of the bounding box of the characters in the text item, and set the standard value of the page size to 1; For native PDF documents, directly extract low-level text items; S103: After extracting text items from all stored PDF documents, the rectangular area of each character in the text item is calculated using the upper, lower, left, and right coordinates of the bounding box of the characters in the text item, and the rectangular area of each character is standardized to a relative value of the page size, and then the font size Font_Size of each character is calculated. i ; S104. After obtaining the font size of each character, the average font size of all characters is calculated using the average formula, and then the characters are divided into two clusters according to the font size using the discrete clustering algorithm, where cluster 1 is the main text font and cluster 2 is the title font; the font size of cluster 2 is larger than that of cluster 1; after the characters are divided into two clusters, the characters in the two clusters are structurally marked, specifically: according to the clustering results, the text items that are identified as titles are marked with corresponding title tags, and the text items that are identified as main text are marked with ordinary paragraph tags, and a markup file is generated.
3. The method for constructing an enterprise equity map based on multi-dimensional data verification according to claim 2 is characterized in that: The specific steps of determining whether there is a contradiction and resolving the contradiction in S400 are: S401. After generating a preliminary equity graph of an enterprise, take the names of shareholders in the preliminary equity graph as nodes and the equity relationships as edges, extract all relationship paths in the equity graph, and judge the relationship paths. When there is only one node and one edge connecting the enterprise in the relationship path, the corresponding path is judged to be a direct path. When there are more than one node and one edge connecting the enterprise in the relationship path, it is judged to be a multi-hop path. S402, extracting the first nodes in the direct path and the multi-hop path, and using the first nodes to classify the relationship paths, and classifying the same first nodes as the same shareholder equity relationship; In the same shareholder equity relationship, the equity map is used to calculate the shareholding ratio of each shareholder. For the direct path, the formula is: , X represents shareholder G r The shareholding ratio of enterprise Q through the direct path, w represents the equity relationship of the edge in the equity graph; for multi-hop paths, the formula is: ; In the formula, D (G r , Q) represents G r Based on the shareholding ratio of enterprise Q in the multi-hop path, P represents the total number of multi-hop paths in the equity relationship of the same shareholder, u represents the total number of edges in a multi-hop path, and G y represents the yth node in a multi-hop path, G y+1 represents the y+1th node in a multi-hop path; x belongs to 1 to P, y belongs to 1 to u; S403. In the same shareholder equity relationship, the shareholder's shareholding ratio of enterprise Q through the direct path and the shareholding ratio of enterprise Q through the multi-hop path are judged. When X=D(G r , Q), it is determined that there is a contradiction in the shareholder's shareholding ratio of enterprise Q, and the direct path is eliminated; the equity relationship of all shareholders is calculated, and all contradictory direct paths are eliminated.
4. The method for constructing an enterprise equity map based on multi-dimensional data verification according to claim 3 is characterized in that: The specific steps of outputting the healthy equity map in S500 are: S501, using a community partitioning algorithm to partition a sub-community in the preliminary equity graph after contradictions are resolved, extracting enterprises and all nodes in the sub-community, calculating the total number of edges S when all nodes in the sub-community are fully interconnected, and calculating the equity density ED in the sub-community; S502: For each node in the subcommunity, calculate the sum of the edge equity relationships between the node and all other nodes, take the sum of the edge equity relationships as the control contribution CC of the corresponding node, and extract the maximum control contribution CC among all nodes max , calculate the maximum control concentration in the subcommunity, the formula is: ; In the formula, represents the maximum control concentration of the subcommunity; w e represents the equity relationship of the e-th edge, and S represents the total number of edges; S503, extracting the identity information of all nodes in the sub-community, extracting the nodes belonging to the same family, calculating the sum of the equity relationships of the nodes in the same family as FC, and using FC as the family control degree in the sub-community; S504, customizing the control density threshold to α D , the maximum control contribution threshold is α CC , the family control threshold is α FC ; When ED>α D and >α CC When FC>α, the subcommunity is judged as a community of interest and marked; FC When the subcommunity is judged to be family-controlled, a virtual node is constructed for the subcommunity in the database, and the virtual node is associated with the equity relationship of the node controlled by the family. The virtual node is highlighted in the preliminary equity map after the contradiction is resolved to obtain a healthy equity map.
5. The method for constructing an enterprise equity map based on multi-dimensional data verification according to claim 4 is characterized in that: The specific steps of updating the healthy equity map using the verified real-time enterprise entity information in S600 are: S601, extracting enterprise entity information in real time, aligning the extracted enterprise entity information with the existing ID in the healthy equity map, searching for all equity relationships in the extracted enterprise entity information in real time in the healthy equity map, and if no equity relationship is found, judging the corresponding equity relationship as suspicious information; When found, determine whether all equity relationships in the real-time extracted enterprise entity information are the same as the equity relationships in the healthy equity map. If they are not the same, mark them as suspicious information in the healthy equity map. All corporate entity information extracted in real time is verified, suspicious information is marked in the healthy equity map, and the healthy equity map is updated.
6. The method for constructing an enterprise equity map based on multi-dimensional data verification according to claim 5 is characterized in that: The specific steps of outputting the updated healthy equity map in S700 are: S701. Save the healthy equity map after contradiction resolution, interest community annotation, and health map verification in the graph database; record the confidence and data source information of all edges and nodes in the equity map, and visualize and output the equity map through the front-end visualization tool.
7. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the execution steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Enterprise relation knowledge base construction method based on knowledge graph
CN115129879A
Data query method and device, storage medium and computer program product
CN118503454A