Enterprise analysis system based on knowledge graph and retrieval enhancement generation technology

By combining knowledge graphs and large language models, the multi-dimensional information integration problem of enterprise analysis systems in the existing technology is solved, accurate question-and-answer and multi-dimensional display of enterprise information is achieved, the reliability and adaptability of the system is improved, and a variety of application scenarios are supported.

CN120386858APending Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510464284.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing enterprise analysis systems rely on single data retrieval and shallow semantic matching, making it difficult to obtain complete and accurate corporate portraits from multi-dimensional information, resulting in low query efficiency and inaccurate analysis results. Moreover, large language models are prone to hallucinations when they lack high-quality data support, affecting system reliability and interpretability.

Method used

Using an enterprise analysis system based on knowledge graph and retrieval enhancement generation technology, we process original data through OCR, create knowledge graph entities and relationships, use pre-trained vector models to encode entity names, combine large language models for semantic analysis and keyword dictionary generation, and realize accurate question-and-answer and multi-dimensional display of enterprise information.

Benefits of technology

It realizes accurate Q&A and multi-dimensional display of enterprise information, improves the reliability and interpretation of analysis results, supports various application scenarios such as enterprise information Q&A, investment decisions and market insights, and has the ability to efficiently coordinate and flexible adapt.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386858A_ABST
    Figure CN120386858A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise analysis system based on a knowledge graph and a retrieval enhancement generation technology, and the system comprises the steps: original data processing: obtaining enterprise data D; based on the enterprise data D, an entity Ei is created in the knowledge graph, and a relation Ri is added; generating an enterprise portrait document based on the enterprise data D and the related information of the knowledge graph entity Ei and the graph relation Ri; encoding the name information of the knowledge graph entity Ei by using the pre-training vector model, and constructing a vector database based on the name vector code Vi; the method comprises the following steps: semantically analyzing user query by utilizing a large language model, and generating a keyword dictionary DiCt and a relationship importance list List; retrieving the knowledge graph by using the entity name Ni and the relationship importance list List, and obtaining query related data and an enterprise portrait document; and generating a user query answer based on the related information data by using a large language model, and providing an enterprise portrait. According to the method, the enterprise multi-source heterogeneous data is orderly integrated through the knowledge graph, the enterprise portrait is generated, accurate question answering and multi-dimensional display of the enterprise information are realized in combination with the generation capability of the large language model, and the method has the advantages of efficient collaboration, flexible adaptation and the like; and various application scenes such as enterprise information questions and answers, investment decisions and market insight can be effectively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large language model applications, and specifically to an enterprise analysis system based on knowledge graph and retrieval augmented generation technology. Background Art

[0002] Enterprise information is of crucial value in aspects such as business decision-making, risk control, and market insight. However, since existing enterprise analysis systems often rely on single data retrieval methods or shallow semantic matching, it is difficult to obtain a complete and accurate enterprise portrait from multi-dimensional information, resulting in problems such as low query efficiency and inaccurate analysis results frequently occurring in practice, severely restricting the effective release of enterprise data value. With the significant breakthroughs made by large language models in the field of text generation and understanding, the industry has begun to attempt to introduce them into enterprise analysis scenarios. However, traditional large language models are prone to the so-called "hallucination" phenomenon when data is incomplete or lacks high-quality support, that is, the model generates non-existent or false content out of thin air during the answering process, reducing the reliability and interpretability of the system in key operations. At the same time, a simple model reasoning solution is difficult to meet the accurate retrieval requirements of massive, multi-source heterogeneous enterprise data. If the model lacks semantic perception of upstream and downstream information and dynamic association relationships, it is also difficult to provide traceable and interpretable comprehensive analysis results. Facing the above problems, there is an urgent need for an enterprise analysis system that can combine the generation ability of large language models with high-quality data retrieval to eliminate the uncertainty brought by model "hallucinations" and achieve global management and in-depth mining of enterprise information. The present invention proposes an enterprise analysis system based on knowledge graph and retrieval augmented generation technology, which orderly integrates enterprise multi-source heterogeneous data through the knowledge graph, combines the generation ability of the large language model, realizes accurate question answering and multi-dimensional display of enterprise information, has advantages such as efficient collaboration and flexible adaptation, and can effectively support various application scenarios such as enterprise information question answering, investment decision-making, and market insight. Summary of the Invention

[0003] An enterprise analysis system based on knowledge graph and retrieval augmented generation technology, characterized in that: the system involves a large language model, a knowledge graph, and a vector database. The system includes the following steps:

[0004] Step 1, use OCR recognition and special format matching to process the original data, and extract editable text data. Perform format conversion, missing value processing, noise filtering, and standardization processing on the text data to obtain the processed data D.

[0005] Step 2, use the enterprise data D to create a knowledge graph entity E i , and add a knowledge graph relationship R i .

[0006] Step 3: Generate an enterprise portrait document using enterprise data D, knowledge graph entities E i and knowledge graph relationships R i related information.

[0007] Step 4: Use a pre-trained vector model to encode the name information of knowledge graph entities E i to obtain a name vector V i and build a vector database.

[0008] Step 5: Use a large language model to semantically parse user queries to generate a keyword dictionary Dict and a relationship importance list List.

[0009] Step 6: Use the keyword dictionary Dict to retrieve the vector database to obtain the matching entity names N i in the knowledge graph.

[0010] Step 7: Use the knowledge graph entity names N i and the relationship importance list List to retrieve the knowledge graph to obtain query-related data and enterprise portrait documents.

[0011] Step 8: Use a large language model to generate user query answers based on the relevant information data and provide an enterprise portrait.

[0012] Specifically, in Step 1, the use of OCR recognition and special format matching to process the original data, perform format conversion, missing value processing, noise filtering, and standardization processing on the data to obtain the processed data D. Specifically: Obtain enterprise data from channels such as industrial and commercial registration, news sentiment, financial statements, and bidding information. First, perform format discrimination on different data sources. For parsable documents, use document parsing tools to extract text and table content and convert it into editable text. For documents containing images or those where text cannot be directly extracted, first perform denoising, binarization, and skew correction operations on the images, then use OCR to recognize the text and table structure and convert it into editable text. Locate and parse keyword fields by customizing special format matching rules to identify core elements such as enterprise names, registered capital, trade transactions, and judgment documents. Uniformly export the above text or structured data in CSV format, and perform missing value filling, noise and duplicate data filtering, and outlier removal. Perform standardization processing on keyword fields, including unifying the format of enterprise names, converting currency units, unifying the format of time fields, and validating numerical ranges.

[0013] Specifically, in Step 2, the use of enterprise data D to create knowledge graph entities E i based on the basic industrial and commercial information of each enterprise and the trade and investment information between enterprises, and add knowledge graph relationships R i, specifically: in the knowledge graph, create enterprise entity E based on the enterprise information in data D fi , E fi 's attributes include information in the following dimensions: name, legal representative, registration status, enterprise address, enterprise scale, registered capital, paid-in capital, business scope, enterprise profile, operating income. Create address entity E for each province and city di , for E fi and its corresponding E di add a "located in" relationship. For all E di that have a "located in" relationship with E fi , calculate the percentile information of operating income, registered capital, and paid-in capital respectively, and add it to the attribute information of E di . According to the "National Economic Industry Classification" document, create industry category entity E ci , for E fi and its corresponding E ci add a "belongs to" relationship. For all E ci that have a "belongs to" relationship with E fi , calculate the percentile information of operating income, registered capital, and paid-in capital respectively, and add it to the attribute information of E ci . Create trade event entity E based on the trade information in the data ti , E ti includes the following dimension information: name, trade amount, report date. For E fi as the supplier, add a "supplier" relationship between it and the corresponding E ti ; for E fi as the customer, add a "customer" relationship between it and the corresponding E ti . Create investment event entity E based on the investment information in the data ni , E ni includes the following dimension information: name, transaction amount, disclosure date, financing round, shareholding ratio. For E fi as the investor, add an "investment" relationship between it and the corresponding E ni ; for E fi as the investee, add a "financing" relationship between it and the corresponding E ni .

[0014] Specifically, in step 3, the relevant information of enterprise data D, knowledge graph entity E i and knowledge graph relationship R i is used to generate an enterprise portrait document, specifically: create an enterprise portrait document for each enterprise. Based on the knowledge graph created in step 2, obtain the attribute information of all dimensions of the corresponding entity E fi as basic business information and add it to the enterprise portrait document. Retrieve and Efi Entity E with a "located in" relationship di , obtain E di 's statistical attribute information and add it as regional dimension information to the enterprise portrait document. Retrieve and E fi Entity E with a "belongs to" relationship ci , obtain E ci 's statistical attribute information and add it as industry dimension information to the enterprise portrait document. Retrieve and E fi Associated entity E ti , obtain trade amount and report date dimension information, analyze indicators such as the distribution of the enterprise's suppliers and customers and the trade amount, and add it as supply chain dimension information to the enterprise portrait document. Retrieve and E fi Associated entity E ni , obtain transaction amount, disclosure date, financing round, shareholding ratio dimension information, analyze the enterprise's investment behavior and financing process, and add it as investment and financing dimension information to the enterprise portrait document. Based on the enterprise data obtained in step 1, retrieve relevant information such as judgment documents, judicial cases, and administrative penalties related to the enterprise, and add it as risk dimension information to the enterprise portrait document.

[0015] Specifically, in step 4, the pre-trained vector model is used to encode the name information of the knowledge graph entity E i to obtain the name vector V i , and construct a vector database, specifically: use an open-source vectorization model to vectorize and encode the names of each entity in the knowledge graph to obtain the name vector V i . Create corresponding metadata tags for each enterprise entity E fi Is the vector encoding of the name of the enterprise's affiliated address, Is the vector encoding of the name of the enterprise's affiliated industry. Use a database tool to create a database and add vector encoding, original text, and metadata tags.

[0016] Specifically, in step 5, the large language model is used to semantically parse the user query to generate a keyword dictionary Dict and a relationship importance list List, specifically: call the large language model API to analyze the user input query. Use the large language model to extract the enterprise name, address name, and industry name included in the user query, and return the extraction result as the keyword dictionary Dict. Obtain all relationship types in the knowledge graph, use the large language model to analyze the correlation between each relationship type and the user query, sort the list according to the degree of correlation, and return the result as the relationship importance list List.

[0017] ​Specifically, in step 6, the vector database is retrieved using the keyword dictionary Dict to obtain the matching entity name N in the knowledge graph i , specifically: extract the enterprise name, address, and industry name fields in Dict, and use an open-source vectorization model to perform vectorization encoding on each field name. If the enterprise name field is not empty, calculate the address vector V di and the industry name vector V ci with the similarity of the metadata tags where if the similarities are not all 0, only retain the enterprise name data of the top 25% of the similarities as the search space; otherwise, retain all enterprise name data as the search space. Calculate the similarity between E fi and other vectors in the search space , and return the enterprise name N with the highest similarity fi . If the enterprise name field is empty and the address name or industry name field is not empty, use all address name data as the search space, calculate the similarity between V di and other vectors in the search space , and return the address name N with the highest similarity di ; use all industry name data as the search space, calculate the similarity between E ci and other vectors in the search space , and return the industry name N with the highest similarity ci .

[0018] Specifically, in step 7, the knowledge graph is retrieved using the knowledge graph entity name N i and the relationship importance list List to obtain query-related data and enterprise portrait documents. Specifically: create a relevant information document, use the name data in step 5 to retrieve the knowledge graph, and obtain the entity E corresponding to each name i , and add all attribute information of E i to the relevant information document. If E i is an enterprise entity, obtain its corresponding enterprise portrait document. Then use the List in step 4 to sequentially traverse the relationship element R, and retrieve all entities E i that have a relationship R with the entity E ri , and add all attribute information of E ri to the relevant information document until the relevant information document reaches the maximum word limit

[0019] Specifically, in step 8, the large language model is utilized to generate user query answers based on relevant information documents and provide an enterprise portrait. Specifically: Call the large language model API to answer user queries. Use the relevant information documents in step 7 as "background knowledge". Through prompt engineering, require the large language model to answer questions strictly based on the background knowledge, and return the output content of the large language model as the answer. If an enterprise portrait document is obtained in step 7, generate a visual front-end page based on the static site generator tool and the document content, and return the website address of the front-end page. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0021] Figure 1 It is a flowchart of an enterprise analysis system based on knowledge graph and retrieval augmentation generation technology.

[0022] Figure 2 It is a flowchart of enterprise portrait generation based on knowledge graph and enterprise data. Detailed Embodiments

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0024] Existing enterprise analysis systems generally rely on single data retrieval and shallow semantic matching, which are difficult to meet the requirements of multi-dimensional information integration and user interaction, resulting in incomplete enterprise portrait results and inconvenient user queries, seriously affecting the effective utilization of data value. To address the above problems, the present invention integrates a large language model and retrieval augmentation technology for enterprise analysis. It uses a knowledge graph to orderly integrate multi-source heterogeneous enterprise data, combines the model generation ability, and realizes accurate question answering and multi-dimensional display of enterprise information. By introducing structured knowledge support, the risk of model hallucinations is reduced, and the reliability and interpretability of the analysis results are improved. This solution effectively supports various application scenarios such as enterprise information question answering, investment decision-making, and market insight, and has the capabilities of efficient collaboration and flexible adaptation. The specific steps are as follows:

[0025] S101: Process the original data, perform format conversion, missing value processing, noise filtering, and standardization operations on the original data to obtain data D.

[0026] Specifically, enterprise data is obtained from multiple public channels such as industrial and commercial registration, news sentiment, financial statements, and bidding information. For the obtained various data sources, the format is first discriminated. For directly parsable documents, document parsing tools are used to extract the text and table content and convert it into editable text. For documents containing images or those where text information cannot be directly extracted, preprocessing operations such as denoising, binarization, and skew correction are first performed on the images, and then the OCR algorithm is used to recognize the page text and table structure, and the recognition results are converted into editable text. Further, custom special format matching rules are used to locate and parse keyword fields in the extracted text, and core elements such as enterprise names, registered capital, trade transactions, and judgment documents are identified. The above text or structured data is uniformly exported to a processable CSV format, and on this basis, missing value filling, noise and duplicate data filtering, and outlier removal are performed. Standardization processing is carried out on the keyword fields, including unified formatting of enterprise names, currency unit conversion, format unification of time and date fields, and range verification of numerical fields.

[0027] S102: Create a knowledge graph entity E using enterprise data D i , and add a knowledge graph relationship R i .

[0028] Specifically, in the knowledge graph, enterprise entities E are created based on each enterprise in the data fi , and the attributes of E fi include information in the following dimensions: name, legal representative, registration status, enterprise address, enterprise scale, registered capital, paid-in capital, business scope, enterprise profile, and operating income. Address entities E are created for each province and city di , and according to the geographical location of E fi , a "located in" relationship is added between E fi and its corresponding E di . The attributes of E di include information in the following dimensions: name. For all E di that have a "located in" relationship with E fi , the 5%, 25%, 50%, and 75% quantile values of operating income, registered capital, and paid-in capital are calculated respectively. According to the national standard industry categories in the "National Economic Industry Classification", industry category entities E ci are created, and according to the industry category of E fi , a "belongs to" relationship is added between E fi and its corresponding E ci . The attributes of E ci include information in the following dimensions: name. For all E ci that have a "belongs to" relationship with E fiCalculate the 5%, 25%, 50%, and 75% percentiles of operating income, registered capital, and paid-in capital respectively. Create a trade event entity E based on each piece of trade information in the data. ti , E ti Includes the following dimension information: name, trade amount, report date, E as supplier in each trade event fi , add it and the corresponding E ti The relationship between the "supplier" and E as the customer fi , add it and the corresponding E ti Create an investment event entity E based on each investment information in the data. ni , E ni Includes the following dimensional information: name, transaction amount, disclosure date, financing round, shareholding ratio, and E as the investor in each investment event. fi , add it and the corresponding E ni The “investment” relationship between E as the investee fi , add it and the corresponding E ni The "financing" relationship between them.

[0029] S103: Using enterprise data D and knowledge graph entity E i Relationship with the graph R i Generate enterprise portrait documents based on relevant information.

[0030] Specifically, a corporate portrait document is created for each enterprise. Based on the knowledge graph created in S102, the corresponding enterprise entity E is obtained. fi All dimension attribute information is added to the enterprise portrait document as basic business information. fi E with a "located in" relationship di Entity, get E di The statistical attribute information (percentile values of operating income, registered capital and paid-in capital) of E fi The relevant information in the horizontal comparison is used to reflect the level of the enterprise in its region and is added to the enterprise portrait document as regional dimension information. fi Entity E with a "belongs to" relationship ci , get E ci The statistical attribute information (percentile values of operating income, registered capital and paid-in capital) of E fi The relevant information in the horizontal comparison is used to reflect the level of the enterprise in its industry field, and is added to the enterprise portrait document as industry dimension information. fi Associated E tiAn entity that obtains information on trade amounts and reporting date dimensions, identifies the position of the enterprise in the upstream and downstream of the supply chain based on the "supplier" or "customer" relationship, and counts the composition of the enterprise's main suppliers, the distribution of supplier industries, and the allocation of procurement amounts; counts the composition of the enterprise's main customers, the distribution of customer industries, and the allocation of sales amounts, and adds them as supply chain dimension information to the enterprise portrait document. Retrieve and E fi associated E ni An entity that obtains information on transaction amounts, disclosure dates, financing rounds, and shareholding ratios, and analyzes the enterprise's activities in the capital market in combination with the "investment" or "financing" relationship. Analyze the enterprise's investment behavior and financing process, reflecting the enterprise's growth process and capital preference, and add them as investment and financing dimension information to the enterprise portrait document. Based on the enterprise data obtained from S101, retrieve relevant information such as judgment documents, judicial cases, and administrative penalties related to the enterprise, and add them as risk dimension information to the enterprise portrait document.

[0031] S104: Use a pre-trained vector model to obtain the name vector encoding V i of the knowledge graph entity E i and build a vector database.

[0032] Specifically, use an open-source vectorization model to vectorize and encode the names of each entity in the knowledge graph to obtain the name vector encoding V i . Create corresponding metadata tags for each enterprise entity E fi , including fi the vector encoding of the name of the affiliated address E fi the vector encoding of the name of the affiliated industry Use a database tool to create a database, add each name vector encoding and its corresponding original text to the database, and add metadata tags to the enterprise name data.

[0033] S105: Use a large language model to semantically parse the user query and generate a keyword dictionary Dict and a relationship importance list.

[0034] Specifically, call the large language model API, require the large language model to identify the enterprise names, address names, and industry names included in the user query, and create a dictionary Dict containing fields for enterprise names, address names, and industry names, and return the extraction results as the keyword dictionary. Obtain all relationship types in the knowledge graph, including "located in, belongs to, invests in, finances, supplier, customer", to get the relationship list List, use the large language model to analyze the correlation between each relationship type and the user query, sort the relationship list according to the degree of correlation, and return the sorting result as the relationship importance list.

[0035] S106: Use the keyword dictionary Dict to retrieve the vector database and obtain the matching entity name N in the knowledge graph. i .

[0036] Specifically, extract the enterprise name, address, and industry name fields in the keyword dictionary, and use an open-source vectorization model to perform vectorization encoding on each field name. Then, determine whether the enterprise name field is empty. If it is not empty, first calculate the address vector V di and the industry name vector V ci similarity with the metadata label, where if the similarities are not all 0, sort the metadata labels according to the similarities, and only retain the enterprise name data with the top 25% similarities as the search space. Otherwise, retain all enterprise name data as the search space. Calculate the similarity between the enterprise name vector E fi and other vectors in the search space, and return the enterprise name N with the highest similarity. fi . If the enterprise name field is empty and the address field is not empty, use all address name data as the search space, calculate the address name vector V di similarity with other vectors in the search space, and return the address name N with the highest similarity. di ; if the enterprise name field is empty and the industry name field is not empty, use all industry name data as the search space, calculate the industry name vector E ci similarity with other vectors in the search space, and return the industry name N with the highest similarity. ci .

[0037] S107: Use the entity name N i and the relationship importance list List to retrieve the knowledge graph and obtain query-related data and enterprise portrait documents.

[0038] Specifically, create a relevant information document, use the name data in S106 to retrieve the knowledge graph, and obtain the entity E corresponding to each name i , and add all the attribute information of E i to the relevant information document. If E i is an enterprise entity, obtain its corresponding enterprise portrait document. Then, use the relationship importance list in S105 to sequentially traverse the relationship elements R, and retrieve all entities E i that have a relationship R with the entity E ri , and add all the attribute information of E ri to the relevant information document until the relevant information document reaches the maximum word limit.

[0039] S108: If a large language model is used to generate user query answers based on relevant information data and provide an enterprise profile.

[0040] Specifically, by calling the large language model API to answer user queries, taking the relevant information documents obtained in S107 as "background knowledge", through prompt engineering, requiring the large language model to answer user queries strictly based on the background knowledge to avoid hallucination problems of the large language model. Then, conduct risk control review on the output result of the large language model to avoid prohibited information, and return the output that passes the risk control review as the final answer to the user. If an enterprise profile document is obtained in S107, then parse the enterprise profile document, use the SSG tool to generate a visual front-end page based on the document content, and return the website address of the front-end page.

[0041] To facilitate the understanding of the present invention, Figure 2 shows the enterprise profile generation flowchart based on the knowledge graph and enterprise data, specifically including:

[0042] S201: For the basic industrial and commercial information dimension, add all attribute information of enterprise entity E in the knowledge graph, including information such as the legal representative of the enterprise, registration status, enterprise address, enterprise scale, registered capital, paid-in capital, business scope, enterprise profile, operating income, etc. fi

[0043] S202: For the geographical dimension, retrieve the address entity E in the knowledge graph that has a "located in" relationship with entity E. fi di Obtain the statistical attribute information of entity E, including the quantile values of the operating income, registered capital, and paid-in capital of the enterprise group in this region. Compare the relevant information of entity E di horizontally with the above statistical data to quantitatively evaluate and display the development level of the target enterprise in the economic group of its region. fi

[0044] S203: For the industry dimension, retrieve the industry entity E in the knowledge graph that has a "belongs to" relationship with entity E. fi ci Obtain the statistical attribute information of entity E, including the quantile data of indicators such as the overall operating income, registered capital, and paid-in capital of the industry. Through horizontal comparison with the relevant information of entity E ci fi form a comprehensive analysis of the industry status of the enterprise, and intuitively and quantitatively reflect the competitiveness and development level of the enterprise in its industry.

[0045] S204: For the supply chain dimension, retrieve the entity E that has a "supplier" or "customer" relationship with entity E. fi ti, extract relevant information such as trade amount and reporting date, and clarify the role and status of the target enterprise in the upstream and downstream of the supply chain based on the transaction relationship. Further retrieve and entity E ti Enterprise entities E with relationships fj , statistically analyze the composition of the target enterprise's main suppliers, the distribution of supplier industries, and the allocation of procurement amounts, as well as the composition of main customers, the distribution of customer industries, and the distribution characteristics of sales amounts, and form an in-depth analysis of the supply chain dimension from the perspective of trade transactions.

[0046] S205: Investment and financing dimension, retrieve entity E related to entity E fi Enterprise entities E with "investment" or "financing" relationships ni , extract information such as transaction amount, disclosure date, financing round, and shareholding ratio, and outline the operation trajectory of the enterprise in the capital market by analyzing the enterprise's own investment behavior and financing process, clarify the enterprise's growth stage and capital preference characteristics, and form a multi-dimensional portrait of the enterprise in the capital field.

[0047] S205: Risk dimension, retrieve relevant information such as judgment documents, judicial cases, and administrative penalty records related to the target enterprise in enterprise data D, and present the legal risks and compliance issues faced or potential by the enterprise in a systematic and comprehensive manner. Summarize, organize, and structure the above risk dimension information for risk analysis and warning.

[0048] The present invention proposes an enterprise analysis system based on knowledge graph and retrieval enhancement technology. Through the knowledge graph, the multi-source heterogeneous data of the enterprise is orderly integrated, combined with the generation ability of the large language model, to achieve accurate question answering and multi-dimensional display of enterprise information, with high-efficiency collaboration and flexible adaptation capabilities, and can effectively support various application scenarios such as enterprise information question answering, investment decision-making, and market insight.

Claims

1. An enterprise analysis system based on knowledge graph and retrieval augmented generation technology, characterized in that, The system involves large language models, knowledge graphs, and vector databases. The system includes the following steps: Step 1, use OCR recognition and special format matching to process the original data and extract editable text data. Perform format conversion, missing value processing, noise filtering, and standardization on the text data to obtain the processed data D; Step 2: Using enterprise data D, create knowledge graph entity E based on the basic industrial and commercial information of each enterprise and the trade and investment information between enterprises, i and add knowledge graph relationship R. i ; Step 3: Generate an enterprise portrait document using the enterprise data D, the knowledge graph entity E i and the relevant information of the knowledge graph relationship R i ; Step 4: Use the pre-trained vector model to encode the name information of the knowledge graph entity E i to obtain the name vector V i , and construct a vector database; Step 5, use the large language model to semantically parse the user query, generating a keyword dictionary Dict and a relationship importance list List; Step 6, use the keyword dictionary Dict to retrieve the vector database and obtain the matching entity name N in the knowledge graph i ; Step 7, utilize the knowledge graph entity name N i and the relationship importance list List to retrieve the knowledge graph and obtain query-related data and enterprise portrait documents; Step 8, use the large language model to generate a user query answer based on the relevant information data and provide an enterprise portrait.

2. The enterprise analysis system based on the knowledge graph and retrieval-augmented generation technology according to claim 1, wherein: In Step 1, the use of OCR recognition and special format matching to process the original data and extract editable text data. Perform format conversion, missing value processing, noise filtering, and standardization on the text data to obtain the processed data D, specifically: Obtain enterprise data from multiple public channels such as industrial and commercial registration, news sentiment, financial statements, and bidding information. First, perform format discrimination on the obtained data sources of various types. For directly parsable documents (such as editable PDFs, Word, Excel, etc.), use document parsing tools to extract the text and table content in the body and convert it into editable text. For documents containing pictures or unable to directly extract text information (such as scanned PDFs, image format reports, etc.), first perform preprocessing operations such as denoising, binarization, and skew correction on the images, and then use the OCR algorithm to recognize the text and table structure on the page and convert the recognition results into editable text. For the above-mentioned extracted editable text, further use custom special format matching rules to locate and parse the key fields in the extracted text, identify core elements such as enterprise names, registered capital, trade transactions, and judgment documents, uniformly export the above text or structured data as a processable CSV format, and on this basis, fill in missing values, filter noise and duplicate data, and eliminate outliers. Perform standardization processing on the key fields, including unified formatting of enterprise names, currency unit conversion, format unification of time and date fields, and range verification of numerical fields.

3. The enterprise analysis system based on knowledge graph and retrieval augmented generation technology according to claim 1, characterized in that: In step 2, using the data D of each enterprise, create the knowledge graph entity E based on the basic industrial and commercial information of each enterprise and the trade and investment information between enterprises i , and add the knowledge graph relationship R i , specifically: in the knowledge graph, create enterprise entities E based on each enterprise in the data fi , and the attributes of E fi include information in the following dimensions: name, legal representative, registration status, enterprise address, enterprise scale, registered capital, paid-in capital, business scope, enterprise profile, operating income. Create address entities E for each province and city di , and according to the geographical location of E fi , add a "located in" relationship between E fi and its corresponding E di . The attributes of E di include information in the following dimensions: name. For all E di that have a "located in" relationship with E fi , calculate the 5%, 25%, 50%, and 75% quantile values of operating income, registered capital, and paid-in capital respectively. Create industry category entities E according to the national standard industry categories in the "National Economic Industry Classification" ci ; according to the industry category of E fi , add a "belongs to" relationship between E Fi and its corresponding E ci . The attributes of E ci include information in the following dimensions: name. For all E Ci that have a "belongs to" relationship with E fi , calculate the 5%, 25%, 50%, and 75% quantile values of operating income, registered capital, and paid-in capital respectively. Create trade event entities E based on each piece of trade information in the data ti , and E ti includes information in the following dimensions: name, trade amount, report date. For the E fi that is the supplier in each trade event, add a "supplier" relationship between it and the corresponding E ti ; for the E fi that is the customer, add a "customer" relationship between it and the corresponding E ti . Create investment event entities E based on each piece of investment information in the data ni , and E ni includes information in the following dimensions: name, transaction amount, disclosure date, financing round, shareholding ratio. For the E fi that is the investor in each investment event, add an "investment" relationship between it and the corresponding E ni ; For E as the investee fi , add the "financing" relationship between it and the corresponding E ni .

4. An enterprise analysis system based on knowledge graph and retrieval augmented generation technology according to claim 1, characterized in that: In step 3, the enterprise data D and knowledge graph entity E are used i Relationship with knowledge graph R i Generate enterprise portrait documents based on the relevant information of the enterprise, specifically: create an enterprise portrait document for each enterprise. Based on the knowledge graph created in step 2, obtain the corresponding entity E fi All dimension attribute information is added to the enterprise portrait document as basic business information. fi Entity E with a "is-at-" relationship di , get E di The statistical attribute information of the enterprise is added to the enterprise portrait document as regional dimension information. fi Entity E with a "belongs to" relationship ci , get E ci The statistical attribute information of the enterprise is added to the enterprise portrait document as industry dimension information. fi Associated entity E ti , obtain the trade amount and report date dimension information, analyze the company's supplier and customer distribution and trade amount and other indicators, and add them to the company portrait document as supply chain dimension information. fi Associated entity E ni , obtain information on transaction amounts, disclosure dates, financing rounds, and shareholding ratios, analyze the company's investment behavior and financing history, and add this information to the company profile document as investment and financing information. Based on the company data obtained in step 1, retrieve relevant information such as judgment documents, judicial cases, and administrative penalties, and add this information to the company profile document as risk information.

5. An enterprise analysis system based on knowledge graph and retrieval augmented generation technology according to claim 1, characterized in that: In step 4, using the pre-trained vector model, encode the name information of the knowledge graph entity E i to obtain the name vector V i , and construct a vector database, specifically: use an open-source vectorization model to vectorize and encode the names of each entity in the knowledge graph. Create corresponding metadata tags for each enterprise entity E fi , including the vector encoding of the name of the address where E fi is located E fi and the vector encoding of the name of the industry to which it belongs Use a database tool to create a database, add each name vector encoding and its corresponding original text to the database, and add metadata tags to the enterprise name data.

6. The enterprise analysis system based on knowledge graph and retrieval-augmented generation technology according to claim 1, characterized in that: In Step 5, the use of the large language model to semantically parse the user query and generate a keyword dictionary and a relationship importance list, specifically: Call the large language model API to analyze the user input query. Use the large language model to extract the enterprise name, address name, and industry name included in the user query, and create a dictionary Dict containing the fields of enterprise name, address name, and industry name. Add the extraction results to the corresponding fields of the dictionary and return it as the keyword dictionary. Obtain all relationship types in the knowledge graph, including "located in, belongs to, invests, finances, supplier, customer", to get the relationship list List. Use the large language model to analyze the correlation between each relationship type and the user query, sort the relationship list according to the degree of correlation, and return the sorting result as the relationship importance list.

7. The enterprise analysis system based on the knowledge graph and retrieval-augmented generation technology according to claim 1, characterized in that: In step 6, the keyword dictionary Dict is used to retrieve the vector database to obtain the matching entity name N in the knowledge graph i , specifically: extract the enterprise name, address, and industry name fields in the keyword dictionary, and use an open-source vectorization model to perform vectorized encoding on each field name. Then, it is judged whether the enterprise name field is empty. If it is not empty, first calculate the address vector V di and the industry name vector V ci with the similarity of the metadata tags, where if the similarities are not all 0, sort the metadata tags according to the similarities, and only retain the enterprise name data with the top 25% similarities as the search space. Otherwise, retain all enterprise name data as the search space. Calculate the similarity between the enterprise name vector E fi and other vectors in the search space, and return the enterprise name with the highest similarity. If the enterprise name field is empty and the address field is not empty, use all address name data as the search space, calculate the similarity between the address name vector V di and other vectors in the search space, and return the address name with the highest similarity; if the enterprise name field is empty and the industry name field is not empty, use all industry name data as the search space, calculate the similarity between the industry name vector E ci and other vectors in the search space, and return the industry name with the highest similarity.

8. The enterprise analysis system based on knowledge graph and retrieval-augmented generation technology according to claim 1, characterized in that: In step 7, using the knowledge graph entity name N i and the relationship importance list List, retrieve the knowledge graph to obtain query-related data and enterprise portrait documents. Specifically: create a relevant information document, and set the maximum word limit of the relevant information document to the maximum number of words in the large model input window. Use the name data obtained in step 5 to retrieve in the knowledge graph to obtain the entity E corresponding to each name i , and add all the attribute information of E i to the relevant information document. If the entity E i is an enterprise entity, obtain the enterprise portrait document corresponding to this entity. Then use the relationship importance list obtained in step 5 to traverse the relationship elements R in the list order, and retrieve all entities E i that have a relationship R with the entity E ri , and add all the attribute information of E ri to the relevant information document until the relevant information document reaches the maximum word limit.

9. The enterprise analysis system based on the knowledge graph and retrieval augmented generation technology according to claim 1, characterized in that: In step 8, the large language model is utilized to generate user query answers based on relevant information data and provide an enterprise portrait. Specifically: Call the large language model API to answer user queries. Use the user's query as the "question" and the relevant information documents obtained in step 7 as the "background knowledge", and input them into the large language model together. Through prompt engineering, require the large language model to answer questions strictly based on the "background knowledge" to avoid hallucination problems of the large language model. Then conduct a risk control review on the output result of the large language model to avoid the occurrence of prohibited information, and return the output that passes the risk control review as the final answer to the user. If an enterprise portrait document is obtained in step 7, then parse the enterprise portrait document and extract the document content. Use the SSG tool (Static Site Generator) to generate a visual front-end page based on the document content and return the URL of the front-end page.

Citation Information

Cited By

  • Enterprise digital transformation scheme generation method and device

    CN120822496A

  • Automobile channel intelligent question and answer method and system based on large language model

    CN121524196A

  • Matching management method and management system between enterprises based on knowledge graph

    CN121563273A

  • Multi-modal engineering design intelligent generation method and system based on knowledge graph and RAG

    CN121901462A

  • Enterprise consultation data retrieval method and system based on graph enhanced retrieval generation

    CN122019759A