An investment BP risk control decision method and system based on a knowledge graph and multi-modal analysis
Patent Information
- Application Number
- CN202610998189.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-28
AI Technical Summary
当前的技术方案要么侧重于某一环节的单一功能实现,要么各环节之间缺乏有机衔接,难以形成从商业计划书解析到风险评估再到决策建议的完整技术闭环
1、本发明通过构建图数据库、向量数据库、结构化数据库三库联动的投资决策知识图谱作为知识底座,结合商业计划书多模态解析技术与多维度事实核查机制,对市场规模、技术可行性、竞品陈述、财务预测等核心陈述开展外部数据交叉验证与文本、数值、多模态自洽性校验,实现商业计划书关键信息的自动化真实性核查与异常点量化标注,突破传统规则评分与关键词匹配方案语义理解能力不足、无法验证数据真实性的技术局限,同时大幅提升信息核查的全面性与准确性,显著降低人工尽调的信息检索与比对成本。
Smart Images

Figure CN122656740A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an investment business plan risk control decision-making method and system based on knowledge graphs and multimodal analysis. Background Technology
[0002] Equity investment, as a crucial financing method in the current process of technological innovation and enterprise development, directly impacts the risk control capabilities and return levels of investment institutions through the professionalism and scientific rigor of its decision-making process. In the actual operations of investment institutions, in-depth analysis and risk assessment of business plans submitted by startups are fundamental steps in the entire investment decision-making process. A typical business plan usually covers core elements such as market analysis, product and technology solutions, business model, financial forecasts, team introduction, and financing plans. This information constitutes the original basis for investment institutions to judge the value of a project. Investment institutions have long faced numerous pain points when analyzing business plans, including difficulty in verifying the authenticity of information, insufficient risk assessment dimensions, inadequate competitor analysis, low efficiency and inconsistent standards, and difficulty in reusing historical data. These problems severely restrict the improvement of the quality and efficiency of investment decisions.
[0003] From a technological evolution perspective, investment institutions have explored various technical solutions in the field of business plan analysis. In the early stages, rule-based automated scoring systems were the mainstream approach. These systems automatically scored business plans using pre-set rules such as bonus points for team background and deductions for market size. However, this rule-based method has significant limitations. Its highly rigid rule system makes it difficult to handle open-ended descriptions in business plans, such as assessing the innovativeness of technological approaches. Furthermore, this type of solution lacks true semantic understanding capabilities, failing to effectively identify logical contradictions and exaggerated statements in business plans. In addition, it cannot utilize external data sources (such as competitor financing history and industry benchmark data) for cross-validation, significantly reducing the accuracy and credibility of the evaluation results. With the development of natural language processing technology, keyword matching-based solutions have gradually become a research hotspot. These solutions extract and match keywords in business plans to identify whether they contain key elements such as patents and market size. However, this approach also has fundamental flaws. It can only extract features at a superficial level and cannot deeply understand the underlying logical structure and argumentation chain of a business plan. More importantly, it cannot verify the authenticity of the data stated in the business plan, such as whether the "market size of 100 billion yuan" claimed in the business plan is consistent with the actual market situation. At the same time, this type of approach lacks systematic risk assessment capabilities, cannot output a structured risk list, and is difficult to meet the actual needs of investment institutions for in-depth risk analysis.
[0004] When the aforementioned technical solutions fail to meet actual business needs, manual due diligence analysis becomes the mainstream choice for investment institutions. Teams composed of investment managers and industry experts manually read business plans and combine this with external research results to conduct a comprehensive risk assessment. While this approach has certain advantages in terms of analytical depth, its drawbacks are equally significant. First, manual analysis is extremely inefficient; a thorough analysis of a single business plan typically requires two to three days of work for an investment manager, sharply contradicting the investment institution's pursuit of efficient decision-making. Second, manual analysis is highly subjective; the judgment standards of different investment managers and industry experts vary significantly, leading to a lack of consistency and comparability in the assessment results. Third, manual analysis struggles to systematically accumulate analytical methods and historical case studies; knowledge cannot be effectively reused, and each analysis must begin from scratch, greatly increasing labor costs. Finally, manual analysis is highly dependent on the industry experience of investment managers; the long training period for new talent and the difficulty in building a talent pipeline also place a heavy burden on investment institutions in terms of human resources.
[0005] In recent years, academia and industry have begun exploring the application of artificial intelligence technology in investment business plan (BP) analysis. Patent CN121860779A discloses a method for recommending investment and financing projects based on dynamic knowledge graphs and large language models. This method generates a project recommendation list by aligning investment institution profiles with the semantics of a differentiable knowledge graph. While this technological advancement introduces knowledge graph technology into the investment and financing field, its solution focuses on recommending and matching investment and financing projects, rather than analyzing and diagnosing the content of the business plan itself. Furthermore, it lacks the ability to verify and validate the facts stated in the business plan; moreover, it does not incorporate competitor knowledge graphs for competitive landscape analysis, thus exhibiting significant limitations in its application scenarios.
[0006] With the continuous advancement of artificial intelligence technology and the increasing demands of investment institutions for scientific decision-making, the existing technological solutions are gradually revealing deep-seated inherent contradictions when dealing with more complex investment decision-making scenarios. Specifically, while addressing the issue of information processing efficiency, the core challenge of existing technologies lies in how to construct a comprehensive technological system capable of conducting all-round fact-checking, achieving multi-dimensional risk quantification, and outputting actionable investment recommendations. Current technological solutions either focus on the implementation of a single function in a particular stage or lack organic connections between different stages, making it difficult to form a complete technological closed loop from business plan analysis to risk assessment to decision-making recommendations. As market competition intensifies and investment risks become increasingly complex, investment institutions urgently need an intelligent decision support system that can systematically integrate multi-source knowledge, perform deep semantic understanding, achieve quantitative risk assessment, and generate actionable recommendations. Existing technological solutions have significant capability gaps when facing this comprehensive demand. Therefore, how to construct an investment business plan risk control decision-making method and system based on knowledge graphs and multimodal analysis, so as to balance the accuracy of fact-checking, the comprehensiveness of risk assessment, and the operability of decision recommendations, has become a key challenge and a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0007] To address existing technical problems, this invention provides an investment business plan (BP) risk control decision-making method and system based on knowledge graphs and multimodal analysis. By constructing an investment decision-making knowledge graph that links three databases, and combining multimodal analysis of business plans, multi-dimensional fact verification, quantitative risk profiling, and competitor benchmarking analysis, a complete technical closed loop from content analysis to decision output is formed. This enables accurate verification of business plan information, standardized quantitative assessment of investment risks, and actionable decision support, comprehensively improving the scientific nature, accuracy, and operational efficiency of investment decisions.
[0008] Firstly, this invention provides an investment business plan (BP) risk control decision-making method based on knowledge graphs and multimodal analysis, including... S100, multi-source knowledge graph construction, is used to integrate industry data, enterprise data, technical data, market data, and competitor data to build an investment decision knowledge graph; S200 and BP multimodal parsing and structuring are used to parse unstructured business plan documents into a structured list of elements. S300, Fact Check and Reasonableness Verification: Based on the investment decision knowledge graph and external data sources, the key data and arguments presented in the business plan are fact-checked and reasonableness verified, and anomalies and exaggerated statements are marked. The S400 multi-dimensional risk profile construction, based on the fact-checking results of S300 and investment decision knowledge graph data, quantifies the investment risk of a project from five dimensions: market risk, technology risk, team risk, financial risk, and compliance risk, and generates a structured risk list. S500, competitor benchmarking and differentiation analysis, based on competitor knowledge graphs, to verify the competitive landscape stated in the business plan and to evaluate the project's differentiated competitive advantages; S600: Investment Decision Recommendation Generation and Output. Based on the analysis results of S300 to S500, structured investment decision recommendations and risk control plans are generated.
[0009] Preferably, the construction of the multi-source knowledge graph includes: Multi-source data preprocessing is used to perform format parsing, data cleaning, entity linking and conflict resolution on raw data, and output standardized data records; Entity recognition and classification is used to identify entities in text using large language models and rule engines, and classify entities into industry nodes, enterprise nodes, technology nodes, product nodes, market nodes, and event nodes. Relation extraction is used to extract relationships between entities from data using large language models and rule engines, generating attribution relationships, competition relationships, technology usage relationships, target market relationships, financing relationships, and referencing relationships; The graph assembly and storage is used to write entities and relationships into a graph database to support graph reasoning, to write vectorized results of text descriptions into a vector database to support semantic retrieval, and to write numerical indicators into a structured database to support precise queries, forming an investment decision knowledge graph that links the three databases.
[0010] Preferably, step S100 further includes competitor knowledge graph construction, used to construct a subset of the investment decision knowledge graph. The competitor knowledge graph includes a competitor identification module, a competitor profiling module, a competitor comparison matrix module, and a competitive relationship reasoning module. Specifically, the competitor identification module identifies a list of competitors of the target company based on industry classification and product function similarity; the competitor profiling module extracts the competitor's technical roadmap, product functions, market share, financing history, and team background; the competitor comparison matrix module constructs a company-competitor comparison matrix covering technology, product, market, team, and financing dimensions; and the competitive relationship reasoning module infers the intensity of competition based on product function overlap, target market overlap, and technical roadmap similarity.
[0011] Preferably, the BP multimodal analysis and structuring includes: The document parsing process supports several business plan document formats, extracts the main text, tables and charts, identifies the table of contents of the business plan, and locates the core sections such as market analysis, product plan, business model, financial forecast, and team introduction. The multimodal content extraction step is used to extract text content, extract structured data from tables, and extract data from charts through OCR and image recognition; The BP element structuring step is used to perform semantic understanding of the extracted business plan content using a large language model, and to convert the business plan into an object containing structured data elements such as project name, company name, and each chapter. The BP element classification step is used to classify the elements of a business plan into five categories: market, technology, business, financial, and team. These correspond to subsequent fact-checking, technical feasibility verification, hypothesis stress testing, rationality verification, and capability profile assessment and analysis strategies, respectively.
[0012] Preferably, the fact-checking and reasonableness verification include: The market size verification step is used to extract market size statements from business plans, retrieve market size data for the corresponding industry from the knowledge graph, calculate the deviation rate between the business plan statements and the knowledge graph data, and mark the market size as exaggerated when the deviation rate exceeds a preset threshold. The technical feasibility verification step is used to extract the technical route statement from the business plan, retrieve the maturity, application cases and technical bottlenecks of the technology in the knowledge graph, and use the large language model to evaluate the feasibility and implementation difficulty of the technical route. When the technical maturity is insufficient or there are major technical bottlenecks, it is marked as technical feasibility questionable. The competitor statement verification step is used to extract competitor statements from the business plan, retrieve the feature list of similar competitors in the competitor knowledge graph, verify whether the competitor statements in the business plan are accurate, and mark the competitor analysis as inaccurate when the competitor knowledge graph shows that other competitors support the same function. The financial forecast reasonableness verification step is used to extract financial forecasts from the business plan, retrieve revenue growth benchmark data of similar companies in the knowledge graph, calculate the deviation between the financial forecasts in the business plan and the industry benchmark, and mark the financial forecast as too optimistic when the growth rate is significantly higher than the industry benchmark.
[0013] Preferably, the construction of the multi-dimensional risk profile includes: Market risk scores are calculated based on penalties for exaggerating market size, intense competition, and low barriers to entry. A technology risk score is calculated based on technology maturity score, technology barrier score, and intellectual property score. The team risk score is calculated based on background matching score, complementarity score, and execution score; A financial risk score is calculated based on optimistic financial forecast penalties, cash flow pressure penalties, and valuation bubble penalties. A compliance risk score is calculated based on legal compliance scores, data privacy scores, and industry regulatory scores.
[0014] Preferably, the competitor benchmarking and differentiation analysis includes: The competitor list of a target company is identified based on four dimensions: industry classification matching, product function similarity calculation, target market overlap calculation, and technology route similarity calculation. The comprehensive score calculation formula is: the comprehensive score equals the industry classification matching degree multiplied by 0.3 plus the function similarity multiplied by 0.4 plus the market overlap multiplied by 0.2 plus the technology similarity multiplied by 0.1. When the comprehensive score is greater than 0.6, it is judged as a competitor. Construct a target company-competitor benchmarking matrix covering technology, product, market, team, and financing dimensions; The competitive advantages of a target company are evaluated based on four dimensions: functional differentiation, technological differentiation, market differentiation, and team differentiation. Generate a structured competitor benchmarking report that includes competitor identification results, competitor benchmarking matrix, differentiated advantage assessment, and competitive strategy recommendations.
[0015] Preferably, the generation and output of the investment decision recommendations includes: Based on the risk profile vector and competitor benchmarking results, a weighted scoring model is used to calculate the investment recommendation score. The calculation formula is as follows: the investment recommendation score is equal to the sum of the products of the five risk dimension scores and their respective weights, plus the differentiation advantage bonus, minus the fact-checking deduction. The differentiation advantage bonus is equal to the differentiation advantage score multiplied by 0.2, and the fact-checking deduction is equal to the number of anomalies multiplied by 5. Generate a structured investment decision proposal containing six parts: project overview, fact-checking results, risk profile analysis, competitor benchmarking analysis, investment decision recommendations, and risk mitigation solutions. Also, generate targeted risk mitigation solutions for the identified key risk points.
[0016] Secondly, the present invention also provides an investment business plan (BP) risk control decision-making system based on knowledge graphs and multimodal analysis, which is applied to the investment business plan (BP) risk control decision-making method based on knowledge graphs and multimodal analysis as described above. The system includes a multi-source knowledge graph construction module, a BP multimodal parsing and structuring module, a fact-checking and rationality verification module, a multi-dimensional risk profile construction module, a competitor benchmarking and differentiation analysis module, an investment decision suggestion generation and output module, and a knowledge graph dynamic update module. The multi-source knowledge graph construction module is used to integrate multi-source heterogeneous data such as industry data, enterprise data, technical data, market data, and competitor data to construct an investment decision knowledge graph. The BP multimodal parsing and structuring module is used to parse unstructured business plan documents into a structured list of elements. The fact-checking and rationality verification module is used to perform fact-checking and rationality verification on the key data and arguments presented in the business plan based on the investment decision knowledge graph and external data sources. The multi-dimensional risk profile building module is used to quantitatively assess the investment risks of a project from five dimensions: market risk, technology risk, team risk, financial risk, and compliance risk. The competitor benchmarking and differentiation analysis module is used to verify the competitive landscape and evaluate differentiated competitive advantages based on the competitor knowledge graph. The investment decision recommendation generation and output module is used to generate structured investment decision recommendations (49) and risk control schemes; The knowledge graph dynamic update module is used to incrementally update the knowledge graph when the triggering conditions occur, and retain historical version snapshots for backtracking queries.
[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. This invention constructs an investment decision knowledge graph that integrates three databases—graph database, vector database, and structured database—as its knowledge foundation. Combined with multimodal parsing technology and a multi-dimensional fact-checking mechanism for business plans, it conducts external data cross-validation and textual, numerical, and multimodal self-consistency checks on core statements such as market size, technical feasibility, competitor statements, and financial forecasts. This enables automated verification of the authenticity of key information in business plans and quantitative labeling of anomalies. It overcomes the limitations of traditional rule-based scoring and keyword matching schemes, which lack semantic understanding and cannot verify data authenticity. Simultaneously, it significantly improves the comprehensiveness and accuracy of information verification and substantially reduces the cost of information retrieval and comparison in manual due diligence.
[0018] 2. By constructing a quantitative risk scoring model covering five dimensions—market, technology, team, finance, and compliance—and combining it with a specialized competitor knowledge graph, we can achieve automatic competitor identification, multi-dimensional benchmarking matrix construction, and differentiated advantage assessment. This generates standardized risk profile vectors and structured risk lists, enabling multi-dimensional quantifiable assessment of investment risks and systematic analysis of the competitive landscape. This addresses the issues of strong subjectivity, inconsistent standards, and incomplete competitor analysis in existing technology risk assessments, ensuring the comparability of assessment results and the depth of risk judgment across different projects, and providing quantifiable risk basis for investment decisions.
[0019] 3. By establishing a closed-loop technology chain that encompasses multimodal structured analysis of business plans, fact-checking and verification, risk profiling, competitor benchmarking analysis, and the generation of tiered investment decision recommendations, and by outputting corresponding investment conclusions based on a weighted scoring model and providing corresponding multidimensional risk mitigation solutions, a complete technology chain of "content analysis - verification and evaluation - decision output" is formed. This overcomes the shortcomings of existing technical solutions, such as single functionality, fragmented processes, and difficulty in outputting actionable decision recommendations. It not only enables the systematic accumulation and reuse of analytical methods and historical cases, but also outputs standardized and implementable investment decision support results, comprehensively improving the scientific nature of investment decisions and the overall operational efficiency. Attached Figure Description
[0020] Figure 1 This is a flowchart of an investment business plan (BP) risk control decision-making method based on knowledge graphs and multimodal analysis.
[0021] Figure 2 This is a schematic diagram of the structure of an investment business plan (BP) risk control decision-making system based on knowledge graphs and multimodal analysis. Detailed Implementation
[0022] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention. It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0023] Please see Figures 1-2 This embodiment proposes an investment business plan risk control decision-making method based on knowledge graphs and multimodal analysis. The following will describe in detail the construction of the investment decision knowledge graph, the multimodal parsing of the business plan, the fact-checking and rationality verification, the construction of multi-dimensional risk profiles, the competitor benchmarking and differentiation analysis, and the implementation details of the investment decision suggestion generation method, so as to ensure that those skilled in the art can implement the technical solution of this invention based on the following technical description.
[0024] In constructing the investment decision knowledge graph, the system first receives raw data from multiple external data sources, including industry research reports, corporate registration data, patent databases, financing databases, academic paper databases, news report databases, historical business plan databases, and investment decision databases. Industry research reports are typically in PDF or Word document format. A format parsing module extracts industry node information, including industry name, classification level, market size, and growth rate data. Market size data is measured in RMB 100 million, and growth rate data is stored as a percentage. Corporate registration data originates from structured data exported from the Enterprise Credit Information Publicity System, containing fields such as company name, unified social credit code, establishment date, registered capital, business scope, and equity structure. Entity links are established using company name and unified social credit code, merging different data records for the same company. The patent database uses a patent search interface provided by the State Intellectual Property Office to extract technical information such as patent number, patent name, applicant, application date, publication date, patent type, abstract, and claims. The system maps this information to technical nodes and patent relationships. The financing database integrates data from multiple investment and financing information platforms, including fields such as financing event ID, financing round, financing amount, investor list, invested companies, and financing date. The system extracts financing event nodes and investor relationships from these data. The academic paper database uses API interfaces from academic databases such as Web of Science and arXiv to obtain information such as paper titles, authors, institutions, journals, publication years, abstracts, and keywords, used to construct technical nodes and academic relationships. The news report database crawls news content from mainstream financial media and industry-specific media, extracting fields such as event titles, event content, involved companies, and publication time to construct event nodes and the relationship between companies and events. The historical business plan database stores anonymized historical business plan documents. The system extracts elements such as project name, industry classification, business model, target market, financial forecasts, and team background to form business plan element nodes and the relationship between projects and elements. The investment decision database records case data of historical investment decisions, including project ID, due diligence results, investment review committee decisions, and post-investment performance, used to construct decision nodes and the relationship between projects and decisions.
[0025] Next, the raw data undergoes multi-source data preprocessing, which includes format parsing, data cleaning, entity linking, and conflict resolution. In the format parsing stage, the system calls the appropriate parser based on the data source's format type: PDF format uses the PDFMiner library for text extraction, Word format uses the python-docx library to parse the document structure, Excel format uses the pandas library to read table data, JSON format uses the standard json library for parsing, and CSV format data exported from structured databases is read using the pandas library. The data cleaning stage includes removing duplicate records, filling in missing values, correcting format errors, and standardizing field names. The system uses a fuzzy matching algorithm to identify potentially duplicate entity records. When two records have an entity name similarity exceeding 0.85 and other key fields overlap, they are considered duplicate records and merged. The entity linking stage associates the same entity from different data sources. The system calculates an entity link score based on text similarity, address similarity, and other attribute similarity. When the link score exceeds 0.8, an entity equivalence relation is established. The conflict resolution process handles conflicting information from different data sources. The system adopts a confidence-weighted strategy, calculating weights based on the authority and timeliness of each data source. When multiple values of the same attribute conflict, the value from the data source with the highest weight is selected as the primary value, and the other values are recorded as reference values in the conflict resolution log.
[0026] After multi-source data preprocessing, a large language model and rule engine are used to identify entities in the text, classifying them into six types: industry nodes, enterprise nodes, technology nodes, product nodes, market nodes, and event nodes. The large language model employs a vertical domain model fine-tuned with investment data. The input is the text paragraph to be identified, and the output is a list of entities in JSON format. Each entity includes fields for entity name, entity type, confidence level, and evidence source. The rule engine, supplementing the large language model, identifies entities conforming to fixed patterns, such as suffixes like "Limited Company" or "Joint-Stock Company" in company names, standard formats for patent numbers, and standard codes for industry classifications. Industry nodes store information such as industry name, classification level (according to the first, second, and third levels of the National Economic Industry Classification Standard GB / T 4754-2017), industry size, growth rate, and life cycle stage. Enterprise nodes store basic information such as company name, unified social credit code, establishment date, registered capital, legal representative, shareholder structure, business scope, and number of insured persons. The technology node stores information such as technology name, technology category, technology maturity level (classified into 1-9 levels according to TRL technology maturity levels), technology principle, applicable scenarios, and technical bottlenecks. The product node stores information such as product name, product category, core functions, technical architecture, pricing strategy, and target users. The market node stores information such as market name, market size, market growth rate, competitive landscape, and market trends. The event node stores information such as event name, event type (financing, product launch, industry policy, major events, etc.), involved companies, time of occurrence, and scope of impact.
[0027] This project utilizes a large language model and a rule engine to extract relationships between entities from data, generating six types of edges: Attribution relationships (connecting companies with industries, technologies, and products, indicating an entity belongs to a certain category); Competition relationships (connecting companies with competitors, indicating two companies compete in the same market); Technology usage relationships (connecting companies with technologies, indicating a company uses a certain technology); Target market relationships (connecting companies with markets, indicating a company's products or services target a specific market); Financing relationships (connecting companies with investors, indicating a company receives funding from an investor); and Citation relationships (connecting companies with patents, papers, etc., indicating a company cites a technological achievement). The large language model employs a few-shot learning method for relation extraction, providing 5-10 example relations in the prompts. The model understands the task requirements and outputs a structured list of relations. The rule engine identifies relationships with clear patterns; for example, equity relationships in corporate business data are directly established through the association between shareholders and invested companies, and the applicant field in patent data is directly mapped to a technology attribution relationship.
[0028] Entities and relationships are written into a graph database, a vector database, and a structured database to form a three-database interconnected investment decision-making knowledge graph. The graph database uses Neo4j as its storage engine, storing node and edge attributes in node attribute tables and edge attribute tables respectively, supporting graph traversal and path analysis using the Cypher query language. The vector database uses Milvus as its storage engine, converting textual descriptions of entities and relationships into dense vectors using pre-trained language models (such as BAAI / bge-base-en-v1.5), supporting vector similarity retrieval for querying related entities based on semantic similarity. The structured database uses MySQL as its storage engine, storing structured attributes and numerical indicators of entities, supporting precise queries and aggregation analysis. The three databases are linked by entity IDs. Each entity node in the graph database has a corresponding vector ID and structured record ID, enabling cross-database joint retrieval. For example, when querying a company's competitive relationships, the system first obtains the company ID by precise matching based on the company name in the structured database, then queries the competitive relationship edges in the graph database based on the company ID to obtain a list of competitors, and finally performs semantic retrieval based on the competitor names in the vector database to verify the accuracy of the competitors.
[0029] As a subset of the investment decision knowledge graph, the competitor knowledge graph is constructed through dedicated sub-modules, including a competitor identification module, a competitor profiling module, a competitor comparison matrix module, and a competitive relationship reasoning module. The competitor identification module identifies a list of competitors for a target company based on industry classification and product function similarity. The system first retrieves enterprise nodes in the knowledge graph that share the same industry classification as the target company, then calculates the product function similarity between these enterprises and the target company. Functional similarity is obtained by comparing the vector similarity of product description texts; when the similarity exceeds 0.6, the enterprise is added to the candidate competitor list. The competitor profiling module extracts information such as the competitor's technology roadmap, product functions, market share, financing history, and team background. This information is retrieved from the investment decision knowledge graph, and supplementary retrieval from external data sources is triggered when information is missing. The competitor comparison matrix module constructs a company-competitor comparison matrix covering five dimensions: technology, product, market, team, and financing. Each dimension includes 3-5 specific indicators; for example, the technology dimension includes indicators such as the number of patents, core technology TRL level, and technical team size. The competitive relationship reasoning module infers competitive intensity based on three dimensions: product function overlap, target market overlap, and technical route similarity. The comprehensive scoring formula is: competitive intensity score equals product function overlap multiplied by 0.4 plus target market overlap multiplied by 0.4 plus technical route similarity multiplied by 0.2. When the score exceeds 0.7, it is determined to be a direct competitive relationship.
[0030] The knowledge graph dynamic update mechanism is used to incrementally update the knowledge graph when trigger conditions occur. These trigger conditions include new financing events, patent announcements, product launches, and industry report releases. New financing events are obtained by monitoring the RSS feeds and API interfaces of multiple investment and financing information platforms. When a new financing event is detected, the system automatically extracts the financing event information and creates a new financing event node in the graph database, while simultaneously updating the financing round and investor relationships of the invested company. Patent announcements are obtained by monitoring the patent announcement interface of the State Intellectual Property Office. When a new patent is granted or published, the system automatically extracts the patent information and creates or updates the technology node in the graph database. Product launches are obtained by monitoring the official news channels, social media accounts, and industry media reports of the target company. When a new product launch is detected, the system automatically extracts the product information and updates the product node. Industry reports are obtained by monitoring the release of reports by professional research institutions. When a new industry report is detected, the system automatically extracts industry data and updates the industry and market nodes. The version management function retains historical version snapshots. A new version snapshot is created with each incremental update, and a version timestamp is recorded. Backtracking queries based on timestamps are supported, such as querying industry market size data at a specific point in time. The consistency check function automatically detects isolated nodes (nodes without any edge connections) and contradictory relationships (such as two contradictory equity relationships of the same enterprise) in the graph after each update, and outputs the detection results to the check log for manual review by the administrator.
[0031] In BP multimodal parsing and structuring, business plan documents are parsed in three mainstream formats: PDF, Word, and PPT. PDF business plan documents are parsed using the PyMuPDF library, extracting text content, table structure, tables, and charts. Text content is stored row-by-row, tables are stored as two-dimensional arrays, and charts are stored as images with captions and axis meanings. Word business plan documents are parsed using the python-docx library, extracting body text, paragraph formatting, tables, and embedded images. Chapter titles are determined by recognizing paragraph styles (Heading 1, Heading 2, etc.), and a table of contents structure is constructed accordingly. PPT business plan documents are parsed using the python-pptx library, extracting the title, body text, tables, and images for each slide. The slide order serves as a reference for page order. After content extraction, the document parsing sub-step identifies the business plan's table of contents, locating the five core chapters: Market Analysis, Product Proposal, Business Model, Financial Forecast, and Team Introduction, providing chapter location information for subsequent multimodal content extraction.
[0032] Multimodal content extraction extracts text content, structured data from tables, and data from charts. The text content extraction module cleans the parsed text, removing headers, footers, watermarks, page numbers, and other distracting elements, retaining the main text and labeling the chapter to which each text block belongs. The table extraction module identifies table elements in the document, extracts table headers and bodies, converts the tables into structured data in JSON format, and records the row index, column index, and cell value for each cell. The chart extraction module performs OCR recognition and image analysis on images in the document, extracting data points and trend information from charts such as line charts, bar charts, and pie charts. When the chart is in image format, a deep learning-based chart recognition model (such as PaddleOCR) is used for data extraction; when the chart is an embedded Excel object, the underlying data is extracted directly.
[0033] BP element structuring utilizes a large language model to perform semantic understanding on the extracted business plan content, transforming the business plan into an object containing structured data elements such as project name, company name, and each chapter. The large language model takes the complete text of the business plan as input and outputs a structured object in JSON format. The object's structure is as follows: The Project Name field extracts the project name from the business plan's title or project overview; the Company Name field extracts the full company name described in the business plan; the Market Analysis section includes structured data such as target market size (numerical value + unit), target market growth rate (percentage), target user profile (age group, income level, consumption preferences, etc.), and market entry strategy; the Product Solution section includes structured data such as a list of core product features, a description of the product's technical architecture, a description of competitive advantages, and a product iteration roadmap; the Business Model section includes structured data such as revenue sources (categorized), pricing strategy, customer lifetime value, and customer acquisition cost; the Financial Forecast section includes structured data such as revenue forecast (by year), cost structure (categorized), profit forecast, cash flow forecast, and financing needs; and the Team Introduction section includes structured data such as a list of core team members (each including name, position, background introduction, and years of relevant experience), team size, and equity allocation.
[0034] The BP element classification sub-step categorizes business plan elements into five categories: market, technology, business, financial, and team. These correspond to subsequent fact-checking, technical feasibility verification, hypothesis stress testing, rationality verification, and capability profile assessment strategies. Market elements include target market size, target market growth rate, market penetration forecast, competitive landscape description, and entry barriers description. Technology elements include core technology roadmap, technology maturity, technological barriers, and intellectual property status. Business elements include business model, revenue sources, pricing strategy, channel strategy, and customer base. Financial elements include revenue forecast, cost forecast, profit forecast, cash flow forecast, and valuation. Team elements include team member background, team size, organizational structure, and equity structure. Each element category has a clear classification label in the structured object, facilitating the subsequent analysis modules to call the appropriate verification methods.
[0035] In fact-checking and rationality verification, the system uses an investment decision knowledge graph and external data sources to conduct fact-checking and rationality verification on the key data and arguments presented in the business plan, and marks outliers and exaggerated statements.
[0036] Market size verification extracts the market size statement from the market analysis section of the business plan, including the absolute size and relative share of the target market. The system first locates the target market size field in the structured BP element object, then retrieves corresponding industry market size data from the investment decision knowledge graph based on industry classification. The search scope includes historical market size data for the past three years and market forecast data published by authoritative research institutions. The deviation rate is calculated as follows: the deviation rate equals the absolute value of the business plan statement minus the knowledge graph data value, divided by the knowledge graph data value. When the deviation rate exceeds a preset threshold of 30%, the system marks the market size statement as "market size may be exaggerated" and records the deviation rate value and the source of evidence. When the corresponding industry data is not available in the knowledge graph, the system calls external data sources (such as data from the National Bureau of Statistics or industry association reports) for supplementary retrieval. If no data is available from external data sources, the statement is marked as "unverifiable" and the reason is recorded as "lack of industry benchmark data."
[0037] The technical feasibility assessment extracts the technical roadmap statement from the product section of the business plan, including the core technology name, implementation principle, maturity level, and expected implementation time. The system retrieves the technology's maturity information from the investment decision knowledge graph, derived from a combination of patent analysis, academic paper citation analysis, and industry expert evaluation. When the technology's TRL level in the knowledge graph is below 6 (from the laboratory verification stage to the system simulation stage), the system labels the technical roadmap as "insufficient technology maturity." When the knowledge graph indicates that the technology has known major technical bottlenecks (such as material limitations, energy consumption issues, and yield problems) and the business plan does not mention corresponding solutions, the system labels the technical roadmap as "having major technical bottlenecks." The system further evaluates the feasibility and implementation difficulty of the technical roadmap using a large language model. Evaluation dimensions include the scientific validity of the technical principles, the rationality of the implementation path, the sufficiency of required resources, and the impact of competing technologies. The evaluation results are output in a structured scoring format.
[0038] The competitive statement verification extracts competitive statements from the competitive analysis section of the business plan, including a list of claimed main competitors and a description of each competitor's relative weaknesses. The system retrieves a list of features from similar competitors in the competitive knowledge graph and compares the claimed competitive features in the business plan with the actual features recorded in the knowledge graph. When the knowledge graph shows that other competitors support the unique features claimed in the business plan, the system marks the competitive statement as "Inaccurate Competitive Analysis" and lists the names of competitors supporting that feature in the knowledge graph. When the business plan omits important potential competitors (competitors with a competition intensity score exceeding 0.7 in the knowledge graph), the system marks the statement as "Incomplete Competitive Coverage" and lists the names of the omitted competitors.
[0039] The financial forecast rationality check extracts revenue growth projections from the financial forecast section of the business plan, including annual revenue and growth rates for the next three years. The system retrieves benchmark revenue growth data for similar companies (same industry, same stage, same business model) from the investment decision knowledge graph and calculates the deviation between the business plan's financial forecast and the industry benchmark. The industry benchmark data uses the median and 80th percentile of revenue growth rates for similar companies. When the business plan's projected growth rate exceeds the 80th percentile of the industry benchmark, the system marks the financial forecast as "potentially overly optimistic" and records the extent of the excess. When the business plan's financial forecast contains obvious logical contradictions (such as revenue growth coupled with cost reductions that do not conform to economies of scale), the system also marks it as "financial logic questionable."
[0040] Multimodal validation employs three methods for cross-validation. The textual semantic validation method uses a large language model to analyze the argumentation logic in the business plan, identifying logical contradictions and insufficient evidence. For example, if a business plan claims a disruptive advantage but provides no empirical data to support it, the large language model marks it as "insufficiently argued." The numerical consistency validation method checks whether the numerical indicators in the business plan are self-consistent, including the recursive relationships between data from different years (e.g., revenue = average order value × customer traffic), the consistency between data for each category and the summary data (e.g., the sum of all cost items equals total cost), and the consistency of the same data cited in different chapters (e.g., whether the market size data cited in the market analysis chapter and the financial forecast chapter are consistent). The external data source cross-validation method connects to third-party data sources (e.g., enterprise registration information, patent databases, court judgment documents) for cross-validation. When significant discrepancies are found between the statements in the business plan and the enterprise registration information, they are marked as "information inconsistency" and the details of the discrepancies are listed.
[0041] The fact-check results are output in the form of a structured report, which includes a list of check items, the check conclusion for each item (pass / questionable / fail), the source of evidence, the anomaly score, and a detailed explanation. The anomaly score is based on a 100-point scale and is calculated comprehensively based on the magnitude of the deviation, the sufficiency of evidence, and the degree of impact.
[0042] In constructing a multi-dimensional risk profile, the system quantifies and assesses the investment risks of a project from five dimensions—market risk, technological risk, team risk, financial risk, and compliance risk—based on fact-checking results and investment decision knowledge graph data, and generates a structured risk list.
[0043] The market risk score is calculated based on three penalty factors: **Market Size Exaggeration Penalty:** This penalty is calculated based on the deviation rate of the market size verification. When the deviation rate exceeds 30%, the penalty value equals the deviation rate multiplied by 100; otherwise, the penalty value is 0. **Intense Competitive Landscape Penalty:** This penalty is calculated based on the number of direct competitors in the competitor knowledge graph. Each direct competitor incurs a penalty value of 5. **Low Barriers to Entry Penalty:** This penalty is calculated based on the entry barrier score recorded in the knowledge graph. The entry barrier score is a comprehensive assessment based on factors such as industry qualification requirements, capital thresholds, technological barriers, and brand barriers. The penalty value equals 1 minus the entry barrier score multiplied by 50. The market risk score calculation formula is: Market Risk Score = 100 minus the sum of the market size exaggeration penalty, intense competitive landscape penalty, and low entry barrier penalty. The score range is 0-100, with a higher score indicating lower market risk. For example, when the market size deviation rate is 45%, there are 3 direct competitors, and the entry barrier score is 0.4, the penalty for market size exaggeration is 45, the penalty for intense competition is 15, the penalty for low entry barriers is 30, and the market risk score is 100 minus 90 equals 10 points, indicating that the market risk is extremely high.
[0044] The technology risk score is calculated based on three factors: Technology Maturity Level (TML) score, determined by the technology lifecycle stage (TML 1-3, 20 points; 4-6, 50 points; 7-9, 80 points); Technology Barrier Score, determined by the quantity and quality of patents (0-5 patents, 40 points; 10+, 60 points), and patent quality score, calculated by weighting the number of claims, citations, and legal status (0-30 points); and Intellectual Property Score, determined by the strength of IP protection, including patent portfolio completeness (whether there are core and peripheral patents), trade secret protection measures (whether there are confidentiality systems and non-compete agreements), and software copyright registration status, with a comprehensive score of 0-100 points. The formula for calculating the technology risk score is: the technology risk score equals the technology maturity score multiplied by 0.4, the technology barrier score multiplied by 0.3, and the intellectual property score multiplied by 0.3. The score range is 0-100 points, and the higher the score, the lower the technology risk.
[0045] The team risk score is calculated based on three factors: Background Matching Score, which assesses the match between team members' technical backgrounds and industry experience and project requirements. Evaluation dimensions include years of industry experience, technical expertise, and successful experiences in relevant fields. Each member is scored from 0 to 100, and the overall team score is a weighted average of these scores, with weights determined by each member's importance in the project. Complementarity Score, which assesses the complementarity of team members' capabilities. Evaluation dimensions include complementarity between technical and business capabilities, between the technical and operational teams, and between core and executive members. Higher complementarity results in a higher score. Execution Score, which assesses the team's historical project execution record, including project completion, timeline control, and cost control. Evaluation data is derived from historical investment cases recorded in the investment decision knowledge graph. The team risk score calculation formula is: Team Risk Score = Background Matching Score multiplied by 0.3 + Complementarity Score multiplied by 0.3 + Execution Score multiplied by 0.4, with a score range of 0-100. A higher score indicates lower team risk.
[0046] The financial risk score is calculated based on three penalty factors: Optimistic Financial Forecast Penalty: Calculated based on the results of the financial forecast reasonableness verification sub-step. When a financial forecast is marked as "potentially overly optimistic," the penalty value is determined by the degree to which it exceeds the industry benchmark: 20 points for exceeding the industry benchmark's 80th percentile or less, 40 points for exceeding it by 50%-100%, and 60 points for exceeding it by more than 100%. Cash Flow Pressure Penalty: Calculated based on the cash flow forecast in the financial forecast. When the number of years with negative cash flow exceeds 50% of the total forecast years, the penalty value increases by 10 points for each year exceeding this threshold, up to a maximum of 30 points. Valuation Bubble Penalty: Calculated based on a comparison of the valuation proposed in the business plan with the valuations of similar companies. When the valuation exceeds the average valuation of similar companies by 200%, the penalty value is 30 points; when it exceeds 300%, the penalty value is 50 points. The financial risk score calculation formula is: Financial Risk Score = 100 minus the sum of the optimistic financial forecast penalty, cash flow pressure penalty, and valuation bubble penalty. The score range is 0-100 points, with a higher score indicating lower financial risk.
[0047] The compliance risk score is calculated based on three factors: Legal compliance score, assessed based on the legal compliance requirements of the target company's industry, including the completeness of industry qualifications and licenses, the existence of legal litigation risks, and whether franchising is involved; the score is calculated using a weighted average of compliance requirement fulfillment. Data privacy score, assessed based on the target company's data processing practices, including whether it involves the collection of users' personal information, whether it has a data security management system, and whether it has passed the information security compliance assessment; the score is calculated using a weighted average of the completeness of compliance measures. Industry regulatory score, assessed based on the regulatory intensity of the target company's industry, is divided into high, medium, and low levels, corresponding to baseline scores of 30, 60, and 90 points respectively. The compliance risk score calculation formula is: Compliance risk score = Legal compliance score multiplied by 0.5 + Data privacy score multiplied by 0.3 + Industry regulatory score multiplied by 0.2, with a score range of 0-100. A higher score indicates lower compliance risk.
[0048] The risk profile vector generation step combines the scores of five risk dimensions into a risk profile vector. The vector format is a five-dimensional vector containing market risk score, technology risk score, team risk score, financial risk score, and compliance risk score. The order of the vector dimensions is fixed to facilitate subsequent investment decision model calculations. The risk list generation sub-step generates a structured risk list based on the risk profile vector. The list includes risk dimensions (market risk, technology risk, team risk, financial risk, and compliance risk), scores (specific score values for each dimension), risk levels (levels are divided according to the scores: 90-100 points for low risk, 70-89 points for medium risk, 50-69 points for relatively high risk, and below 50 points for high risk), key risk points (the specific risk item with the lowest score in each dimension), and mitigation suggestions (general mitigation suggestions for each key risk point).
[0049] In competitor benchmarking and differentiation analysis, the system verifies the competitive landscape described in the business plan based on the competitor knowledge graph and evaluates the project's differentiated competitive advantages.
[0050] The competitor list for a target company is identified based on four dimensions: Industry Classification Matching: All companies with the same industry classification as the target company are retrieved from the investment decision knowledge graph. Matching is calculated using the industry classification hierarchy similarity: 1 point for a first-level similarity and 0.5 points for a second-level similarity. Product Function Similarity: The vector similarity between the target company and candidate companies in their product function descriptions is calculated using cosine similarity, and the similarity value is directly used as the function similarity score. Target Market Overlap: The degree of overlap between the target company and candidate companies in their target markets (geographical regions or user groups) is calculated using the Jaccard coefficient, and the overlap value is directly used as the market overlap score. Technology Route Similarity: The similarity between the target company and candidate companies in their core technology routes is calculated using a combination of technical keyword matching and patent citation network analysis. The comprehensive score is calculated as follows: Comprehensive score = Industry Classification Matching score multiplied by 0.3 + Functional Similarity score multiplied by 0.4 + Market Overlap score multiplied by 0.2 + Technology Similarity score multiplied by 0.1. A comprehensive score greater than 0.6 indicates a competitor and is added to the competitor list.
[0051] We constructed a benchmarking matrix of target companies and competitors, encompassing technology, product, market, team, and financing dimensions. Comparison metrics for the technology dimension include the number of patents, the quality of core patents, the size of the technical team, and the number of technical partners. Comparison metrics for the product dimension include product feature coverage, user experience rating, product stability rating, and product iteration speed. Comparison metrics for the market dimension include market share, geographical coverage, number of customers, and customer retention rate. Comparison metrics for the team dimension include team size, average years of experience of core members, team integrity rating, and historical project success rate. Comparison metrics for the financing dimension include cumulative financing amount, investor lineup rating, valuation of the most recent financing round, and financing pace. Data for each metric is derived from cross-validation using an investment decision knowledge graph and external data sources. Data gaps are marked as "unknown," and the weight of that metric is reduced in subsequent analyses.
[0052] The differentiation advantage assessment process evaluates a target company's competitive advantage based on four dimensions: Functional Differentiation: This dimension assesses the degree of difference between the target company's product functions and those of competitors. The functional differentiation score equals the number of unique functions of the target company divided by the total number of functions of the target company, multiplied by 100, with a score range of 0-100. Technological Differentiation: This dimension assesses the degree of difference between the target company's technological approach and that of competitors. The technological differentiation score is calculated based on the uniqueness of its technology patents, the leading position of its technological approach, and the height of its technological barriers. Market Differentiation: This dimension assesses the degree of difference between the target company's market positioning and that of competitors. The market differentiation score is calculated based on the exclusivity of its target market, the uniqueness of its user profile, and the innovation of its market strategy. Team Differentiation: This dimension assesses the degree of difference between the target company's team and that of competitors. The team differentiation score is calculated based on the uniqueness of its team background, the complementarity of its team capabilities, and the leading performance of its team. The differentiation advantage score is calculated as follows: Differentiation advantage score equals the sum of the functional differentiation score, technological differentiation score, market differentiation score, and team differentiation score, divided by 4, with a score range of 0-100. A higher score indicates a more significant differentiation advantage.
[0053] The competitor benchmarking report generation process generates a structured competitor benchmarking report that includes competitor identification results, a competitor benchmarking matrix, a differentiation advantage assessment, and competitive strategy recommendations. The first part of the report lists all identified competitors and their overall scores; the second part presents a tabular comparison of the target company and each competitor across various dimensions; the third part displays the scores for each dimension of the differentiation advantage assessment and the overall score; and the fourth part generates competitive strategy recommendations based on the differentiation advantage assessment results. Recommendation types include technology leadership strategy, market focus strategy, differentiation positioning strategy, and defensive strategy, with specific recommendations determined based on the dimension with the highest differentiation advantage score.
[0054] In the process of generating and outputting investment decision recommendations, the system generates structured investment decision recommendations and risk control plans based on the analysis results of the aforementioned steps.
[0055] The investment decision model is calculated based on risk profile vectors and competitor benchmarking results, using a weighted scoring model to calculate the investment recommendation score. The formula for calculating the investment recommendation score is: Investment Recommendation Score = Market Risk Score multiplied by 0.25 + Technical Risk Score multiplied by 0.20 + Team Risk Score multiplied by 0.20 + Financial Risk Score multiplied by 0.20 + Compliance Risk Score multiplied by 0.15 + Differentiation Advantage Bonus - Fact Check Deduction. The Differentiation Advantage Bonus is equal to the Differentiation Advantage Score multiplied by 0.2, and the Fact Check Deduction is equal to the number of anomalies multiplied by 5. The number of anomalies is the total number of verification items marked "Questionable" and "Not Passed" in the fact check report. A strong buy recommendation is output when the investment recommendation score is greater than or equal to 80; a conditional buy recommendation with risk mitigation conditions is output when the investment recommendation score is between 60 and 79; a cautious buy recommendation with significant adjustments is output when the investment recommendation score is between 40 and 59; and a no-buy recommendation is output when the investment recommendation score is below 40.
[0056] The investment decision proposal generation process generates a structured investment decision proposal comprising six parts: project overview, fact-checking results, risk profile analysis, competitor benchmarking analysis, investment decision recommendations, and risk mitigation plans. The project overview extracts basic information from the structured business plan (BP) elements, such as project name, company name, industry, funding round, and funding amount. The fact-checking results section summarizes the conclusions and scores of each verification item in the fact-checking report. The risk profile analysis section displays a five-dimensional risk profile radar chart and a list of key risk points. The competitor benchmarking analysis section presents the core content and differentiated advantage assessment results of the competitor benchmarking report. The investment decision recommendations section outputs the corresponding investment recommendation conclusions based on the investment recommendation score. The risk mitigation plans section lists mitigation suggestions for each key risk point.
[0057] The risk mitigation plan generation process involves developing targeted risk mitigation plans for identified key risk points. Market risk mitigation plans include requiring third-party market research reports for data verification, reducing the optimism of market growth assumptions, and providing detailed explanations of market entry strategies. Technology risk mitigation plans include requiring third-party feasibility assessment reports, setting phased investment conditions for technology milestones, and clarifying intellectual property ownership. Team risk mitigation plans include requiring detailed background checks on core team members, setting equity vesting clauses for team members, and providing team expansion plans. Financial risk mitigation plans include requiring detailed explanations of financial forecast assumptions, adjusting valuations to a reasonable industry range, and setting performance-based clauses. Compliance risk mitigation plans include requiring plans for obtaining relevant licenses and permits, providing data security compliance certificates, and providing contingency plans for industry regulatory responses.
[0058] The system outputs documents in PDF or Word format for investment decision recommendations, Excel format for risk lists, PDF or Excel format for competitor benchmarking reports, and PDF format for fact-checking reports.
[0059] The technical solution of the present invention will be further explained below with reference to a specific case.
[0060] An analysis of the business plan for a Series B funding round of an AI startup, Company A, reveals that the company develops intelligent customer service systems. It claims its core product, based on self-developed natural language processing technology, achieves over 95% accuracy in intent recognition. Its target market is the Chinese SaaS intelligent customer service market, which is projected to reach 50 billion RMB by 2025, with an annual growth rate of 60%. Founded in 2020, Company A's core team comes from a major internet company. The CEO has 10 years of experience in the AI field, and the CTO has 8 years of NLP technology development experience. The business plan indicates a planned funding round of 50 million RMB, with a pre-investment valuation of 300 million RMB.
[0061] The system analyzed Company A's business plan following the aforementioned steps. In the multi-source knowledge graph construction step, the system integrated data sources such as AI industry research reports, SaaS industry reports, enterprise registration data, patent databases, and financing event databases to construct an investment decision knowledge graph containing 1200 industry nodes, 800 enterprise nodes, and 500 technology nodes. For the intelligent customer service sub-sector, the system additionally constructed a competitor knowledge graph containing 15 competing products, identifying 8 direct competitors.
[0062] In the BP multimodal parsing and structuring process, the system successfully parsed Company A's PDF business plan, extracting structured elements from five core sections: market analysis, product plan, business model, financial forecast, and team introduction. The market analysis section extracted two key data points: a target market size of 50 billion yuan and a growth rate of 60%. The product plan section extracted two key statements: the core technology route is a pre-trained Transformer-based model, and the intent recognition accuracy is 95%. The financial forecast section extracted three-year projections: 5 million yuan in revenue in 2023, 20 million yuan in 2024, and 80 million yuan in 2025.
[0063] In the fact-checking and reasonableness verification steps, the system conducted checks in four aspects. Market size verification showed that the knowledge graph recorded the Chinese intelligent customer service market size in 2025 as 28 billion yuan (source: iResearch Consulting 2024 Research Report), a deviation of 78.6% from Company A's claimed 50 billion yuan, exceeding the 30% threshold, and was marked as "Market size may be exaggerated." Technical feasibility verification showed that Company A's claimed Transformer technology route had a TRL level of 7 (commercially used) and a technology maturity score of 80 in the knowledge graph, but Company A only held 2 utility model patents, insufficient patent quantity, and was marked as "Technical feasibility questionable – weak intellectual property layout." Competitor statement verification showed that Company A listed only 3 competitors in its business plan, but the competitor knowledge graph showed 5 other direct competitors supporting its claimed unique functions, and was marked as "Inaccurate competitor analysis – omission of important competitors." The financial forecast reasonableness check shows that the median revenue growth rate of similar Series B AI companies is 200%, and the 80th percentile is 250%. Company A's forecast of a 300% growth rate in 2024 (from 5 million to 20 million) exceeds the 80th percentile and is marked as "the financial forecast may be too optimistic".
[0064] In the multi-dimensional risk profile construction process, the system calculated risk scores across five dimensions. The market risk score was 10 (78.6 for market size exaggeration penalty, 40 for intense competition penalty, and 30 for low entry barriers penalty), indicating extremely high market risk. The technology risk score was 56 (80×0.4=32 for technology maturity score, 9.75 for technology barriers score (5×4 for number of patents + 15 for patent quality score), and 21 for intellectual property score), indicating moderate to high technology risk. The team risk score was 72 (85×0.3=25.5 for background matching score, 22.5 for complementarity score, and 32 for execution score), indicating moderate team risk. The financial risk score was 50 (50 for optimistic financial forecast penalty, 0 for cash flow pressure penalty, and 0 for valuation bubble penalty), indicating relatively high financial risk. The compliance risk score is 75 (legal compliance score 80×0.5=40, data privacy score 70×0.3=21, industry regulatory score 70×0.2=14), indicating a moderate compliance risk. The comprehensive risk profile vector is [10,56,72,50,75].
[0065] In the competitor benchmarking and differentiation analysis steps, the system identified 8 direct competitors and constructed a 5-dimensional competitor benchmarking matrix. The differentiation advantage assessment showed a score of 65 for functional differentiation, 55 for technological differentiation, 60 for market differentiation, and 70 for team differentiation, resulting in a comprehensive differentiation advantage score of 62.5.
[0066] In the investment decision recommendation generation and output step, the system calculates the investment recommendation score. The investment recommendation score equals 10×0.25+56×0.20+72×0.20+50×0.20+75×0.15+62.5×0.2-20×5, with 4 outliers (exaggerated market size, questionable technical feasibility, inaccurate competitor analysis, and overly optimistic financial forecasts), and a deduction of 20 points for fact-checking. The calculated result is 2.5+11.2+14.4+10+11.25+12.5-20=41.35 points, corresponding to the decision conclusion of "prudent investment recommendation with a requirement for significant adjustments".
[0067] The following example illustrates the effect of the technical solution of the present invention compared to using only a single verification method.
[0068] In the comparative analysis, traditional manual due diligence methods were used to analyze the business plan of the same company, A. These methods primarily include due diligence interviews, public information retrieval, and expert consultation. Regarding market size verification, manual due diligence typically only obtains 1-2 industry reports for reference, making it difficult to cover detailed information on all competitors. The accuracy of the verification depends heavily on the industry knowledge accumulated by the due diligence personnel. In terms of technical feasibility verification, manual due diligence relies heavily on expert interviews, and the assessment results are significantly influenced by the experts' subjective judgment, making it difficult to systematically quantify the intellectual property layout. Regarding competitor analysis, manual due diligence is limited by time constraints, typically only verifying competitors mentioned in the business plan, making it difficult to systematically identify overlooked important competitors. In terms of financial forecasting, manual due diligence lacks systematic benchmark data from similar companies, relying mainly on the experience and judgment of the due diligence personnel. Regarding risk assessment, manual due diligence lacks unified quantitative standards, and the assessment results of different projects lack comparability.
[0069] The table below shows the comparison data between the examples and comparative examples of this invention: The aforementioned case studies and comparative data demonstrate that the technical solution of this invention can achieve comprehensive fact-checking and multi-dimensional risk quantification assessment of business plans, offering significant advantages over traditional manual due diligence methods. By constructing an investment decision knowledge graph as its knowledge foundation, the system can quickly retrieve and compare large amounts of multi-source data; through multimodal parsing technology, the system can automatically extract structured elements from business plans; through a quantitative scoring model, the system can generate comparable investment recommendation scores; and through a structured risk list and risk mitigation solutions, the system can output actionable investment decision support.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A risk control decision-making method for investment business plans based on knowledge graphs and multimodal analysis, characterized in that, Includes the following steps; S100, multi-source knowledge graph construction, is used to integrate industry data, enterprise data, technical data, market data, and competitor data to build an investment decision knowledge graph; S200 and BP multimodal parsing and structuring are used to parse unstructured business plan documents into a structured list of elements. S300, Fact Check and Reasonableness Verification: Based on the investment decision knowledge graph and external data sources, the key data and arguments presented in the business plan are fact-checked and reasonableness verified, and anomalies and exaggerated statements are marked. The S400 multi-dimensional risk profile construction, based on the fact-checking results of S300 and investment decision knowledge graph data, quantifies the investment risk of a project from five dimensions: market risk, technology risk, team risk, financial risk, and compliance risk, and generates a structured risk list. S500, competitor benchmarking and differentiation analysis, based on competitor knowledge graphs, to verify the competitive landscape stated in the business plan and to evaluate the project's differentiated competitive advantages; S600: Investment Decision Recommendation Generation and Output. Based on the analysis results of S300 to S500, structured investment decision recommendations and risk control plans are generated.
2. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 1, characterized in that, The construction of the multi-source knowledge graph includes: Multi-source data preprocessing is used to perform format parsing, data cleaning, entity linking and conflict resolution on raw data, and output standardized data records; Entity recognition and classification is used to identify entities in text using large language models and rule engines, and classify entities into industry nodes, enterprise nodes, technology nodes, product nodes, market nodes, and event nodes. Relation extraction is used to extract relationships between entities from data using large language models and rule engines, generating attribution relationships, competition relationships, technology usage relationships, target market relationships, financing relationships, and referencing relationships; The graph assembly and storage is used to write entities and relationships into a graph database to support graph reasoning, to write vectorized results of text descriptions into a vector database to support semantic retrieval, and to write numerical indicators into a structured database to support precise queries, forming an investment decision knowledge graph that links the three databases.
3. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 1, characterized in that, S100 further includes competitor knowledge graph construction, used to construct a subset of the investment decision knowledge graph. The competitor knowledge graph includes a competitor identification module, a competitor profiling module, a competitor comparison matrix module, and a competitive relationship reasoning module. Specifically, the competitor identification module identifies a list of competitors of the target company based on industry classification and product function similarity; the competitor profiling module extracts the competitor's technology roadmap, product functions, market share, financing history, and team background; the competitor comparison matrix module constructs a company-competitor comparison matrix covering technology, product, market, team, and financing dimensions; and the competitive relationship reasoning module infers the intensity of competition based on product function overlap, target market overlap, and technology roadmap similarity.
4. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 2, characterized in that, The BP multimodal parsing and structuring includes: The document parsing process supports several business plan document formats, extracts the main text, tables and charts, identifies the table of contents of the business plan, and locates the core sections such as market analysis, product plan, business model, financial forecast, and team introduction. The multimodal content extraction step is used to extract text content, extract structured data from tables, and extract data from charts through OCR and image recognition; The BP element structuring step is used to perform semantic understanding of the extracted business plan content using a large language model, and to convert the business plan into an object containing structured data elements such as project name, company name, and each chapter. The BP element classification step is used to classify the elements of a business plan into five categories: market, technology, business, financial, and team. These correspond to subsequent fact-checking, technical feasibility verification, hypothesis stress testing, rationality verification, and capability profile assessment and analysis strategies, respectively.
5. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 4, characterized in that, The fact-checking and reasonableness verification include: The market size verification step is used to extract market size statements from business plans, retrieve market size data for the corresponding industry from the knowledge graph, calculate the deviation rate between the business plan statements and the knowledge graph data, and mark the market size as exaggerated when the deviation rate exceeds a preset threshold. The technical feasibility verification step is used to extract the technical route statement from the business plan, retrieve the maturity, application cases and technical bottlenecks of the technology in the knowledge graph, and use the large language model to evaluate the feasibility and implementation difficulty of the technical route. When the technical maturity is insufficient or there are major technical bottlenecks, it is marked as technical feasibility questionable. The competitor statement verification step is used to extract competitor statements from the business plan, retrieve the feature list of similar competitors in the competitor knowledge graph, verify whether the competitor statements in the business plan are accurate, and mark the competitor analysis as inaccurate when the competitor knowledge graph shows that other competitors support the same function. The financial forecast reasonableness verification step is used to extract financial forecasts from the business plan, retrieve revenue growth benchmark data of similar companies in the knowledge graph, calculate the deviation between the financial forecasts in the business plan and the industry benchmark, and mark the financial forecast as too optimistic when the growth rate is significantly higher than the industry benchmark.
6. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 5, characterized in that, The construction of the multi-dimensional risk profile includes: Market risk scores are calculated based on penalties for exaggerating market size, intense competition, and low barriers to entry. A technology risk score is calculated based on technology maturity score, technology barrier score, and intellectual property score. The team risk score is calculated based on background matching score, complementarity score, and execution score; A financial risk score is calculated based on optimistic financial forecast penalties, cash flow pressure penalties, and valuation bubble penalties. A compliance risk score is calculated based on legal compliance scores, data privacy scores, and industry regulatory scores.
7. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 6, characterized in that, The competitor benchmarking and differentiation analysis includes: The competitor list of a target company is identified based on four dimensions: industry classification matching, product function similarity calculation, target market overlap calculation, and technology route similarity calculation. The comprehensive score calculation formula is: the comprehensive score equals the industry classification matching degree multiplied by 0.3 plus the function similarity multiplied by 0.4 plus the market overlap multiplied by 0.2 plus the technology similarity multiplied by 0.
1. When the comprehensive score is greater than 0.6, it is judged as a competitor. Construct a target company-competitor benchmarking matrix covering technology, product, market, team, and financing dimensions; The competitive advantages of a target company are evaluated based on four dimensions: functional differentiation, technological differentiation, market differentiation, and team differentiation. Generate a structured competitor benchmarking report that includes competitor identification results, competitor benchmarking matrix, differentiated advantage assessment, and competitive strategy recommendations.
8. The investment business plan risk control decision-making method based on knowledge graph and multimodal analysis according to claim 7, characterized in that, The generation and output of the investment decision recommendations include: Based on the risk profile vector and competitor benchmarking results, a weighted scoring model is used to calculate the investment recommendation score. The calculation formula is as follows: the investment recommendation score is equal to the sum of the products of the five risk dimension scores and their respective weights, plus the differentiation advantage bonus, minus the fact-checking deduction. The differentiation advantage bonus is equal to the differentiation advantage score multiplied by 0.2, and the fact-checking deduction is equal to the number of anomalies multiplied by 5. Generate a structured investment decision proposal containing six parts: project overview, fact-checking results, risk profile analysis, competitor benchmarking analysis, investment decision recommendations, and risk mitigation solutions. Also, generate targeted risk mitigation solutions for the identified key risk points.
9. An investment business plan (BP) risk control decision-making system based on knowledge graphs and multimodal analysis, characterized in that: The system is applied to an investment business plan (BP) risk control decision-making method based on knowledge graph and multimodal analysis as described in any one of claims 1-8. The system includes a multi-source knowledge graph construction module, a BP multimodal parsing and structuring module, a fact-checking and rationality verification module, a multi-dimensional risk profile construction module, a competitor benchmarking and differentiation analysis module, an investment decision suggestion generation and output module, and a knowledge graph dynamic update module. The multi-source knowledge graph construction module is used to integrate multi-source heterogeneous data such as industry data, enterprise data, technical data, market data, and competitor data to construct an investment decision knowledge graph. The BP multimodal parsing and structuring module is used to parse unstructured business plan documents into a structured list of elements. The fact-checking and rationality verification module is used to perform fact-checking and rationality verification on the key data and arguments presented in the business plan based on the investment decision knowledge graph and external data sources. The multi-dimensional risk profile building module is used to quantitatively assess the investment risks of a project from five dimensions: market risk, technology risk, team risk, financial risk, and compliance risk. The competitor benchmarking and differentiation analysis module is used to verify the competitive landscape and evaluate differentiated competitive advantages based on the competitor knowledge graph. The investment decision recommendation generation and output module is used to generate structured investment decision recommendations (49) and risk control schemes; The knowledge graph dynamic update module is used to incrementally update the knowledge graph when the triggering conditions occur, and retain historical version snapshots for backtracking queries.
Citation Information
Patent Citations
Investment and financing project recommendation method based on dynamic knowledge graph and large language model
CN121860779A