Systems and methods for data-driven identification of entities in early-stage venture capital
A data-driven system with rule-based models and expert review addresses the challenge of identifying early-stage companies for venture capital by consolidating and filtering noisy data, improving investment decision-making.
Patent Information
- Application Number
- PCT/IB2024/060937
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2024-11-05
- Publication Date
- 2025-12-04
AI Technical Summary
Investors face challenges in identifying promising early-stage companies for venture capital due to limited and noisy data, lack of historical information, and operational inefficiencies, making it difficult to make informed investment decisions.
A data-driven system using rule-based models and large language models to consolidate and filter data, combined with expert review, to identify and prioritize investment opportunities.
Enhances the ability to make informed decisions on early-stage companies by providing accurate insights and predictions, leveraging diverse data sources and expert validation.
Smart Images

Figure IB2024060937_04122025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR DATA-DRIVEN IDENTIFICATION OF ENTITIES IN EARLY-STAGE VENTURE CAPITALCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of domestic priority under 35 U.S.C. § 119(e) from Provisional U.S. Patent Application No. 63 / 654,326, filed May 31, 2024, the disclosures of which are incorporated herein by reference in their entirety for all purposes.TECHNICAL FIELD
[0002] The embodiments described herein are generally directed to systems and methods for data-driven identification of entities, and, more particularly, to systems and methods for data-driven identification of entities in early-stage venture capital.BACKGROUND
[0003] Most investors targeting early-stage companies for potential venture capital investment face significant challenges in identifying promising opportunities. The primary difficulty lies in detecting undiscovered companies through early signals in a scalable manner. This task is complicated by the nature of early-stage company data, which often contains unstructured data mixed with irrelevant information and lacks meaningful historical information. This makes it challenging to process information and make decisions using traditional machine learning models effectively. In addition, early-stage entities typically lack a proven track record and may still be developing or refining their business models. This uncertainty, combined with limited financial transparency and potential operational inefficiencies, adds to the complexity of analyzing signals. Given these challenges, there is a need for improved methods to analyze early signals and filter through noisy data to make informed decisions about which companies to engage with for potential investment.
[0004] Further, the agility of the system is crucial in adapting to a changing industry landscape and the emergence of new sectors. Current data-driven systems struggle to identify and evaluate information about early-stage entities because these startups typically lack comprehensive and publicly available data. Unlike established firms, early-stage entities often have minimal historical financial records, limited market presence, and little to no detailed performance metrics that can be easily analyzed. Therefore, the absence of standardizedinformation makes it challenging for data-driven systems to generate accurate insights or predict future performance to guide investor’s decisions for venture capital funding.
[0005] U.S. Patent Publication No. 20220114518 Al describes a system and method for determining the likelihood of future success of a business to inform investment decisions. However, this approach is less effective for analyzing early-stage companies, which typically have limited operational data and sparse public information. Early-stage entities often have only minimal operational data and company descriptions that may be brief, contain noise, or not be available in public sources. To address these issues, a data-driven system and method have been developed for identifying early-stage entities for venture capital engagement, utilizing a combination of rule-based models and large language models (LLMs) instead of traditional deep neural networks (DNNs). This allows us to extract meaningful signals from limited and diverse data sources, including noisy or incomplete information. The goal is to make informed decisions about engaging with early-stage companies, rather than making direct investment decisions as in the U.S. Patent No. 20220114518 Al. The present disclosure is directed toward overcoming one or more of the problems discovered by the inventors.SUMMARY
[0006] In an embodiment, a non-transitory computer-readable medium having instructions stored to identify entities to receive a resource, wherein the instructions, when executed by a processor, cause the processor to: obtain data information of entities from more than one source; consolidate disparate data points into unified company profiles using rule-based matching, fuzzy matching, and manual matching; apply a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; provide the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subject the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and use the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
[0007] In an embodiment, a method comprising using at least one hardware processor to: obtaining data information of entities from more than one source; consolidating disparate data points into unified company profiles using rule-based matching, fuzzy matching, and manualmatching; applying a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; providing the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subjecting the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and using the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
[0008] In an embodiment, a system to identify entities for venture capital investment, the system comprises: at least one hardware processor; and software that is configured to, when executed by the at least one hardware processor, obtain data information of entities from more than one source; consolidate disparate data points into unified company profiles using rulebased matching, fuzzy matching, and manual matching; apply a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; provide the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subject the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and use the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
[0009] It should be understood that any of the features in the methods and systems above may be implemented individually or with any subset of the other features in any combination. Thus, to the extent that the appended claims would suggest particular dependencies between features, disclosed embodiments are not limited to these particular dependencies. Rather, any of the features described herein may be combined with any other feature described herein, or implemented without any one or more other features described herein, in any combination of features whatsoever. In addition, any of the methods, described above and elsewhere herein, may be embodied, individually or in any combination, in executable software modules of a processor-based system, such as a server, and / or in executable instructions stored in a non- transitory computer-readable medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The details of the present invention, both as to its structure and operation, may be gleaned in part by study of the accompanying drawings, in which like reference numerals refer to like parts, and in which:
[0011] FIG. 1 illustrates an example infrastructure in which one or more of the disclosed processes may be implemented, according to an embodiment;
[0012] FIG. 2 illustrates a processing system that may be used to implement processes described herein, according to an embodiment;
[0013] FIG. 3 illustrates a conceptual architecture of a machine learning model module in a system for data-driven identification of entities in early-stage venture capital, according to an embodiment;
[0014] FIG. 4 illustrates a sequential arrangement of modules in a system for data-driven identification of entities in early-stage venture capital, according to an embodiment;
[0015] FIG. 5 illustrates an example of a data acquisition subsystem focused on job posts of an entity, according to an embodiment;
[0016] FIG. 6 illustrates an example of a data acquisition subsystem focused on website traffic of an entity, according to an embodiment;
[0017] FIG. 7 an illustrates an example of a data acquisition subsystem focused on mobile application metrics of an entity, according to an embodiment;
[0018] FIG. 8 illustrates an example of a data acquisition subsystem focused on new product information of an entity, according to an embodiment;
[0019] FIG. 9 illustrates an example of a data acquisition subsystem focused on news information of an entity, according to an embodiment;
[0020] FIG. 10 illustrates an example of an analytical tool dashboard used in a system for data-driven identification of entities in early-stage venture capital, according to an embodiment;
[0021] FIG. 11 illustrates an infrastructure of an aggregator module, according to an embodiment;
[0022] FIG. 12 illustrates an example of a rule-based model module in a system for data- driven identification of entities in early-stage venture capital, according to an embodiment;
[0023] FIG. 13 illustrates an example of a fuzzy matching module in a system for data- driven identification of entities in early-stage venture capital, according to an embodiment;
[0024] FIG. 14 illustrates a decision-making framework within a system for data-driven identification of entities in early-stage venture capital, according to an embodiment;
[0025] FIG. 15 illustrates a large language model module in a system for data-driven identification of entities in early-stage venture capital, according to an embodiment;
[0026] FIG. 16 illustrates an example of a reviewer module interface dashboard, according to an embodiment;
[0027] FIG. 17 illustrates an example of a user interface to collect reviewer’s feedback; according to an embodiment; and
[0028] FIG. 18 illustrates an example of a dashboard with priority results of recommendations for entities in early-stage venture capital.DETAILED DESCRIPTION
[0029] The detailed description set forth below, in connection with the accompanying drawings, is intended as a description of various embodiments, and is not intended to represent the only embodiments in which the disclosure may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that embodiments of the invention can be practiced without these specific details.
[0030] In some instances, well-known structures and components are shown in simplified form for brevity of description. For clarity and ease of explanation, some surfaces and details may be omitted in the present description and figures. It should also be understood that the various components illustrated herein are not necessarily drawn to scale. In other words, thefeatures disclosed in various embodiments may be implemented using different relative dimensions within and between components than those illustrated in the drawings.
[0031] After reading this description, it will become apparent to one skilled in the art how to implement the invention in various alternative embodiments and alternative applications. However, although various embodiments of the present invention will be described herein, it is understood that these embodiments are presented by way of example and illustration only, and not limitation. As such, this detailed description of various embodiments should not be construed to limit the scope or breadth of the present invention as set forth in the appended claims.
[0032] FIG. 1 illustrates an example infrastructure in which one or more of the disclosed processes may be implemented, according to an embodiment. The infrastructure may comprise a platform 110 (e.g., one or more servers) which hosts and / or executes one or more of the various processes (e.g., methods or functions, implemented as software modules) described herein. Platform 110 may comprise dedicated servers, or may instead be implemented in a computing cloud, in which the resources of one or more servers are dynamically and elastically allocated to multiple tenants based on demand. In either case, the servers may be collocated and / or geographically distributed. Platform 110 may also comprise or be communicatively connected to a server application 112 and / or one or more databases 114. In addition, platform 110 may be communicatively connected to one or more user systems 130 via one or more networks 120. Platform 110 may also be communicatively connected to one or more external systems 140 (e.g., other platforms, websites, etc.) via one or more networks 120.
[0033] Network(s) 120 may comprise the Internet, and platform 110 may communicate with user system(s) 130 through the Internet using standard transmission protocols, such as HyperText Transfer Protocol (HTTP), HTTP Secure (HTTPS), File Transfer Protocol (FTP), FTP Secure (FTPS), Secure Shell FTP (SFTP), and the like, as well as proprietary protocols. While platform 110 is illustrated as being connected to various systems through a single set of network(s) 120, it should be understood that platform 110 may be connected to the various systems via different sets of one or more networks. For example, platform 110 may be connected to a subset of user systems 130 and / or external systems 140 via the Internet, but may be connected to one or more other user systems 130 and / or external systems 140 via an intranet. Furthermore, while only a few user systems 130 and external systems 140, one server application 112, and one set of database(s) 114 are illustrated, it should be understood that theinfrastructure may comprise any number of user systems, external systems, server applications, and databases.
[0034] User system(s) 130 may comprise any type or types of computing devices capable of wired and / or wireless communication, including without limitation, desktop computers, laptop computers, tablet computers, smart phones or other mobile phones, servers, game consoles, televisions, set-top boxes, electronic kiosks, point-of-sale terminals, and / or the like. Each user system 130 may comprise or be communicatively connected to a client application 132 and / or one or more local databases 134.
[0035] Platform 110 may comprise web servers which host one or more websites and / or web services. In embodiments in which a website is provided, the website may comprise a graphical user interface, including, for example, one or more screens (e.g., webpages) generated in HyperText Markup Language (HTML) or other language. Platform 110 transmits or serves one or more screens of the graphical user interface in response to requests from user system(s) 130. In some embodiments, these screens may be served in the form of a wizard, in which case two or more screens may be served in a sequential manner, and one or more of the sequential screens may depend on an interaction of the user or user system 130 with one or more preceding screens. The requests to platform 110 and the responses from platform 110, including the screens of the graphical user interface, may both be communicated through network(s) 120, which may include the Internet, using standard communication protocols (e.g., HTTP, HTTPS, etc.). These screens (e.g., webpages) may comprise a combination of content and elements, such as text, images, videos, animations, references (e.g., hyperlinks), frames, inputs (e.g., textboxes, text areas, checkboxes, radio buttons, drop-down menus, buttons, forms, etc.), scripts (e.g., JavaScript), and the like, including elements comprising or derived from data stored in one or more databases (e.g., database(s) 114) that are locally and / or remotely accessible to platform 110. It should be understood that platform 110 may also respond to other requests from user system(s) 130.
[0036] Platform 110 may comprise, be communicatively coupled with, or otherwise have access to one or more database(s) 114. For example, platform 110 may comprise one or more database servers which manage one or more databases 114. Server application 112 executing on platform 110 and / or client application 132 executing on user system 130 may submit data (e.g., user data, form data, etc.) to be stored in database(s) 114, and / or request access to data stored in database(s) 114. Any suitable database may be utilized, including without limitationMySQL™, Oracle™, IBM™, Microsoft SQL™, Access™, PostgreSQL™, MongoDB™, and the like, including cloud-based databases and proprietary databases. Data may be sent to platform 110, for instance, using the well-known POST request supported by HTTP, via FTP, and / or the like. This data, as well as other requests, may be handled, for example, by serverside web technology, such as a servlet or other software module (e.g., comprised in server application 112), executed by platform 110.
[0037] In embodiments in which a web service is provided, platform 110 may receive requests from user system(s) 130 and / or external system(s) 140, and provide responses in extensible Markup Language (XML), JavaScript Object Notation (JSON), and / or any other suitable or desired format. In such embodiments, platform 110 may provide an application programming interface (API) which defines the manner in which user system(s) 130 and / or external system(s) 140 may interact with the web service. Thus, user system(s) 130 and / or external system(s) 140 (which may themselves be servers), can define their own user interfaces, and rely on the web service to implement or otherwise provide the backend processes (e.g., methods and functionality), storage, and / or the like, described herein. For example, in such an embodiment, a client application 132, executing on one or more user system(s) 130, may interact with a server application 112 executing on platform 110 to execute one or more or a portion of one or more of the various process(es) described herein.
[0038] Client application 132 may be “thin,” in which case processing is primarily carried out server-side by server application 112 on platform 110. A basic example of a thin client application 132 is a browser application, which simply requests, receives, and renders webpages at user system(s) 130, while server application 112 on platform 110 is responsible for generating the webpages and managing database functions. Alternatively, the client application may be “thick,” in which case processing is primarily carried out client-side by user system(s) 130. It should be understood that client application 132 may perform an amount of processing, relative to server application 112 on platform 110, at any point along this spectrum between “thin” and “thick,” depending on the design goals of the particular implementation. In any case, the software described herein, which may wholly reside on either platform 110 (e.g., in which case server application 112 performs all processing) or user system(s) 130 (e.g., in which case client application 132 performs all processing) or be distributed between platform 110 and user system(s) 130 (e.g., in which case server application 112 and client application 132 both perform processing), can comprise one or more executable software modulescomprising instructions that implement one or more of the processes (e.g., methods or functions) described herein.
[0039] FIG. 2 illustrates a processing system that may be used to implement processes described herein, according to an embodiment. System 200 may comprise one or more processors 210. Processor(s) 210 may comprise a central processing unit (CPU). Additional processors may be provided, such as a graphics processing unit (GPU), an auxiliary processor to manage input / output, an auxiliary processor to perform floating-point mathematical operations, a special-purpose microprocessor having an architecture suitable for fast execution of signal-processing algorithms (e.g., digital-signal processor), a subordinate processor (e.g., back-end processor), an additional microprocessor or controller for dual or multiple processor systems, and / or a coprocessor. Such auxiliary processors may be discrete processors or may be integrated with a main processor 210. Examples of processors which may be used with system 200 include, without limitation, any of the processors (e.g., Pentium™, Core i7™, Xeon™, etc.) available from Intel Corporation of Santa Clara, California, any of the processors available from Advanced Micro Devices, Incorporated (AMD) of Santa Clara, California, any of the processors (e.g., A series, M series, etc.) available from Apple Inc. of Cupertino, any of the processors (e.g., Exynos™) available from Samsung Electronics Co., Ltd., of Seoul, South Korea, any of the processors available from NXP Semiconductors N.V. of Eindhoven, Netherlands, and / or the like.
[0040] Processor 210 may be connected to a communication bus 205. Communication bus 205 may include a data channel for facilitating information transfer between storage and other peripheral components of system 200. Furthermore, communication bus 205 may provide a set of signals used for communication with processor 210, including a data bus, address bus, and / or control bus (not shown). Communication bus 205 may comprise any standard or nonstandard bus architecture such as, for example, bus architectures compliant with industry standard architecture (ISA), extended industry standard architecture (EISA), Micro Channel Architecture (MCA), peripheral component interconnect (PCI) local bus, standards promulgated by the Institute of Electrical and Electronics Engineers (IEEE) including IEEE 488 general-purpose interface bus (GPIB), IEEE 696 / S-100, and / or the like.
[0041] System 200 may comprise main memory 215. Main memory 215 provides storage of instructions and data for programs executing on processor 210, such as one or more of the functions and / or modules discussed herein. It should be understood that programs stored inthe memory and executed by processor 210 may be written and / or compiled according to any suitable language, including without limitation C / C++, Java, JavaScript, Perl, Python, Visual Basic, .NET, and the like. Main memory 215 is typically semiconductor-based memory such as dynamic random access memory (DRAM) and / or static random access memory (SRAM). Other semiconductor-based memory types include, for example, synchronous dynamic random access memory (SDRAM), Rambus dynamic random access memory (RDRAM), ferroelectric random access memory (FRAM), and the like, including read only memory (ROM).
[0042] System 200 may comprise secondary memory 220. Secondary memory 220 is a non-transitory computer-readable medium having computer-executable code and / or other data (e.g., any of the software disclosed herein) stored thereon. In this description, the term “computer-readable medium” is used to refer to any non-transitory computer-readable storage media used to provide computer-executable code and / or other data to or within system 200. The computer software stored on secondary memory 220 is read into main memory 215 for execution by processor 210. Secondary memory 220 may include, for example, semiconductor-based memory, such as programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), and flash memory (block-oriented memory similar to EEPROM).
[0043] System 200 may comprise an input / output (I / O) interface 235. I / O interface 235 provides an interface between one or more components of system 200 and one or more input and / or output devices.
[0044] System 200 may comprise a communication interface 240. Communication interface 240 allows software to be transferred between system 200 and external devices, networks, or other information sources. For example, computer-executable code and / or data may be transferred to system 200, over one or more networks, from a network server via communication interface 240. Examples of communication interface 240 include a built-in network adapter, network interface card (NIC), Personal Computer Memory Card International Association (PCMCIA) network card, card bus network adapter, wireless network adapter, Universal Serial Bus (USB) network adapter, modem, a wireless data card, a communications port, an infrared interface, an IEEE 1394 fire-wire, and any other device capable of interfacing system 200 with a network or another computing device. Communication interface 240 preferably implements industry-promulgated protocol standards, such as Ethernet IEEE 802 standards, Fiber Channel, digital subscriber line (DSL), asynchronous digital subscriber line(ADSL), frame relay, asynchronous transfer mode (ATM), integrated digital services network (ISDN), personal communications services (PCS), transmission control protocol / Internet protocol (TCP / IP), serial line Internet protocol / point to point protocol (SLIP / PPP), and so on, but may also implement customized or non-standard interface protocols as well.
[0045] Software transferred via communication interface 240 is generally in the form of electrical communication signals 255. These signals 255 may be provided to communication interface 240 via a communication channel 250 between communication interface 240 and an external system 245. In an embodiment, communication channel 250 may be a wired or wireless network, or any variety of other communication links. Communication channel 250 carries signals 255 and can be implemented using a variety of wired or wireless communication means including wire or cable, fiber optics, conventional phone line, cellular phone link, wireless data communication link, radio frequency (“RF”) link, or infrared link, just to name a few.
[0046] Computer-executable code is stored in main memory 215 and / or secondary memory 220. Computer-executable code can also be received from an external system 245 via communication interface 240 and stored in main memory 215 and / or secondary memory 220. Such computer-executable code, when executed by processor(s) 210, enable system 200 to perform the various functions of the disclosed embodiments as described elsewhere herein.
[0047] FIG. 3 illustrates a conceptual architecture of a machine learning model module in a system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. In an embodiment, the application may utilize one or more machinelearning models or other artificial intelligence (Al) to facilitate one or more aspects of project management in a system for data-driven identification of entities in early-stage venture capital 400. FIG. 3 illustrates a high-level diagram of machine learning, according to an embodiment. In general, each machine-learning model 300 is trained to make predictions and / or recommendations in a training stage 310 and operated to make predictions in an operation stage 320. It should be understood that each machine-learning model 300 used by the application may undergo its own training stage 310 and operation stage 320, and that two or more machinelearning models 300 may operate independently from each other to perform different predictive tasks or may operate in combination with each other to perform a single predictive task.
[0048] In training stage 310, a machine-learning (ML) model 300 is trained using a dataset 312. Machine-learning model 300 may be trained using supervised or unsupervised learning. In supervised learning, dataset 312 may comprise vectors of features, with each feature vector labeled or annotated with the desired output and comprising a plurality of features that may be relevant to the determination of that output. In a case in which machine-learning model 300 is intended to perform image recognition or classification, dataset 312 may instead comprise images that have been labeled or annotated with the desired recognition or classification output. In either case, dataset 312 may be cleaned and augmented in any known manner.
[0049] In subprocess 314, feature engineering may be used to identify the features to be used in the feature vectors in dataset 312. The feature engineering may utilize any known manner of identifying relevant features that may correlate to an output. Features that are determined to be irrelevant may be removed from the feature vectors of dataset 312. In an alternative embodiment or embodiments which do not use feature vectors (e.g., image recognition or classification), subprocess 314 may be omitted.
[0050] In subprocess 316, machine-learning model 300 is trained using dataset 312. Specifically, machine-learning model 300 is applied to dataset 312 (e.g., which may be divided into training and validation datasets) and updates its internal structure to minimize the error between the desired output, represented by the labels in dataset 312, and its actual output. Machine-learning model 300 may comprise any type of machine-learning algorithm, including, without limitation, an artificial neural network (e.g., a convolutional neural network, a deep neural network, etc.), a linear regression, a logistic regression, a decision tree, a random forest algorithm, a support vector machine (SVM), a naive Bayesian classifier, a k-Nearest Neighbors (kNN) algorithm, a K-Means algorithm, gradient boosting algorithms (e.g., XGBoost, LightGBM, CatBoost), and the like. It should be understood that the particular machinelearning algorithm that is used will depend on the problem being solved, and that different machine-learning algorithms may be used by platform 110 for different tasks.
[0051] In subprocess 318, machine-learning model 300 may be evaluated to determine its accuracy in performing the predictive task for which it was designed. If the accuracy is not sufficient, the training stage 310 may continue. For example, a different set of features may be used for training, a different dataset 312 may be used for training, a different machine-learning algorithm may be used, and / or the like, until the evaluation in subprocess 318 demonstrates that machine-learning model 300 is suitably accurate. It should be understood that thenecessary accuracy may depend on the predictive task for which machine-learning model 300 was designed. For example, a machine-learning model 300 used for a predictive task that may affect a client interaction may require more accuracy than a machine-learning model 300 used for a predictive task that only affects internal interactions between team members.
[0052] Once machine-learning model 300 has been trained to a sufficient accuracy, machine-learning model 300 may be moved to operation stage 320 to perform its predictive task on data 322 in a production environment of platform 110. Data 322 may comprise feature vectors derived from any of the data discussed herein and / or image data derived from any of the media discussed herein (e.g., photographs, video, etc.). In subprocess 324, machinelearning model 300 is applied to data 322 to produce a prediction 326. Prediction 326 may comprise a classification (e.g., a single most likely classification, a probability vector comprising confidences for each of a plurality of possible classifications, etc.), a recommendation (e.g., recommended next action), and / or the like.
[0053] FIG. 4 illustrates a sequential arrangement of modules in a system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. Data source module 410 serves as the repository for collecting data. Data source module’s 410 data may encompass a wide array of information, including company information, recruiting jobs, website performance metrics, App performance, people in organization, website, number of peoples. The diversity and depth of the data sourced enhance the subsequent analytical processes.
[0054] The collected data from data source module 410 is then routed to aggregator module 420. The primary function of aggregator module 420 is to consolidate disparate data streams into a singular, coherent data set. The process in aggregator module 420 ensures that the data fed into the following stages is representative, comprehensive, and structured for advanced analysis.
[0055] Decision engine 430 is the next critical component of system for data-driven identification of entities in early-stage venture capital 400. Utilizing sophisticated algorithms, machine learning techniques, large language model and historical data patterns, decision engine 430 performs a dynamic analysis of the aggregated data from aggregator module 420. Decision engine 430 assesses potential investment opportunities against a multitude of factors to predict success and viability, thus generating preliminary investment recommendations. However,recognizing the complexity and nuance of venture capital investment decisions, system for data-driven identification of entities in early-stage venture capital 400 incorporates reviewer module 440. Reviewer module 440 facilitates a human-in-the-loop review process, wherein experts can weigh in on the recommendations provided by decision engine 430.
[0056] Finally, the process in system for data-driven identification of entities in early-stage venture capital 400 culminates with a decision module 450. Here, the outcome of decision engine's 430 analysis, augmented by Reviewer module’s 440 insights, is synthesized into a final decision in decision module 450 on whether the early-stage company represents a favorable investment opportunity. Decision module 450 stands as the definitive gateway, ensuring that only the most promising ventures, as determined by rigorous data-driven analysis and expert review, are selected for further scrutiny by the investment team. System for data- driven identification of entities in early-stage venture capital 400 leverages a variety of data sources, including recruitment information, website traffic statistics, app store rankings, Product Hunt, and news articles. Below, the detailed approach is described by to utilizing each listed source.
[0057] FIG. 5 illustrates an example of a data acquisition subsystem 500 focused on job posts of an entity, according to an embodiment. Data source module 410, as a critical component of system for data-driven identification of entities in early-stage venture capital 400, is tasked with the extraction and organization of data relevant to potential pre-defined investment targets. Data acquisition subsystem 500 contains the extraction of job posts from various digital mediums which may serve as a leading indicator of a company's growth and hiring trends. Upon the collection of job postings, data acquisition subsystem 500 initiates a crawl company information sequence, wherein it employs web crawling technologies to gather additional data about the companies associated with the job postings.
[0058] Subsequently, data acquisition subsystem 500 performs a verification step to ascertain the presence of the company within the existing database. If the company name is already present in the database, this indicates prior recognition and potentially accumulated data, which warrants further analysis. When a job post is confirmed, the data from the post is saved to a jobs table in the database, ensuring that each opportunity is logged and time-stamped, providing a temporal data point for analysis. Concurrently, if the company is not recognized in the database, its information is inserted into a companies table in the database, which enables a rich historical data profile for each company to be constructed over time.
[0059] FIG. 6 illustrates an example of a data acquisition subsystem 600 focused on website traffic of an entity, according to an embodiment. Data source module 410 encapsulated area is tasked with capturing website traffic data, which is a parameter indicative of consumer interest and product market fit. The data in data acquisition subsystem 600 is not indiscriminately captured; parameters such as country and category are applied to filter the traffic relevant to the geographic market and industry sector of interest. Additionally, data acquisition subsystem 600 is configured to select data from the most recent period, ensuring that the analysis is based on current trends and user engagement. Once the website traffic data is obtained, it is processed by the crawl company information function.
[0060] Upon extraction, the website traffic data is systematically saved to a traffic table in the database, creating a historical data point for each company that can be analyzed over time to identify growth patterns and spikes in user interest. Data acquisition subsystem 600 then queries the database to determine if the company name already exists within the stored data (e.g. “Is company name in database?”). If the company is previously known to the database, this indicates ongoing monitoring and data collection, which can be helpful for tracking progress and development over time.
[0061] When a company's presence in the database is confirmed, the new information is inserted into a companies table in the database. This step updates the company's profile with the latest engagement data, allowing for a dynamic assessment of the company's trajectory and potential for success. By integrating website traffic data with a robust set of parameters and historical information, Data acquisition subsystem 600 enriches the overall analytical capability of system for data-driven identification of entities in early-stage venture capital 400, enabling more informed and nuanced decisions regarding early-stage venture capital investments.
[0062] FIG. 7 an illustrates an example of a data acquisition subsystem 700 focused on mobile application metrics of an entity, according to an embodiment. In this example of data acquisition subsystem 700, data source module 410 of system for data-driven identification of entities in early-stage venture capital 400 specifically targets an Appstore ranking as a key indicator of an application's popularity and user adoption. This metric enhances data acquisition subsystem 700, as it reflects the competitive positioning of an app within the crowded marketplace of mobile applications. To ensure a comprehensive understanding of market presence, data acquisition subsystem 700 is engineered to crawl app information fromboth the Apple App Store and Google Play Store across each category, harnessing a broad spectrum of user interaction data.
[0063] Once the rankings and associated app details are retrieved, this information is saved to an Apps table in the database. This dedicated table within the database can be structured to store various data points about the apps, including their rankings, category, developer information, and any other relevant metadata that could signal the company's potential for growth and scalability. Parallel to this, data acquisition subsystem 700 can execute a “Find Related Company” operation, whereby it matches the applications to their corresponding developing companies. This step can enhance app association and performance metrics with the companies responsible for their development.
[0064] Subsequently, a crawl company information step is initiated. This step involves an extensive search for additional data on the company, potentially including financials, team size, and market presence. The crawl company information step can broaden the context of the app's performance by providing insights into the company's overall operations and market strategy. After, a verification process can follow where data acquisition subsystem 700 checks (e.g. “Is company name in database?”) to determine whether the company is already being tracked. If the check returns negative, indicating the company is not present in the database, data acquisition subsystem 700 proceeds to insert into a companies table in the database. This action adds the new company to the database, thus beginning the process of monitoring and collecting data over time. By incorporating app performance metrics from both major app stores and directly linking them to the operational data of the companies, this subsystem provides a data- rich foundation for assessing the market traction and growth potential of early-stage companies, which be helpful for making informed venture capital investment decisions.
[0065] FIG. 8 illustrates an example of a data acquisition subsystem 800 focused on new product information of an entity, according to an embodiment. Data acquisition subsystem 800 begins by monitoring new products in a server (e.g. “New product in Product Hunt”). For example, Product Hunt is a website that features new startups and tech products. To ensure the timely capture of data, data acquisition subsystem 800 employs RSS feeds, which are particularly adept at real-time updates, allowing for the immediate detection of product listings as they are published.
[0066] Upon identifying a new product, the crawl product information operation is executed. This operation is meticulously designed to extract a comprehensive set of data points from the product listing. The information crawled can include the product name, which provides a direct identifier for the offering; a short description, offering a concise summary of the product; a long description, which provides an in-depth look at the product's features, use cases, and benefits; and details regarding the “teams” behind the product, including team size, key personnel, and their roles. This breadth of information provides a holistic view of the product and the company's ability to articulate and market their value proposition, as well as the strength and composition of the team responsible for the product's development — factors that are often indicative of a company's potential for success.
[0067] Finally, the data extracted can be inserted into a products table in the database. This database table can be structured to store detailed records of each product, maintaining a historical dataset that can be analyzed for patterns indicative of company growth, market acceptance, and innovation trends.
[0068] FIG. 9 illustrates an example of a data acquisition subsystem 900 focused on news information of an entity, according to an embodiment. Commencing with the “news” component, data acquisition subsystem 900 can be programmed to monitor and capture news articles using RSS feeds. This enables the system to swiftly identify news pieces concerning startups as they are released, ensuring an up-to-date stream of information. Data acquisition subsystem 900 is fine-tuned to recognize keywords associated with startup activity, such as "start-up" and "khdi nghiep," which are indicative of relevant content for the target analysis.
[0069] Once news related to startups is detected, the crawl news content process is initiated, which involves the systematic retrieval of the content of news articles. This content provides a rich source of information, containing updates on startup achievements, funding rounds, product launches, and other newsworthy events. Following the retrieval of content, the system employs an “Extract Company Name” procedure. This procedure utilizes entity extraction algorithms to parse the news content and accurately identify company names mentioned within the text.
[0070] Subsequent to the extraction of company names, the data is saved to a news table in the database. This action creates a record of each news item associated with startups, contributing to a dataset that can be analyzed to discern industry trends, market sentiment, andstartup activity levels. Simultaneously, the crawl company information step can take place. This step further delves into the details of the identified companies, gathering more in-depth information that can include the company's market segment, operational status, leadership team, and historical performance. Data acquisition subsystem 900 then queries the database with to question if the company name is in the database (e.g. “Is company name in database?”). If the company is not already present, signifying a new entity within the scope of data acquisition subsystem’s 900 monitoring, the process moves to insert into the companies table in the database. This inclusion ensures that newly surfaced companies are tracked and their progress is recorded for future analysis.
[0071] FIG. 10 illustrates an example of an analytical tool dashboard 1000 used in system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. Analytical tool dashboard 1000 serves as an analytical tool, providing real-time insights into the frequency and volume of data harvested from various sources that are integral to the identification process. The left section of analytical tool dashboard 1000 displays a breakdown of data types 1010 along with their statuses. Each source name corresponds to a source from which system for data-driven identification of entities in early-stage venture capital 400 collects information. The statuses, including “active,” “Added to pipeline,” and “Rejected,” represent the current engagement of the venture capital firm with the identified opportunities, indicating the flow and outcomes of the deals.
[0072] The table within data types 1010 quantifies the number of leads from each source that are currently active, those that have been added to the investment pipeline, and those that have been rejected. This tabulation aids monitoring of the performance and relevance of each data source in contributing to the pipeline of potential investments. On the right, graph 1020 visualizes the aggregate lead created by date and type, offering a temporal perspective of data accumulation from each source over a given period. The peaks and troughs in graph 1020 signal the activity volume, with an evident spike representing a significant influx of data at a particular point in time.
[0073] This graphical representation in graph 1020 allows for the immediate assessment of data consistency and reliability across sources, as well as the identification of anomalies or trends in data acquisition, for example, as the pronounced spike observed in November 2023. Such insights can direct attention to potential market movements, seasonal trends, or the effectiveness of data crawling algorithms. Collectively, dashboard 1000 facilitates continuousmonitoring and evaluation of the data sources, ensuring that the system remains robust and that the data input streams are operational and yielding valuable information. This level of oversight can help maintain the integrity of the VC firm's deal flow and to make informed decisions about where to focus efforts for sourcing early-stage investment opportunities.
[0074] FIG. 11 illustrates an infrastructure of an aggregator module 420, according to an embodiment. As previously discussed in FIGS. 5-9, data source module 410 gathers information from various streams including jobs, website ranking, App ranking, and / or news. Each of these data sources in data source module 410 provide a different facet of information on companies, which when combined, offer a comprehensive view of a company's market presence and potential for growth. Building upon the initial data gathering framework, system for data-driven identification of entities in early-stage venture capital 400 is designed to continuously evolve by integrating new data sources, thereby enhancing its analytical capabilities and ensuring relevance in a rapidly changing market environment. The process of adding new data sources operates through two primary mechanisms.
[0075] Firstly, any novel deals or investment opportunities identified by an investment team, which are not yet recorded in system for data-driven identification of entities in early- stage venture capital 400, serve as a trigger for further investigation. This involves tracing the origins of such information back to the investment team to pinpoint the new data source. This method ensures that all pertinent data influencing investment decisions is captured and integrated into system for data-driven identification of entities in early-stage venture capital 400. Secondly, an investment team can actively collaborate with identified data partners to assess and integrate new data streams in system for data-driven identification of entities in early-stage venture capital 400. This collaboration can enhance the discovering of data sources that align with any given investment objectives, providing a targeted approach to data collection. Upon identifying a promising new data source, a dedicated programming team can step in to analyze its structure and relevance. Subsequently, they can develop and deploy specialized crawlers tasked with the periodic retrieval of this data. These crawlers are programmed to efficiently harvest data without compromising the integrity and performance of the data source. By incorporating these new streams, system for data-driven identification of entities in early-stage venture capital 400 continually adapts to include the most relevant and impactful data, thereby refining our understanding of market dynamics and enhancing our investment strategies.
[0076] The objective of the aggregator module 420 is to centralize the information related to a company into one single identifier, making it easier to manage and reference against the additional information in the source data. This is achieved through a multi-stage matching process to ensure the accuracy and integrity of the company profiles. The initial step in this matching process is rule-base matching stage 1110, which employs a set of predetermined rules, such as the consistency of company names and website URLs, to automatically link data points to the corresponding company profiles.
[0077] Following this, fuzzy matching stage 1120 is applied to instances where rule-based matching stage 1110 is inconclusive. Fuzzy matching stage 1120 method utilizes an algorithmic approach to measure the similarity between data points, automatically associating them with company profiles if they achieve a similarity score greater than 90%. This threshold ensures high confidence in the automated matches while allowing for minor discrepancies in the data. Fuzzy matching stage 1120 stands as a pivotal technique in data preparation and analysis, particularly within the domain of systems for the identification of early-stage venture capital investment opportunities. Fuzzy matching stage 1120 plays an integral role in discerning and associating disparate data points to a singular company entity, which can be helpful for the consolidation of information within system for data-driven identification of entities in early-stage venture capital 400.
[0078] At the core of fuzzy matching stage 1120 is the generation of a similarity score between two strings of text. This score is a quantifiable measure of the likeness between the strings, encapsulating various dimensions of similarity. The methodology accounts for character overlap, which evaluates the strings based on the number and sequence of shared characters. Edit distance, often associated with the Levenshtein Distance algorithm, gauges the number of insertions, deletions, or substitutions required to transform one string into another. Additionally, phonetic similarity be assessed, which is particularly useful for matching strings that may not be identically spelled but sound alike, thus accounting for variances in linguistic representation.
[0079] In cases where fuzzy matching stage 1120 yields a score below the confidence threshold, the process advances to manually matching stage 1130. During manually matching stage 1130, users are presented with suggestions for potential matches, prioritized by the highest fuzzy match score from fuzzy matching stage 1120. This allows for human discernment in cases where the algorithm is uncertain, thereby providing a safeguard against erroneousmatches. Upon successful identification, whether through rule-based matching stage 1110, fuzzy matching stage 1120, or manual matching stage 1130, the data is then associated with a single company id 1140. This unique identifier acts as the linchpin in system for data-driven identification of entities in early-stage venture capital 400, enabling users to quickly and efficiently access a centralized repository of information for each company. Aggregator module 420 allows system for data-driven identification of entities in early-stage venture capital 400 ability to provide a unified and accurate database of company profiles, which is critical for the effective evaluation and tracking of potential investment opportunities in early- stage companies.
[0080] FIG. 12 illustrates an example of a rule-based matching stage 1110 in a system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. As described in FIG. 11, the initial step in this matching process is rule-base matching stage 1110, which employs a set of predetermined rules, such as the consistency of company names and website URLs, to automatically link data points to the corresponding company profiles. FIG. exemplifies the different rules that dictate the filtering through rulebase matching stage 1110.
[0081] FIG. 13 illustrates an example of a fuzzy matching stage 1120 in a system for data- driven identification of entities in early-stage venture capital 400, according to an embodiment. As described in FIG. 11, fuzzy matching stage 1120 is the generation of a similarity score between two strings of text. This score is a quantifiable measure of the likeness between the strings, encapsulating various dimensions of similarity. The methodology accounts for character overlap, which evaluates the strings based on the number and sequence of shared characters. Edit distance, often associated with the Levenshtein Distance algorithm, gauges the number of insertions, deletions, or substitutions required to transform one string into another. FIG. 13 exemplifies fuzzy matching stage 1120 process results.
[0082] FIG. 14 illustrates a decision-making framework within a system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. The process initiates in aggregator module 420, where data concerning a company is centralized into a single identifier. This data encompasses a wide array of metrics, including company fundamentals, website description, business model description, positive signals like job postings, app rankings, website traffic, and other pertinent indicators that together offer a comprehensive profile of the company.
[0083] This consolidated data in aggregator module 420 is then transmitted to decision engine 430, which initiates the evaluation process with rule-based filtering. This filtering applies a dynamically updated set of predefined criteria to systematically exclude companies that fall outside the investment scope of the venture capital firm. The criteria for exclusion, which are regularly refined and expanded upon in review meetings, include state organizations, gambling sites, large corporations, financial institutions, video game companies, adult content producers, religious organizations, nonprofit entities, agencies, consulting, and outsourcing companies. Such rigorous and regularly updated filtering criteria are helpful for maintaining the investment firm's focus on its core objective of identifying emerging and scalable ventures. If a company's profile does not pass this initial rule-based filtering stage 1110 — meaning it falls into one of the excluded categories — it is promptly set to disapprove, and no further analysis is conducted on this entity.
[0084] Conversely, for profiles that successfully navigate through the rule-based filter, the process advances to the LLMs model in decision engine 430. Traditional models, such as supervised learning algorithms and neural networks, typically require extensive labeled datasets and feature explicit programming of rules for data processing. In contrast, the LLM used in this invention is an advanced form of generative artificial intelligence that leverages deep learning techniques to understand human-like text based on the input it receives. The model is pre-trained on a diverse range of internet text and fine-tuned in venture capital area, enabling it to generate accurate, context-relevant content with minimal supervision. This sophisticated Al-driven model assesses the company profile in-depth, leveraging vast amounts of data and advanced pattern recognition capabilities to predict the potential success and growth trajectory of the company.
[0085] The technical advantages of integrating a LLM in system for data-driven identification of entities in early-stage venture capital 400 are manifold. Firstly, the LLM exhibits exceptional adaptability and flexibility; unlike traditional models that necessitate retraining for new data or applications, the LLM can dynamically adjust to varied inputs without additional programming of system for data-driven identification of entities in early- stage venture capital 400. This capability significantly diminishes the time and resources required for model updates and maintenance. Secondly, the LLM enhances data handling by efficiently understanding and generating responses from large volumes of unstructured data — a considerable improvement over traditional models, which typically struggle with such dataformats and require extensive preprocessing. Finally, the LLM's enhanced learning capabilities allow our invention to benefit from the model's continual learning and improvement based on new data inputs. This ongoing learning process starkly contrasts with traditional models, which remain static in their learning capabilities once deployed unless explicitly updated.
[0086] The model result in reviewer module 440 is the output of the LLMs model's analysis from decision engine 430, which provides a decision on which company to engage with and provides a rationale for each selection. This result is critical as it forms the basis for the final decision-making phase.
[0087] Reviewer module 440 is the concluding stage of the process. Here, human experts analyze decision engine 430 results, adding layers of strategic and subjective assessment to the data-driven findings of the LLMs model from decision engine 430. This human-in-the-loop approach ensures that the final investment decisions are informed not only by quantitative analysis but also by industry knowledge, investor experience, and nuanced judgment.
[0088] FIG. 15 illustrates a large language model module 1500 in a system for data-driven identification of entities in early-stage venture capital 400, according to an embodiment. Decision engine 430 is fueled by two primary streams of input data: Company information from aggregator module 420 and historical decision in database 1510. The company information from aggregator module 420 consists of a vectorized representation of current company data, which includes quantitative and qualitative aspects such as financial health, market positioning, and growth metrics. This vectorization facilitates a machine-readable format for comparison and analysis.
[0089] To augment the analysis, the engine leverages a decision of similar companies, which utilizes cosine similarity measures to draw parallels between the prospective investment company and those previously evaluated. This comparative analysis is critical, as it provides a contextual backdrop for the LLM by highlighting historical investment outcomes in similar cases. The LLM model stands at the core of decision engine module 430, operating on a complex prompt structure that integrates several components:Instruction: A specific directive that outlines the analytical task for the model, such as evaluating investment potential or growth prospects.Context: This includes the decision data regarding similar companies, providing the LLM with reference points that enrich its predictive capabilities.Input Data: Detailed company information that has been processed by the Aggregator, serving as the direct subject matter for the model's analysis.Thesis: Here, the model is informed of the investment thesis, which may include criteria such as innovation, scalability, and market disruption potential.Industry Context: Expert insights into the industry sector of focus are provided to give the LLM a deeper understanding of market dynamics and sector-specific considerations.Output Indicator: This defines the expected format and structure of the model's output, ensuring that the results are compatible with subsequent review and action steps.
[0090] To ensure that the LLM remains closely aligned with the latest market conditions and strategic considerations, system for data-driven identification of entities in early-stage venture capital 400 has a rigorous process for continuously updating and refining information, particularly regarding the investment thesis and industry context. This process includes an ongoing research initiative where system for data-driven identification of entities in early-stage venture capital 400 or a team can actively review recent industry publications, white papers, and analytical reports. Additionally, system for data-driven identification of entities in early- stage venture capital 400 can enhance understanding through direct interactions with industry experts, founders, and key stakeholders. These engagements offer a firsthand, nuanced perspective on emerging trends, challenges, and opportunities across various industries, critically informing system for data-driven identification of entities in early-stage venture capital’s 400 investment thesis refinement. Insights are methodically evaluated and integrated, ensuring they meet a given strategic investment criteria, which focus on innovation potential, scalability of the business models, and the capacity for market disruption.
[0091] Once this updated information is curated and validated, it is methodically integrated into the LLM model's prompt under the thesis' and industry context components. This integration not only enhances the model's predictive accuracy but also ensures that the assessments it generates are reflective of the latest industry landscapes and investment paradigms. Consequently, this dynamic updating process maintains the relevance and efficacy of decision engine module 430, empowering it to deliver robust, informed, and timely investment analyses.
[0092] System for data-driven identification of entities in early-stage venture capital 400 as depicted integrates reviewer module 440 that enhances the predictive accuracy of the LLM models within decision engine module 430. Following the generation of initial model resultsin reviewer module 440 by the LLM models, these outputs undergo a meticulous review by a panel of investment experts. These experts can possess extensive knowledge of the venture capital industry and understand complex market dynamics intricately. Their primary responsibility is to evaluate the initial model results against actual investment outcomes and grounded expert insights, pinpointing discrepancies, biases, and precision gaps in the model's predictions to.
[0093] After this initial review, the experts also provide their independent decision inputs, which are then compared with the initial model results. This comparative analysis can help for identifying areas where the model's predictions diverge from expert decisions. Based on this comparison, the reviewers make detailed annotations and adjustments, which culminate in revised model results by reviewer 1520. This output incorporates corrections for any identified biases, enhancements in interpreting complex data points, and adjustments informed by emerging trends that were not initially captured by the model.
[0094] These meticulously revised results are then reintegrated into system for data-driven identification of entities in early-stage venture capital 400 through a finetune process, which involves technically sophisticated methods of retraining or recalibrating the LLM models. This recalibration is based on the enriched dataset comprising original model outputs, expert decisions, and the subsequent revisions. The fine-tuning process leverages advanced machine learning techniques, potentially including transfer learning and gradient descent adjustments, to align the model more closely with the nuanced realities of venture capital decision-making.
[0095] Through this iterative process of comparison, revision, and refinement, the LLM models evolve to reflect a deeper and more accurate understanding of the investment landscape. This continuous feedback loop significantly enhances the models' sensitivity to the subtleties of venture capital decisions, thereby improving both the accuracy and reliability of the decision-making process within the venture capital firm. The integration of both quantitative corrections and qualitative insights ensures that system for data-driven identification of entities in early-stage venture capital 400 adapts over time, maintaining relevance and efficacy in a dynamic market environment.
[0096] The resultant data flow, from company information aggregation to the application of historical decision context in database 1510, and the iterative refinement of the LLM model, constitutes a dynamic and self-improving decision-making framework. This framework is notstatic; it evolves with each investment analysis cycle, thereby increasing the precision and reliability of the system in identifying promising early-stage companies for venture capital investments in system for data-driven identification of entities in early-stage venture capital 400.
[0097] FIG. 16 illustrates an example of a reviewer module 440 interface dashboard 1600, according to an embodiment. Interface dashboard 1600 is designed with a clear and uncluttered layout to facilitate the review process in reviewer module 440. Interface dashboard 1600 presents a table where each row represents a different company. For each company, the following information is displayed:Action: Buttons such as 'Reject' and 'Add Positive' allow the reviewer to take immediate action based on their assessment of the company.Title: The name of the company or the product is prominently displayed for easy identification.Type: This field indicates the company's industry or the type of product it offers, providing quick context at a glance.Short Description: A brief overview of the company's operations or product offering is provided. This concise format ensures reviewers are not overwhelmed with information upfront. Should the reviewer need more details, they have the option to delve deeper into a full profile by selecting the specific company entry.Iggy Reason: This column articulates the rationale behind the model's decision. It encapsulates the reasoning from the LLM model, offering insights into why the company was preliminarily approved or rejected.Iggy Decision: The decision output from the model is shown here, with a binary indicator where T' signifies an approval recommendation by the model.
[0098] Interface dashboard 1600 is pivotal in streamlining the review process. It allows the reviewer to quickly gauge the model's preliminary assessment and the reasons behind it, providing a basis for their expert evaluation. This initial filtering by the model enables reviewers to focus their attention on potentially suitable investment opportunities while allowing for the swift dismissal of those that do not meet the required criteria.
[0099] FIG. 17 illustrates an example of a user interface 1700 to collect reviewer’s feedback; according to an embodiment. User interface 1700 comprises two primary components: a dropdown menu labeled “Reason” and a text input field labeled “Comment.” The “Reason” dropdown menu, as shown to the left of user interface 1700, allows users to select from a list of pre-defined reasons that delineate why they accept or reject a model'soutput. These predefined reasons are periodically updated based on operational feedback and discussions in weekly working group meetings. Adjacent to this, the “Comment” section offers a space for users to articulate detailed thoughts, opinions, or arguments supporting their decision. This dual-input mechanism ensures that each submission is accompanied by both a categorical reason and a qualitative comment, enhancing the data's utility for subsequent training and refinement of LLM in system for data-driven identification of entities in early- stage venture capital 400. Additionally, user interface 1700 incorporates a system to generate statistical reports on user submissions, aimed at identifying and addressing patterns of poor or repetitive data entry, thereby maintaining the integrity and effectiveness of the feedback process. Buttons labeled “Cancel” and “Submit” enable users to either abandon or complete their feedback submission, respectively. Some examples of reviewer’s feedback:
[0100] FIG. 18 illustrates an example of a dashboard with priority results 1800 of recommendations for entities in early-stage venture capital. Dashboard with priority results 1800 is meticulously designed to facilitate a detailed comparison and analysis of each company that has been processed by the decision engine. Dashboard with priority results 1800 includes a set of filters at the top for “Type,” which allows selection among sources ensuring that the data can be segmented and reviewed according to source type and the individual responsible for the review. Each row in the “Deal detail” section of the dashboard represents a unique company, displaying:ID and Title: These fields provide a unique identifier for the company and its name or the product title, respectively.Website: This column lists the company's web address.Short Description: A succinct description of the company’s main business or service offering is given here for quick reference.Comment: Reviewers can add notes here for context or reminders for future reviews.Iggy Decision: The decision suggested by the model is indicated here — 'approve,' 'active,' 'unsure,' or left blank if no decision could be determined.Human Decision: After reviewing, the human reviewer's decision is recorded, which could be 'active,' 'added to pipeline,' or 'rejected.'Iggy Reason: This important column elucidates the rationale provided by the model for its decision.Pipeline at date: This field indicates when the company was added to the pipeline, providing a temporal context for the review process.
[0101] The contrasting fields of “Human decision” and “Iggy decision” are central to this interface, as they allow for a direct comparison between the model's output and the reviewer's judgment. Discrepancies between these decisions are critical data points, as they highlight instances where the model's assessment may not align with the human reviewer's expertise and understanding.
[0102] When such differences are identified, they prompt a careful review of the model's prompt structure and the resulting output. The insights gained from this comparative analysis are then used to revise the model's prompts and outputs, which subsequently serve as refined data for fine-tuning the model. This iterative process of review, comparison, and refinement enhances the accuracy and reliability of the model, ensuring that it continuously evolves to meet the complex decision-making requirements of venture capital investment evaluations.
[0103] Furthermore, while the processes, described herein, are illustrated with a certain arrangement and ordering of subprocesses, each process may be implemented with fewer, more, or different subprocesses and a different arrangement and / or ordering of subprocesses. In addition, it should be understood that any subprocess, which does not depend on the completion of another subprocess, may be executed before, after, or in parallel with that other independent subprocess, even if the subprocesses are described or illustrated in a particular order.
[0104] The above description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles described herein can be applied to other embodiments without departing from the spirit or scope of the invention. Thus, it is to be understood that the description and drawings presented herein represent a presently preferred embodiment of the invention and are therefore representative of the subject matter which is broadly contemplated by the present invention. It is further understood that the scope of the present invention fully encompasses other embodiments that may become obvious to those skilled in the art and that the scope of the present invention is accordingly not limited.
[0105] As used herein, the terms “comprising,” “comprise,” and “comprises” are open- ended. For instance, “A comprises B” means that A may include either: (i) only B; or (ii) B in combination with one or a plurality, and potentially any number, of other components. In contrast, the terms “consisting of,” “consist of,” and “consists of’ are closed-ended. For instance, “A consists of B” means that A only includes B with no other component in the same context.
[0106] Combinations, described herein, such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more ofA, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A andB, A and C, B and C, or A and B and C, and any such combination may contain one or more members of its constituents A, B, and / or C. For example, a combination of A and B may comprise one A and multiple B’s, multiple A’s and one B, or multiple A’s and multiple B’s.
Claims
CLAIMSWhat is claimed is:
1. A non-transitory computer-readable medium having instructions stored to identify entities to receive a resource, wherein the instructions, when executed by a processor, cause the processor to: obtain data information of entities from more than one source; consolidate disparate data points into unified company profiles using rule-based matching, fuzzy matching, and manual matching; apply a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; provide the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subject the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and use the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
2. The processor of claim 1, wherein the one or more sources include job postings, website traffic, mobile application metrics, news, or and other indicators that together offer a comprehensive profile of the entities.
3. The processor of claim 1, wherein the processor obtains data information by initiating a crawl company information sequence, wherein the processor employs web crawling technologies to gather data information about the early-stage entities associated with the early - stage entities.
4. The processor of claim 1, wherein the sources can be assessed and integrate new source streams.
5. The processor of claim 1, wherein the fuzzy matching comprises: measuring the similarity between data points; and automatically associating the data points with an entity if the data points achieve a similarity score of 90% or more.
6. The processor of claim 5, wherein a manual matching is available for data points with a similarity score of 89% or less.
7. The processor of claim 1, wherein the entity is an early-stage company.
8. The processor of claim 1, wherein the rule-based matching includes matching company names and website URLs to automatically link data points to the corresponding company profiles.
9. The processor of claim 1, wherein the rule-based filtering excludes entities that fall outside the investment scope of a venture capital firm.
10. The processor of claim 9, wherein the rule-based filtering excludes state organizations, gambling sites, large corporations, financial institutions, video game companies, adult content producers, religious organizations, nonprofit entities, agencies, consulting, and outsourcing companies.
11. The processor of claim 1, wherein the large language model module is driven by entity information and a historical results database.
12. The processor of claim 1, further comprising a recalibration process driven by a dataset including original model outputs, expert decisions, and subsequent revisions.
13. The processor of claim 1, wherein the recommendations by the panel of investment experts comprises a user interface with a predefined answer dropdown and a text field; wherein the predefined answer dropdown is configured to allow the panel of investment experts to select from a list of pre-defined reasons that delineate why they accept or reject a model's output, wherein the text field is configured to allow the panel of investment experts to articulate detailed comments.
14. A method comprising using at least one hardware processor to: obtaining data information of entities from more than one source; consolidating disparate data points into unified company profiles using rule-based matching, fuzzy matching, and manual matching; applying a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; providing the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subjecting the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and using the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
15. The method of claim 14, wherein the one or more sources include job postings, website traffic, mobile application metrics, news, or and other indicators that together offer a comprehensive profile of the entities.
16. The method of claim 14, wherein the rule-based matching includes matching company names and website URLs to automatically link data points to the corresponding company profiles.
17. The method of claim 14, wherein the recommendations by the panel of investment experts comprises a user interface with a predefined answer dropdown and a text field; wherein the predefined answer dropdown is configured to allow the panel of investment experts to select from a list of pre-defined reasons that delineate why they accept or reject a model's output, wherein the text field is configured to allow the panel of investment experts to articulate detailed comments.
18. A system to identify entities for venture capital investment, the system comprising: at least one hardware processor; and software that is configured to, when executed by the at least one hardware processor, obtain data information of entities from more than one source; consolidate disparate data points into unified company profiles using rule-based matching, fuzzy matching, and manual matching; apply a rule-based filtering to the consolidated data to rule out irrelevant entities and produce a list of candidate entities; provide the list of candidate entities to a large language model module configured to perform an analysis and produce a recommendation to prioritize resource allocation to the candidate entities; subject the recommendations to a review by a panel of investment experts to provide a decision on whether to follow the recommendations; and use the decision from the panel of investment experts to train and improve the large language model module for future recommendations.
19. The system of claim 18, wherein the rule-based matching includes matching company names and website URLs to automatically link data points to the corresponding company profiles.
20. The system of claim 18, wherein the recommendations by the panel of investment experts comprises a user interface with a predefined answer dropdown and a text field; wherein the predefined answer dropdown is configured to allow the panel of investment experts to select from a list of pre-defined reasons that delineate why they accept or reject a model's output, wherein the text field is configured to allow the panel of investment experts to articulate detailed comments.
Citation Information
Patent Citations
Venture capital investment information tracking and collection method
CN106709806A
Fund investment risk control method and device, medium and terminal
CN116485552A
Computing system for adaptive investment recommendations
US11803915B1
Methods and systems for generating composite index using social media sourced data and sentiment analysis
US20120296845A1
System and method for screening entities using multi-level rules and financial information
US20210012424A1