An information management system and method
The information management system addresses the inefficiencies of existing search engines by using a taxonomy-based approach with cascaded relevancy and categorization modules, ensuring accurate and secure data delivery, thereby enhancing decision-making and data accessibility.
Patent Information
- Application Number
- PCT/IN2025/051088
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-26
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing database search engines face challenges with low relevancy rates, high processing costs, and complexity in delivering accurate and efficient search results, particularly in patent search scenarios, due to the reliance on classification systems and semantic analysis, which often require specialized knowledge and substantial computational resources.
An information management system comprising a delivery subsystem, curation subsystem, and database, utilizing a taxonomy library and cascaded relevancy and categorization modules to ensure accurate, relevant, and well-organized information delivery, with role-based access control and secure memory resources.
The system enables efficient, secure, and accurate organization and delivery of data, improving decision-making and data accessibility by reducing latency and computational costs, while maintaining high accuracy and adaptability across different domains.
Smart Images

Figure IN2025051088_29012026_PF_FP_ABST
Abstract
Description
AN INFORMATION MANAGEMENT SYSTEM AND METHOD DESCRIPTIONFIELD OF THE INVENTION
[0001] The present invention relates to the field of data processing technologies, and more particularly to an information management system and method designed for efficient, secure, and accurate organization, curation, and delivery of data, thereby enhancing decision-making, collaboration, and data accessibility across organizations.BACKGROUNDINTERPRETATION CONSIDERATIONS
[0002] This section describes the technical field in detail and discusses problems encountered in the technical field. Therefore, statements in the section are not to be construed as prior art.DISCUSSION
[0003] In today's digital landscape, database search engines are extensively used for research and document exploration within databases. These engines typically rely on a matching strategy, comparing user-provided search terms, such as keywords, document titles, or numbers, to locate relevant documents. While this method ensures that all documents containing the search terms are displayed, it often results in a low relevancy rate among the listed documents, forcing users to sift through numerous irrelevant results.
[0004] The backend of these search engines translates user queries into machine-level instructions, requiring specific knowledge of the database's structure, network environment, server architecture, or other interaction elements of the network architecture. This translation process usually involves SQL queries or complex machine learning algorithms, demanding specialized knowledge of the database and network systems.
[0005] In patent search scenarios, database search engines may utilize classification systems such as the Cooperative Patent Classification (CPC) or the International Patent Classification (IPC) to organize and retrieve patents efficiently. While these classifications standardize and expedite document retrieval, they introduce complexity and may be time-consuming for users unfamiliar with the systems.
[0006] Alternatively, some database search engines employ semantic analysis, utilizing word-sense disambiguation to interpret the meaning of queries and map them to relevant categories. Although semantic analysis reduces the risk of missed results, it struggles with complex queries, often requiring significant processing on the front end. This negatively impacts the speed, accuracy, and delivery of search results.
[0007] Both classification-based and semantic-based approaches generate an overwhelming number of relevant documents, making it challenging to analyze and effectively display the results. Furthermore, the processing time required for these methods is often substantial, leading to higher power consumption and a poor user experience.
[0008] Processing costs for search engines continue to rise, especially for Al-powered products that require real-time, high-accuracy results. Predicting processing requirements in advance is difficult due to fluctuating user activity, and while loadbalancing tools exist, high costs and uncertainty about active user numbers remain significant concerns. Limiting search results to control costs may lead to incomplete analytics and ambiguous decision-making.
[0009] PRIOR ART
[0010] One such system and method, disclosed in the CN publication 115221881A, relates to a text processing method and device, a storage medium, and electronic equipment. The method introduces text processing that uses word vector sets to identify entities corresponding to a target organization name, enhancing processing efficiency and enabling more accurate and faster search results by leveraging the semantic relationships captured in the vectors. While this vector-based approach improves theretrieval of relevant information even for short or ambiguous organization names, it also introduces significant challenges: the implementation is complex, requiring specialized expertise and substantial computational resources, the cost of storing and querying high-dimensional vectors is high, and the process of deploying and fine- tuning such systems is time-consuming. Additionally, accuracy is affected by the quality of embedding and the context-dependence of vector models, making it difficult to guarantee consistent performance across different tasks and domains.
[0011] Another US publication, 10942948 B2, relates to a cloud-based pluggable classification system and method. A source system sends a text term to be classified over a communication network to the classification system, which includes multiple classifiers. Upon receiving the request, the classification system accesses a rule associated with the text term, which specifies which classifier or classifiers to use from the available set. The selected classifier then generates classification information by mapping the text term to a category within a taxonomy, and this information is sent back to the source system via the network. However, this approach presents several problems, including accuracy concerns due to the reliance on predefined rules and classifier selection, increased operational costs associated with maintaining and scaling multiple classifiers in the cloud, the system complexity from managing and updating classifier rules and integrations, and potential delays or time-consuming processes due to network latency and the overhead of coordinating multiple components in real time.
[0012] Therefore, there is a need for a system and method that can deliver highly accurate outputs suitable for informed decision-making without requiring additional processing, while also maintaining reasonable costs, reducing latency, and significantly improving the user experience.SUMMARY
[0013] An objective of the present invention is to provide an information management system and method that enables efficient, secure, and accurate organization, curation,and delivery of data, thereby improving decision-making, collaboration, and data accessibility across the organization.
[0014] Another objective of the present invention is to facilitate effective, safe, and precise management, refinement, and distribution of data, ensuring that data is handled with high accuracy, security, and reliability throughout its lifecycle to support robust organizational operations and informed decision-making.
[0015] This and other objectives are achieved by providing an information management system and method as defined in the features of the independent claims. Additional advantageous embodiments and improvements of the invention are listed in the dependent claims. The use of expressions like “... aspect according to the invention” or “in one embodiment” or similar terminology is intended to refer to examples or embodiments consistent with the broadest scope of the invention as defined by the independent claims.
[0016] According to a first aspect, the present invention discloses an information management system comprising a delivery subsystem, a curation subsystem, and a database. The database is connected with the delivery subsystem and the curation subsystem for storing a taxonomy library. The delivery subsystem retrieves a taxonomy from the taxonomy library and a portion of curated data corresponding to the retrieved taxonomy from the database. Further, the delivery subsystem provides the retrieved taxonomy and the portion of the curated data. This integrated system ensures users receive accurate, relevant, and well-organized information efficiently.
[0017] In an embodiment of the present invention, an interface 2 receives and directs input to the delivery subsystem, enabling efficient processing and retrieval of relevant information.
[0018] In another embodiment of the present invention, the delivery subsystem retrieves a taxonomy from the taxonomy library stored in the database based on the input from the interface 2, ensuring that the information provided is tailored to the user's specific input.
[0019] In another embodiment of the present invention, the delivery subsystem is connected to the database comprising a shared memory resource allocated to a shared portion of retrieved taxonomy and a shared portion of curated data, and at least one private memory resource allocated to a private portion of retrieved taxonomy and a private portion of curated data. The shared and private memory resources ensure both collaborative and private access to data.
[0020] In yet another embodiment of the present invention, the system comprises an access validation subsystem to assign the at least one shared memory resource or private memory resource to a user, ensuring appropriate role-based access control and data security based on user permissions.
[0021] In yet another embodiment of the present invention, the delivery subsystem provides, through the interface 2, a selective combination of the shared and private portions of curated data and a selective combination of the shared and private portions of retrieved taxonomy based on the assignment by the access validation subsystem, thereby enabling personalized and secure information delivery tailored to each user’s role-based access rights.
[0022] In yet another embodiment of the present invention, the taxonomy comprises a hierarchical structure and relationship data of multiple levels of the hierarchy, enabling an organized and meaningful categorization of information for a defined purpose.
[0023] According to a second aspect of the present invention, the present invention discloses a method for information management. The method comprises the steps of: a) receiving a dataset through an interface 1 ; b) selecting a taxonomy from a taxonomy library; c) implementing a relevancy module in a cascaded structure to determine a relevancy level and a confidence score of each data point of the dataset based on the selected taxonomy; d) sorting a data point based on the determined relevancy level and the confidence score; and e) implementing a categorization module in a cascaded structure to classify the sorted data point based on the selected taxonomy; wherein the categorization module is selected based on the determined relevancy level and the confidence score.
[0024] This method ensures efficient, attribute-driven organization and classification of information, enabling an integrated system to deliver accurate, relevant, and well- organized information to the users with optimal efficiency.
[0025] In an embodiment of the present invention, implementing the relevancy module in a cascaded structure comprises processing the data point selectively through a plurality of connected relevancy identifiers. This method enhances the precision and adaptability of relevancy assessment for each data point.
[0026] In another embodiment of the present invention, processing the data point selectively through the plurality of connected relevancy identifiers comprises processing each data point through a first relevancy identifier to determine the relevancy level and the confidence score and finalizing the relevancy level or processing the data point through a second relevancy identifier based on the determined relevancy level and the confidence score. This method ensures a more accurate and dynamic determination of relevancy levels for improved data selection.
[0027] In still another embodiment of the present invention, the plurality of connected relevancy identifiers exchanges the determined relevancy level and the confidence score of each data point or stores the determined relevancy level and the confidence score in the database to generate an automated training dataset. This facilitates continuous learning and improvement through the generation of high-quality training datasets.
[0028] In still another embodiment of the present invention, implementing the categorization module in a cascaded structure comprises processing the data point through a multi-class categorization classifier to determine a tentative category and associated confidence score and processing the data point selectively through a plurality of connected single-class categorization classifiers based on the tentative category and the associated confidence score to determine a finalized category and the confidence score. This method increases the accuracy and confidence of data categorization outcomes.
[0029] In yet another embodiment of the present invention, implementing the categorization module in a cascaded structure comprises generating and exchanging feedback value by interconnected single-class categorization classifiers or multi-class categorization classifiers to perform automated reinforcement learning or feedforward learning. This enables the method to refine and optimize its categorization performance through continuous feedback-driven learning.
[0030] In still another embodiment of the present invention, the multi-class categorization classifier and the plurality of connected single-class categorization classifiers exchange the determined category and the confidence score with each other or store the determined category and the confidence score in the database to generate an automated training dataset. This supports ongoing model enhancement and accuracy through the accumulation of a valuable training dataset.
[0031] In yet another embodiment of the present invention, the method is implemented in a curation subsystem. This ensures that the dataset curation processes are systematic, scalable, and easily integrated within the overall system.
[0032] In still another embodiment of the present invention, the relevancy identifier, the multi-class categorization classifier, or the single-class categorization classifier includes a rule engine, artificial intelligence (Al) or machine learning models, a statistical model, a heuristic model, natural language processing and understanding engines, artificial intelligence (Al) generative artificial intelligence (Al) models or an ensemble model. By incorporating rule-based logic, advanced artificial intelligence or machine learning algorithms, statistical analysis, heuristic approaches, or natural language processing and understanding capabilities, the method is tailored to handle diverse types of datasets and complex classification tasks. Additionally, the use of ensemble models enables the combination of multiple techniques to improve accuracy and robustness. This versatility ensures that the information management method is optimized for different domains, dataset characteristics, and evolving organizational requirements.BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Various aspects, as well as embodiments of the present invention, are better understood by referring to the following detailed description. To better understand the invention, the detailed description should be read in conjunction with the drawings.
[0034] FIG. 1(A) illustrates an information management system in accordance with an embodiment of the present invention;
[0035] FIG. 1(B) illustrates an information management system in accordance with another exemplary embodiment of the present invention;
[0036] FIG. 1(C) illustrates an information management system in accordance with another exemplary embodiment of the present invention;
[0037] FIG. 2 illustrates a curation subsystem in accordance with an exemplary embodiment of the invention;
[0038] FIG. 3 illustrates a schematic of an apparatus in accordance with an embodiment of the present invention; and
[0039] FIG. 4 illustrates a method for information management in accordance with an embodiment of the present invention.DETAILED DESCRIPTION
[0040] FIG. 1 (A) illustrates an information management system 100 in accordance with an embodiment of the present invention. The information management system 100 comprises a delivery subsystem 102, a curation subsystem 104, a database 106, an interface 2 108, a user 110 (representing one or more users, 110-1, 110-2,... ,110-n), an access validation subsystem 112, and an interface 1 114. The system 100 is designed to provide efficient, attribute-driven organization, secure access control, and delivery of accurate, relevant, and well-organized information to the user 110 with varying access privileges. The system 100 is specifically designed for environments that require the curation, classification, and delivery of large volumes of data or datasets collected from multi-disciplinary unstructured datasets.
[0041] The database 106 is the core of the system 100 and serves as a central repository for storing a taxonomy library, curated data, or datasets across various domains and contexts. Alternatively, the curated data is stored within the curation subsystem 104. The database 106 is a shared storage between the curation subsystem 104 and the delivery subsystem 102. In one example, the curated data may include, but is not limited to, organization data, patent curation per industry, bibliographic data per industry, microservices, or client-specific private data or client private data. The database 106 stores curated data, which is not mandatorily accessed or utilized on an immediate or real-time basis but can be instead maintained for potential future inputs from the user 110. Such data or dataset is commonly referred to as "resting data" or "data at rest”. The data or dataset stored in the database 106 is indexed for fast and efficient retrieval. For example, the data or dataset related to “LIDAR” may be assigned an index number, enabling the system 100 to retrieve the data or dataset rapidly based on user inputs. This indexing supports compliance with data handling policies and minimizes risks associated with real-time processing. The databasel06 may be connected to APIs, including but not limited to Rest API, Google Remote Procedure Call API (gRPC API), Simple Object Access Protocol API (SOAP API), or web crawlers to detect and update changes in the stored dataset or data points. For instance, if the organization shifts focus from smartphones to semiconductor chips, database 106 updates relevant datasets using information obtained via the APIs or web crawlers.
[0042] The taxonomy library comprises a hierarchical structure, enabling the categorization of information into multiple levels and relationships, thus supporting complex classification schemes. The taxonomy library may be related to one or more industries For example, in the communication industry, the taxonomy library is structured or organized into top-level categories such as "wired" and "wireless" communication, each further subdivided into relevant subcategories (e.g., short-range, long-range). This hierarchical arrangement enables precise classification and retrieval of data or datasets within the system 100. The taxonomy library itself may also beindexed to enhance retrieval speed. In one example, the taxonomy library may also be indexed to enhance retrieval speed, as explained above.
[0043] In an example implementation for the patent classification, the database 106 is designed to include a comprehensive hierarchical structure to structure the patent and the organization data or dataset, including product information and technology-specific details. This structure enables categorization by technology area, sub-area, and application domain, with each application area containing detailed product listings and specifications. This system not only facilitates intuitive navigation but also enables users to analyze relationships between technology domains, application areas, and products. The hierarchical approach supports scalable management of complex data, enhances data integrity, and improves decision-making by providing a comprehensive view of an organization's technological landscape.
[0044] The curation subsystem 104 is operatively connected to the database 106 and is responsible for processing, organizing, and updating the curated data or dataset. The curation subsystem 104 receives raw or unstructured data or datasets from various sources, including patent numbers, organization data, product details, market or technology datasets, and data points, via the interface 1 114, which may be APIs or web crawlers for real-time data access. In one example, the curation subsystem 104 may receive data from Bloomberg or Factiva. In another example, the curation subsystem 104 may use real-time web crawlers based on rules defined in the rule engine, Robotic Process Automation (RPA), or generative artificial intelligence (Al) or large language models. Alternatively, the curation subsystem 104 is connected to the interface 2 108, enabling the user 110 to upload the data or dataset (backend team for the first time). As used herein, “data,” “dataset,” and “data point” are used interchangeably throughout the description. In one scenario, a “data point” may refer to a specific, narrow part or subset of the dataset. The curation subsystem 104 applies the taxonomy structure to classify, annotate, and index incoming data, using automated algorithms such as rule engines, machine learning models, statistical models, heuristic models, natural language processing and understanding engines, generative Al models,ensemble models, or combinations thereof. This ensures accurate, consistent relevancy determination and categorization of datasets or data points. The curated data or dataset is systematically indexed, easily retrievable, and updated in real time to maintain data integrity and accuracy.
[0045] As part of its advanced curation capabilities, the system 100 implements a cascaded processing pipeline within the curation subsystem 104. This pipeline is specifically adapted to receive datasets, or data points, through the interface 1 114. In another scenario, this pipeline receives datasets or data points through the interface 2 108 (uploaded by the user 110 for curation) or by selecting relevant taxonomies from the taxonomy library stored in the database 106.
[0046] The curation subsystem 104 or the pipeline includes a relevancy module arranged in a cascaded structure of one or more connected relevancy identifiers to process each data point in the retrieved dataset and a categorization module. The first relevancy identifier determines an initial relevancy level for the data point along with a confidence score. Depending on this determination, the data point may be finalized or forwarded to subsequent relevancy identifiers for further relevancy determination along with their confidence score. This multi-stage approach ensures robust and exact relevancy determination. The relevancy and confidence score may be exchanged between the relevancy identifiers or stored in database 106, supporting automated training dataset generation for continuous improvement. In one example, the relevancy identifiers may use feedforward learning to generate the training datasets, allowing the system 100 to iteratively refine the models and improve accuracy over time by leveraging the stored scores as feedback for future learning cycles. This approach supports a self-improving curation process, where the system's 100 performance, accuracy, and reliability are continuously enhanced through automated data-driven feedback. In one scenario, the system 100 may receive input with marked relevancy from the interface 1 114 or the interface 2 108. In one example, the user 110 may upload or select the dataset and initiate a curation process on the selected or uploaded dataset via the interface 2 108. Different types of relevancy identifiers are implemented withinthe system to balance the advantages and disadvantages of each, optimizing performance and cost. The overall cost of using the generative Al or large language models (LLMs) is reduced by selectively applying these models by the system 100 only where the most effective relevancy identifiers are needed. This approach enables the system 100 to maintain high-quality data or dataset curation and relevance while managing computational resources efficiently, ensuring that LLMs are utilized in a targeted manner rather than across all processes indiscriminately. The system 100 features a three-level management approach for relevancy determination, operating at the user level, domain level, and entity level. This structure allows the system 100 to determine and tailor dataset relevance based on individual user needs, the specific requirements of different domains, and the unique characteristics of entities within the dataset, ensuring precise and context-aware information retrieval and delivery.
[0047] For example, when a batch of patent documents is processed, the first relevancy identifier may use a rule engine to determine the documents relevant to a specific technology domain, such as communication technology, along with an associated confidence score. Documents not conclusively determined are forwarded to subsequent identifiers employing techniques such as machine learning, generative Al, or large language models, for further analysis and relevancy determination (including both relevancy level and confidence score) into other domains (e.g., computing or mechanical technology). This staged approach enables multi-domain classification and efficient filtering, reducing manual effort and computational overhead. The same process applies when assessing organizations or companies, where domain-specific rules, pattern recognition algorithms, or the above-defined models are used to determine the relevancy level or confidence score of datasets or data points by their alignment with target technology domains.
[0048] The system 100 of the present invention is not limited to any specific kind of dataset or data points. The examples in the description illustrate possible applications, including patent, product, scientific, non-patent literature, image data, news articles, orother online information. These illustrations are provided for clarity and do not limit the scope of the invention.
[0049] This staged approach ensures comprehensive evaluation, enabling accurate determination of relevant datasets or data points even if missed in earlier stages.
[0050] The relevancy determination process is adaptable and may be performed by a single or multiple relevancy identifiers, depending on the complexity and format of the raw or unstructured data, data point, or dataset received from the various sources. In one scenario, for heterogeneous datasets containing various formats (e.g., text, structured metadata, user-generated content), multiple relevancy identifiers may be deployed, each specializing in specific data types or domains to enhance accuracy and efficiency.
[0051] An advantage of this multi-stage approach is that only datasets or data points not determined by earlier relevancy identifiers (e.g., rule engine, machine learning models) are forwarded to subsequent relevancy identifiers, such as those employing large language or generative Al models. The large language or generative Al models are used only when necessary, optimizing computational resources and reducing costs.
[0052] Ultimately, this approach enables scalable and automated processing of large datasets or data points, regardless of the source or structure. By employing the rule engines and above-mentioned models individually or in combination, the system 100 quickly and accurately filters and determines the relevancy and confidence score of datasets according to their relevance to specific technological domains, streamlining workflows and supporting informed strategic decisions. Following relevancy determination, a categorization module, arranged in a cascade of one or more classifiers, assigns each data point to an appropriate category within the taxonomy. The categorization module comprises a multi-class categorization classifier and a proposal classifier.
[0053] The multi-class categorization classifier assigns a tentative category and confidence score to each data point. If the confidence score is below a threshold or further refinement is needed, then single-class classifiers or the proposal classifier re-examine the data point to assign a more precise category along with the confidence score. The classifiers may exchange or feedforward feedback values to enable reinforcement or feedforward learning, and may store category and the confidence score in the database 106 to facilitate automated training datasets generation. This feedback-driven approach ensures that the categorization process remains adaptive and self-improving over time.
[0054] For example, the curation subsystem 104 receives a patent dataset related to automotive technology after the relevancy determination. The multi-class categorization classifier assigns it a tentative category "communication" with a confidence score of 65% using the rule engine, as defined above. As the confidence score is below the predefined threshold (for example, 80%), the dataset is then passed to a single-class categorization classifier, which re-examines the dataset by applying more focused and granular criteria (models (GAI or large language) as defined above) and assign it to the “communication” with a confidence score of 92%. This refined categorization and the updated confidence score are stored in the database 106. The exact process applies when assessing organizations or companies, where domainspecific rules, pattern recognition algorithms, or large language models are used to determine the tentative category and corresponding confidence score of datasets or data points by their alignment with target technology domains. The curation subsystem is explained in detail in FIG. 2.
[0055] Apart from the patent or organization dataset curation, the curation subsystem 104 may calculate quality (Q) and market (M) factors by leveraging advanced computational methods such as microservices architectures or big data analytics platforms. The microservice architecture may be integrated with the curation subsystem 104 and the delivery subsystem 102 through the interface 1 114. This curation subsystem 104 receives and processes unstructured or pre-curated data, extracts relevant attributes necessary for commercialization analysis. By integrating scalable analytics, the curation subsystem 104 efficiently handles large datasets and delivers real-time or near-real-time assessments of innovation and invention indexesfor strategic decision-making in intellectual property management. The system 100 may utilize different microservices for calculating the quality factor, with each microservice tailored to specific parameters or quality attributes relevant to the dataset, domain request, or input by the user 110. The microservices focus on aspects such as performance, scalability, reliability, maintainability, security, or availability, depending on the requirements of the data being curated or the needs of the application. By leveraging a modular microservices architecture, the system 100 may dynamically select and deploy the most appropriate quality assessment methods, ensuring that the calculated Q factor accurately reflects the relevant characteristics for each context. This approach allows for flexible, scalable, and precise quality evaluation, supporting continuous improvement and adaptation as system 100 priorities or data types evolve.
[0056] In one example, the calculation of the quality factor (Q) involves the analysis of several patent-centric attributes. These may include, but are not limited to, the number of patent citations, jurisdictions filed, claim complexity and scope (including whether they are generic or highly specific), or the age of the patent. Each of these attributes contributes to a composite quality score, often through weighted algorithms or standardized scoring models. For instance, a patent with numerous citations, broad jurisdictional coverage, and broader claims is likely to receive a higher quality score, indicating a stronger position in terms of technological innovation and legal robustness. The Q factor may be calculated using quantitative and qualitative parameters.
[0057] Similarly, the market factor (M) is calculated using organization-specific attributes that reflect the commercial potential of innovation. These attributes may include, but are not limited to, organization type, industry sector, or market presence. By quantifying these factors, the curation subsystem 104 generates a market score that, combined with the quality score, provides a comprehensive view of both the technical merit and commercial viability of the innovation. This dual-index approach enables stakeholders to prioritize innovations with high potential for successful commercialization and strategic impact. The calculated Q and M factors may be storedin the database 106. The calculated quality and market factors are stored in the database 106 for future use and automated training dataset curation.
[0058] The system 100 may further include a relationship classifier to establish associations between products and subcategories by analyzing parameters such as product name, type, technical details, user manuals, and other relevant data, constructing a comprehensive hierarchy for each product, and storing the same in the database 106. Additionally, the system 100 may leverage artificial intelligence or machine learning algorithms to automatically identify and list one or more relevant patents corresponding to each product or subcategory, enabling the users 110 to quickly retrieve intellectual property information, streamlining research and development processes. In this scenario, the system 100 generates a knowledge graph that maps the relationships between patents and products, providing a visual and dataset or data- driven representation of how specific innovations are connected to tangible market offerings. This knowledge graph enables the users 110 to trace the lineage of a product back to its foundational patents, uncovering insights into technology transfer, intellectual property utilization, and innovation pathways. The system 100 is flexible and not limited to patents and products, and may also generate knowledge graphs for different types of datasets, adapting to various domains and use cases. Whether linking research publications to organizations, mapping technologies to market sectors or components, or connecting inventors to their contributions, the system’s 100 knowledge graph capabilities support comprehensive data or dataset exploration and discovery across a wide range of scenarios.
[0059] The system 100 may include dataset or datapoint redundancy cleaners at each stage of the dataset or data point processing workflow to ensure that the dataset or individual data points are thoroughly cleansed of duplicate or unnecessary information. By systematically identifying and removing redundant entries, the dataset or datapoint redundancy cleaners help to maintain the integrity and quality of the curated data or data points. This process not only optimizes storage and processing efficiency but also enhances the accuracy and reliability of the dataset used for analysis, knowledge graphgeneration, and automated training dataset generation. As a result, the system 100 consistently delivers clean, high-quality datasets that are ready for downstream applications and decision-making.
[0060] The delivery subsystem 102 is also connected to the database 106 to retrieve or receive a selected taxonomy from the taxonomy library or the taxonomy library permitted, as well as a corresponding portion of curated data from the curation subsystem 104 stored in the database 106. Unlike conventional systems, which directly connect the delivery subsystem 102 to the curation subsystem 104, causing a lot of processing time. In the present invention, the database 106, which is indexed for each industry, is placed between the curation subsystem 104 and the delivery subsystem 102. This allows the fast retrieval of the datasets. The delivery subsystem 102 is connected to the database 106 comprising both shared memory resources (shown in the FIG. 1(A) and the FIG. 1 (B)) and private memory resources (shown in the FIG. 1 (A) and the FIG. 1 (B). The shared memory resources are allocated to store shared portions of retrieved taxonomy and curated data, supporting collaborative access among the multiple users (110-1, 110-2,... 110-n), while the private memory resources store private portions of retrieved taxonomy and curated data, enabling individualized access and privacy. This allows the users (110-1, 110-2,... 110-n) to access information that is specific to their roles or permissions and protected from unauthorized access. The shared memory resources are visible to all authorized users, while the private memory resources may be available to a single user or a defined group. The private memory resource (also known as “least privileged data access”) may be implemented as instances or microservices hosted remotely, locally, or on client-dedicated systems, protected by a firewall and encryption protocols. The data here may include API access, an internal dataset, or datapoints related to patents or products. Access to the data or dataset is managed by the access validation system 112. The system 100 enables the user 110 to modify curated data, with a copy securely stored in the private memory resource. Only upon explicit approval is the modified copy written back to database 106, ensuring unauthorized updates are prevented. Isolation between the shared and private memoryresources preserves data confidentiality and integrity. In one example, when the user 110 subsequently requests the dataset or data points, the system 100 retrieves the dataset or data points from the database 106 and presents the most recent modified version, overwriting previous data with the updated content to the user 110, thereby ensuring that only the latest authorized changes are reflected in their view.
[0061] For example, the user 110 edits a "LIDAR technology" dataset (the private memory resource) in database 106. The system 100 saves the edited version in the user’s private memory resource, while the original remains unchanged. The private copy is inaccessible to others, maintaining confidentiality. Changes are committed to the database 106 only upon explicit approval, protecting data integrity and confidentiality by isolating private edits until formally approved.
[0062] The system 100 offers users 110 flexible options for data or dataset delivery by supporting both asynchronous and synchronous data access modes. In asynchronous data access mode, the system 100 delivers only the specific shared dataset or shared memory resource required for a particular use case, such as providing only the bibliographic dataset to a user working on patent landscapes, excluding unrelated information like the litigation dataset. In contrast, the synchronous data access mode enables the delivery of both private and shared datasets or memory resources to the user 110 in a continuous, time-synchronized stream, making it ideal for applications that require real-time, high-speed, and comprehensive data transfer. By accommodating both data access modes, the system 100 ensures efficient, targeted, and user-specific dataset access while optimizing performance and resource utilization.
[0063] The relevancy identifiers and categorization classifiers within the curation subsystem 104 may be implemented using a variety of technologies, including rule engines, artificial intelligence (Al) or machine learning models, statistical models, heuristic models, natural language processing or understanding engines, generative artificial intelligence (Al) or large language models or ensemble models. This flexibility allows the system 100 to adapt to different data types and domains,leveraging the most appropriate analytical techniques for each use case. The words “curated” and “categorized” are used interchangeably throughout the description.
[0064] For example, curation subsystem 104 may use a combination of natural language processing (NLP), natural language understanding (NLU), and a machine learning model to implement the relevancy identifiers and classifiers. When new technical documents are received, NLP extracts key concepts and context, and machine learning classifies documents according to taxonomy. Alternatively, rule-based engines and statistical models may process structured datasets for precise, explainable categorization.
[0065] The delivery subsystem 102 presents the real-time dataset or data points based on the user 110 input or selection, managing the transfer of data or data points between the database 106, the curation subsystem 104, and the interface 2 108, ensuring that only the appropriate portions of the taxonomy and the curated data are delivered according to the user 110 inputs and access rights.
[0066] The access validation subsystem 112 manages the user 110 authentication, authorization, assignment of memory resources, and role-based access. The role-based access in the system 100 is determined by two primary factors: hierarchy and user profile. The hierarchy reflects the user’s 110 position within the organization’s structure, such that higher-level roles automatically inherit all permissions granted to subordinate roles. For instance, a project manager would have all the access rights of a lead engineer, a project engineer, and a technician. The user profile, on the other hand, takes into account the specific attributes or job functions of each user 110, allowing for further refinement of access within each hierarchical level. This means that even the users 110 at the same level can have different permissions based on their technical responsibilities or areas of expertise. This dual-factor model ensures that access is both structured according to organizational hierarchy and tailored to technical roles, providing precise and secure control over engineering resources. In one example, the user 110-1 can only access the taxonomy of the automotive and related curated data, while the other user 110-2 may access the taxonomy of all the industries and relatedcurated data. That means different user profiles can see different workspaces based on the role-based access.
[0067] Upon receiving the user's 110 request, the access validation subsystem 112 verifies the user’s 110 credentials and determines the appropriate level of access. Based on this assessment, the access validation subsystem 112 allows access to view, edit, or share either the shared or private memory resources to the user 110 via the delivery subsystem 102. This mechanism enforces strict access control policies, ensuring that the users (110-1, 110-2,... 110-n) may only retrieve information they are authorized to access, thereby enhancing the security and privacy of the system 100.
[0068] For example, a normal user (110-1) requesting a communication patent dataset is authenticated and granted access only to non-sensitive taxonomy and curated dataset via the shared memory resource. A privileged user (110-2), such as an administrator or lead researcher, is authorized to retrieve both shared (non-sensitive) and sensitive (private) datasets, delivered through a private memory resource. This ensures robust security and privacy controls. An exemplary implementation of the private and shared memory resources is illustrated and described in detail above. However, it is understood that the disclosed implementation is equally applicable to other forms of private and shared data, memory architectures, or equivalent resource allocation mechanisms. This flexible design ensures that data security and the user 110 privacy are maintained across various deployment scenarios.
[0069] The interface 2 108 may receive a single input, a batch of inputs, or a dataset from the user 110. The interface 1 114 may receive a dataset through the API call or generate an API call based on the stored instructions or algorithms.
[0070] The interface 2 108 is operatively coupled to the delivery subsystem 102 and serves as the primary point of interaction between the users (110-1, 110-2,... 110-n) and the system 100. The interface 2 108 receives input, directs requests to the delivery subsystem 102, which processes the user 110 input, receives the relevant taxonomy and curated data from the curation subsystem 106, and returns the results through the interface 2 108. The interface 2 108 may comprise a range of human-machine interface(HMI) technologies designed to facilitate seamless interaction between the users 110 and the system 100. The interface 1 114 and the interface 2 108 may include, but are not limited to, display devices such as liquid crystal displays (LCD), light-emitting diode (LED) screens, and organic light- emitting diode (OLED) screens. The interface 2 108 may also be integrated into mobile phones or other portable devices for remote access. In one example, the interface 2 108 may be integrated into the delivery subsystem 102.
[0071] In addition to visual displays, the interface 1 114 and the interface 2 108 may include touchscreens, voice input interfaces, or application interfaces for integration with external software systems or mobile applications, enhancing user interaction and interoperability.
[0072] The design of the interface 1 114 and the interface 2 108 supports programmable keys, keypads, stylus input, and connectivity options (Ethernet, USB, or wireless protocols) for communication with peripheral devices and networks. This ensures robust, user-friendly, and adaptable interaction with the system 100 for various use cases.
[0073] The interface 2 108 presents a selective combination of the shared and private portions of curated data and taxonomy, as determined by the access validation subsystem 112, enabling personalized and secure information delivery, tailored to each user’s (110-1, 110-2,... 110-n) specific needs, preferences, and access rights. The interfaces 1 114 and 2 108 may be implemented as a web portal, mobile application, desktop client, or any other suitable form, depending on the deployment environment.
[0074] Alternatively, the interface 2 108 may be operatively coupled to the curation subsystem 104, facilitating seamless interaction between the user 110 and the underlying data processing components. In one implementation, interface 2 108 serves as a shared access point for both the curation subsystem 104 and the delivery subsystem 102, simplifying the system 100 architecture by providing a single-entry point for input and output.
[0075] Alternatively, the system 100 architecture may be designed such that the curation subsystem 104 and the delivery subsystem 102 each have their dedicated interfaces (the interface 1 114 and the interface 2 108), allowing specialized the user 110 interactions or API calls tailored to the unique requirements of curation and delivery processes. This enhances modularity and flexibility, as changes to one interface do not directly impact the other.
[0076] Both shared and separate interface architectures offer advantages, shared interfaces promote simplicity and consistency, while separate interfaces allow for customization and isolation. The choice depends on the system's 100 requirements, the user's 110 needs, and maintainability, and may be guided by design patterns such as the frontage pattern.
[0077] During operation, the user 110 accesses the system 100 via the interface 2 108 and selects a taxonomy. The access validation subsystem 112 authenticates the user 110 and determines the appropriate access level. The delivery subsystem 102 then retrieves the relevant taxonomy and the corresponding curated data from the database 106, utilizing assigned shared and / or private memory resources, and delivers through the interface 2 108.
[0078] The system 100 supports multiple simultaneous users (110-1, 110-2,... 110-n), each with distinct access rights and information needs. By leveraging the hierarchical taxonomy structure and integrating the shared and private memory resources, it allows both collaborative and individualized data access for a wide range of use cases.
[0079] For example, the user 110 may select a taxonomy such as “automotive,” prompting the delivery subsystem 102 to retrieve relevant curated data from the database 106 and present it. This enables efficient, taxonomy-based access to curated information for future use.
[0080] The interface 2 108 of the system 100 enables the user 110 to perform a “SMART SEARCH” that automatically retrieves relevant datasets or data points in response to the user 110 input. For example, when the user 110 enters a keyword such as “LIDAR”, the system 100 searches the database 106 to locate correspondingdatasets. If the keyword “LIDAR” is not found within the database 106, the system 100 constructs a knowledge graph (KG) to expand the search scope by identifying relevant synonyms, related technologies, or contextual associations using generative Al or other models as defined above. The backend repository for the knowledge graph is maintained either locally within the system 100 (in the database 106) or remotely on a server (not shown). Additionally, the system 100 may leverage contextual knowledge through generative Al models or Retrieval- Augmented Generation (RAG) techniques, thereby enhancing the accuracy and comprehensiveness of data retrieval by grounding search results in authoritative sources and minimizing information gaps or inaccuracies. This approach ensures that the user's 110 queries yield the most pertinent and reliable datasets, even when the exact keyword is not present in the database 106.
[0081] The architecture of system 100 supports efficient processing and retrieval of relevant information by leveraging attribute-driven organization, hierarchical taxonomy, and robust access control. The interconnected components of the system 100 work together to deliver well-organized, accurate, and relevant information to the users (110-1, 110-2,... 110-n), ensuring both collaborative and individualized dataset access as required. The system 100 is implemented in various embodiments, including different configurations of memory resource allocation and taxonomy structures, to suit specific operational or security requirements.
[0082] The system 100 may include a processing unit and memory unit as essential hardware components to carry out the overall operation of all the subsystems described above. The processing unit executes instructions, algorithms, and manages the flow of data between the delivery subsystem 102, curation subsystem 104, access validation subsystem 112, the interface 1 114, and interface 2 108. The processing unit and memory unit may be integrated within one or more of the aforementioned subsystems and interfaces, thereby enabling localized processing and storage. Alternatively, the processing unit and memory unit may be implemented as standalone components within the system 100, facilitating centralized control and resource management. Thisflexible configuration allows for optimized performance and scalability of the system 100 in accordance with specific application requirements.
[0083] The processing unit may be integrated within the system 100 or located remotely, and may comprise a single or multi-core processor, cloud server, digital signal processor (DSP), graphics processing unit (GPU) microcontroller, system on a chip (SoC), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or a combination thereof. The processor unit may use co-processors for complex tasks and include specialized hardware, software, or firmware modules. The processing unit may be a general-purpose or special-purpose processor, or any device capable of implementing system 100 operations. The processing unit may include one or more specialized hardware, software, and / or firmware modules (not shown) specially configured with particular circuitry, instructions, algorithms, or data to perform functions of the disclosed system 100 and methods. In some scenarios, resource-intensive computing, such as for large language models or generative artificial intelligence models, is implemented on a separate server to handle intensive computational tasks, while resource-efficient or low- resource computing is deployed locally on the client computers for lighter processing needs. This setup leverages the client-server architecture, where the server manages heavy data processing and complex algorithms, and the client focuses on user interaction and basic tasks, optimizing both performance and resource utilization across the system 100.
[0084] The processing unit may be coupled to the memory unit, which may include both volatile (e.g., RAM) and non-volatile (e.g., storage devices) types, temporarily holds data and instructions required for immediate processing by the processing unit. The volatile memory elements may include, random access memory, such as DRAM, SRAM, SDRAM, non-volatile memory elements (for example, ROM, hard drive, etc.), magnetic, semiconductor, tape, optical, removable, non-removable, or other types of storage devices or tangible and combinations thereof. Typical forms of non-transitory media include, for example, a flash drive, a flexible disk, a hard disk, a solid state drive, magnetic tape or other magnetic data storage medium, a CD-ROM or other optical datastorage medium, any physical medium with patterns of holes, a non-transitory computer-readable medium, RAM, a PROM, and EPROM, a FLASH-EPROM, other flash memory, NVRAM, a cache, a register, other memory chip or cartridge, or networked versions of the same. The memory unit may have a distributed architecture or local architecture, store software programs or algorithms, historical curated data, confidence score, and thresholds for datasets. In one scenario, a dedicated cache is implemented within the delivery subsystem 102 to achieve optimal speed. This ensures faster data retrieval and improved system performance for the users 110. The cache allocation for each user 110 can be predicted based on the size of the curated dataset relevant to the workspace. Hence, load balancing is not dependent on unpredictable variables like existing systems.
[0085] In one implementation, the database 106 may be part of the memory unit or managed as a distinct resource, depending on the system 100 design and deployment requirements.
[0086] FIG. 1(B) illustrates an information management system 100 in accordance with another exemplary embodiment of the present invention. The functionality and structure shown in FIG. 1(B) are the same as in the FIG. 1(A), but FIG. 1(B) provides a more detailed visualization of the components and operations of the system 100. This detailed depiction allows for a clearer understanding of how each part of the system 100 interacts and functions within the information management architecture.FIG. 1(C) illustrates an information management system 100 according to another exemplary embodiment, where the overall functionality and structure remain consistent with those depicted in FIGS. 1(A) and 1(B). However, FIG. 1(C) specifically highlights the exemplary possible components within each subsystem and demonstrates how these subsystems and interfaces (102-114) can be remotely distributed across different servers or locations. This arrangement provides a clearer understanding of the connectivity, interactions, and operational flow between subsystems and interfaces (102-114), showcasing the flexibility and scalability of the information managementsystem 100 in accommodating distributed deployments and diverse operational requirements.
[0087] FIG. 2 illustrates a curation subsystem 200 in accordance with an exemplary embodiment of the invention. The curation subsystem 200 is configured to receive a dataset from an API, web-crawlers (as shown and described in FIG. 1(A) and FIG. 1(B) or an interface (as shown in FIG. 1(A). The curation subsystem 200 is composed of two primary modules, a relevancy module 202 and a categorization module 204. The purpose of the curation subsystem 200 is to systematically filter, evaluate, and organize large datasets, including batches of patent documents, industry or company-related datasets, or other datasets, to facilitate downstream analysis and classification.
[0088] The relevancy module 202 is central to the initial processing of the dataset. The relevancy module 202 comprises a plurality of relevancy identifiers (RI1 202-1, RI2 202-2, RI3 202-3, ... , RIN 202 -N), each connected in a cascaded manner, either serially or in parallel. This arrangement allows each relevancy identifier to process only the subset of data points that have not met the threshold criteria (which is the confidence score, as explained in FIG. 1(A)) set by the preceding identifier in the sequence. For example, the first relevancy identifier RI1 202-1 may receive a batch of 500 patent data points, such as those related to the automotive industry. RI1 202-1 determines each data point by retrieving a taxonomy, which is retrieved from a database (as discussed and shown in FIGS.l (A) and (B). Alternatively, the determination may be done using a rule engine, an artificial intelligence (Al) or machine learning models, a statistical model, a heuristic model, a natural language processing and understanding engine, generative artificial intelligence (Al) or large language models, or an ensemble model. The rule engine may be trained or augmented using generative artificial intelligence (Al) models or large language models (LLMs) to automatically generate, refine, and optimize rules that govern the operation of the system as described in the FIG. 1 (A), FIG. 1(B), FIG. 1(C) and FIG. 2. In this approach, the generative Al models analyze historical data, domain knowledge, and operational requirements to dynamically create or update rule sets, potentially leveraging techniques such as Retrieval-AugmentedGeneration (RAG), which involves fetching pertinent information from trusted sources or documents before generating rules. This ensures that the rules created are well- supported by accurate, up-to-date information and reduces the likelihood that the generative Al will provide incorrect or misleading content, often referred to as “hallucinations”. While the rule engine itself enforces logical consistency, traceability, and compliance by applying these rules in a deterministic and explainable manner, thereby combining the adaptability and creative capabilities of the generative Al model with the precision and reliability of traditional rule-based decision-making systems. The relevancy of each data point is determined by comparing the data point with a prestored threshold criterion (which is the confidence score), such as greater than 95% for RI1 202-1. In another example, the threshold criterion may be defined in terms of performance metrics, including false negatives, true positives, false positives, and true negatives. In one scenario, if the two relevancy identifiers (RI1 and RI2 either connected in serial or parallel) have the same output, such as a relevancy level with a 100% confidence score, then the same (final output) is sent to the categorization module 204. In another scenario, if the output from the first relevancy identifier is 100% and the other one has a 30% confidence score, or a conflicting situation, the rule engine’s weightages or further Al processing resolve the relevancy. Each relevancy identifier may include one or more rule engines.
[0089] The data points meeting the threshold at any stage are finalized, while those that are not passed to the next relevancy identifier, which is the second relevancy identifier RI2 202-2. RI2 202-2 may, for example, receive 300 data points from RI1 202-1 and apply a different threshold, such as greater than 70%. This identifier may also utilize a specialized set of keywords, stored in the database or rule engine, to further refine the relevancy determination. The data points not meeting the second threshold are then forwarded to the third relevancy identifier RI3 202-3, which may receive, for example, 100 data points for further assessment. The third relevancy identifier, RI3 202-3, employs a human-intervention interface, a generative Al model,the analytical hierarchy process (AHP), or other advanced filtering mechanisms for quality checks, with threshold criteria, such as equal to 0%.
[0090] Throughout this process, each relevancy identifier may operate according to instructions from the rule engine or the models discussed above. The rule engine may dynamically update or enrich the curation rules based on the evolving nature of the dataset and the taxonomy selected. The taxonomy library provides a structured schema for determining relevancy, selecting, and retrieving the appropriate taxonomy based on attributes of the dataset, such as industry, company, or semantic features of the dataset.
[0091] Once the relevancy levels and the confidence score or threshold are determined by the connected relevancy identifiers (RI1 202-1, RI2202-2, RI3 202-3, ... , RIN 202- N), the output is forwarded to the categorization module 204. The categorization module 204 organizes curated data points into predefined categories, enhancing the value of the curated dataset or data points for downstream applications such as analytics, reporting, or machine learning training sets.
[0092] Additionally, the relevancy identifiers (RI1 202-1, RI2 202-2, RI3 202-3, ... , RIN 202-N) may exchange or feedforward identified relevancy levels and the confidence score with each other or store these determinations in the database, facilitating the generation of enriched automated training datasets and supporting iterative improvements to the curation process. This modular and hierarchical approach to data curation ensures high accuracy and scalability, making the curation subsystem 200 suitable for large-scale, dynamic data environments such as patent analysis or realtime electronic data streams.
[0093] The categorization module 204 includes a multi-class classifier 204-1 and a proposal classifier 204-2. The proposal classifier 204-2 includes one or more cascaded categorization classifiers (CC 204-21, CC 204-22,... 204-2N).
[0094] The categorization module 204 is a sophisticated component designed to efficiently organize and classify curated data points into predefined categories, ensuring both accuracy and adaptability in the system. The categorization module 204operates in close integration with the relevancy module 202, forming a comprehensive framework for data curation.
[0095] Upon receiving data points from the relevancy module 202, the categorization module 204 initiates a multi-class categorization process. The multi-class categorization classifier 204-1 serves as the initial decision-making entity, analyzing each data point or data and assigning a tentative category with an associated confidence score. For example, a patent data point may be classified as “communication” (100% confidence score), “computing” (80-90%), or “automotive” (50-70%). These results are forwarded to the proposal classifier 204-2 for further refinement. In one example, the categorization module 204 may include one or more multi-class categorization classifiers corresponding to each domain. In another example, rejected data points may be sent to a human intervention interface (manual error checking) for quality assessment, after which they are returned to the proposal classifier 204-2. False positives may be discarded or used to retrain the system, supporting continuous improvement. The architecture of the proposal classifier 204-2 comprises a sequence (connected serially or parallelly) of single-class categorization classifiers (CC 204-21, CC 204-22, ..., 204-2N), each specializing in the identification of a specific category. As the data point progresses through each classifier in the sequence, the data point is determined for the respective category and assigned a finalized category assignment along with an updated confidence score. For instance, the first single-class classifier CC 204-21 may determine that the data points definitively belong to the communication category with a 100% confidence score. If the confidence score is 80- 90% and the category is automotive, the data points proceed to the next classifier, such as CC 204-22. If the confidence score is 50-70% and the category is computing, the data points proceed to the next classifier, such as CC 204-23. In one example, the categorization module 204 may either use the multi-class categorization classifier 204- 1 or single-class categorization classifiers (CC 204-21, CC 204-22, ...204-2N). In one example, if the two single-class categorization classifiers give the same output for two categories (communication and computing), the rule engine identifies this as anautomated error according to predefined rules and resolves it by applying a priority or weightage technique to select the appropriate category.
[0096] Once a finalized category, along with the confidence score, is determined by one of the single-class categorization classifiers (CC 204-21, CC 204-22, ...204-2N), this result is passed to the rule engine or other models (discussed in FIG.1) for validation and verification. The rule engine may employ a combination of rule-based logic, statistical algorithms, machine learning models, and, where necessary, utilize a human intervention interface review process to ensure the accuracy and appropriateness of the category assignment. For example, high-confidence score data points (100%) may be validated automatically, and intermediate or low-confidence score data points (80-90% and 50-70%) may require an advanced artificial intelligence (Al) model, a generative Al model, or human intervention for review checks.
[0097] A key feature of the categorization module 204 is the ability to support automated reinforcement or feedforward learning and continuous improvement. The categorization module 204 facilitates the exchange of feedback between the multi-class and single-class classifiers, allowing them to share determined categories along with the confidence score (or threshold) and outcomes in the database, generating an automated dynamic training dataset that is used to retrain and optimize both the relevancy module 202 and the categorization module 204. The training datasets generated through this process may be stored in memory or a dedicated database, depending on system requirements and scalability considerations.
[0098] The entire operation of the curation subsystem 200, including all interactions and dataset exchange, is orchestrated by a processing unit, as explained in FIGS. 1(A)- 1(C). This processing unit manages task sequencing, data delivery, the execution of rule-based, machine learning algorithms, or models discussed in the above figures, and the storage and retrieval of the training dataset. The modular and extensible design of the categorization module 204 allows for the adaptation to new categories, updated models, and evolving rules or models, thereby providing a future-proof solution for complex dataset categorization challenges.
[0099] FIG. 3 illustrates a schematic of an apparatus 300 in accordance with an embodiment of the present invention. The apparatus 300 may be a relevancy module, a curation module, a chip, a chip system, or a system (shown and discussed in FIGS. 1(A)- 1(C) a processor, or other hardware component that supports the implementation of the methods described in the preceding embodiments. The apparatus 300 is specifically configured to implement the functionalities and processes outlined in the foregoing method embodiments.
[0100] The apparatus 300 comprises one or more processing units 302. These processing unit 302 are responsible for executing the core functions of the system, as detailed in the previous figures. The processing unit 302 may be realized as a general- purpose processor, a dedicated processor, or a specialized processing component such as a baseband processor, graphics processing unit (GPU), or a central processing unit (CPU). For instance, a baseband processor within the apparatus 300 may be configured to process data points to determine the relevancy and category of the data points with a specified confidence score. Alternatively, the CPU may oversee the overall control of the apparatus 300, execute software programs, and process data associated with those programs.
[0101] In some configurations, the processing unit 302 executes instructions and processes data that enable the apparatus 300 to perform the methods described in the preceding embodiments. Additionally, the apparatus 300 may optionally include a transceiver unit designed to facilitate the reception and transmission of information or data points between various modules. The transceiver unit may be implemented as a transceiver circuit, interface, or interface circuit, and may be configured for either separate or integrated sending and receiving functionalities. The transceiver unit is also capable of reading and writing code or data, as well as transmitting or transferring signals as required by the system.
[0102] Further, the apparatus 300 may include a circuit dedicated to implementing the communication functions, such as sending and receiving datasets, as described in the method embodiments. Optionally, the apparatus 300 may be equipped with one or morememory units 304. These memory units 304 are utilized to store instructions and algorithms 306, which, when executed by the processing unit 302, enable the apparatus 300 to carry out the described methods. The memory unit 304 may also serve as storage for datasets, and the arrangement of the processing and memory units (302, 304) may be either discrete or integrated. For example, a correspondence described in the foregoing method embodiments may be stored in the memory unit 304.
[0103] The apparatus 300 may also include a transceiver unit 308 and / or an antenna 310 to support wireless communication. The transceiver unit, which may be referred to as a transceiver machine, circuit, apparatus, or module, is configured to provide transceiver functionality within the system.
[0104] The apparatus 300 is adaptable and may be configured to perform the methods described in the embodiments corresponding to FIG. 1 (A) to FIG. 2, or any combination thereof. The processing unit 302 and the transceiver unit 308 may be implemented using a variety of integrated circuit technologies, including but not limited to complementary metal oxide semiconductor (CMOS), n-type or p-type metal oxide semiconductor (NMOS / PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), and gallium arsenide (GaAs).
[0105] The apparatus 300, as described, may serve as either the relevancy module or the categorization module, but its scope is not limited to these functions or the structural representations in FIG. 1 (A) to FIG. 2. The apparatus 300 may be realized as an independent device or as a component within a larger system. Examples include, but are not limited to, an independent integrated circuit (IC), a chip or chip system, a set of ICs (optionally including storage components), an application-specific integrated circuit (ASIC), a modem (MSM), an embedded module, a receiver, a cellular phone, a wireless device, a handheld device, a mobile unit, a cloud device, an artificial intelligence device, or a machine device.
[0106] FIG. 4 illustrates a method 400 for information management in accordance with an embodiment of the present invention. The method 400 comprises the steps of a) receiving 402, a dataset through an interface; b) selecting 404, a taxonomy from ataxonomy library; c) implementing 406, a relevancy module in a cascaded structure to determine a relevancy level and a confidence score of each data point of the dataset based on the selected taxonomy; d) sorting 408, a data point based on the determined relevancy level and the confidence score; and e) implementing 410, a categorization module in a cascaded structure to classify the sorted data point based on the selected taxonomy; wherein the categorization module is selected based on the determined relevancy level and the confidence score.
[0107] The method 400 enables highly accurate and efficient information management by combining context-aware taxonomy selection with cascaded relevancy and categorization modules. This structured approach ensures that only the most relevant data points are processed and classified according to the most appropriate taxonomy, significantly reducing manual effort, minimizing errors, and improving the overall quality and usefulness of the managed information. As a result, faster and more informed decisions are made while optimizing resource utilization.
[0108] In step c), implementing 406 the relevancy module in a cascaded structure involves processing each data point selectively through a plurality of connected relevancy identifiers. This means that a first relevancy identifier first processes each data point to determine the relevancy level and the confidence score. Based on the determination, the relevancy level may either be finalized at this stage or, if further evaluation is needed, the data point is passed to a second relevancy identifier for processing. This selective processing allows for a more refined and accurate determination of relevancy and the confidence score, as each identifier applies different criteria or models to determine the relevance and the confidence score of the data point. Furthermore, the plurality of connected relevancy identifiers exchanges or feedforward the determined relevancy levels and the confidence score of each data point or stores the determined relevancy level in the database to generate an automated training dataset. This method enhances the accuracy of relevancy assessment and enables the generation of a comprehensive automated training dataset, supporting continuous improvement of the relevancy module through machine learning and analytics.
[0109] In step e), implementing 410, the categorization module in a cascaded structure involves a two-stage classification process for each selected data point. First, the data point is processed through a multi-class categorization classifier, which assigns a tentative category along with an associated confidence score. This preliminary classification provides an initial understanding of where the data point may best fit within the taxonomy. Next, based on the tentative category and the confidence score determined by the multi-class classifier, the data point is then selectively processed through a plurality of connected single-class categorization classifiers (proposal classifier). Each of these single-class classifiers further evaluates the data point, focusing on specific categories to refine and determines a finalized category and the confidence score. This cascaded approach ensures that the final category and the confidence score assigned to the data point are both accurate and reliable, leveraging the strengths of both broad and focused classification techniques to enhance overall categorization performance.
[0110] In step e), implementing 410, the categorization module in a cascaded structure further enhances the classification capabilities by generating and exchanging feedback values between the interconnected single-class categorization classifiers and the multiclass categorization classifier. This exchange of feedback enables the system to perform reinforcement or feedforward learning, allowing the classifiers to continuously learn from each other's outputs and improve their accuracy over time. The multi-class categorization classifier and the plurality of connected single-class categorization classifiers also exchange or feedforward the determined category information and the confidence score with each other or store the finalized category and the confidence score in a database to generate an automated training dataset. By doing so, the method 400 not only refines its current classification decisions but also generates a comprehensive automated training dataset. This dataset is leveraged for ongoing training and optimization of the classification models, resulting in a more adaptive, intelligent, and accurate categorization process that allows the system to process more datasets.
[0111] The method 400 described is implemented in a curation subsystem. This means that all the steps involved in receiving datasets, selecting and applying taxonomies, determining relevancy, and categorizing data points are carried out within a dedicated component of the overall information management system. By utilizing a curation subsystem, the method 400 ensures that the dataset is systematically processed, organized, and maintained according to predefined rules and intelligent algorithms. This centralized approach streamlines the management of large and complex datasets, enhances data quality, and supports efficient retrieval and use of information across the organization. The curation subsystem thus plays a crucial role in enabling automated, scalable, and consistent information management.
[0112] The relevancy identifier, the multi-class categorization classifier, or the singleclass categorization classifier, each include a rule engine, an artificial intelligence (Al) or machine learning models, a statistical model, a heuristic model, natural language processing and understanding engines, large language or generative artificial intelligence (Al) models or an ensemble model. This flexible architecture allows the method 400 to leverage a wide range of analytical and computational techniques for determining relevancy and categorizing data points. By incorporating rule-based logic, advanced Al or machine learning algorithms, statistical analysis, heuristic approaches, or natural language processing and understanding capabilities, the method 400 is tailored to handle diverse types of datasets and complex classification tasks. Additionally, the use of ensemble models enables the combination of multiple methods to improve accuracy and robustness. This versatility ensures that the information management method 400 is optimized for different domains, dataset characteristics, and evolving organizational requirements.
Claims
CLAIMSWe Claim1. An information management system comprising: a delivery subsystem; a curation subsystem; and a database, connected with the delivery subsystem and the curation subsystem, for storing a taxonomy library; wherein the delivery subsystem is configured to retrieve a taxonomy from the taxonomy library, and a portion of curated data corresponding to the retrieved taxonomy from the database; and wherein the delivery subsystem is configured to provide the retrieved taxonomy and the portion of the curated data.
2. The information management system according to claim 1, wherein an interface 2 is configured to receive and direct input to the delivery subsystem.
3. The information management system according to claim 2, wherein the delivery subsystem is configured to retrieve a taxonomy from the taxonomy library stored in the database based on the input from the interface 2.
4. The information management system according to claim 1, wherein the delivery subsystem is connected to the database comprising a shared memory resource allocated to a shared portion of retrieved taxonomy and a shared portion of curated data, and at least one private memory resource allocated to a private portion of retrieved taxonomy and a private portion of curated data.
5. The information management system according to claim 4, wherein the system comprises an access validation subsystem to assign the at least one shared memory resource or private memory resource to a user.
6. The information management system according to claim 5, wherein the delivery subsystem is configured to provide, through the interface 2, a selective combination of the shared and private portions of curated data and a selective combination of the shared and private portions of retrieved taxonomy based on the assignment by the access validation subsystem.
7. The information management system according to claim 1, wherein the taxonomy comprises a hierarchical structure and relationship data of multiple levels of the hierarchy.
8. A method for information management comprising: receiving a dataset through an interface 1 ; selecting a taxonomy from a taxonomy library; implementing a relevancy module in a cascaded structure to determine a relevancy level and a confidence score of each data point of the dataset based on the selected taxonomy; sorting a data point based on the determined relevancy level and the confidence score; and implementing a categorization module in a cascaded structure to classify the sorted data points based on the selected taxonomy, wherein the categorization module is selected based on the determined relevancy level and the confidence score.
9. The method according to claim 8, wherein implementing the relevancy module in a cascaded structure comprises processing the data point selectively through a plurality of connected relevancy identifiers.
10. The method according to claim 9, wherein processing the data point selectively through the plurality of connected relevancy identifiers comprises: processing each data point through a first relevancy identifier to determine the relevancy level and the confidence score; and finalizing the relevancy level or processing the data point through a second relevancy identifier based on the determined relevancy level and the confidence score.
11. The method according to claim 10, wherein the plurality of connected relevancy identifiers exchanges the determined relevancy level and the confidence score of each data point, or stores the determined relevancy level and the confidence score in the database to generate an automated training dataset.
12. The method according to claim 8, wherein implementing the categorization module in a cascaded structure comprises: processing the data point through a multi-class categorization classifier to determine a tentative category and associated confidence score; and processing the data point selectively through a plurality of connected singleclass categorization classifiers based on the tentative category and the associated confidence score to determine a finalized category and the confidence score.
13. The method according to claim 8, wherein implementing the categorization module in a cascaded structure comprises generating and exchanging feedback value byinterconnected single-class categorization classifiers or multi-class categorization classifiers to perform automated reinforcement learning or feedforward learning.
14. The method according to claim 13, wherein the multi-class categorization classifier and the plurality of connected single-class categorization classifiers exchange the determined category and the confidence score with each other or store the determined category and the confidence score in the database to generate an automated training dataset.
15. The method according to claim 8, wherein the method is implemented in a curation subsystem.
16. The method according to claim 8, wherein the relevancy identifier, the multi-class categorization classifier, or the single-class categorization classifier includes a rule engine, artificial intelligence (Al) or machine learning models, a statistical model, a heuristic model, natural language processing and understanding engines, large language or generative artificial intelligence (Al) models, or an ensemble model.
Citation Information
Patent Citations
Methods for screening infections
CN110546157A
NLP and AIS of I / O, prompts, and collaborations of data, content, and correlations for evaluating, predicting, and ascertaining metrics for IP, creations, publishing, and communications ontologies
US12094018B1
Method and Apparatus for Coupling the Internet, Environment and Intrinsic Memory to Users
US20160267187A1
Audio content processing systems and methods
US20200105274A1