Standardized Project Data Templates Using Service Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional project management systems face challenges in managing and standardizing project data across multiple participants, leading to inefficiencies in data sharing, inconsistent updates, and inaccurate benchmarking and analysis, which hinders informed decision-making and resource optimization.
Innovation Solution
A system utilizing machine learning models to extract services from text data, cluster them based on similarity, and generate standardized data storage templates, enabling efficient conversion and storage of project data in a unified format, with real-time updates and alerts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If project data is stored in non-standardized formats by individual participants, then each participant can use their preferred hardware or software platform, but data consistency and comparability across projects deteriorate
Solution Approach 1:
The patent introduces an intermediary processing layer that includes text extraction, embedding generation, and clustering components. This intermediary system converts diverse non-standardized project data into standardized representations without requiring changes to the original participant platforms, thus maintaining platform flexibility while achieving data standardization.
Solution Approach 2:
The patent transforms project data by changing its representation parameters through text embedding and clustering. The raw text data is converted into numerical embeddings, then grouped into clusters that represent standardized service categories, effectively transforming unstandardized data into standardized form through parameter transformation.
2Measurement precision
If manual methods are used for bid management and data comparison, then detailed evaluation can be performed, but processing time and resource consumption increase
Solution Approach 1:
The patent replaces manual mechanical processes of bid evaluation with an automated system using machine learning models. The clustering model automatically compares bids against historical data and extracts service information, substituting human manual analysis with computational processes that maintain accuracy while dramatically reducing time consumption.
Solution Approach 2:
The system performs preliminary actions by pre-processing project data through text extraction and embedding generation before bid evaluation. Historical project data is预先 clustered and stored as reference standards, enabling rapid comparison during bid management without performing full analysis from scratch each time.
3Productivity
If project data is not standardized, then data can be collected from multiple sources, but the ability to benchmark and analyze data across projects deteriorates
Solution Approach 1:
The patent extracts essential information from unstandardized project data through text extraction and embedding generation. By taking out only the critical service-related information and representing it in a standardized embedding format, the system maintains data collection efficiency while preserving analytical value for benchmarking and comparison.
4Ease of manufacture
If conventional data storage methods are used, then implementation is simple, but file sizes and memory usage increase
Solution Approach 1:
The patent changes the parameter representation of project data from raw text format to compressed embedding vectors. This parameter transformation significantly reduces the storage volume required while maintaining the essential information content, as numerical embeddings are more compact than original text descriptions.
Data Source
AI summary
Techniques for improving project data storage by generating standardized data storage templates are disclosed. An example system aggregates project data corresponding to at least two projects that is in a non-standardized format, extracts text data from the project data, and inputs the text data and an input prompt into a first machine learning (ML) model configured to extract one or more services indicated by the text data. The example system further generates a text embedding for each service and applies a second ML model to (i) the services and (ii) the text embeddings. Applying the second ML model includes: clustering the services and the text embeddings into a set of clusters, and generating, based on the clusters, at least one data storage template indicating a standardized set of services in a standardized format for projects with associated project data included in one or more clusters.


