AI/ML Asset Discovery in Enterprise Code Repositories
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a need for a method and system to automatically discover and identify Artificial Intelligence/Machine Learning (AI/ML) source code, models, parameters, data input and output specifications, and data transforms within a production code repository, as existing methods lack the capability to do so effectively.
Innovation Solution
A computer-implemented method and system that uses AI/ML to automatically analyze source code from various sources, including open-source AI/ML libraries, non-open-source AI/ML libraries, and tagged/pre-classified code, to perform semantic matching and identify AI/ML models and their associated parameters and data specifications within a production code repository.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If AI/ML approaches are used to analyze source code, then bug detection and prediction capabilities are improved, but the quality of analysis depends on the model and training data which are difficult to control and standardize
Solution Approach 1:
The system segments the code analysis task into multiple independent modules: syntax analysis, semantic analysis, and AI/ML model analysis. Each module processes specific aspects of the code separately, allowing for standardized evaluation of each component while maintaining overall system reliability.
Solution Approach 2:
The patent introduces an intermediary layer of standardized interfaces and protocols between the code repository and the AI/ML analysis models. This intermediary ensures consistent data formatting and communication, improving measurement precision while allowing flexible model selection for reliability.
2Measurement precision
If manual identification and classification of AI/ML source code is performed, then accuracy and governance are improved, but time consumption and labor requirements increase significantly
Solution Approach 1:
The system implements self-service automation where the AI/ML analysis tools automatically scan, identify, and classify code without human intervention. The models self-evaluate and self-correct, reducing time loss while maintaining accuracy through continuous learning and validation processes.
Solution Approach 2:
The patent applies preliminary action by pre-processing and pre-classifying code segments before final AI/ML analysis. The system prepares code data in advance, organizing it into standardized formats that accelerate subsequent analysis while ensuring accuracy through pre-validation checks.
3Loss of information
If comprehensive analysis of all source code is performed to ensure complete discovery of AI/ML assets, then visibility and governance are improved, but system complexity and processing overhead increase
Solution Approach 1:
The system applies local quality by analyzing different portions of the code base with appropriate levels of detail. Critical AI/ML components receive comprehensive analysis, while standard code receives lighter analysis. This selective approach ensures complete visibility of AI/ML assets without unnecessarily increasing overall system complexity.
Solution Approach 2:
The patent implements dynamic analysis where the system adjusts its complexity level based on the code being analyzed. The analysis depth, processing speed, and resource allocation dynamically adapt to the specific requirements of each code segment, maintaining comprehensive visibility while optimizing system complexity for each task.
Data Source
AI summary
A method and a system for the automatic discovery of AI/ML models, their parameters, data input and output specifications, and data transforms in a production code repository using Artificial Intelligence/Machine Learning are disclosed. A method and system for automatic discovery of the location, identification, classification, and definition of the AI/ML models, their parameters, data input and output specifications, and data transforms in the production code repository using Artificial Intelligence/Machine Learning are also disclosed. The method and system utilize a plurality of source codes from a plurality of sources, such as open-source AI/ML libraries with the source codes, non-open-source AI/ML libraries, and tagged/pre-classified code, in conjunction with a production code repository, to identify the method of working on the plurality of source codes using Artificial Intelligence/Machine Learning.


