System and method for enabling development of an end-to-end no code artificial intelligence vision software with ai and computer vision frameworks
Patent Information
- Application Number
- PCT/MY2025/050019
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-07
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies face challenges in enabling users to develop their own artificial intelligence and computer vision models due to the need for programming expertise and tedious manual coding, which limits accessibility and efficiency.
A web-based platform providing image processing and deep learning model training services in a code-free manner, allowing users to develop end-to-end artificial intelligence vision software through intuitive graphical interfaces and modules for image processing, annotation, augmentation, model training, and fine-tuning, with support for both traditional and deep learning algorithms.
Empowers users to build robust computer vision models efficiently, enhancing accessibility, accuracy, and operational efficiency, while reducing the need for specialized skills and lowering costs by leveraging cloud and local computing resources.
Smart Images

Figure MY2025050019_02102025_PF_FP_ABST
Abstract
Description
[0001]
[0002] SYSTEM AND METHOD FOR ENABLING DEVELOPMENT OF AN END- TO-END NO CODE ARTIFICIAL INTELLIGENCE VISION SOFTWARE WITH Al AND COMPUTER VISION FRAMEWORKS
[0003] FIELD OF INVENTION
[0004] The invention broadly relates to artificial intelligence and computer vision. More particularly, the invention relates to a system, or a platform, to enable users to develop an end-to-end artificial intelligence and computer vision models, and a corresponding method.
[0005] BACKGROUND OF THE INVENTION
[0006] Artificial intelligence is a rapidly evolving field within computer science that aims to enable machines to perform tasks that may include learning, problem-solving, image perception, reasoning, and language understanding. Generally, artificial intelligence may be achieved by the use of machine learning algorithms and models that enable patterns in data to be analysed for making future predictions on the data. To utilise artificial intelligence in performing a task, a suitable machine learning model has to be trained on a large amount of relevant data to achieve reasonable or satisfactory accuracy in performing the said tasks. However, due to tedious and repetitive manual coding, time factor remains to be the biggest challenge in training the models. Furthermore, there is a barrier of entry for users to train their own artificial intelligence and computer vision model, as they are required to have proper skills and expertise, especially in the field of programming and software engineering to appropriately train a machine learning model for a desired task. Among the prior arts that may relate to breaking the barrier of entry for users to train their own machine learning model may include US20220405587, which merely relates to the generation of a synthetic dataset for machine learning. As such, this prior art fails to enable its users to develop their own artificial intelligence model. Accordingly, a system and method to provide such enablement is desirable.
[0007] SUMMARY OF INVENTION
[0008] The present invention intends to provide a web-based system that operates on a cloud environment. The system may generally comprise one or more devices, each having at least one processor for operating one or more modules to provide a platform. Some of these devices may be located remotely away from the user, whereby they may be networked devices that are configured to share computing, storage, and network resources between each other, which allows them to form, or be part of, or is part of, a cloud computing infrastructure (i.e. cloud). Some of these devices may be located close to the user, whereby they may be one or more on-premise or local devices that form, or is part of, a local computing infrastructure, which is under ownership and / or control of the user. The user may interact with any one or both the cloud or the local computing infrastructure via their own end-user device over a network such as the internet.
[0009] The platform includes a first part or component that relates to an image processing in computer vision service, and a second part or component that relates to a deep learning model training service. The platform enables users to develop an end-to-end artificial intelligence vision software with artificial intelligence and computer vision frameworks in a code-free manner through the services provided by each part of the platform, providing enhanced user accessibility from data acquisition, image processing, annotation, augmentation, model training, model exportation to analysis and fine tuning of model.
[0010] Preferably, the first part involves modules that include an image acquisition, object detection, object recognition, image filtering, image recognition, image transformation, custom solution and image iteration to enable users to identify, locate and describe the desired elements within a visual data. Preferably, the first part involves further modules that include a feature detection and description module for pin-pointing specific areas within the images, and an image segmentation module for performing segmentation on images.
[0011] Preferably, the first part further involves modules that include the advanced functions of image processing to enable image enhancement from any of circle detection, colour specs, comer detection, edge detection, image segmentation, foreground extraction, hough line, morphological and feature matching.
[0012] Preferably, the first part further involves modules that include a machine learning toolbox module that enables machine learning support for any one of a combination of tasks that include object detection, tracking, and instance segmentation through utilizing the traditional computer vision algorithm.
[0013] Preferably, the second part involves modules that utilise the advanced deep learning algorithm to allow the development of artificial intelligence models in several steps, which consist of data acquisition, preprocessing, annotation, augmentation, training, model export, and analysis and fine tuning of the models.
[0014] Preferably, the second part further involved modules that include capability to fine tune the parameters for training the machine learning models.
[0015] Preferably, the second part further involves modules that are a comparison module that is configured to perform a comparison between a plurality of machine learning models.
[0016] Preferably, the second part further involves modules that include a testing module that is configured to perform testing of at least one machine learning module.
[0017] Preferably, the second part further involves modules that include a first web page module that is configured to display one or more web pages pertaining to the image processing service, and the second part further involves modules that include a second web page module that is configured to display one or more web pages pertaining to the deep learning model training service.
[0018] One skilled in the art will readily appreciate that the invention is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. The embodiments described herein are not intended as limitations on the scope of the invention.
[0019] BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To facilitate an understanding of the invention, there are illustrated in the accompanying drawings the preferred embodiments from an inspection of which when considered in connection with the following description, the invention, its construction and operation and many of its advantages would be readily understood and appreciated.
[0021] FIG. 1 is a diagram illustrating a flowchart describing interactions between the first part of the platform and a user.
[0022] FIG. 2 is a diagram illustrating a flowchart describing interactions between the second part of the platform and a user.
[0023] FIG. 3 is a diagram illustrating a flowchart that is a continuation of FIG. 2.
[0024] FIG. 4 is a flowchart describing a sequence of display of web pages that is shown to the user for the first part of the platform and the second part platform as provided by their respective web page modules.
[0025] FIG. 5 is a flowchart describing the steps involved within a training workflow of the first embodiment.
[0026] FIG. 6 is a schematic diagram illustrating an example implementation of the second embodiment.
[0027] FIG. 7 is a flowchart describing the steps involved within a training workflow of the second embodiment.
[0028] FIG. 8 is a diagram illustrating an example user interface that may be displayed to one or more users of the platform by the cloud.
[0029] DETAILED DESCRIPTION OF THE INVENTION
[0030] The present invention relates to a platform that provides a user-friendly online environment for developing robust and functional vision detection systems. It intends to streamline the development process by offering diverse functions of computer vision and providing easy access to users through a website. The present invention aims to develop an intuitive, web-based, and no-code computer vision development application that seamlessly caters to both technical and non-technical users. A corresponding method for implementing such a system is further described.
[0031] From hereon, one or more steps pertaining to the method of the present invention may be described. It is to be noted that the described steps are not to be interpreted as nonlimiting, and minor modifications to the steps (e.g. additions, omissions, repetitions, swaps, or the like) are permissible by a skilled person without substantial deviation from as described.
[0032] The present invention intends to provide a web-based system that generally operates on a cloud environment. The system may generally comprise one or more devices that may be in the form of computers, each having at least one processor for operating one or more modules to provide a platform. The processors of these devices may be, but shall not be limited to, a conventional processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), or a combination thereof. Some of these devices may be located remotely away from the user, whereby they may be networked devices that are configured to share computing, storage, and network resources between each other, which allows them to form, or be part of, or is part of, a cloud computing infrastructure (i.e. cloud). Some of these devices may be located close to the user, whereby they may be one or more on-premise or local devices that form, or is part of, a local computing infrastructure, which is under ownership and / or control of the user. The user may interact with any one or both the cloud or the local computing infrastructure via their own end-user device over a network such as the internet. The end-user device may be, by way of example, a desktop, a laptop, a smartphone, or the like.
[0033] The platform includes a first part or component that relates to image processing in computer vision service, and a second part or component that relates to a deep learning model training service. The platform enables users to develop an end-to-end artificial intelligence vision software with artificial intelligence and computer vision frameworks in a code-free manner through the services provided by each part of the platform, providing enhanced user accessibility to the entire model development life cycle, spanning from data acquisition, image processing, annotation, augmentation, training, model exportation to analysis and fine tuning of model.
[0034] Advantageously, the present invention shall empower clients in the manufacturing, healthcare, agriculture and other industries to develop a robust vision solution and achieve continuous performance improvement of artificial intelligence. The platform comprehensively caters to the requirements, encompassing image processing, data annotation, model development, evaluation, comparison, prediction, and deployment. The provided invention is a comprehensive solution that revolutionizes efficiency and drives innovation across industries by streamlining computer vision model development, enabling artificial intelligence developers, computer vision engineers, and students to rapidly build and deploy cutting-edge models. With the provided invention, users shall gain a competitive edge through enhanced accuracy, improved operational efficiency, and optimized decision-making. By leveraging the power of Artificial Intelligence, the provided invention empowers its users to unlock the full potential of artificial intelligence, driving transformative outcomes in the industries. Progress and innovation in artificial intelligence computer vision are currently limited by issues like time-consuming manual coding and restricted accessibility. Therefore, the present invention provides simplified approaches to overcome these obstacles and pave the way for a more fruitful and inclusive future in the field. Compared to the prior art, the present invention shall allow users to build and develop their own computer vision-based machine learning model through both traditional computer vision algorithm and advance deep learning algorithm under one single platform that includes data acquisition, image processing, data annotation, model development, evaluation to comparison, prediction, and deployment, eliminating the need of exporting data from / to any external platform.
[0035] The present invention may include a first part or component that relates to image processing in computer vision. This first part or component may be separated into four sections, which may include image filtering, object detection, feature description, and segmentation. These may be provided by one or more relevant modules.
[0036] Users may use the functions of the first part or component through the following sequence of steps. In a first step, the user may upload at least one image dataset. The image dataset may comprise a plurality of images. In a second step, the user may select regions of interest (ROI) within said images. In a third step, the user may select basic functions. In a fourth step, the user may provide input parameters. In a fifth step, the user may initiate the processing of the image dataset. In a sixth step, the platform may provide an output. In a seventh step, the user may select advanced functions. In an eighth step, the user may provide input parameters. In a ninth step, the user may initiate the advance processing of the image dataset. In a tenth step, the platform may provide an output. In an eleventh step, the user may initiate means to provide custom solutions.
[0037] In a twelfth step, the user may provide input variables. In a thirteenth step, the platform may provide calculation results. In a fourteenth step, the user may provide or input custom logic. In a fifteenth step, the platform may provide a logic assessment. In a sixteenth step, the platform may display an iteration page. In a seventeenth step, the user may upload a dataset folder. In an eighteenth step, the user may initiate iteration of the dataset on the platform. In a nineteenth step, the platform may provide an output. Finally, in a twentieth step, the user may download the output dataset.
[0038] The present invention may include a second part or component that relates to a deep learning - artificial intelligence training service. More specifically, the second part relates to a service that provides customized training of computer vision-based artificial intelligence models, which may include training of at least one machine learning model, transfer learning of at least one machine learning model, fine-turning of at least one machine learning model, machine learning model performance analysis, machine learning model comparison, and testing of the model.
[0039] The second part may involve six sections in sequence, which include a data collection section, model design section, model training section, validation section, testing and deployment section, which may be provided in the form of one or more web pages. These may be provided by one or more relevant modules.
[0040] The build section may comprise one or more steps in sequence. In one sequence of steps, there is included a first step that involves building the machine learning model, a second step that involves starting a new experiment, a third step that involves starting training of the model, a fourth step that involves uploading at least one dataset, and a fifth step the involves selecting a new model. In one alternative sequence of steps, there is included a first step that involves initiating transfer learning of a pre-built machine learning model, a second step that involves configuring parameters of the machine learning mode, a third step that involves validating the machine learning model, a fourth step that involves input values to the machine learning model, a fifth step that involves training the machine learning model, a sixth step that involves providing a report on the training, and a seventh step the involves generating files related to the machine learning model.
[0041] The track section may comprise one or more steps in sequence. In one sequence of steps, there is included a first step that involves tracking the machine learning model, a second step that involves viewing experiment details, a third step that involves viewing the report page, a fourth step that involves a re-validation process, a fifth step that involves uploading a dataset, a sixth step that involves validation, a seventh step that involves configuring parameters related to the machine learning mode, an eighth step that involves validation, a ninth step that involves inputting dataset and parameter values, a tenth step that involves validation, and a tenth step that involves providing a report.
[0042] The predict section may comprise one or more steps in sequence. In one sequence of steps, there is included a first step that involves starting the prediction, a second step that involves uploading the images and the files related to the machine learning model, a third step that involves validation, a fourth step that involves carrying out the prediction process, and a fifth step the involves providing prediction results.
[0043] The general operational flow between the user and the platform may follow one or more steps in sequence. In particular, generally, there is a first step that involves registration of the user on the platform, a second step that involves the user logging into the platform, a third step that involves the selection of services offered by the platform that image processing and / or machine learning model training services, a fourth step that involves the user creating a new project on the platform, a fifth step that involves the user uploading a dataset onto the platform, a sixth step that involves the user uploading a dataset onto the platform, a seventh step that involves the user selecting functions that relate to the selected service, an eighth step that involves the user selecting activities that relate to the selected service, a ninth step that involves the generation of output results, and a tenth step that involves the user logging out from the platform.
[0044] Advantageously, the present invention provides enhanced accessibility to users. More specifically, the platform offered by the present invention introduces an intuitive graphical user interface (GUI), thereby simplifying user interactions even for those without technical expertise in machine learning or coding.
[0045] Advantageously as well, the present invention also offers customization and enablement of precise configurations, as users can define precise parameters, align functions precisely and customize their pipeline based on their unique requirements.
[0046] Advantageously as well, the present invention also provides a comprehensive testing environment. More specifically, the present invention offers a sandbox-like space for hands-on experimentation and refinement of functions related to the Open Computer Vision Library (OpenCV), and effective optimization of settings for specific use cases.
[0047] Advantageously as well, the present invention also provides streamlined training and validation of artificial intelligence and computer vision models. More specifically, the present invention provides real-time reporting during training, re-validation with new datasets, and robust metrics to ensure continual model improvement.
[0048] Advantageously as well, the present invention also provides expanded model capabilities as there is provided further integration with machine learning model providers such as Ultralytics “You Only Loop Once” (“YOLO”), which provides a range of cutting-edge models for various computer vision tasks.
[0049] Advantageously as well, the present invention also provides variable and logic customization upon the machine learning models. More specifically, users can input custom conditional statements and rules to enable nuanced decision-making based on specific criteria.
[0050] Advantageously as well, the present invention also provides an iteration page for bulk analysis. More specifically, the present invention enables swift processing and categorization of results across multiple images to expedite assessments of extensive image collections.
[0051] Advantageously as well, the present invention also provides real-time assessment of machine learning models to determine whether they passes or fails in their tasks. More specifically, the present invention provides an instant evaluation that categorizes detections as 'pass' or 'fail,' thereby offering immediate feedback on image datasets.
[0052] Advantageously as well, the present invention also provides real-time reporting. More specifically, the present invention provides detailed metrics like Mean Average Precision (mAP) during the training phase to aid users in better decision-making and model adjustment.
[0053] Advantageously, the present invention also provides a unified training and prediction framework. More specifically, the present invention enables seamless integration of workflows to ensure a cohesive and streamlined user experience from training to realtime predictions on uploaded images.
[0054] The first part of the platform of the present invention, which relates to an image processing service, shall now be further described. In particular, the objective of the first part is to (i) simplify present OpenCV Library usage by providing an easy-to-use user interface, (ii) enable non-coders and anyone to apply and perform image processing effortlessly, and (iii) expand accessibility by making OpenCV Library usage more user-friendly for diverse applications.
[0055] The OpenCV library is a versatile and powerful open-source library that offers a comprehensive set of tools and algorithms for computer vision, image processing, and machine learning tasks. It provides a framework to perform a wide range of operations, from basic filtering and transformation to complex tasks like object detection, recognition, deep learning-based analysis, and the like.
[0056] The first part of the platform of the present invention relates to a user-friendly interface that abstracts the complexities of OpenCV's functionalities. Rather than writing intricate code to execute operations, users interact with an intuitive interface where they can input parameters and settings directly. This abstraction layer simplifies the utilization of OpenCV Library’s diverse functions, allowing individuals without extensive programming knowledge to harness the power of image processing, object detection, and other computer vision tasks. By streamlining the process of defining parameters within the software's interface, users can perform various operations within the OpenCV Library efficiently, minimizing the need for in-depth coding expertise and accelerating the development and implementation of computer vision solutions.
[0057] It is to be stated that the first part is configured to support tasks that include any one or a combination of (i) image filtering and transformation, which includes operations like blurring, sharpening, resizing, and colour space conversions, (ii) object detection and recognition, which includes identifying and locating objects in images or videos, (iii) feature detection and description, which involves locating and describing specific points or areas in images, and (iv) image segmentation, which involves dividing an image into segments to simplify analysis.
[0058] With this, the first part of the present application involves the usage of computer vision tools, provided by the OpenCV Library to detect and analyse objects in images and videos. The first part may further involve modules that enable the implementation of image filtering, object detection, feature description, and segmentation techniques. This shall enable elements within visual data to be identified, located, and described, thereby enabling robust analysis and recognition.
[0059] For implementing the first part or component of the platform, the present invention may comprise an image filtering and transformation module. The image filtering and transformation module may be configured to perform the step of enabling image enhancement to be performed in a simplified manner through a user-friendly interface, and optimizing usage of the OpenCV Library for tasks like blurring, resizing, and color conversions.
[0060] For implementing the first part or component of the platform, the present invention may further comprise an object detection and recognition module. The object detection and recognition module may be configured to perform the step of implementing precise object recognition using the OpenCV Library, and ensure accurate detection in images.
[0061] For implementing the first part or component of the platform, the present invention may further comprise a feature detection and description module. The feature detection and description module may be configured to perform the step of pin-pointing specific areas within images accurately by leveraging on the functionalities of the OpenCV Library, thereby facilitating detailed analysis.
[0062] For implementing the first part or component of the platform, the present invention may further comprise an image segmentation module. The image segmentation module is configured to perform the step of segmenting images for focused analysis by utilizing the OpenCV Library to simplify complex data interpretation.
[0063] The process flow for the interactions between the first part or component of the platform and the user is illustrated in the flowchart of FIG. 1.
[0064] For implementing the first part or component of the platform, the present invention may further comprise a first web page module configured to perform the step of facilitating the display of one or more web pages that include (i) a basic functions webpage, (ii) an advanced functions webpage, (iii) a custom solution webpage, and (iv) an iteration webpage. The sequence of display of these web pages may be according to as shown in FIG. 4
[0065] Within the first part or component of the platform, the basic functions webpage and the advanced functions webpage may display functions that shall enable users to key in the parameters based on their criteria and preferences. Users have the flexibility to input specific parameters aligned with their unique requirements, enabling customization and precision in their desired operations. By aligning these parameters with their unique requirements, individuals can customize their pipeline precisely, ensuring that the outcomes match their intended objectives. Moreover, the comprehensive testing environment as provided by the platform may serve as a playground for exploration and refinement. It offers a safe space for users to experiment with a multitude of functions housed within the OpenCV Library. This hands-on approach enables individuals to finetune operations, adjusting parameters and settings according to their criteria and preferences. This process may be done iteratively, and shall empower users to not only understand the nuances of the OpenCV Library and its capabilities, but also to optimize these functions to suit their specific use cases.
[0066] Within the first part or component of the platform, the custom solution webpage offers users the freedom to shape their own unique processes. It enables the present invention to serve as a versatile platform where users can define and refine their specific requirements for various operations. This solution stands out for its adaptability, allowing users to personalize their interactions with the system based on their distinct needs and preferences by providing their own variables and logic parameters. More specifically, users may define variables that directly influence the functionality of the platform. Here, they may input parameters pertinent to their tasks, such as specifying calculations like determining the area of detected circles or computing distances between points along identified lines. This customization ensures precision, enabling users to fine-tune the system's behaviour according to their exact requirements.
[0067] Within the first part or component of the platform, the iteration webpage enables the users to gain a powerful tool to expedite bulk analysis of image datasets. By uploading a JSON file housing their predetermined parameters and subsequent logic, coupled with a folder of images, users can seamlessly extend their analyses across multiple images. For instance, if parameters were initially set to identify screw-related features, such as detecting circles within screws, applying this configuration to an entire folder of images instantly generates detection results for each image. Moreover, the platform evaluates these outcomes against pre-set logic, and promptly categorizes detections as 'pass' or 'fair based on the user-defined criteria. This capability significantly accelerates the assessment process, ensuring consistent and efficient analysis of extensive image collections.
[0068] Furthermore, for implementing the first part or component of the platform, the present invention may further comprise a variable definer module. More specifically, within the platform, users may access a dedicated section for defining variables that directly influence the platform’s functionality. Here, users may input parameters pertinent to their tasks, such as specifying calculations like determining the area of detected circles or computing distances between points along identified lines. This customization ensures precision, enabling users to fine-tune the system's behaviour according to their exact requirements.
[0069] Furthermore, for implementing the first part or component of the platform, the present invention may further comprise a logic segment module. The logic segment module provides users with a canvas to input their own set of conditional statements and rules. This empowers users to create sophisticated logic sequences that drive the platform’s responses. For instance, users might articulate conditions like triggering a 'FAIL' outcome if a detected circle's area falls below a predefined threshold. This feature allows for nuanced decision-making within the system, enhancing its adaptability and responsiveness to user-defined criteria.
[0070] Finally, it is to be noted that the first part may be implemented using programming languages known in the art, which may include, but shall not be limited to Python, FastAPI framework, TypeScript, Next JS framework, Tailwind CSS, MongoDB Atlas as database, or the like.
[0071] The second part or component of the platform of the present invention, which relates to a deep learning and artificial intelligence training service, shall now be further described. In particular, the objective of the second part is to (i) simplify and accelerate the computer vision development process, (ii) enable non-coders and anyone to apply and perform image processing effortlessly, and (iii) expand accessibility by making OpenCV Library usage more user-friendly for diverse applications.
[0072] The second part or component of the platform of the present invention relates to means to provide a no-code computer vision training service for users to develop and maintain their own computer vision machine learning model. The second part of the platform may be a web application that is hosted in a cloud database. The second part is configured to provide one or more features that include (i) enablement of seamless adoption of artificial intelligence by unlocking artificial intelligence computer vision capabilities for all users, (ii) simplifying artificial intelligence development lifecycle management through an agile software development lifecycle, (iii) enablement of hardware-free artificial intelligence solutions, which empowers user innovation while minimizing costs, and (iv) enable users to have access to cutting-edge development, whereby they can create state-of-the-art computer vision machine learning models with ease.
[0073] Within the second part or component of the platform of the present application, the platform operates within a bifurcated architecture. There is a frontend that primarily serves a graphical user interface (GUI) function, with its core purpose being user interaction. There is a backend in which all data processing requests are directed thereto for execution and subsequent response delivery to the frontend. The backend may function as an application protocol interface (API) for facilitating seamless communication and data exchange between the frontend and itself.
[0074] For example, consider a scenario where a user intends to initiate a machine learning model training process. The user may be prompted to upload a dataset and configure specific parameters. Upon completion, the user triggers the process by clicking the "Start" button. Subsequently, the frontend transmits the request, inclusive of the requisite data, to the backend. The backend, upon receipt of the request, undertakes the responsibility of executing the machine learning model training process in accordance with the provided data and parameters. This process exemplifies the clear demarcation of roles, with the frontend orchestrating user interactions and the backend orchestrating substantive computational tasks.
[0075] The second part or component of the platform of the present application may comprise a machine learning toolbox module, which is preferably provided by Ultralytics Y OLO. This enables the second part of the platform to incorporate features that support any one or a combination of object detection, tracking, instance segmentation, image classification, and pose estimation tasks.
[0076] In particular, the machine learning model toolbox module includes or provides a spectrum of models, which may include, but shall not be limited to, Y0L0v3, Y0L0v5, Y0L0v6, Y0L0v8, RT-DETR, or the like. Each of these models may be integrated into the platform for augmentation of each of their capabilities.
[0077] Within the second part or component of the platform, the training process of at least one machine learning model may be performed by invoking a training application protocol (API) interface of the machine learning toolbox module (i.e. an API provided by Ultralytics YOLO. This API call will involve fitting the provided data and parameters by the user, thereby initiating and overseeing the training process.
[0078] It is to be stated that the second part or component of the platform is configured to support tasks that relate to training, transfer learning, and fine-tuning of at least one computer vision machine learning model. More specifically, for implementing the second part, there may further include one or a combination of modules that include (i) an analysis module that is configured to perform one or more steps that relate to the task of analysing the performance of the machine learning model, (ii) a comparison module that is configured to perform one or more steps that relate to the task of making a comparison between a plurality of machine learning models, (iii) a tracking module that is configured to perform one or more steps that relate to the task of tracking the training of the machine learning model, and (iv) a testing module that is configured to perform one or more steps that relate to the task of testing the machine learning model.
[0079] The process flow for the interactions between the second part of the platform and the user is illustrated in the flowchart of FIGS 2 - 3.
[0080] For implementing the second part or component of the platform, the present invention may further comprise a second web page module configured to display one or more web pages that include (i) a build webpage, (ii) a track webpage, (iii) a report webpage, and (iv) a validation webpage, and (v) a predict webpage. The sequence of display of these web pages may be according to as shown in FIG. 4.
[0081] Within the second part or component of the platform, the report webpage may enable the platform to provide enhanced transparency and monitoring capabilities during the training process. As the machine learning model undergoes training, detailed metrics, which may include the Mean Average Precision (mAP) metric, may be dynamically presented on the report webpage in real-time. This real-time visualization empowers users to closely monitor the progress of the training process, facilitating informed decision-making.
[0082] Furthermore, within the report webpage, users may retain control throughout the training phase with the ability to cancel the process at any point, ensuring flexibility and efficiency in model development. Upon completion of the training, a comprehensive summary of the training results is generated and made available on the report webpage. This comprehensive report enables users to evaluate the performance of the machine learning mode, thereby providing insights into key metrics and facilitating informed adjustments for subsequent training iterations.
[0083] Within the second part, the tracking webpage may further streamline user engagement and facilitate comprehensive oversight, by incorporating experiment tracking. This webpage may consolidate information pertaining to all training experiments conducted by a particular user. Users can conveniently review various experiment details and assess model performance at a glance. This fosters efficient experiment management, allowing users to track and compare multiple models and configurations effortlessly. In essence, the tracking webpage may aid and enrich user experience by providing realtime insights during training, empowering users to make data-driven decisions, and offering efficient experiment tracking and model evaluation in a centralized manner.
[0084] Within the second part or component of the platform, the validation webpage provides continuous model validation for robust performance assessment. To address potential limitations associated with small or unrepresentative validation sets, users can initiate a re-validation process directly from the report page. Clicking a "Re-validate" button may initiate a smooth transition to the validation webpage. On this page, users may be prompted to upload a new dataset and select the previously trained model that they intend to revalidate. Upon commencing the re-validation process, the system may process the newly provided data and parameters. The results of the re-validation are generated and promptly displayed on the report page, denoted as "Version 2". This dualversion approach affords users a comparative analysis, fostering a deeper understanding of the adaptability of a machine learning model to diverse datasets and configurations.
[0085] Within the second part or component of the platform, the prediction webpage may enable users to intuitively train machine learning models. The prediction webpage may facilitate uploading of the images alongside a machine learning model file that is preferably saved in Py Torch format. This simplifies the prediction workflow, enabling users to promptly initiate the prediction process. Upon initiating prediction on the prediction webpage, the backend portion invokes Ultralytics YOLO's prediction API. This API call leverages the user-provided images and the pre-trained model in the Py Torch format to generate predictions. The results of the prediction are then seamlessly integrated and displayed directly on the uploaded images, offering users a comprehensive and visually accessible interpretation of the insights of the machine learning model.
[0086] The unified framework provided by the second part or component of the platform encompasses both training and prediction functionalities that leverage state-of-the-art Al architectures. Users are enabled to navigate the entire process effortlessly, from training sophisticated machine learning models to obtaining real-time predictions, fostering a cohesive and streamlined experience within the ambit of computer vision development.
[0087] Finally, it is to be noted that the second part or component of the platform may be implemented using programming languages known in the art, which may include, but shall not be limited to Python, FastAPI framework, TypeScript, Next JS framework,
[0088] Tailwind CSS, MongoDB Atlas for database, or the like.
[0089] It should be noted that the modules described so far may be operated under a cloud computing infrastructure. Eliminating the needs of having high-performance hardware, these modules can empower users to perform model training anytime and anywhere, streamlining the model training process and reducing time constraints.
[0090] The descriptions of the present application, so far, had described a first embodiment, wherein platform provides a workflow involving annotation to training, which is intended to be fully implemented within the cloud computing infrastructure (i.e. cloud).
[0091] FIG. 5 illustrates a flowchart that describes the steps involved within a training workflow of the first embodiment. First, there may be a step 510, which involves uploading images and annotations by at least one user via their end-user device. Next, there may be a step 520, which involves the user interacting with the cloud for image annotation and storage via an interface provided by the cloud, which corresponds to the first part or component of the platform. Next, there may be a step 530, which involves performing training of at least one machine learning model on the cloud, wherein the cloud is to handle all training tasks, which corresponds the second part or component of the platform. The model may be trained based on the YOLO algorithm. Next, there may be a step 540, which involves storing the trained model in the cloud. Next, there may be a step 550, which involves retrieving the trained model by the user, wherein the user may download the trained model to their end-user device.
[0092] While providing such a workflow within the cloud can facilitate collaboration and scalability, modem machine learning models for image analysis such as object detection, classification, and segmentation require large amounts of data and complex training processes to achieve high accuracy. Such training often involves specialised hardware, such as graphics processing units (GPUs), and can be computationally and financially expensive. This significantly increases resource consumption and associated costs for subscribing, hosting and / or maintaining the cloud, for supporting the resources of the machine learning models. Furthermore, the cloud may become overloaded, especially when multiple users simultaneously interface with the cloud to request and perform training sessions that involve models which have a large learning capacity or a large number of learning parameters.
[0093] The present invention may be configured to be of a second embodiment, which similarly relates to a system and method that provides a platform. The platform of the second embodiment may include a first part or component that relates to image processing in computer vision service, and a second part or component that relates to a deep learning model training service. However, the platform of the second embodiment is implemented within a cloud computing infrastructure (i.e. cloud) and a local computing infrastructure.
[0094] More specifically, within the second embodiment of the platform, the first part or component that relates to image processing and / or image annotation may be implemented via the cloud. However, second part or component that relates to a deep learning model training service may be implemented via the local computing infrastructure, whereby on-premises or local training of machine learning models is performed on the local computing infrastructure though deep learning frameworks (e.g., Ultralytics YOLO, etc.). The cloud and the local computing infrastructure may be configured to interact and / or be interfaced with each other for the storage and retrieval of machine learning models. As such, the second embodiment may be regarded one that provides a hybrid workflow for computer vision tasks such as object detection, as this second embodiment combines (i) the processing capabilities of the cloud, and (ii) processing capabilities of at least one on-premise or local device of a local computing infrastructure, for training of at least one machine learning model.
[0095] The need for the second embodiment stems from the fact that there exists a certain group of users, such as organisations or power users, which have robust on-premise or local devices with high performance hardware, which may be, by way of example, high- end desktops or dedicated servers with GPUs. By providing a platform according to the first embodiment, these robust on-premise or local devices of these users may be underutilised. Furthermore, by only using the cloud for collaborative annotation and central data storage, and offloading the resource-intensive training step to on-premise or local devices, cloud computing costs may be reduced and concurrent use between one or more users may be managed more efficiently. Hence, the hybrid workflow of the second embodiment may continue to offer the benefits of cloud-based annotation and data management, while delegating computationally expensive training tasks to onpremises or local devices
[0096] It is to be noted that the second embodiment of the present invention may be regarded as an extension of the first embodiment of the present invention. Thus, it is to be understood that the descriptions for the first embodiment of the present invention (e.g., modules, etc.) may similarly be applicable for the second embodiment.
[0097] The second embodiment may be configured to provide cloud-based annotation, which may be implemented within a first part or component of the platform. In particular, users may be allowed to interact with a web-based interface to upload and annotate one or more images via the cloud. The annotations for each of these images may then be stored within a blob container, along with its corresponding metadata, in a cloud database.
[0098] The second embodiment may be further configured to provide trigger and metadata storage. In particular, once users finalize annotations and request model training, which includes hyperparameters for higher-accuracy demands, this request may be logged in a database of the cloud (i.e. cloud database), wherein the request may be linked with a user ID number or a project ID number.
[0099] The second embodiment may be further configured to provide orchestration of localised training. More specifically, there may be a local agent module that is deployed within the on-premise or local devices of the local computing infrastructure, which periodically check the cloud database for any pending requests that were made by the users of the platform. When the local agent module finds a pending request, it will download the relevant images and their corresponding labels from the cloud database. The labels downloaded from the cloud database may be YOLO-compatible.
[0100] The second embodiment may be further configured to provide training of machine learning models on on-premise or local devices of the local computing infrastructure, such as a local server having a GPU, through a machine learning framework, such as the Ultralytics YOLO framework. This shall reduce cloud resource consumption and address the cost and concurrency constraints inherent in cloud-only solutions.
[0101] The second embodiment may be further configured to provide export and storage of trained machine learning models. In particular, a trained model may be exported into a desired format after training finishes, and the trained model of the desired format may then be uploaded back into the blob container. The format in which the trained model may be exported into may be, by way of example, the Open Neural Network Exchange (ONNX) format, or the like. The cloud database may then be updated, and the cloud may alert the user and / or other users that a freshly trained model is now available.
[0102] By delegating resource-intensive training to local hardware within local computing infrastructure provided by the users of the platform, the second embodiment of the present invention dramatically lowers cloud expenses, supports multiple concurrent training requests without excessive cloud scaling costs, and preserves the convenience and collaboration features of cloud-based annotation and storage.
[0103] FIG. 6 illustrates a schematic diagram of the second embodiment. In the second embodiment, there may still be a cloud 600 and a cloud database 610. The cloud 600 may include modules for image annotation and storage. Users may still connect to and interact with the cloud 600 for annotating images in the cloud 600, and storing annotated images and their metadata thereon. The cloud database 610 may include one or more data storage mediums that store training requests, hyperparameters, user identifications, or the like. Further included in this second embodiment is a local agent module 621 that may be operated by an on-premise or local device 620 of the local computing infrastructure. The local agent module 621 may be configured to poll the cloud database 610 for pending jobs. In the event that a job is found, the local agent module 621 may download the images and labels into the on-premise or local device, and may prompt its local training module 622 to facilitate training of a machine learning model locally on the on-premise or local device 620. When local training of the machine learning model is completed, the local agent module 621 may then upload the trained machine learning model to a cloud storage 630, and update the cloud database 610 with information regarding the trained machine learning model.
[0104] The second embodiment may comprise modules that may be operated by processors of one or more devices within a cloud computing infrastructure, i.e. a cloud 600, which include an annotation interface module, at least one cloud database 610, and at least one cloud storage 630.
[0105] The annotation interface module may provide a browser-based or API-driven environment that allows users to interact with the cloud 600 for uploading and labelling one or more images that are to be provided to a machine learning model for its training later on.
[0106] The cloud database 610 may be provided by one or more data storage mediums that may be part of the cloud 600. The cloud database 610 may store data in a structured format, which may be in the form of tables. The cloud database 610 may be configured to manage accounts of one or more users of the platform, identification (ID) codes of the users of the platform (e.g. user ID numbers), identification (ID) codes of the projects on the platform (e.g. project ID numbers, etc.), and logs of training requests, which may include hyperparameters of the machine learning models.
[0107] The cloud storage 630 may be implemented as a blob container, which is provided and supported by one or more data storage mediums that may be part of the cloud 600. The cloud storage 630 may store and / or retain both (i) the raw image files used for training of the machine learning models, and (ii) the final annotation files that may correspond to a relevant algorithm used for training of the machine learning models, such as the YOLO algorithm.
[0108] The second embodiment may comprise modules that may be operated by processors of one or more on-premise or local devices 620 of a local computing infrastructure, which include the local agent module 621 and its local training module 622, for executing a machine learning model training framework so that a machine learning model is trained. The local agent module 621 may be a daemon or software program that is configured to periodically checks the cloud database for new training tasks. The local agent module 621 may also be further configured to (i) fetch at least one dataset relevant to training of a machine learning model from the cloud storage 630 (i.e. blob container), (ii) facilitate training execution of the machine learning model via an algorithm (i.e. the YOLO algorithm), (iii) convert the trained machine learning model into a suitable format such as the ONNX format, and (iv) upload the format-converted trained machine learning model back to the cloud storage 630 (i.e. the blob container). Meanwhile, the local training module 622 may be part of the local agent module 621. The local training module 622 may be a machine learning model training framework, such as the Ultralytics YOLO framework, which is installed on an on-premise or local device 620 that may be, by way of example, in the form of a GPU-equipped local server.
[0109] FIG. 7 illustrates a flowchart that describes the steps involved within a training workflow of the second embodiment.
[0110] First, there may be a step 710, which involves performing data annotation and data conversion. More specifically, the cloud 600 may provide an interface that may facilitate users to create a new project for uploading one or more images to the cloud 600. The cloud 600 may also enable and allow users to annotate the images uploaded to the cloud. Optionally, the cloud 600 may support multiple annotation formats and may convert them into YOLO-compatible label files. By way of example, high-level metadata, such as classes and bounding box coordinates, may be stored in the cloud storage 630 (i.e. blob container), thereby ensuring all annotated files remain in a centralised location.
[0111] Next, there may be a step 720, which involves logging at least one training request. More specifically, the user may provide a training request, which specifies the images that are to be used for training of the machine learning model, and hyperparameters of the machine learning model such as batch size and epoch count. These hyperparameters may be tuned for higher accuracy. As the training request is logged, a new row or record may be created in the cloud database 610, which associates the request with an ID number of the user, an ID number of the project, a timestamp of the request, and specified training parameters.
[0112] Next, there may be a step 730, which involves initialising or starting the local agent module 621 for it to operate on the on-premise or local device 620 of the local computing infrastructure.
[0113] Next, there may be a step 740, which involves the local agent module 621 polling the cloud database 610 at pre-set intervals for jobs that relate to pending training requests for at least one machine learning model to be trained. In the event that there is no pending training request, the local agent module 620 may be configured to repeat step 720 after a pre-set interval. In the event that the local agent module 620 finds at least one pending training request for at least one machine learning model to be trained, step 720 may proceed to step 730.
[0114] Step 750 may involve the local agent module 621 retrieving data that includes any one or a combination of images, annotated data, and / or hyperparameters relevant to the machine learning model to be trained, which may have been provided per steps 710 and 720, as well as the relevant project ID number and user ID number. These data may be retrieved from the cloud 600, or more specifically, the cloud storage 630 (i.e. the blob container). The data retrieved by the local agent module 621 may be downloaded onto the on-premise or local device 620.
[0115] Following step 750 may be step 760, which involves the local agent module 621 executing training of the machine learning model on the on-premise or local device 620 based on the retrieved data. More specifically, the local agent module 621 may launch the training job on the on-premise or local device 620, for training of the machine learning model to be facilitated by its local training module 622. This may result in a trained machine learning model and its corresponding artifacts.
[0116] Following step 760 may be step 770, which involves the local agent module 621 converting and / or exporting the trained machine learning model into a suitable format such as the ONNX format, or other formats known in the art, for the trained machine learning model to become a converted and trained machine learning model. Alternatively, the local agent module 620 may immediately convert and export the trained machine learning model into a standardised format, which is the ONNX format.
[0117] Following step 770 may be step 780, which involves the local agent module 621 uploading the converted and trained machine learning model, as well as its corresponding artifacts, to the cloud storage 630 (i.e. the blob container) for storage. The local agent module 621 may be further configured to store the converted trained machine learning model and its artifacts within the cloud storage 630 under a same project, alongside with corresponding images and annotations that were used for its training.
[0118] Following step 780 may be step 790, which involves the local agent module 620 updating the cloud database 610 to inform that the machine learning model, which was pending for training, has already been trained on the on-premise device of the user. As such, the record of the training request is updated to mark the request as completed. The user that had submitted the training request may then be notified of the completion of the training via the cloud interface, via an automated message sent to email of the user, and / or via a software operating on the end-user device of the user (e.g., a mobile application, etc.). With that, step 790 may loop back to step 740 to check for more jobs, and repeat therefrom.
[0119] In certain alternative implementations, the local agent module 621 may have been initialised or started prior to steps 710 and 720. In these certain alternative implementations, step 720 may directly proceed to step 740.
[0120] FIG. 8 illustrates an example user interface (UI) 800 that may be displayed to one or more users of the platform as a cloud web interface, when the users interact with the cloud via their end-user device. The example UI 800 may also be part of a dashboard interface that lists down current projects of the users and training requests of the users. In particular, the example UI 800 indicates one or more users of the platform that may be requesting training sessions in a concurrent manner. The example UI 800 as per FIG. 8 may display a first user A, a second user B, and at least one third user C. Each user A - C may be requesting training of a machine learning models for different projects or projects of the same. The example UI 800 as per FIG. 8 indicates a request for training a machine learning model made by the first user A is completed, and a link to download the trained model from the blob container may be made available to the first user A. The example UI 800 as per FIG. 8 further indicates a request for training a machine learning model made by the second user B, which is in progress, wherein the machine learning model is currently being trained. The example UI 800 as per FIG. 8 further indicates a request for training a machine learning model made by the third user C is currently being queued for waiting for local and / or cloud resources to become available, so that training of the machine learning mode as requested by the user C may be performed.
[0121] Advantageously, the second embodiment of the present invention may provide reduced cloud costs. By shifting the training process, which is the most resource-intensive step, to local hardware, the subscription and usage fees for cloud resources are greatly minimize. This may be particularly advantageous when there are multiple users or clients that intend to train machine learning models in a concurrent manner.
[0122] Advantageously as well, the second embodiment of the present invention may also provide scalability and flexibility. In particular, organisations can adjust their local hardware capacity according to anticipated workloads. Furthermore, upgrading or adding specialised hardware (e.g. graphics processing units (GPUs)) onto on-premise or local devices may be more cost-effective than scaling cloud instances.
[0123] Advantageously as well, the second embodiment of the present invention may also provide concurrent user support. In particular, there may be multiple on-premise or local devices, each operating separate local agent modules. This shall enable support of simultaneous training jobs without incurring excessive cloud costs.
[0124] Advantageously as well, the second embodiment of the present invention may also provide collaboration and centralised data storage within for an organisation. In particular, the annotation and results may remain in a centralised cloud storage (i.e. blob container), which is accessible via an interface of the cloud, thereby preserving the benefits of remote and / or shared work on large datasets.
[0125] The second embodiment of the present invention may be implemented in one or more exemplary settings.
[0126] In one exemplary setting, a small company having at least one on-premise or local device, such as a local server, may be configured to implement the aforementioned second embodiment. The local server may have at least one specialised hardware (e.g., GPU), and may operate the local agent module. The local server may be configured to handle training request and perform training of machine learning models occasionally, or for a certain period of time (e.g. during night-time), thereby reducing associated costs for maintaining one or more cloud instances.
[0127] In another exemplary setting, an enterprise having a plurality of on-premise or local devices may be configured to implement the aforementioned second embodiment. The on-premise or local devices of the enterprise may each have one or more specialised hardware (e.g., GPUs), and may each operate at least one local agent module. The second embodiment may enable different teams within the enterprise to share a single annotation workflow for a machine learning model, but train the machine learning model on separate on-premise or local devices in a simultaneous manner.
[0128] In yet another exemplary setting, an academic or research team having multiple researchers may utilise the second embodiment of the present application. In particular, each researcher may store all annotated datasets in the cloud for easy access, but each researcher may deploy at least one local agent module on a high-performance computing (HPC) cluster of their research institution, thereby reducing associated costs for maintaining one or more cloud instances and saving budget for other research tasks.
[0129] The present invention may also be configured to be of a third embodiment, which similarly relates to a system and method that provides a platform. The platform of the third embodiment may include a first part or component that relates to image processing in computer vision service, and a second part or component that relates to a deep learning model training service. However, the platform of the third embodiment may be fully implemented within a local computing infrastructure. It is to be understood that the implementation of the third embodiment may be based on the descriptions of any one or both the first embodiment and the second embodiment.
[0130] The present disclosure includes as contained in the appended claims, as well as that of the foregoing description. Although this invention has been described in its preferred form, it is understood that the present disclosure of the preferred form has been made only by way of example and numerous changes in the details of the construction, combination and arrangements of parts may be resorted to without departing from the scope of the invention.
Claims
CLAIMS1. A sy stem compri sing one or more devices that operate on cloud environment, locally, or both, to provide a platform that includes a first part that relates to an image processing in computer vision service which utilises Computer Vision algorithms; and a second part that relates to a deep learning model training service which utilises a deep learning algorithm; wherein the platform enables users to develop an end-to-end artificial intelligence vision software with artificial intelligence and computer vision frameworks in a code-free manner through the services provided by each part of the platform.
2. The system according to claim 1, wherein the first part involves modules that include an image acquisition, object detection, object recognition, image filtering, image recognition, image transformation, custom solution and image iteration to enable users to identify, locate and describe the desired elements within a visual data.
3. The system according to claim 2, wherein the first part involves further modules that include a feature detection and description module for pin-pointing specific areas within the images, and an image segmentation module for performing segmentation on images. the advance functions of image processing to enable image enhancement from any of circle detection, colour specs, comer detection, edge detection, feature matching, image segmentation, foreground extraction, hough line, morphological and matching feature.
4. The system according to any one of the preceding claims, wherein the first partinvolves modules that include a machine learning toolbox module that enables machine learning support for any one of a combination of tasks that include object detection, tracking, and instance segmentation through utilizing the traditional computer vision algorithm.
5. The system according to claim 4, wherein the second part further involves modules that utilise the deep learning algorithms to allow the development of artificial intelligence models in several steps, which consists of data acquisition, pre-processing, annotation, augmentation, training, model export, and analysis and fine tuning of the models.
6. The system according to any one of claims 4 or 5, wherein the second part further involves modules that include capability to fine tune the parameters for training the machine learning models.
7. The system according to any one of claims 4 to 6, wherein the second part further involves modules that are a comparison module that is configured to perform a comparison between a plurality of machine learning models. a testing module that is configured to perform testing of at least one machine learning module.
8. The system according to any one of claims 4 to 7, wherein the first part further involves modules that include a first web page module that is configured to display one or more web pages pertaining to the image processing service; and the second part further involves modules that include a second web page module that is configured to display one or more web pages pertaining to the machine learning model training service.
9. A method comprising the steps of providing a one or more devices that operate on a cloud environment, locally, or both, to provide a platform that includes a first part that relates to an image processing service in computer vision; and a second part that relates to a deep learning model training service; wherein the platform enables users to develop an end-to-end artificial intelligence vision software with artificial intelligence and computer vision frameworks in a code-free manner through the services provided by each part of the platform, providing enhanced user accessibility from data acquisition, image processing, annotation, augmentation, training, model exportation to analysis and fine turning of model.