System and method for a machine learning evaluation pipeline

A unified pipeline automates the ML evaluation process by integrating a requirement management layer, execution layer, and user interfaces, addressing inefficiencies in fragmented ML evaluation systems and enhancing automation.

JP7702016B2Active Publication Date: 2025-07-02WOVEN BY TOYOTA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024066930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-15
Filing Date
2024-04-17
Publication Date
2025-07-02
Estimated Expiration
2044-04-17

AI Technical Summary

Technical Problem

Existing machine learning (ML) model evaluation processes are fragmented, requiring separate systems for obtaining test data, executing the model, and interpreting results, which are inefficient and labor-intensive.

Method used

A unified pipeline integrates a requirement management layer, execution layer, and user interfaces to automate the ML evaluation process, allowing for seamless interpretation and execution of evaluation requirements, and displaying results through a user interface.

Benefits of technology

The integrated pipeline rationalizes and automates the ML evaluation process, improving efficiency and reducing human intervention by encapsulating the entire evaluation within a single, automated workflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702016000001
    Figure 0007702016000001
  • Figure 0007702016000002
    Figure 0007702016000002
  • Figure 0007702016000003
    Figure 0007702016000003
Patent Text Reader

Abstract

To provide a method, system, and device for evaluating a machine learning (ML) model.SOLUTION: A method may include: receiving, by a requirements management layer, at least one requirement obtained from a storage layer; interpreting, by the requirements management layer, the at least one requirement; and transmitting, by the requirements management layer, instructions to perform an ML evaluation process to an execution layer based on the interpreted requirements, where the execution layer transmits an output signal with the results of the ML evaluation process upon completing the ML evaluation process.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary embodiments consistent with the present disclosure relate to providing a pipeline for evaluating a machine learning model.

Background Art

[0002] Machine learning (ML) models can be used to automate various tasks. When developing an ML model, a developer may have certain criteria or parameters that the ML model needs to meet. For example, if an ML model is intended to automate a safety-critical task, the ML model may need to achieve a certain reliability assessment. Therefore, in order to ensure that the ML model can meet the standards, the developer may need to test and evaluate the ML model by performing ML evaluation.

[0003] In related technologies, simple metrics can be used to automatically determine whether a model is good. For example, the mean average precision (mAP) can be used as a simple number indicating whether an ML model is good, bad, or excellent.

[0004] The process of evaluating an ML model used by conventional systems and methods can be limited and slow. In particular, usually, each component used in the ML evaluation process is separate. For example, the means for obtaining test data, the ML model itself, and the ML evaluation test unit can all be in separate systems. Further, in related technologies, each step in the evaluation process may require a human user to interpret the results of the ML evaluation and consider how to repeat the evaluation process to obtain a more optimal ML model.

[0005] Therefore, there is a need for a more rationalized and automated method for performing ML evaluation.

Summary of the Invention

[0006] According to one or more exemplary embodiments, an apparatus and a method for evaluating a machine learning (ML) model are provided. In particular, the apparatus and method according to the exemplary embodiments receive at least one requirement obtained from a storage layer in a requirement management layer, interpret the requirement in the requirement management layer, and send, by the requirement management layer, an instruction for performing an ML evaluation process to an execution layer based on the interpreted test parameters. When the ML evaluation process is completed, an output signal having the result of the ML evaluation process may be sent by execution. Based on the output signal, the result information may be displayed on a user interface (UI). Accordingly, the entire process of configuring and performing the ML evaluation process can be rationalized / capsulated in a single pipeline and, optionally, presented to the user using a user interface connected to the layers of the pipeline, thus improving the automation of the evaluation process.

[0007] According to an embodiment, a method for evaluating an ML model may be provided. The method may include receiving, by a requirement management layer, at least one requirement obtained from a storage layer, interpreting, by the requirement management layer, the at least one requirement, and sending, by the requirement management layer, an instruction for performing an ML evaluation process to an execution layer based on the interpreted test parameters, wherein the execution layer sends an output signal having the result of the ML evaluation process when the ML evaluation process is completed.

[0008] The at least one test parameter may be in the form of a requirement code (RaC) file.

[0009] The storage layer may communicate with a first user interface configured to allow a user to edit at least one test parameter.

[0010] The execution layer may be configured to send the output signal to a second user interface configured to display the result of the ML evaluation process.

[0011] The first user interface and the second user interface may be displayed simultaneously.

[0012] The execution layer may include an inference component and a unit test component. When receiving instructions for performing an ML evaluation process, the inference component is configured to receive test data and perform an inference process based on the instructions for obtaining an output from the test data and the ML model. The unit test component is configured to perform an evaluation process based on the instructions for obtaining the output from the ML model and metrics.

[0013] The inference component may be configured to receive test data from a test data storage layer.

[0014] According to an embodiment, an apparatus for evaluating a machine learning (ML) model may be provided. The apparatus may include at least one memory storing computer-executable instructions and at least one processor. The at least one processor is to receive at least one requirement obtained from a storage layer by a requirement management layer, interpret the at least one requirement by the requirement management layer, and transmit, by the requirement management layer, instructions for performing an ML evaluation process based on the interpreted requirement to an execution layer. The execution layer is to transmit an output signal having the result of the ML evaluation process when the ML evaluation process is completed. The at least one processor is configured to execute the computer-executable instructions to perform the above.

[0015] Additional aspects are described in part below, become apparent in part from the description, or may be realized by the practice of the presented embodiments of the disclosure.

Brief Description of the Drawings

[0016] Features, aspects, and advantages of specific preferred embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numerals indicate like elements.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0017] The following detailed description of exemplary embodiments refers to the accompanying drawings. The present disclosure provides examples and explanations, but is not intended to be comprehensive or to limit one or more exemplary embodiments to the exact forms disclosed. Modifications and variations are possible in view of the present disclosure or can be learned from the implementation of one or more exemplary embodiments. Further, one or more features or components of an exemplary embodiment can be incorporated into or combined with one or more features of another exemplary embodiment (or another exemplary embodiment). Additionally, in the flowchart diagrams and descriptions of operations provided herein, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed (at least partially) simultaneously, and the order of one or more operations may be switched.

[0018] Exemplary embodiments of the systems and / or methods and / or non-transitory computer-readable storage media described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. It will be apparent that the actual specific control hardware or software code used to implement the systems and / or methods is not a limitation of one or more exemplary embodiments. Accordingly, the operation and behavior of the systems and / or methods and / or non-transitory computer-readable storage media are described herein without reference to specific software code. It is understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.

[0019] Even if a particular combination of features is recited in the claims and / or disclosed herein, the combination is not intended to limit the disclosure of possible exemplary embodiments. Indeed, many of the features may be combined in ways that are not specifically recited in the claims and / or not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible exemplary embodiments includes each dependent claim combined with all the other claims in the set of claims.

[0020] Elements, acts, or instructions used in this specification should not be construed as important or essential unless specifically described otherwise. Also, the articles "a" and "an" used in this specification are intended to include one or more matters and may be used interchangeably with "one or more". If only one matter is intended, the term "one" or a similar term is used. Also, the terms "has", "have", "having", "include", "including", or the like used in this specification are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless specifically stated otherwise. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0021] Figure 1 is a diagram of exemplary components of a machine learning (ML) evaluation device 100. As shown in Figure 1, the ML evaluation device 100 may include a bus 110, a processor 120, a memory 130, a storage component 140, an input component 150, an output component 160, and a communication interface 170.

[0022] Bus 110 includes components that enable communication among the components of ML evaluation device 100. Processor 120 may be implemented in hardware, firmware, or a combination of hardware and software. Processor 120 may be a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In one or more exemplary embodiments, processor 120 includes one or more processors that are programmable to perform functions. Memory 130 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 220.

[0023] The storage component 140 stores information and / or software related to the operation and use of the ML evaluation device 100. For example, the storage component 140, together with a corresponding drive, may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium. The input component 150 includes components (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone) that enable the ML evaluation device 100 to receive information via user input or the like. Additionally or alternatively, the input component 150 may include sensors (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator) that detect information. The output component 160 includes components (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)) that provide output information from the ML evaluation device 100.

[0024] The communication interface 170 includes transceiver-like components (e.g., a transceiver and / or separate receivers and transmitters) that enable the ML evaluation device 100 to communicate with other devices via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 170 may enable the ML evaluation device 100 to receive information from and / or provide information to other devices. For example, the communication interface 170 may include, but is not limited to, an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, or the like.

[0025] The ML evaluation device 100 can perform one or more of the exemplary processes described herein. According to one or more exemplary embodiments, the ML evaluation device 100 can perform the process in response to the processor 120 executing software instructions stored by a non-transitory computer-readable medium such as the memory 130 and / or the storage component 140. The computer-readable medium is defined herein as a non-transitory memory device. The memory device includes a memory space within a single physical storage device or a memory space spanning multiple physical storage devices.

[0026] The software instructions can be read into the memory 130 and / or the storage component 140 from another computer-readable medium or from another device via the communication interface 170. The software instructions stored in the memory 130 and / or the storage component 140, when executed, can cause the processor 120 to perform one or more of the processes described herein.

[0027] Additionally or alternatively, hardwired circuitry can be used in place of or in combination with the software instructions to perform one or more of the processes described herein. Accordingly, one or more exemplary embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0028] The number and arrangement of components shown in FIG. 1 are provided as an example. In fact, the ML evaluation device 100 can include additional components, fewer components, different components, or differently arranged components compared to those shown in FIG. 1. Additionally or alternatively, a set of components (e.g., one or more components) of the ML evaluation device 100 can perform one or more of the functions described as being performed by another set of components of the ML evaluation device 100.

[0029] FIG. 2 is a system architecture diagram showing layers and components according to one or more exemplary embodiments. Layers (e.g., execution layer 200, requirements management layer 210, and storage layer 220) may be combined to be regarded as a machine learning (ML) evaluation pipeline.

[0030] According to an embodiment, an execution layer 200 may be provided. The execution layer 200 may include an inference component 201 and a unit test component 202. The execution layer 200 may be responsible for executing an ML evaluation process for the ML model 240. It should be understood that the execution layer 200 may implement any suitable means for performing the ML evaluation. The execution layer 200 may be configured to receive instructions from the requirements management layer 210. The execution layer 200 may also interact with a display results UI 260 to output the results of the evaluation. In particular, the execution layer 200 may transmit an output signal including the results of the evaluation upon completion of the evaluation, and according to some embodiments, the display results UI 260 may receive the output signal and display the results of the evaluation. The display results UI 260 may be implemented using any suitable UI. For example, the display results UI 260 may have a tabular format for listing the results of the evaluation tests, and the display results UI 260 may highlight specific portions of the test results based on whether the evaluation test passes or fails. Nevertheless, it should be understood that the display UI 260 may depend on a particular implementation manner as determined by those skilled in the art.

[0031] The execution layer 200 may also receive test data from the test data storage 222, either via the requirements management layer 210 or directly from the test data storage 222.

[0032] According to some embodiments, the inference component 201 may be responsible for receiving data from the test data storage 222 and inputting (i.e., inferring) the data to the ML model 240 for calculating an output using the ML model 240.

[0033] According to some embodiments, the unit test component 202 may be responsible for executing unit tests on the ML model 240. Specifically, the unit test may evaluate the output results from the ML model 240 obtained during the inference process performed on the ML model 240 using the inference component 201. Metrics regarding the performance of the ML model 240 may be obtained using the unit test component 202.

[0034] A requirements management layer 210 may be provided. The requirements management layer 210 may be responsible for interpreting files stored in the storage layer (e.g., requirements code (RAC) files 1, 2,..., N (221-1, 221-2, 221-N,...)) or test data from the test data storage 222. Based on the interpretation of the RAC files, the requirements management layer 210 may send instructions to the execution layer 200 to perform ML evaluations.

[0035] The storage layer 220 may include all data used by the execution layer 200 and the requirements management layer 210. The storage layer 220 may be implemented by any suitable storage means (e.g., a database, cloud storage, etc.). The storage layer 220 may include any number of RAC files 221-1, 221-2, 221-N and a test data storage 222. It should be understood that each RAC file and the test data storage may be stored on the same storage medium or different storage media.

[0036] Requirement Code (RAC) files 221-1, 221-2, 221-N may bear the memory of the expected behavior of the ML model 240. The expected behavior may have requirements (which may include test parameters). In particular, the RAC file may be in a human-readable format that specifies the performance metrics of the ML model to be tested, based on, for example, specific metric criteria that need to be achieved, specifications that need to be achieved, types of tests, etc. The RAC file may include requirements, pass criteria, test objectives, file paths of test data, and test conditions used during the ML evaluation process in the execution layer 200. Thus, such requirements, pass criteria, test objectives, test data paths, and test conditions can be easily interpreted by the requirement management layer 210 to determine how the ML test and evaluation of the ML model 240 should be performed by the execution layer 200.

[0037] According to some embodiments, the RAC files 221-1, 221-2, 221-N may be editable by the editing prompt UI 250. In particular, the editing prompt UI 250 may be a general-purpose command-line interface used to edit source files, or the editing prompt UI 250 may be a graphical user interface (GUI) having interactive elements that allow a user to drag and drop or select a predetermined configuration. Nevertheless, it should be understood that any suitable user interface may be implemented for the editing prompt UI 250.

[0038] The test data storage 222 may contain data used by the inference component 201 and input into the ML model 240. According to some embodiments, the test data storage 222 may be stored separately from the RAC file. The specific format of the data file for the test data may be any suitable format.

[0039] It should be understood that the editing prompt UI 250 and the display result UI 260 can be implemented in any suitable environment, for example, in a web interface or only on a local user device. According to some embodiments, it should also be understood that the editing prompt UI 250 and the display result UI 260 can be displayed separately or simultaneously.

[0040] According to an embodiment, a UI (which may include features of the editing prompt UI 250 and the display result UI 260) may be provided, which may include additional features for streamlining the evaluation process. For example, the user interface may communicate with the requirement management layer 210 along with an interface that enables the user to select which set of requirements from the storage layer 220 should be used and the ordering thereof, for example, whether the first set and the second set of requirements should be evaluated sequentially or in parallel. The user interface may also enable the user to select a test data set (e.g., from the test data storage 222). It is also envisioned that the user may operate across different candidate ML models to configure the same set of requirements to automatically indicate which candidate ML model has the best metrics.

[0041] FIG. 3 is a flowchart diagram showing a method 300 for evaluating an ML model according to one or more exemplary embodiments.

[0042] Referring to FIG. 3, in operation S310, requirements including test parameters are received by the requirement management layer 210 from the storage layer 220. According to some embodiments, the requirements (such as test parameters) may be one or more of the RAC files 221-1, 221-2, 221-N.

[0043] After receiving requirements from the storage layer 220, in operation S320, the requirements can be interpreted by the requirement management layer 210 to determine how to instruct the execution layer 200 to perform an ML evaluation on the ML model 240. This may include interpreting model adjustment parameters such as requirements, acceptance criteria, test objectives, test data paths, confidence thresholds, etc., interpreting test conditions based on the requirements, and determining appropriate tests and test parameters to be performed using the execution layer 200. Therefore, the requirements to be interpreted can be obtained by the requirement management layer 210.

[0044] In operation S330, the requirement management layer 210 sends an instruction to the execution layer 200 based on the test parameters interpreted from operation S320 to cause the execution layer 200 to perform an ML evaluation on the ML model 240. According to some embodiments, when the execution layer 200 finishes performing the ML evaluation, the execution layer 200 may send an output signal including the result of the evaluation. According to some embodiments, this output signal may be sent to the display result UI260.

[0045] FIG. 4 is a flowchart diagram showing a method 400 for evaluating an ML model using an inference component and a unit test component according to one or more exemplary embodiments. Layers and components similar to those shown in FIG. 2 can be used to implement method 400, and a complete description of steps similar to method 300 as shown in FIG. 3 can be omitted for readability.

[0046] Referring to FIG. 4, in operation S410, an instruction for performing an ML evaluation can be sent from the requirement management layer 210 and received by the execution layer 200 to cause the execution layer 200 to perform an ML evaluation on the ML model 240. The instruction received in operation S410 can be similar to that sent in operation S330 as described above with respect to FIG. 3.

[0047] In operation S420, the execution layer 200 may send an instruction to the inference component 201 to execute an inference operation using the ML model 240 together with test data (e.g., test data obtained from the test data storage 222). The output of this operation may include an inference log, and the inference log may include data from the execution of the inference operation on the ML model 240 using the test data.

[0048] In operation S430, the execution layer 200 may send the inference log from the inference component 201 to the unit test component 202.

[0049] In operation S440, the execution layer 200 may send an instruction to the unit test component 202 to evaluate metrics (e.g., by calculating and comparing metrics regarding test criteria according to requirements).

[0050] In operation S450, the execution layer 200 may output the evaluated metrics (which may include comparison results). According to some embodiments, this may include sending the metrics and comparison results to a UI such as the display result UI 260.

[0051] FIG. 5 is a flowchart diagram showing a method for evaluating an ML model including adding or updating requirements according to one or more exemplary embodiments. Layers and components similar to those shown in FIG. 2 may be used to implement method 500, and a complete description of steps similar to method 300 as shown in FIG. 3 and method 400 as shown in FIG. 4 above may be omitted for readability.

[0052] Referring to FIG. 5, in operation S510, requirements including test parameters are received by the requirement management layer 210 from the storage layer 220. This may be similar to operation S310 described above with respect to FIG. 3.

[0053] After receiving requirements from the storage layer 220, in operation S520, the requirements can be interpreted by the requirement management layer 210 to determine how to instruct the execution layer 200 to perform an ML evaluation on the ML model 240. This can be similar to operation S320 described above with reference to FIG. 3.

[0054] In operation S530, the requirement management layer 210 sends an instruction to the execution layer 200 based on the test parameters interpreted from operation S520 to cause the execution layer 200 to perform an ML evaluation on the ML model 240. This can be similar to operation S330 described above with reference to FIG. 3.

[0055] In operation S540, the instruction for performing the ML evaluation can be received by the execution layer 200 to cause the execution layer 200 to perform an ML evaluation on the ML model 240. This can be similar to operation S410 described above with reference to FIG. 4.

[0056] In operation S550, the instruction execution layer 200 can send an instruction to the unit test component 202 to evaluate metrics (e.g., by calculating and comparing metrics regarding test criteria according to the requirements). This can be similar to operation S440 described above with reference to FIG. 4. It should be understood that although not explicitly shown, according to some embodiments, operations similar to S420 - S430 may also be included before operation S550.

[0057] In operation S560, the execution layer 200 can output the evaluated metrics (which may include comparison results) and send the metrics and comparison results to a UI such as the display result UI 260. This can be similar to operation S450 described above with reference to FIG. 4.

[0058] In operation S570, the execution layer 200 may send a command to the requirement management layer 210 to add requirements to the storage layer 220 (e.g., by adding a new RAC file to the storage layer 220) or to update existing requirements (e.g., by editing an existing RAC file in the storage layer 220). The command may be sent to update ML model adjustment parameters such as, but not limited to, trust thresholds and test parameters for optimizing the overall performance of the system including the ML model 240.

[0059] In operation S580, the requirement management layer 210 receives the command sent in operation S570 and accordingly adds or updates requirements in the storage layer 220. Thereafter, the steps may be repeated from S510. According to this embodiment, an iterative process for exploring requirements and parameters that may not yet have been discovered can be implemented.

[0060] Based on the above embodiments, it can be understood that the entire process of configuring and executing the ML evaluation process can be rationalized / capsulated into a single pipeline and optionally presented to the user using a user interface connected to the layers of the pipeline, thus improving the automation of the evaluation process.

[0061] The above disclosure provides examples and explanations, but is not intended to be exhaustive or to limit one or more exemplary embodiments to the exact form disclosed. Modifications and variations are possible in view of the present disclosure or can be learned from the practice of one or more exemplary embodiments.

[0062] One or more exemplary embodiments may relate to systems, methods, and / or computer-readable media at any possible technical detail level of integration. Further, one or more of the components described above may be implemented as instructions stored in a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to perform operations.

[0063] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0064] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0065] The computer-readable program code / instructions for performing the operations can be in the form of assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, where the programming languages include object-oriented programming languages such as Smalltalk, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed fully on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or fully on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection to an external computer can be made (e.g., through the Internet using an Internet service provider). In one or more exemplary embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions to perform the aspects or operations.

[0066] The computer-readable program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or block(s) of the block diagram. The computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium storing instructions therein comprises an article of manufacture including instructions which implement the aspects of the functions / acts specified in the flowchart and / or block(s) of the block diagram.

[0067] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block(s) of the block diagram.

[0068] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible exemplary embodiments of systems, methods, and computer-readable media according to one or more exemplary embodiments. In this regard, each block in the flowchart or block diagram may represent a micro-service, module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function. The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks arranged differently compared to those depicted in the drawings. In one or more alternative exemplary embodiments, the functions recited in the blocks may occur without regard to the order and relationship depicted in the figures. For example, two blocks shown in succession may actually be executed simultaneously or substantially simultaneously, or the blocks may be executed in the reverse order depending on the related functions. It should also be noted that each block of the block diagram and / or flowchart diagram, and combinations of blocks of the block diagram and / or flowchart diagram, may be implemented by a dedicated hardware-based system that performs the specified function or action or by a combination of dedicated hardware and computer instructions.

[0069] It will be apparent that the systems and / or methods described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement the systems and / or methods is not a limitation of one or more exemplary embodiments. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.

Claims

1. 1. A method for evaluating a machine learning (ML) model, comprising: The method comprises: receiving, by a requirements management layer, at least one requirement retrieved from the storage layer; interpreting said at least one requirement by said requirement management layer; sending instructions by the requirements management layer to an execution layer to perform an ML evaluation process based on the interpreted requirements, the execution layer sending an output signal having a result of the ML evaluation process upon completion of the ML evaluation process; A method comprising:

2. The method of claim 1 , wherein the at least one requirement is in the form of a Requirements Code (RaC) file.

3. The method of claim 1 or 2, wherein the storage layer is in communication with a first user interface configured to enable a user to edit the at least one requirement.

4. The method of claim 3 , wherein the execution layer is configured to send the output signal to a second user interface configured to display the results of the ML evaluation process.

5. 5. The method of claim 4, wherein the execution layer comprises an inference component and a unit test component, and upon receiving the instructions to perform an ML evaluation process, the inference component is configured to receive test data and perform an inference process based on the test data and instructions to obtain outputs from the ML model, and the unit test component is configured to perform an evaluation process based on the outputs from the ML model and the instructions to obtain indicators.

6. 6. The method of claim 5, wherein the output from the ML model comprises an inference log, and upon completing the inference process, the inference component is configured to forward the inference log to the unit testing component, and the evaluation process includes evaluating indicators from the inference log, and the evaluated indicators are displayed in the second user interface.

7. 7. The method of claim 6, further comprising receiving, by the requirements management layer, an instruction to add or update at least one requirement in the storage layer, and wherein the output signal is sent from the execution layer upon completing the ML evaluation process based on the evaluated metrics.

8. 1. An apparatus for evaluating a machine learning (ML) model, comprising: The apparatus comprises: at least one memory storing computer executable instructions; At least one processor; Equipped with The at least one processor receiving, by a requirements management layer, at least one requirement retrieved from the storage layer; interpreting said at least one requirement by said requirement management layer; sending instructions by the requirements management layer to an execution layer to perform an ML evaluation process based on the interpreted requirements, the execution layer sending an output signal having a result of the ML evaluation process upon completion of the ML evaluation process; 23. An apparatus configured to execute the computer-executable instructions to:

9. The apparatus of claim 8 , wherein the at least one requirement is in the form of a Requirements Code (RaC) file.

10. The apparatus of claim 8 or 9, wherein the storage layer is in communication with a first user interface configured to enable a user to edit the at least one requirement.

11. The apparatus of claim 10 , wherein the execution layer is configured to send the output signal to a second user interface configured to display the results of the ML evaluation process.

12. 12. The apparatus of claim 11, wherein the execution layer comprises an inference component and a unit test component, wherein upon receiving the instructions to perform an ML evaluation process, the inference component is configured to receive test data and perform an inference process based on the test data and instructions to obtain outputs from the ML model, and the unit test component is configured to perform an evaluation process based on the outputs from the ML model and the instructions to obtain indicators.

13. 13. The apparatus of claim 12, wherein the output from the ML model comprises an inference log, and wherein upon completing the inference process, the inference component is configured to forward the inference log to the unit testing component, and wherein the evaluation process includes evaluating indicators from the inference log, and the evaluated indicators are displayed in the second user interface.

14. 14. The apparatus of claim 13, wherein the processor is further configured to execute the computer-executable instructions to receive, by the requirements management layer, instructions to add or update at least one requirement in the storage layer, and wherein the output signal is sent from the execution layer upon completing the ML evaluation process based on the evaluated metrics.

15. 1. A non-transitory computer-readable storage medium having instructions executable by at least one processor to cause the processor to perform a method, comprising: The method comprises: receiving, by a requirements management layer, at least one requirement retrieved from the storage layer; interpreting said at least one requirement by said requirement management layer; sending instructions by the requirements management layer to an execution layer to perform an ML evaluation process based on the interpreted requirements, the execution layer sending an output signal having a result of the ML evaluation process upon completion of the ML evaluation process; A non-transitory computer-readable recording medium comprising:

16. 20. The non-transitory computer-readable medium of claim 15, wherein the at least one requirement is in the form of a Requirements Code (RaC) file.

17. 17. The non-transitory computer-readable storage medium of claim 15 or 16, wherein the storage layer is in communication with a first user interface configured to enable a user to edit the at least one requirement.

18. 20. The non-transitory computer-readable storage medium of claim 17, wherein the execution layer is configured to send the output signal to a second user interface configured to display the results of the ML evaluation process.

19. 20. The non-transitory computer-readable storage medium of claim 18, wherein the execution layer comprises an inference component and a unit test component, wherein upon receiving the instructions to perform an ML evaluation process, the inference component is configured to receive test data and perform an inference process based on the test data and instructions to obtain outputs from an ML model, and the unit test component is configured to perform an evaluation process based on the outputs from the ML model and the instructions to obtain indicators.

20. 20. The non-transitory computer-readable storage medium of claim 19, wherein the output from the ML model comprises an inference log, and wherein upon completing the inference process, the inference component is configured to forward the inference log to the unit testing component, and wherein the evaluation process includes evaluating metrics from the inference log, and the evaluated metrics are displayed in the second user interface.

Citation Information

Patent Citations

  • Information processing method and information processing system

    JP2019096285A

  • Data analyzer, data analysis method, and data analysis program

    JP2021018508A

  • Allocating Processing Resources To Concurrently-Executing Neural Networks

    US20220194423A1

  • User acceptance test system for machine learning systems

    US20220300754A1