Vision inspection control system based on interactive artificial intelligence agent, and operating method thereof

WO2026205889A1PCT designated stage Publication Date: 2026-10-01LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004442
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-03-17
Filing Date
2026-03-19
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004442_01102026_PF_FP_ABST
    Figure KR2026004442_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an artificial intelligence-based vision inspection technology and, more specifically, to a vision inspection technology using an interactive artificial intelligence agent, which enables, through natural language conversation with a user, the entire process of generating, training, evaluating, and distributing a vision inspection model to be automated and managed in an integrated manner.
Need to check novelty before this filing date? Find Prior Art

Description

Conversational AI Agent-Based Vision Inspection Control System and Method of Operation

[0001] The present disclosure relates to artificial intelligence-based vision inspection technology, and more specifically, to vision inspection technology using a conversational artificial intelligence agent that enables the entire process from the creation, training, evaluation, and deployment of a vision inspection model to the creation, training, evaluation, and deployment of a vision inspection model through natural language conversation with a user to be automated and managed in an integrated manner.

[0002] Specifically, the present disclosure relates to a vision inspection control system based on a conversational artificial intelligence agent and a method of operation thereof. More specifically, the present disclosure relates to a vision inspection control system based on a conversational artificial intelligence agent and a method of operation thereof that enables the control of a vision inspection model by simulating a user's ambiguous natural language input based on multiple agents and replacing it with specific and quantitative parameters.

[0003] Furthermore, the present disclosure relates to an integrated vision inspection system using a conversational artificial intelligence agent and a method for constructing the same. More specifically, the invention relates to an integrated vision inspection system using a conversational artificial intelligence agent and a method for constructing the same, which acquires and analyzes various environmental information, including constraints of an inspection target and external linked data, from external resources, and performs active conversational reverse questioning to the user in the event of missing information, thereby enabling the integrated construction of a vision inspection model optimized for the actual application environment.

[0004] Recently, the importance of AI-based vision inspection systems for inspecting product quality and preventing the leakage of defective products is growing day by day in various manufacturing industries, such as smart factories.

[0005] Conventional vision inspection systems primarily adopt a method in which users directly manipulate specific tools provided through a graphical user interface (GUI) to develop and train inspection models.

[0006] However, this approach requires a high level of expertise in deep learning and repetitive manual intervention at each stage of model building, such as data preparation and labeling, model architecture design, and hyperparameter tuning.

[0007] As a result, there was a very high barrier to entry for general field workers, who are not AI experts, to directly build and efficiently utilize the system.

[0008] In particular, in actual manufacturing sites, subjective and qualitative requirements from workers frequently arise, such as "inspect a little more strictly" or "filter out appropriately without lowering the yield."

[0009] Conventional technology failed to automatically convert such ambiguous natural language instructions into quantitative control variables (e.g., defect detection thresholds) that the model could understand, resulting in inefficiency where operators had to manually manipulate parameters through countless trials and errors.

[0010] Furthermore, conventional model building methods have limitations in that they focus primarily on generating general-purpose models in ideal data environments at the laboratory level.

[0011] As a result, when deploying the constructed model to target equipment on an actual production line (e.g., edge devices), a problem arises in that it fails to properly reflect the unique physical constraints of the target infrastructure (e.g., inference processing speed limits, memory capacity, etc.) or the specific quality standards of the particular process.

[0012] In other words, due to the discrepancy between the laboratory environment and the actual mass production site, compatibility conflicts or a sharp decline in inference performance (Performance Drop) frequently occurred during model deployment.

[0013] Furthermore, even after model deployment, existing systems faced limitations in terms of post-management and maintenance (MLOps), as they consumed enormous time and quality costs when unexpected environmental changes occurred—such as shifts in lighting conditions at the manufacturing site, camera lens contamination, or the emergence of new types of defect patterns—requiring manual resetting of equipment hardware parameters or retraining the entire model from scratch.

[0014] Therefore, there is an urgent need for a new technical solution that supports even AI non-experts in easily building the entire lifecycle of vision inspection models, from planning to deployment, using only everyday natural language conversation. Furthermore, it can autonomously provide a robust vision inspection pipeline optimized for actual operating environments by independently recognizing ambiguous operator requirements and various constraints of the target environment and reflecting them in perfect quantitative figures.

[0015] One embodiment of the present disclosure is devised to solve the problems of the prior art as described above, and aims to provide a vision inspection technology using a conversational artificial intelligence agent that enables the entire process from the creation, training, evaluation, and distribution of a vision inspection model to the creation, training, evaluation, and distribution of a vision inspection model through natural language conversation with a user to be automated and managed in an integrated manner.

[0016] Specifically, one embodiment of the present disclosure aims to provide a conversational artificial intelligence agent-based vision inspection control system and a method of operation thereof, which enables the control of a vision inspection model by simulating ambiguous natural language input from a user based on multiple agents and replacing it with specific and quantitative parameters.

[0017] In addition, one embodiment of the present disclosure aims to provide an integrated vision inspection system using a conversational artificial intelligence agent and a method for constructing the same, which acquires and analyzes various environmental information, including constraints of an inspection target and external linkage data, from external resources, and performs active conversational reverse questioning to the user in the event of missing information, thereby enabling the integrated construction of a vision inspection model optimized for the actual application environment.

[0018] However, the technical problems to be solved by the present disclosure and the embodiments thereof are not limited to the technical problems described above, and other technical problems may exist.

[0019] One embodiment of the present disclosure comprises a method executed by a computer, wherein at least one processor of the computer receives at least one input data through at least one interface, wherein the input data includes at least one natural language input indicating a task objective; wherein the at least one processor detects at least one qualitative requirement embedded in the natural language input; wherein the at least one processor generates a plurality of agents having mutually conflicting objective functions based on the detected qualitative requirement, wherein the plurality of agents are each assigned at least one of a different persona or optimization goal and constitute a multi-agent collaboration structure that interprets the qualitative requirement in a multidimensional manner; wherein the at least one processor executes a simulation based on the generated plurality of agents; and wherein the at least one processor determines at least one optimal parameter value to substitute the qualitative requirement based on the results of the executed simulation. The above-mentioned at least one processor takes the received input data as an input (ingest) of at least one artificial intelligence model and reflects the determined optimal parameter value as at least one control variable to generate a task plan required for building at least one target model; wherein the artificial intelligence model is characterized by including an architecture that interprets the intent of the natural language input based on at least one language model and dynamically calls at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model to orchestrate a workflow;The above at least one processor comprises the step of constructing the above at least one target model based on the generated task plan, wherein the target model is an artificial intelligence-based task execution model optimized to perform a task corresponding to the task objective by being trained through a function tool called according to the task plan; and the above at least one processor comprises the step of applying the constructed target model to at least one target environment, wherein the target environment comprises at least one of at least one infrastructure or edge device to which the constructed target model is deployed to perform at least one target task.

[0020] In another aspect, the step of detecting the qualitative requirements comprises: a step of parsing the natural language input into semantic units by utilizing at least one of a natural language processing (NLP) algorithm or a pre-configured system prompt; a step of identifying the remaining phrases among the text phrases of the parsed natural language input, excluding the phrases corresponding to the hyperparameters of the target model; and a step of extracting and labeling the identified remaining phrases as the qualitative requirements.

[0021] In another aspect, the step of executing the simulation includes executing a parallel simulation based on the plurality of agents, which simultaneously performs at least one of a prediction or classification operation on the same validation dataset based on at least one of a multi-thread or distributed processing environment.

[0022] In another aspect, the step of executing the simulation comprises: each of the plurality of agents interpreting the qualitative requirements based on at least one of the different personas or optimization goals to set at least one of different thresholds or hyperparameters; and each of the plurality of agents applying at least one of the set thresholds or hyperparameters to at least one validation dataset to perform at least one of prediction or classification operations.

[0023] In another aspect, the step of determining the optimal parameter value comprises: a step of aggregating the simulation results to derive a quantitative trade-off analysis result regarding performance evaluation indicators among the plurality of agents; and a step of determining the optimal parameter value to be applied to the target model based on the derived quantitative trade-off analysis result.

[0024] In another aspect, the step of determining the optimal parameter value comprises the step of displaying the trade-off quantitative analysis result on the interface, the step of receiving user selection data for selecting an optimal balance point corresponding to the displayed trade-off quantitative analysis result, and the step of determining the parameter corresponding to the received user selection data as the optimal parameter value.

[0025] In another aspect, the step of determining the optimal parameter value includes a step of performing a comparative operation between the trade-off quantitative analysis result and a pre-set optimal parameter value judgment criterion, and a step of determining the optimal parameter value based on the result of the comparative operation.

[0026] In another aspect, one embodiment of the present disclosure further comprises: a step in which the at least one processor maps the determined optimal parameter value to at least one user identification information or at least one of the qualitative requirements and stores it in at least one memory; and a step in which the at least one processor, when a natural language input identical to the qualitative requirements is received through the interface, applies the optimal parameter value based on the data stored in the memory.

[0027] In another aspect, the step of generating the work plan includes the step of injecting the determined optimal parameter value as a control variable for building the target model, and the step of specifying the workflow of the work plan by updating the setting value for at least one functional tool based on the injected control variable.

[0028] In another aspect, an embodiment of the present disclosure further comprises: a step in which the at least one processor obtains target-specific information from at least one external storage system, the data including at least one of constraints or quality standard data for the target environment; a step in which the at least one processor cross-validates the obtained target-specific information and the natural language input to identify at least one missing data, which is missing information required for building the target model but is absent; and a step in which the at least one processor updates the work plan by supplementing the identified missing data.

[0029] In another aspect, the step of updating the work plan includes generating an interactive clarification question to supplement the identified missing data and outputting it through the interface, and updating the work plan based on a user response received in response to the output interactive clarification question.

[0030] In another aspect, the step of updating the above work plan includes extracting at least one parameter, setting value, or instruction embedded in the user response and mapping it to the area of ​​the identified missing data.

[0031] In another aspect, the step of building the target model comprises the step of designing the architecture of the target model by utilizing at least one of an automated model architecture exploration technique or a pre-trained base model based on control variables reflected in the work plan, and the step of performing learning on the designed architecture using at least one functional tool.

[0032] In another aspect, the step of applying the target model to the target environment includes converting the constructed target model into a format optimized for the target environment based on the target-specific information, and distributing the converted target model to the target environment through at least one communication network.

[0033] In another aspect, the step of applying the target model to the target environment includes collecting real-time log data from the target environment and updating the target model by dynamically updating the work plan based on the collected real-time log data.

[0034] In another aspect, one embodiment of the present disclosure further comprises: the step of the at least one processor calculating at least one performance evaluation metric for the target model using at least one verification dataset; and the step of the at least one processor manifesting the calculated performance evaluation metric through the interface.

[0035] Meanwhile, one embodiment of the present disclosure comprises: at least one processor; and at least one memory storing at least one instruction that performs the following when executed by the at least one processor, wherein the at least one instruction comprises: the step of the at least one processor receiving at least one input data through at least one interface—wherein the input data includes at least one natural language input indicating a task objective; the step of the at least one processor detecting at least one qualitative requirement embedded in the natural language input; the step of the at least one processor generating a plurality of agents having mutually conflicting objective functions based on the detected qualitative requirement—wherein the plurality of agents are assigned at least one of different personas or optimization goals and constitute a Multi-Agent Collaboration Structure that interprets the qualitative requirement in a multidimensional manner; the step of the at least one processor executing a simulation based on the generated plurality of agents; and the step of the at least one processor determining at least one optimal parameter value to substitute the qualitative requirement based on the result of the executed simulation.The above-mentioned at least one processor comprises a step of generating a task plan required for building at least one target model by using the received input data as an input (ingest) of at least one artificial intelligence model and reflecting the determined optimal parameter value as at least one control variable, wherein the artificial intelligence model includes an architecture that orchestrates a workflow by interpreting the intent of the natural language input based on at least one language model and dynamically calling at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model; the above-mentioned at least one processor comprises a step of building at least one target model based on the generated task plan, wherein the target model is an artificial intelligence-based task execution model optimized to perform a task corresponding to the task objective by being trained through the functional tool called according to the task plan; and the above-mentioned at least one processor comprises a step of applying the built target model to at least one target environment, wherein the target environment includes at least one of at least one infrastructure or edge device in which the built target model is deployed to perform at least one target task.

[0036] In another aspect, one embodiment of the present disclosure further comprises a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network comprising: a plurality of neurons arranged in an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.

[0037] In another aspect, one embodiment of the present disclosure further comprises an Application Specific Integrated Circuit (ASIC) for a given artificial neural network, comprising: a plurality of neurons organized into an array including at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synapse circuits.

[0038] In another aspect, one embodiment of the present disclosure comprises: a plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synapse circuits, further comprising a neuromorphic circuit for a predetermined artificial neural network.

[0039] On the other hand, one embodiment of the present disclosure comprises a method executed by a computer, wherein at least one processor of the computer receives at least one input data through at least one interface, wherein the input data comprises at least one natural language input indicating a task objective; wherein the at least one processor generates a task plan that reflects at least one control variable required for building at least one target model, using the received input data as an ingest for at least one artificial intelligence model, wherein the artificial intelligence model comprises an architecture that orchestrates a workflow by interpreting the intent of the natural language input based on at least one language model and dynamically calling at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model; and wherein the at least one processor builds the at least one target model based on the generated task plan, wherein the target model is an artificial intelligence-based task execution model optimized to perform a task corresponding to the task objective by being trained through the functional tool called according to the task plan. and the step of the at least one processor applying the constructed target model to at least one target environment—wherein the target environment comprises at least one of at least one infrastructure or edge device in which the constructed target model is deployed to perform at least one target task;

[0040] In another aspect, the step of generating the work plan comprises: accessing at least one external storage system based on the received input data and loading target-specific information including at least one data among constraints or quality standard data for the target environment to which the target model is to be applied; and generating the work plan by reflecting the loaded target-specific information.

[0041] In another aspect, the step of loading the target-specific information includes the step of parsing and identifying the target environment from the input data, and the step of dynamically accessing the external storage system based on a Machine Control Protocol (MCP)-based communication structure to obtain the target-specific information.

[0042] In another aspect, the step of generating the above-mentioned work plan further includes the step of semantically parsing the loaded target-specific information and the natural language input and cross-validating them to identify at least one missing data point, which is missing information required for building the target model but is absent.

[0043] In another aspect, the step of generating the work plan further includes the step of the artificial intelligence model generating a conversational clarification question to supplement the identified missing data and outputting it through the interface, and the step of updating the work plan based on a user response received in response to the output conversational clarification question.

[0044] In another aspect, the step of updating the above work plan includes extracting at least one parameter, setting value, or instruction embedded in the user response and mapping it to the area of ​​the identified missing data.

[0045] In another aspect, the step of generating the work plan further comprises: detecting at least one qualitative requirement embedded in the natural language input; generating a plurality of agents having mutually conflicting objective functions based on the detected qualitative requirement; executing a simulation based on the generated plurality of agents; determining an optimal parameter value to substitute the qualitative requirement based on the results of the executed simulation; and updating the work plan based on the determined optimal parameter value.

[0046] In another aspect, the step of detecting the qualitative requirements comprises: a step of parsing the natural language input into semantic units by utilizing at least one of a natural language processing (NLP) algorithm or a pre-configured system prompt; a step of identifying the remaining phrases among the text phrases of the parsed natural language input, excluding the phrases corresponding to the hyperparameters of the target model; and a step of extracting and labeling the identified remaining phrases as the qualitative requirements.

[0047] In another aspect, the step of generating the plurality of agents includes the step of dynamically generating the plurality of agents, each assigned at least one of different personas or optimization goals, to form a multi-agent collaboration structure.

[0048] In another aspect, the step of executing the simulation comprises: each of the plurality of agents interpreting the qualitative requirements based on at least one of the different personas or optimization goals to set at least one of different thresholds or hyperparameters; and each of the plurality of agents applying at least one of the set thresholds or hyperparameters to at least one validation dataset to perform at least one of prediction or classification operations.

[0049] In another aspect, the step of determining the optimal parameter value comprises: a step of aggregating the simulation results to derive a quantitative trade-off analysis result regarding performance evaluation indicators among the plurality of agents; and a step of determining the optimal parameter value to be applied to the target model based on the derived quantitative trade-off analysis result.

[0050] In another aspect, the step of building the target model comprises the step of designing the architecture of the target model by utilizing at least one of an automated model architecture exploration technique or a pre-trained base model based on control variables reflected in the work plan, and the step of performing learning on the designed architecture using at least one functional tool.

[0051] In another aspect, the step of applying the target model to the target environment includes converting the constructed target model into a format optimized for the target environment based on the target-specific information, and distributing the converted target model to the target environment through at least one communication network.

[0052] In another aspect, the step of applying the target model to the target environment includes collecting real-time log data from the target environment and updating the target model by dynamically updating the work plan based on the collected real-time log data.

[0053] On the other hand, one embodiment of the present disclosure comprises: at least one processor; and at least one memory storing at least one instruction that performs the following when executed by the at least one processor, wherein the at least one instruction comprises: the step of the at least one processor receiving at least one input data through at least one interface—wherein the input data comprises at least one natural language input indicating a task objective; and the step of the at least one processor generating a task plan that reflects at least one control variable required for building at least one target model by using the received input data as an input (ingest) to at least one artificial intelligence model—wherein the artificial intelligence model comprises an architecture that orchestrates a workflow by interpreting the intent of the natural language input based on at least one language model and dynamically calling at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model. The above-mentioned at least one processor comprises a step of constructing the above-mentioned at least one target model based on the above-mentioned generated task plan—wherein the target model is an artificial intelligence-based task execution model optimized to perform a task corresponding to the task objective by being trained through a function tool called according to the above-mentioned task plan; and the above-mentioned at least one processor comprises a step of applying the constructed target model to at least one target environment—wherein the target environment is characterized by including at least one of at least one infrastructure or edge device to which the constructed target model is deployed to perform at least one target task.

[0054] In another aspect, one embodiment of the present disclosure further comprises a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network comprising: a plurality of neurons arranged in an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.

[0055] In another aspect, one embodiment of the present disclosure further comprises an Application Specific Integrated Circuit (ASIC) for a given artificial neural network, comprising: a plurality of neurons organized into an array including at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synapse circuits.

[0056] In another aspect, one embodiment of the present disclosure comprises: a plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synapse circuits, further comprising a neuromorphic circuit for a predetermined artificial neural network.

[0057] A vision inspection method using a conversational artificial intelligence agent according to one embodiment of the present disclosure supports the automation and integrated management of the entire process from the creation, training, evaluation, and distribution of a vision inspection model through natural language conversation with a user, thereby enabling even field workers lacking coding knowledge or expertise in deep learning modeling to intuitively and easily build and operate an advanced, customized vision inspection pipeline, which has the effect of significantly lowering the barrier to system adoption and drastically reducing the time and cost required for model development.

[0058] In addition, a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure supports controlling a vision inspection model by simulating a user's ambiguous natural language input based on multiple agents and replacing it with specific and quantitative parameters, thereby eliminating unclear inspection criteria that relied on the operator's subjective sense or intuition and implementing parameter optimization based on actual objective data and simulation results, which has the effect of significantly improving the judgment reliability and process stability of the vision inspection model.

[0059] In addition, the vision inspection method using a conversational artificial intelligence agent according to one embodiment of the present disclosure acquires and analyzes various environmental information, including constraints of the inspection target and external integration data, from external resources, and performs active conversational reverse questioning to the user in the event of missing information to support the integrated construction of a vision inspection model optimized for the actual application environment. This allows the system to autonomously fill in information gaps and preemptively reflect the physical characteristics of the actual deployment environment in the model design, even if the user is not fully familiar with all constraints or essential parameters of the complex target infrastructure. Consequently, it has the effect of establishing a customized vision inspection pipeline that fundamentally prevents performance drops and compatibility risks that frequently occur due to the discrepancy between the laboratory environment and the actual mass production site before deployment.

[0060] However, the effects obtainable in this disclosure are not limited to those mentioned above, and other unmentioned effects can be clearly understood from the description below.

[0061] FIG. 1 illustrates an example of a block diagram of a computing system implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0062] FIG. 2 illustrates an example of the structure of a neuromorphic circuit that may be included in a processor according to one embodiment of the present disclosure.

[0063] FIG. 3 illustrates an example of a block diagram of a computing device implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0064] FIG. 4 illustrates an example of a block diagram showing the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.

[0065] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0066] FIG. 6 illustrates an example of a block diagram showing the data flow and system interaction of a vision inspection method service application process using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0067] FIG. 7 illustrates an example of the architecture of a Universal Dynamic Multi-Agent System according to one embodiment of the present disclosure.

[0068] FIG. 8 illustrates an example of a conceptual diagram of an entire workflow that is automated from data preparation to model deployment for building a vision inspection model based on an interactive interface and a vision inspection agent according to one embodiment of the present disclosure.

[0069] FIG. 9 illustrates an example of a block diagram showing the internal components of a vision inspection agent and the interlocking relationships of each functional tool according to one embodiment of the disclosure.

[0070] FIG. 10 illustrates an example of an overall architecture block diagram of a vision inspection agent platform based on a communication interface, comprising a plurality of communication protocol clients (MCP Clients) and servers (MCP Servers) according to one embodiment of the present disclosure.

[0071] FIG. 11 illustrates an example of a flowchart for explaining a method for constructing an integrated vision inspection system using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0072] FIG. 12 illustrates an example of a flowchart for explaining the operation method of a vision inspection control system based on an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0073] As the present disclosure is capable of various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present disclosure, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various forms. In the following embodiments, terms such as "first," "second," etc., are used not in a limiting sense but for the purpose of distinguishing one component from another. Furthermore, singular expressions include plural expressions unless the context clearly indicates otherwise. Additionally, terms such as "include" or "have" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Furthermore, in the drawings, the size of components may be exaggerated or reduced for convenience of explanation. For example, the size and thickness of each component shown in the drawings are arbitrarily depicted for convenience of explanation, so the present disclosure is not necessarily limited to what is depicted.

[0074] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.

[0075]

[0076] [Exemplary system providing a vision inspection method using a conversational AI agent]

[0077] Hereinafter, an exemplary system for implementing a vision inspection method using a conversational artificial intelligence agent, which enables the automation and integrated management of the entire process from the creation, training, evaluation, and deployment of a vision inspection model through natural language conversation with a user, will be described in detail with reference to the attached drawings.

[0078] At this time, the 'vision inspection model' mentioned in this disclosure is merely an example for convenience of explanation, and the scope of the present invention is not limited thereto. That is, the 'vision inspection model' should be interpreted as a broad concept encompassing an 'artificial intelligence-based task execution model' that performs various industrial tasks such as acoustic inspection, anomaly detection, and predictive maintenance.

[0079] FIG. 1 illustrates an example of a block diagram of a computing system implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0080] Referring to FIG. 1, a computing system (1000) implementing a vision inspection method using an interactive artificial intelligence agent of the present disclosure includes a user computing device (110), a server computing system (130), and a training computing system (150), and each device and / or system is connected to communicate via a network (170).

[0081] A vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure may be implemented and provided locally by a user computing device (110), implemented and provided in the form of a web service by a server computing system (130) communicating with the user computing device (110), or implemented and provided by the user computing device (110) and the server computing system (130) in conjunction with each other.

[0082] In this embodiment, the user computing device (110) and / or server computing system (130) can train a machine learning model (120 and / or 140) through interaction with a training computing system (150) that is communicatedly connected via a network (170).

[0083] Additionally, in embodiments of the present disclosure, machine learning models (120, 140) may include various forms of AI agents or agentic architectures beyond simple prediction models.

[0084] In one embodiment, the machine learning model may be an AI agent having a structure that receives system prompts and user prompts, autonomously plans and executes tasks through core components such as planning, memory, reasoning, and tools, and improves itself through feedback.

[0085] In another embodiment, the machine learning model may include a Search Augmented Generative (RAG) architecture that retrieves relevant information from an external database and generates a response based thereon to provide an accurate answer based on the latest information or expertise.

[0086] In another embodiment, the machine learning model may be a multi-agent system in which multiple AI agents cooperate to achieve a specific goal. The multi-agent system may have a supervisory pattern in which a central supervisor agent distributes tasks to subordinate expert agents and aggregates the results. Alternatively, it may have a hierarchical pattern in which a meta-agent acts as an intermediary manager to control and coordinate subordinate agents. Furthermore, it is possible to include a multi-agent debate pattern in which multiple agents present different opinions and derive an optimal conclusion through discussion and evaluation.

[0087] The server computing system (130) can host AI agents such as those mentioned above, particularly multi-agent systems requiring complex computations or large-scale long-term memory. Additionally, the server computing system (130) includes a Multi-Channel Processing (MCP) server to manage integration with various external tools, and can relay communication with cloud APIs, payment services, search engines, etc.

[0088] The training computing system (150) can generate a ToolFormer model that learns how to use a specific tool, or perform iterative learning that gradually improves the performance of the agent through a self-reflection mechanism in which another LLM evaluates and modifies the results generated by the agent.

[0089] At this time, the training computing system (150) may be separate from the server computing system (130) or part of the server computing system (130). Additionally, in some embodiments, the training computing system (150) may be separate from the user computing device (110) or part of the user computing device (110).

[0090] And at this time, the artificial intelligence model can be 1) trained directly locally by a user computing device (110), 2) trained by the server computing system (130) and the user computing device (110) interacting with each other through a network (170), and 3) trained by a separate training computing system (150) using various training and learning techniques. It may also be implemented by transmitting the artificial intelligence model trained by the training computing system (150) to the user computing device (110) and / or the server computing system (130) through the network (170) to provide / update it.

[0091] - User Computing Device (110: User Computing Device)

[0092] The user computing device (110) may include all other types of computing devices, such as a smartphone, a mobile phone, a digital broadcasting device, a PDA (personal digital assistants), a PMP (portable multimedia player), a desktop, a wearable device, an embedded computing device and / or a tablet PC.

[0093] Additionally, in the embodiment, the user computing device (110) may further include a predetermined server computing device that provides a vision inspection method environment using a conversational artificial intelligence agent.

[0094] This user computing device (110) includes at least one processor (111) and memory (112).

[0095] Here, the processor (111) of the user computing device (110) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.

[0096] In particular, according to the embodiment, this processor (111) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which is a hardware technology for implementing a certain digital circuit.

[0097] Here, a Field Programmable Gate Array (FPGA) implementation can refer to a flexible digital circuit that is programmable according to user needs.

[0098] As an example, a field programmable gate array implementation may include a register that temporarily stores data and controls the flow and timing of signals to maintain intermediate results or state information of operations to support synchronized operation of the FPGA, programmable logic that programs operations within the FPGA to perform specific functions or operations as logic circuits configurable according to user needs, and an input interface that receives signals from external devices or sensors and transmits them to internal circuits as a channel for receiving data from outside the FPGA.

[0099] Through the combination of the above components, a field-programmable gate array implementation can provide flexible and various types of digital circuits.

[0100] Meanwhile, an Application-Specific Integrated Circuit (ASIC) can refer to a custom integrated circuit that is fixedly designed to perform a specific use or function.

[0101] As an example, the application-dedicated integrated circuit may include a register, which is a small memory device for temporarily storing and managing data and supports the rapid processing of ASIC operations by storing intermediate calculation results or state information; a microprocessor, which is a central processing unit that performs control and operations within the ASIC and coordinates the operation of the entire system by performing various operations or generating control signals when necessary; and an input block, which is an interface for receiving data from the outside, which receives data to be processed by the ASIC and transmits it internally, and receives various input data through connections with sensors or external devices.

[0102] Through the combination of the components mentioned above, an application-specific integrated circuit can perform specific purpose tasks in an optimized manner.

[0103] Additionally, according to an embodiment, the processor (111) may include a structure of a neuromorphic circuit in the form of an array including a plurality of neuron circuits.

[0104] FIG. 2 briefly illustrates the structure of a neuromorphic circuit (300) that may be included in a processor (111, 131, 151) according to one embodiment.

[0105] Referring to FIG. 2, for example, a neuromorphic circuit (300) may include a plurality of presynaptic neuron circuits (310), a plurality of presynaptic lines (311) extending laterally from the plurality of presynaptic neuron circuits (310), a plurality of postsynaptic neuron circuits (320), a plurality of postsynaptic lines (321) extending longitudinally from the plurality of postsynaptic neuron circuits (320), and a plurality of synaptic circuits (330) provided at the intersection of the plurality of presynaptic lines (311) and the plurality of postsynaptic lines (321).

[0106] A plurality of free synaptic neuron circuits (310) can transmit signals input from the outside in the form of electrical signals to a plurality of synaptic circuits (330) through a plurality of free synaptic lines (311).

[0107] Additionally, a plurality of post-synaptic neuron circuits (320) can receive electrical signals from a plurality of synaptic circuits (330) through a plurality of post-synaptic lines (321).

[0108] Furthermore, multiple post-synaptic neuron circuits (320) may transmit electrical signals to multiple synaptic circuits (330) through multiple post-synaptic lines (321).

[0109] A plurality of synapse circuits (330) can store weights included in layers constituting a neural network system implemented by a neuromorphic circuit (300) and perform a predetermined operation based on the weights and input data.

[0110] For example, each of the plurality of synaptic circuits (330) may include a resistive memory cell having a variable resistance. In this case, the resistance value of the plurality of synaptic circuits (330) changes by a voltage applied through the plurality of presynaptic neuron circuits (310) or the plurality of postsynaptic neuron circuits (320), and can store weight data according to this resistance change.

[0111] The neuromorphic circuit (300) is formed by mimicking the structure of neurons and synapses, which are essential elements of the human brain. When a deep neural network (DNN) is realized using the neuromorphic circuit (300), the data processing speed can be improved and power consumption can be reduced compared to when the existing von Neumann structure is utilized.

[0112] Returning to the point, the memory (112) of the user computing device (110) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs memory storage functions on the internet. This memory (112) may store data (113) and instructions (114) necessary for the at least one processor (111) to perform functional operations, such as training an artificial intelligence model or executing a vision inspection method using a conversational artificial intelligence agent through the artificial intelligence model.

[0113] In one embodiment, the user computing device (110) can store at least one machine learning model (120).

[0114] For example, the user computing device (110) may be composed of various machine learning models, such as multiple neural networks (e.g., deep neural networks) that perform a vision inspection method using a conversational artificial intelligence agent based on structured / quantitative data, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.

[0115] For example, the machine learning model (120) may store a model for a vision inspection method using linear regression, decision tree, random forest, gradient boosting or / and deep learning-based conversational AI agent. And the neural network may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, Transformer or / and other forms of neural networks.

[0116] At this time, according to the embodiment, the machine learning model (120) may be directly installed in the server computing system (130) or may operate as a separate device from the server computing system (130) to perform deep learning, etc. for a vision inspection method using the conversational artificial intelligence agent.

[0117] Additionally, according to an embodiment, the user computing device (110) may store a model to be used in each process and a prompt template that serves as the basis for input to the model in order to perform at least part of the process for a vision inspection method using a conversational artificial intelligence agent through a large language model (LLM).

[0118] In one embodiment, a user computing device (110) receives at least one machine learning model (120) from a server computing system (130) through a network (170), stores it in memory (112), and then executes the stored machine learning model (120) by a processor (111) to perform a vision inspection method using a conversational artificial intelligence agent, etc.

[0119] In another embodiment, the user computing device (110) may provide the user with a vision inspection method using an interactive artificial intelligence agent by performing an operation through a machine learning model (140) including at least one machine learning model (140) in conjunction with a server computing system (130) and communicating related data to the outside.

[0120] For example, a user computing device (110) can perform a vision inspection method using an interactive artificial intelligence agent in such a way that a server computing system (130) provides an output for the user's input using a machine learning model (140) via the web.

[0121] Additionally, the artificial intelligence model can be implemented in such a way that at least some of the machine learning models (120 and / or 140) are executed on a user computing device (110) and the rest are executed on a server computing system (130).

[0122] Additionally, the user computing device (110) may include at least one input component (121) that detects user input.

[0123] For example, the user input component (121) may include a touch sensor (e.g., a touch screen and / or a touch pad, etc.) that detects a touch of a user's input medium (e.g., a finger or a stylus), a vision data sensor (e.g., an image sensor) that detects a user's motion input, a microphone that detects user voice input, a button, a mouse and / or a keyboard, etc.

[0124] Here, the vision data sensor may include an image processing module. Specifically, the vision data sensor may process still images or video obtained by a vision data sensor device (e.g., CMOS or CCD).

[0125] In addition, the vision data sensor can process still images or videos acquired through the vision data sensor device using a vision data recognition process (e.g., OCR, etc.) and / or an image processing module to extract necessary information and transmit the extracted information to a processor.

[0126] Additionally, the input component (121) can receive input from an external controller (e.g., mouse, keyboard, etc.) based on an interface module, and in this case, may include an external output device (e.g., speaker).

[0127] At this time, the interface module may be configured to include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, an earphone port, a power amplifier, an RF circuit, a transceiver, and other communication circuits.

[0128] In addition, the external output device may include a display system that outputs various information related to a vision inspection method using an interactive artificial intelligence agent as graphic vision data (e.g., images).

[0129] Such a display system may be implemented by including at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, a 3D display, and an e-ink display.

[0130] Meanwhile, the user computing device (110) including the above-described components may further perform at least some of the functional operations performed by the server computing system (130) described later.

[0131] -Server Computing System (130: Server Computing System)

[0132] The server computing system (130) can perform a series of processes to provide a vision inspection method using a conversational artificial intelligence agent.

[0133] In detail, in an embodiment, the server computing system (130) can provide a vision inspection method using a conversational artificial intelligence agent by exchanging data necessary to enable the vision inspection method process using a conversational artificial intelligence agent to be executed on an external device such as a user computing device (110).

[0134] More specifically, in an embodiment, the server computing system (130) can provide an environment in which an application can run on a user computing device (110).

[0135] To this end, the server computing system (130) may include an application program, data and / or instructions, etc. for the application to operate, and may transmit and receive various data based thereon with the external device.

[0136] Additionally, the server computing system (130) includes at least one processor (131) and memory (132).

[0137] Here, the processor (131) of the server computing system (130) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.

[0138] In particular, depending on the embodiment, such a processor (131) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which are hardware technologies for implementing a specific digital circuit. A detailed description thereof is omitted by applying the description of the FPGA and ASIC mentioned above.

[0139] Additionally, according to an embodiment, the processor (131) may have a structure of a neuromorphic circuit in the form of an array including a plurality of neuron circuits (see FIG. 2).

[0140] And the memory (132) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (132) may store data (133) and instructions (134) necessary for the processor (131) to perform functional operations, such as training an artificial intelligence model or executing a vision inspection method using a conversational artificial intelligence agent through the artificial intelligence model.

[0141] In one embodiment, the server computing system (130) may be implemented to include at least one computing device. For example, the server computing system (130) may be implemented to operate a plurality of computing devices according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Additionally, the server computing system (130) may include a plurality of computing devices connected to a network (170).

[0142] Additionally, the server computing system (130) may store at least one machine learning model (140). For example, the server computing system (130) may include a neural network and / or other multi-layer non-linear model as the machine learning model (140). Exemplary neural networks may include a feed-forward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.

[0143] In an embodiment, the server computing system (130) may further include a data store computing system (hereinafter, data store) which is a storage for continuously storing and managing raw data that forms the basis of a vision inspection method using an interactive artificial intelligence agent.

[0144] Such data stores may include various forms of data storage, ranging from file systems to cloud storage. For example, a data store may include at least one database among a relational database that uses a structured query language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse optimized for querying and analysis by centralizing large volumes of data from multiple sources as a system used for reporting and data analysis, a data warehouse that stores large volumes of raw data in basic formats such as structured data, semi-structured data, and unstructured data, and a local storage device or Network Attached Storage (NAS) that stores data in files in a format generally accessible by a computer operating system.

[0145] Additionally, in the embodiment, the server computing system (130) may further include a plurality of specialized engines and repositories that are logically and physically separated to perform a vision inspection method using a conversational artificial intelligence agent.

[0146] In one embodiment, the server computing system (130) may include at least one engine among a reasoning engine that processes a user's natural language input or system event to establish an action plan, a membership management engine, and a supervision signal generation engine.

[0147] Here, the term engine may include not only a set of instructions executed by a processor (131) to perform specific logic, but also dedicated hardware circuits to accelerate said logic.

[0148] Such engines may run on hardware accelerators optimized to handle the computational load of large-scale generative artificial intelligence models. The hardware accelerators are processors specialized for matrix operations and vector processing and may include at least one of the aforementioned Tensor Processing Unit (TPU), Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), or Application-Specific Integrated Circuit (ASIC). These hardware accelerators can provide technical improvements that distribute the computational load of large-scale language models (LLM) and / or diffusion models and enable real-time inference.

[0149] Additionally, the data (133) may not be a simple set of data, but may include a structured embedding repository to support search augmentation generation (RAG) of the AI ​​model. The embedding repository stores high-dimensional vector representations of text, code, or audio data, thereby enabling the engine performing a vision inspection method using a conversational AI agent to perform high-speed search based on semantic similarity. This can serve as a technical means to suppress hallucinations in the model and increase the accuracy of the output.

[0150] Specifically, the memory (132) of the server computing system (130) may include a structured Knowledge Base Layer to physically support Search Augmentation Generation (RAG). The Knowledge Base Layer may include an Embedding Repository that stores high-dimensional vector representations of unstructured text, code, or multimodal data, and a Policy Document Repository that stores behavioral constraints and business rules of agents.

[0151] At this time, the vision inspection method engine using a conversational artificial intelligence agent can be configured to vectorize a query received from a user computing device (110) and query the embedding repository to retrieve context information with high semantic similarity in real time. This structure can technically suppress the hallucination phenomenon of the artificial intelligence model by allowing the agent to refer to external verified knowledge rather than relying only on intrinsic parameters.

[0152] Additionally, the user computing device (110) may include a trigger event detector that provides an interface for interaction with an agent and initiates the operation of the agent. The trigger event detector can detect not only user input but also the arrival of a specific time, a change in the state of an external system, etc., and transmit a processing request to a server computing system (130).

[0153] - Training Computing System (150: Training Computing System)

[0154] The training computing system (150) includes at least one processor (151) and memory (152).

[0155] Here, the processor (151) of the training computing system (150) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.

[0156] In particular, depending on the embodiment, this processor (151) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which are hardware technologies for implementing a specific digital circuit. A detailed description thereof is omitted by applying the description of the FPGA and ASIC mentioned above.

[0157] Additionally, according to an embodiment, the processor (131) may have a structure of a neuromorphic circuit in the form of an array including a plurality of neuron circuits (see FIG. 2).

[0158] And the memory (152) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (152) may store data (153) and instructions (154) necessary for the processor (151) to perform learning of an artificial intelligence model, etc.

[0159] For example, the training computing system (150) may include a model trainer (160) that trains a machine learning model (120 and / or 140) stored in a user computing device (110) and / or a server computing system (130) using various training or learning techniques, such as back propagation of error (according to the framework illustrated in FIG. 5).

[0160] For example, such a model trainer (160) can perform updates to one or more parameters of a machine learning model (120 and / or 140) for a vision inspection method using an interactive artificial intelligence agent based on a defined loss function in a backpropagation manner.

[0161] In some embodiments, performing backpropagation of the error may include performing truncated backpropagation through time. The model trainer (160) may perform a number of generalization techniques (e.g., weight devaluation, dropout and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning model (120 and / or 140) being trained.

[0162] Additionally, the model trainer (160) can train a machine learning model (120 and / or 140) based on a series of training data (161). Here, the training data (161) may include data of different forms, such as, for example, vision data, audio samples and / or text. Examples of types of vision data that may be used may include video frames, LiDAR point clouds, X-ray vision data, computed tomography scans, hyperspectral vision data and / or various other forms of vision data.

[0163] These training data (161) may be provided by a user computing device (110) and / or a server computing system (130). When the training computing device trains a machine learning model (120 and / or 140) on specific data of the user computing device (110), the machine learning model (120 and / or 140) may be characterized as a personalized model.

[0164] And the model trainer (160) includes computer logic that is utilized to provide the desired function.

[0165] Additionally, the model trainer (160) may be implemented as hardware, firmware, and / or software that controls a general-purpose processor. In one embodiment, the model trainer (160) may include a program file stored in a storage device, be loaded into memory (152), and be executed by one or more processors (151). In another embodiment, the model trainer (160) includes one or more sets of computer-executable data (153) and instructions (154) stored in a tangible computer-readable storage medium, such as a RAM hard disk or an optical or magnetic medium.

[0166] Additionally, according to an embodiment, the model trainer (160) may include a Supervision Signal Engine. The Supervision Signal Engine may compare a result generated by an agent (e.g., Raw Output) with a verified result obtained through an external tool (e.g., search engine, API) (e.g., Grounded Output) to calculate a difference value, and execute a reinforcement learning process to update a reward model or fine-tune an agent model based on this.

[0167] Meanwhile, a computing system (1000) according to one embodiment of the present disclosure may be connected in a predetermined manner through a wired / wireless network (170) for communication between a user computing device (110) and a server computing system (130).

[0168] Network (170) includes, but is not limited to, 3GPP (3rd Generation Partnership Project) network, LTE (Long Term Evolution) network, WIMAX (World Interoperability for Microwave Access) network, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth network, satellite broadcasting network, analog broadcasting network and / or DMB (Digital Multimedia Broadcasting) network.

[0169] Generally, communication through the network (170) can be performed using any type of wired and / or wireless connection through various communication protocols (e.g., TCP / IP, HTTP, SMTP and / or FTP, etc.), encodings or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, Secure HTTP and / or SSL, etc.).

[0170] As such, in one embodiment, the system (1000) of the present disclosure may be implemented as a technical system in which specialized hardware accelerators, vectorized data storage, and physical engines controlling the same are organically combined, rather than as a simple set of software algorithms.

[0171] FIG. 3 illustrates an example of a block diagram of a computing device implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0172] Including FIG. 3, the computing device (100) included in the user computing device (110), server computing system (130), and training computing system (150) includes a plurality of applications (e.g., Application 1 to Application N). Each application may include a machine learning library and one or more machine learning models. For example, the applications may include a vision data processing application (e.g., Detection, Classification, and / or Segmentation, etc.), a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application and / or a chat-bot application, etc.

[0173] Additionally, for example, the computing device (100) may include a specialized application for [core technology name of the invention], which provides related services to a user, including a model that performs a vision inspection method using a conversational artificial intelligence agent.

[0174] In an embodiment, the computing device (100) may include a model trainer (160) for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, it may provide output data according to a predetermined input data.

[0175] Each application of the computing device (100) can communicate with a number of other components of the computing device (100), such as, for example, at least one sensor, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (e.g., a public Application Programming Interface). In one embodiment, the API used by each application may be specific to that application.

[0176] FIG. 4 illustrates an example of a block diagram showing the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.

[0177] Referring to FIG. 4, a computing device (100) according to one embodiment of the present disclosure may have a pipeline structure that receives input data (402) and generates output data (412) through a series of transformation processes to perform a vision inspection method using an interactive artificial intelligence agent. This process may be performed through a preprocessing module (404), an encoder / embedding model (406), a neural network layer (408), and a decoder / generation head (410).

[0178] First, the preprocessing module (404) receives input data (402) (e.g., a text prompt, an image, or a multimodal signal) from a user or system. The preprocessing module (404) can perform tokenization and normalization on the input data to generate a sequence of tokens, which are the smallest units that the model can process.

[0179] Next, the encoder / embedding model (406) receives the generated token as input and converts it into a vector / embedding mapped to a number in a high-dimensional vector space. At this stage, the discrete information of the input data is converted into a continuous numeric matrix so that it can become a form that the machine learning model can compute.

[0180] Next, the neural network layer (408) receives the vector / embedding and performs deep computation. The neural network layer (408) may have a structure in which a plurality of sub-layers (e.g., layer 1 to layer N) are stacked. Each layer may abstract and refine input features through an attention mechanism or convolution operation.

[0181] In particular, the final output of the neural network layer (408) is defined as a latent representation. This latent representation has a structure different from the original input data (402) and may correspond to an intermediate representation in which the semantic features of the data are highly compressed and abstracted. This may mean that it is not a simple data transmission, but a technical data structure that is valid only within the system.

[0182] Finally, the decoder / generation head (410) receives the potential representation as a conditioning input. The decoder / generation head (410) can interpret the compressed potential representation and reconstruct or generate output data (412) in a form recognizable by the user (e.g., natural language text, image pixels, control codes, etc.).

[0183] This stepwise data transformation structure (token -> vector -> latent representation -> output) can clearly demonstrate that it functions not as a simple sequence of operations, but as a concrete device that technically processes input data to generate useful information.

[0184] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device implementing a vision inspection method using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0185] Referring to FIG. 5, the computing device (200) includes a plurality of applications (e.g., Application 1 to Application N). Each application can communicate with a central intelligence layer. For example, the applications may include a vision data processing application, a text messaging application, an email application, a dictation application, a virtual keyboard application and / or a browser application, etc. In one embodiment, each application can communicate with the central intelligence layer (and a model stored therein) using an API (e.g., a common API across all applications).

[0186] In addition, in one embodiment of the present disclosure, the application may include a vision inspection method application using a conversational artificial intelligence agent, an energy management application, a logging and analysis application, etc.

[0187] And each application can interface with a shared model within the central intelligence layer through a designated API (e.g., a common API).

[0188] Here, the central intelligence layer may include a plurality of machine learning models. For example, as illustrated in FIG. 5, at least some of each machine learning model may be provided for each application and managed by the central intelligence layer. In another embodiment, two or more applications may share a single machine learning model. For example, in some embodiment, the central intelligence layer may provide a single model for all applications. In some embodiment, the central intelligence layer may be included within the operating system of the computing device (200) or otherwise implemented.

[0189] In one embodiment, the central intelligence layer may be integrated as part of the operating system or implemented as a separate logical layer, and may perform the role of transmitting input time series data to the corresponding model to return a prediction result.

[0190] In addition, the central intelligence layer can communicate with the central device data layer.

[0191] Here, the central device data layer may be a centralized data store for the computing device (200).

[0192] For example, the central device data layer can integrate and store sensor data, device state information, external environment information, etc., stored within the computing device (200), and provide this as input data required for a vision inspection method using a conversational artificial intelligence agent. Each device component (e.g., sensor, context, state manager, etc.) can communicate with the corresponding data layer through a private API, etc.

[0193] As illustrated in FIG. 5, the central device data layer can communicate with a number of other components of the computing device (200), such as, for example, one or more sensors, a context manager, a device state component and / or additional components, and in some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0194] The technology described herein may refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from said systems. It will be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, division of tasks, and functionality between and from components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components may operate sequentially or in parallel.

[0195] In addition, the data storage, prediction model, and / or application, etc. of this specification may be configured and operated in a distributed manner locally or over a network, and these configurations may be flexibly applicable to various system architectures.

[0196] FIG. 6 illustrates an example of a block diagram showing the data flow and system interaction of a vision inspection method service application process using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0197] Referring to FIG. 6, the computing system (1000) may be composed of an organic data pipeline between a trigger event detector (510), an inference engine (530), and a mobile device screen (550).

[0198] First, the trigger event detector (510) may be configured to monitor and detect a signal initiating the operation of the system. The trigger event detector (510) may receive at least one of i) a time event indicating the arrival of a specific point in time, ii) a user action resulting from physical input by a user, or iii) a system state indicating a change in internal system data, and transmit an activation signal to the inference engine (530). This may mean that the service can be actively initiated depending on the situation without an explicit request from the user.

[0199] The inference engine (530) can perform a multi-stage operation process to convert raw data into a final result in response to the activation signal. This process can be implemented as a series of logically connected prompt chains.

[0200] For example, the first prompt, contextualization (532), allows the inference engine (530) to receive unstructured raw data (e.g., user logs, channel metadata), analyze and refine it, and generate compressed summary information.

[0201] And the second prompt, Content Generation (534), can generate multiple candidate results that match the user's intent or situation using a generative AI model based on the summary information generated above.

[0202] In addition, the third prompt, Verification (536), can produce a reliable final result by filtering and verifying the generated candidate results against a predefined policy (e.g., safety guidelines, format rules).

[0203] Finally, the final result generated by the inference engine (530) can be transmitted to and implemented on a mobile device screen (550) through an auto pre-filling (554) operation. Specifically, the system changes the state of the interface by directly writing the final result to a memory address of a specific target field (552) (e.g., text input field, setting value) within the mobile device screen (550). Here, the mobile device screen (550) may be an example of a user computing device (110).

[0204] Such a configuration can go beyond the simple display of information and provide specific technical means for data generated by an external trigger to physically control and complete the input interface of the user terminal.

[0205] FIG. 7 illustrates an example of the architecture of a multi-agent system (Universal Dynamic Multi-Agent System, 600) according to one embodiment of the present disclosure.

[0206] Referring to FIG. 7, the multi-agent system (600) can be configured around an orchestration engine (610) that determines and executes an optimal agent collaboration structure in real time according to the nature of the user's request or task.

[0207] Additionally, this multi-agent system (600) can operate in an organically combined manner, including a complexity analyzer (605), an agent pool (620), a shared memory fabric (630), and a tool execution interface (640).

[0208] 1. Input Analysis and Dynamic Topology

[0209] The complexity analyzer (605) can evaluate the complexity of the user query and the required domain expertise. Based on this evaluation result, the orchestration engine (610) can dynamically configure an optimal agent collaboration topology for task resolution.

[0210] For example, in the case of a simple question, the orchestration engine (610) can enable a single agent mode.

[0211] If complex planning is required, the orchestration engine (610) can configure a Supervisor pattern to instantiate one supervisor agent to control subordinate agents.

[0212] When high accuracy is required, the orchestration engine (610) can configure a debate pattern to set up a control path for multiple agents to perform mutual criticism.

[0213] 2. Agent Pool and Instantiation (Agent Pool)

[0214] The agent pool (620) may be a repository of template agents that have prompts and tool sets specialized for specific functions (e.g., web search, code generation, data analysis). The orchestration engine (610) can select and activate the necessary agents at runtime according to the determined topology.

[0215] 3. Shared Memory Fabric

[0216] The shared memory fabric (630) may be a data pipeline that synchronizes the state and context between multiple collaborating agents. This mediates the short-term memory of individual agents and the knowledge base (RAG) of the entire system, ensuring that the output of agent A is passed to agent B as input without loss.

[0217] 4. Tool Execution & Feedback Loop

[0218] Each agent can communicate with an external API (search engine, calculator, AWS, etc.) through a tool execution interface (640). The execution result generated at this time is fed back to the orchestration engine (610), and the engine can determine the consistency of the result and perform self-correction logic to instruct the agent to rework or proceed to the next step.

[0219] This structure enables the implementation of an adaptive artificial intelligence system that flexibly changes the system's processing structure according to the nature of the input problem, rather than a fixed (static) algorithm.

[0220] Meanwhile, the orchestration engine (610) can analyze the characteristics of the task received from the complexity analyzer (605) (e.g., creativity, logic, whether coding is required) and select specialized agents waiting in the agent pool (620) to configure a dynamic collaboration topology. The topology can be reconfigured into various operation modes as follows, depending on the type of task.

[0221] First Operation Mode: Sequential Pattern

[0222] For tasks such as a single-flow creation or report writing, the orchestration engine (610) connects multiple agents in series. For example, a pipeline can be formed in the order of user agent -> writer agent -> style agent so that the output of the previous stage is passed as the input to the next stage.

[0223] Second Operation Mode: Supervisor Pattern

[0224] In cases where complex sub-tasks are mixed, the orchestration engine (610) can form a centralized star topology in which one agent is designated as a supervisor and the remaining agents (e.g., research agent, math agent) are assigned as workers. The supervisor agent can distribute tasks to subordinate agents and aggregate results.

[0225] Third Operation Mode: Hierarchical Pattern

[0226] In cases of high complexity, such as large-scale project management, the orchestration engine (610) can establish a command system by forming a tree structure topology in which a meta-agent is placed as the upper manager and multiple specialized agent groups are placed below it.

[0227] 4th Operation Mode: Debate Pattern

[0228] In cases where the correct answer is unclear or high reliability is required (e.g., social issue analysis), the orchestration engine (610) can derive an optimal conclusion by configuring a competitive topology in which multiple agents present different perspectives on the same topic and perform mutual criticism and voting.

[0229] Fifth Operation Mode: Mixture-of-Agents Pattern

[0230] When parallel processing is required, the orchestration engine (610) can form a parallel processing structure by placing multiple agents in multiple layers to perform tasks simultaneously and finally integrating the results through an aggregator agent.

[0231] Additionally, the agent pool (620) may include various template agents that can be deployed into the topology. Each agent may be instantiated to perform the following core methodologies under the control of the orchestration engine (610).

[0232] For example, at least one agent can execute a loop that repeats reasoning and action to perform reasoning and action (ReAct Paradigm) that refines answers based on results from an external tool (e.g., Google Search).

[0233] In addition, at least one agent can perform code-based behavior (CodeAct Paradigm) that involves complex calculations or data analysis by generating and executing executable code (e.g., Python) instead of natural language.

[0234] In addition, at least one agent can perform self-reflection to improve the quality of the output by carrying out a metacognitive process of self-evaluating (Critique) and modifying the generated output.

[0235] In addition, the system can provide a shared memory fabric (630) to enable individual agents to cooperate organically without disconnection. This can prevent context loss by managing the conversation history (Short-Term Memory) between agents and an external knowledge base in an integrated manner. Furthermore, the tool execution interface (640) can securely connect agents with external APIs, cloud services, and databases through protocols such as Multi-Channel Processing (MCP).

[0236] In this way, the multi-agent system (600) of the present disclosure can implement an adaptive artificial intelligence platform that is not a fixed single model, but rather dynamically changes the system's processing structure and behavior according to the nature of the input problem.

[0237] Specifically, one embodiment of the present disclosure may be implemented as an Agentic RAG (Retrieval-Augmented Generation) system that supports advanced question-answering by including agent functions. This system may be a search-based generation system that retrieves information from external data sources and generates answers based thereon.

[0238] The system's data processing pipeline can consist of a data extraction phase and an Agentic RAG pipeline phase. In the data extraction phase, content in various formats, such as text and images, can be collected from designated websites. Subsequently, text and metadata are extracted from the collected content, the text is chunked into small units, and each chunk is vectorized using an embedding model and stored in a vector database.

[0239] The Agentic RAG pipeline can be divided into search and generation phases. In the search phase, when a user query is input, query rewriting and embedding are performed, and highly relevant documents can be found through similarity searches in a vector database. Subsequently, the search results can be ranked based on relevance to construct context. In the generation phase, the user query and the searched context are combined to construct a prompt, which is then input into a large language model to generate a final response. During this process, the agent can respond to complex queries by utilizing agentic elements such as memory storage, calling external tools, and planning.

[0240] An AI agent according to one embodiment of the present disclosure can support complex decision-making and task execution through a hierarchical memory structure similar to human memory. The agent's memory can be broadly divided into short-term memory and long-term memory.

[0241] Short-term memory is a temporary memory space focused on the currently ongoing workflow and may include working memory, which manages workflow-specific reasoning and task flows, and cache memory, which provides quick access to frequently used data or results.

[0242] Long-term memory is a memory space based on knowledge and experience that is continuously preserved, and may include episodic memory, which records events or incidents manually saved in a specific workflow; semantic memory, which stores conceptual or factual knowledge; and procedural memory, which stores knowledge of how to perform specific tasks or procedural knowledge.

[0243] This memory structure can be integrated with the language model framework through a central memory controller. Additionally, it can leverage external knowledge or support real-time integration by connecting with external vector databases, semantic databases, or third-party APIs via the MCP server. Through this, the agent can generate user-customized responses that comprehensively reflect past experiences and current context.

[0244] One embodiment of the present disclosure may include various agentic workflows to solve complex problems. These workflows may be designed to suit specific business purposes.

[0245] For example, the system may include a Plan and Execute workflow. In this workflow, a Planner can break down a single top-level task into multiple sub-tasks, and after each sub-task is processed by a specialized agent, the results are integrated. This can be utilized for business process automation or data pipeline orchestration.

[0246] As another example, the system may include an Orchestrator-Worker workflow. In this structure, a central orchestrator language model can break down tasks, distribute them to multiple worker language models for processing, and then integrate the results. This can be used in the implementation of Agentic RAGs or coding agents.

[0247] As another example, the system may include a routing workflow. This workflow may be structured to analyze input tasks, classify them into the most suitable ones among several predefined subtasks, and forward them to a specialized language model or path capable of processing the task. This can be applied to customer support agents or multi-agent discussion systems.

[0248] One embodiment of the present disclosure may include a protocol for efficient and secure communication between a plurality of agents or between an agent and an external tool.

[0249] In one embodiment, an Agent2Agent (A2A) protocol may be used for communication between agents. The A2A protocol can enhance security by enabling each agent to communicate without directly sharing their internal data. Through this protocol, multiple agents can share tasks and negotiate, and each agent can operate independently using its own language model, framework, and database.

[0250] In another embodiment, a Multi-Channel Processing (MCP) protocol may be used for communication between an agent and an external function server. MCP may have a structure in which each external function, such as file access, search, and cloud API calls, is separated into a separate server for communication. For example, one agent may communicate with a local file system or a search engine via the MCP protocol, while another agent may communicate with a cloud provider such as AWS or a communication tool such as Slack via the same protocol.

[0251]

[0252] [Vision Inspection Method Using Conversational AI Agents]

[0253] Hereinafter, a vision inspection method using a conversational artificial intelligence agent, in which at least one processor of a computing system (1000) according to one embodiment of the present disclosure can automate and integrally manage the entire process from the creation, training, evaluation, and distribution of a vision inspection model through natural language conversation with a user, will be described in detail with reference to the attached drawings.

[0254] At this time, a computing system (1000) according to one embodiment of the present disclosure can implement a method for constructing an integrated vision inspection system using a conversational artificial intelligence agent, which acquires and analyzes various environmental information including constraints of an inspection target and external linkage data from external resources, and performs active conversational reverse questioning to the user when information is missing, thereby enabling the integrated construction of a vision inspection model optimized for the actual application environment.

[0255] In addition, a computing system (1000) according to one embodiment of the present disclosure can implement a method of operation for a vision inspection control system based on a conversational artificial intelligence agent, which enables controlling a vision inspection model by simulating ambiguous natural language input from a user based on multiple agents and replacing it with specific and quantitative parameters.

[0256] In the following description, for effective explanation, at least one processor of the computing system (1000) performing each step is collectively referred to as the computing system (1000).

[0257] In addition, among the terms used in the embodiments and / or claims of the present disclosure, the process of 'ingesting' data to an artificial intelligence model, etc., may be used to encompass a series of input operations that go beyond the simple one-dimensional transmission of data, in which a specific data set is received by the neural network computation layer and / or normalization pipeline of the model and is mechanically loaded and processed as basic data to perform actual inference, training, and / or probabilistic computation.

[0258] Additionally, the process of ‘manifesting’ the final result, etc. through an interface can be used as a broad concept encompassing all concretization and output operations, such as displaying abstract data (e.g., a sample of probability distribution data, tensor values, weight matrices, etc.) computed within the computing system (1000) on a display screen in a physical form that the user can visually and / or intuitively perceive (e.g., a rendered graph, derivation of quantitative risk indicators in text form, etc.), or constructing and storing it as explicit structured data in a physical / logical database for subsequent simulation and / or data utilization.

[0259] FIG. 8 illustrates an example of a conceptual diagram of an entire workflow that is automated from data preparation to model deployment for building a vision inspection model based on an interactive interface and a vision inspection agent according to one embodiment of the present disclosure, FIG. 9 illustrates an example of a block diagram showing the interoperability of internal components of a vision inspection agent and each functional tool according to one embodiment of the disclosure, and FIG. 10 illustrates an example of an overall architecture block diagram of a vision inspection agent platform based on a communication interface including a plurality of communication protocol clients (MCP Clients) and servers (MCP Servers) according to one embodiment of the present disclosure.

[0260] Before describing the specific methodology of the present disclosure, the configuration and operation structure of a vision inspection agent platform implemented by a computing system (1000) according to one embodiment of the present disclosure will be briefly described with reference to FIGS. 8 to 10.

[0261] Referring to FIG. 8, a computing system (1000) according to one embodiment of the present disclosure can receive vision data, related documents and / or natural language input, etc. from a user through a user interface (hereinafter, conversational interface) including at least one conversational window and / or tool, and can automatically execute a workflow of data preparation, model design and training, evaluation and analysis, and model deployment through at least one artificial intelligence agent.

[0262] In detail, as an example, the computing system (1000) can provide an intuitive interactive window as a user interface that allows the user to input everyday language (e.g., "I have pilot process vision data and I want to make an inspection model") without complex coding or parameter manipulation.

[0263] In addition, in the embodiment, the computing system (1000) can visually display normal (OK) and bad (NG) judgment results through a tool provided with the above-described interactive interface, and visualize data labeling or model training processes to support the user in easily understanding the current system status.

[0264] For example, the computing system (1000) can provide real-time user interface feedback, including label recognition path visualization, error code list, process status icon, etc., by displaying it on the received vision data.

[0265] In addition, in the embodiment, when a user uploads at least one related document, such as a development process document and / or a quality definition document, the computing system (1000) can analyze it in conjunction with at least one artificial intelligence agent and independently plan and execute a sequential processing process from the data preparation stage to the model deployment stage.

[0266] In addition, in the embodiment, the computing system (1000) can collect, store, and manage vision data including a predetermined format (e.g., JPEG, etc.) based on a structured directory system divided into training (train) and verification (val) directories, and sub-directories for good (OK) and defective (NG) products.

[0267] In addition, in the embodiment, the computing system (1000) can control the entire flow to search for an optimal model structure using an external neural architecture search (NAS), etc., after data preparation is complete, or to perform learning based on a pre-trained inspection foundation model, etc., and to distribute the model that has undergone final evaluation to the inspection system at the site.

[0268] Through these configurations, the computing system (1000) in the embodiment has the effect of enabling even field workers lacking expertise in vision inspection to easily build the entire process of a high-performance inspection model and automatically distribute it through interactive interaction.

[0269] Also, referring to FIG. 9, a computing system (1000) according to one embodiment of the present disclosure can run a vision inspection agent including at least one vision language model, context memory, prompt configuration and / or document search module, and can interact with functional tools such as a data loader, out-of-distribution data detection (OOD Detection) and / or model training and external resources (storage system) through at least one planning module.

[0270] In detail, in an embodiment, the computing system (1000) can interpret natural language input from a user received through a conversational interface using at least one vision language model that acts as the brain of the system, and establish a response and task plan that are appropriate to the context by referring to past conversation history, task performance records and / or success or failure experiences, etc., stored in at least one context memory.

[0271] Additionally, in the embodiment, the computing system (1000) can actively search for documents directly uploaded by the user and / or documents stored in an external storage system through at least one document search module to identify constraints and quality standards required for model construction and transmit them to a vision language model.

[0272] In addition, in the embodiment, the computing system (1000) can dynamically call tools including data loading, data labeling, model training, noisy label detection, out-of-distribution data detection and / or model evaluation according to instructions from at least one vision language model, and can perform actual vision data processing and parameter tuning tasks in parallel in conjunction with external systems and inspection equipment based on the called tools.

[0273] At this time, as an example, the computing system (1000) can generate standard evaluation indicators such as accuracy, recall, and / or F1 score through the model evaluation tool described above and generate them in the form of a training report.

[0274] Additionally, as an example, the computing system (1000) can run the aforementioned noise label detection and out-of-distribution data detection tools to identify abnormal vision data and process it to automatically correct or exclude it from learning.

[0275] As a result, the computing system (1000) can provide the effect of maximizing the reliability of the inspection model and data robustness by not only accurately reflecting the user's intention based on the modularized tool call and long-term memory-based architecture as described above, but also filtering abnormal data distributions on its own.

[0276] Meanwhile, referring to FIG. 10, a computing system (1000) according to one embodiment of the present disclosure can perform integrated monitoring by establishing a communication interface including at least one communication protocol client (MCP Client) and / or server (MCP Server), etc., and transmitting and receiving data bidirectionally with at least one target infrastructure such as an external inspection device, database, quality management system, etc.

[0277] In detail, in an embodiment, the computing system (1000) may provide an integrated system layer of a comprehensive vision inspection agent platform that includes at least one monitoring tool in addition to at least one interactive interface and / or model learning tool described above.

[0278] At this time, according to the embodiment, the computing system (1000) can perform detailed functions such as alarm logger, real-time monitoring, and data uncertainty prediction through at least one monitoring tool.

[0279] To this end, in the embodiment, the computing system (1000) can hierarchically connect many-to-many client and server protocols to perform safe and standardized bidirectional communication with heterogeneous inspection equipment or external infrastructure that may have different communication standards.

[0280] As a specific example, the computing system (1000) can perform standard input / output (stdin / stdout) communication based on JSON format messages to smoothly exchange vision data, control commands, and evaluation results collected from the inspection equipment.

[0281] For example, the computing system (1000) can immediately receive vision data or artificial intelligence defect judgment results generated in real time from inspection equipment 1 and 2 physically placed at the site via a communication server to an upper client, and can detect abnormal changes in emission rates or uncertainty in equipment data based on a monitoring tool and report real-time warnings to the user through an alarm logger.

[0282] In addition, in the embodiment, the computing system (1000) can induce retraining to prevent model drift, a phenomenon in which the performance of a model deteriorates over time, by analyzing the emission rate trend over a preset period (e.g., 7 days) through the monitoring tool described above.

[0283] Based on this macroscopic integrated framework, the computing system (1000) can not only simply deploy one-off inspection models but also flexibly combine with various manufacturing site software and hardware systems to provide a significant effect of completing a complete closed-loop intelligent quality management ecosystem ranging from data collection, active learning, multiple deployments, and real-time post-management.

[0284] Meanwhile, a computing system (1000) according to one embodiment of the present disclosure can run at least one artificial intelligence agent based on a hierarchical memory structure similar to a human memory method for performing advanced tasks.

[0285] Specifically, the computing system (1000) may subdivide the memory of the artificial intelligence agent into a short-term memory including a working memory and a cache memory that manage the inference flow for each workflow currently in progress, and a long-term memory including an episodic memory, semantic memory, and / or procedural memory that permanently preserves procedural knowledge of a specific task or past success / failure experiences.

[0286] In addition, in the embodiment, the computing system (1000) can operate the artificial intelligence agent according to the React paradigm, in which a language model repeatedly performs reason and action to solve a problem.

[0287] Furthermore, the computing system (1000) can maximize the flexibility of data processing by controlling the artificial intelligence agent to operate according to the CodeAct paradigm, which directly generates and executes programming code such as Python when complex calculations such as data analysis are required.

[0288] In addition, the computing system (1000) according to the embodiment can dynamically configure various multi-agent system patterns in which multiple agents cooperate according to the type of task in order to overcome the limitations of a single agent.

[0289] According to an embodiment, the computing system (1000) may implement the multi-agent system pattern based on at least one of a sequential pattern in which a plurality of agents inherit the results of a previous step, a supervisor pattern in which a central supervisor model (Supervisor LM) distributes tasks to lower-level professional agents, a hierarchical pattern in which a meta-agent acts as an intermediate manager, a multi-agent debate pattern in which mutually exclusive opinions are discussed, and a mixture-of-AI agents pattern in which results are integrated across multiple levels.

[0290] At this time, according to the embodiment, the computing system (1000) can control the internal data exchange between agents based on an A2A (Agent2Agent) protocol that enhances internal security.

[0291] In addition, the computing system (1000) can ensure a safe and efficient inter-module communication and collaboration ecosystem by separating and controlling the interaction with external systems (inspection equipment, cloud API, etc.) through a server based on the aforementioned MCP (Multi-Channel Processing) protocol.

[0292] FIG. 11 illustrates an example of a flowchart for explaining a method for constructing an integrated vision inspection system using an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0293] Returning to FIG. 11, a computing system (1000) according to one embodiment of the present disclosure may perform the steps of: receiving at least one model building request data (S101); obtaining target-specific information based on the received model building request data (S103); identifying missing data based on the obtained target-specific information (S105); updating a model building work plan based on the identified missing data (S107); detecting ambiguous expression data within the updated model building work plan or the model building request data (S109); updating a model building work plan based on the detected ambiguous expression data (S111); building a model according to the updated model building work plan (S113); and distributing and monitoring the built model (S115).

[0294] In detail, a computing system (1000) according to one embodiment of the present disclosure can receive at least one model building request data. (S101)

[0295] Here, the model building request data according to the embodiment is data input by a user through an interactive interface, and may refer to broad data encompassing at least one of vision data subject to learning or verification, relevant documents associated with the vision inspection process (e.g., quality definition documents, development process documents, equipment drawings, etc.), and qualitative natural language input indicating the purpose of creating the vision inspection model or inspection criteria.

[0296] Specifically, in the embodiment, the computing system (1000) can support various receiving methods and input modalities to receive the model building request data in a multifaceted way, such as a file upload function provided in an interactive interface, direct application programming interface (API) linkage with an external storage system, real-time data streaming through a network, text conversion of voice input using speech-to-text (STT) technology and / or direct typing through a text input window.

[0297] More specifically, as an example, the computing system (1000) may receive, through an interactive interface, an upload command for at least one vision data and / or related document, etc., and / or qualitative text data, etc., spoken in the form of everyday language, as model building request data.

[0298] For example, when a user inputs a vision data set in file or streaming form into an interactive interface and provides natural language input including the target, purpose, and required inspection level of the vision inspection, such as "make the inspection model of the A-line pilot process strict," the computing system (1000) can process various input modalities integrally and recognize and acquire them as valid model building request data for system control.

[0299] Additionally, in an embodiment, the computing system (1000) may operate to transmit at least one received model building request data to an orchestration engine and / or vision language model within the computing system (1000).

[0300] In this embodiment, the vision language model of the computing system (1000) can precisely parse the user's work intent, requirements, and context embedded in the text data in semantic units based on a natural language processing algorithm, and the orchestration engine can dynamically configure an initial computation pipeline to preprocess the received vision data, check the labeling status, and plan subsequent work based on the parsed intent.

[0301] Through this, the computing system (1000) can support users who do not possess professional coding skills or complex hyperparameter tuning knowledge regarding deep learning architecture or vision inspection modeling, to easily access an advanced vision inspection platform and initiate a model building process through everyday and intuitive conversational interactions.

[0302] In addition, a computing system (1000) according to one embodiment of the present disclosure can obtain target-specific information based on received model building request data. (S103)

[0303] Here, the target-specific information according to the embodiment may refer to data indicating physical and environmental constraints and / or unique quality standards of the target environment (e.g., specific infrastructure, production lines of a specific factory and / or individual inspection equipment, etc.) where the vision inspection model to be built will be deployed and operated.

[0304] For example, target-specific information may include quality definitions, development process documents, and / or facility drawings that explicitly state optical constraints of the target environment, processing speed limits (e.g., frames per second throughput), communication protocol specifications, or unique defect criteria for a specific process.

[0305] In detail, in an embodiment, the computing system (1000) can parse and identify a specific target (e.g., 'A-line pilot process') from model building request data entered by a user into an interactive interface through a Planner Agent running within the system, and actively search for at least one relevant information by accessing an external resource associated with the target through a communication interface.

[0306] More specifically, in an embodiment, the computing system (1000) can dynamically access external storage systems such as an in-house database, an enterprise resource management (ERP) system, and / or a programmable logic controller (PLC) by utilizing a Retrieval-Augmented Generation (RAG) architecture and / or a machine control protocol (MCP)-based many-to-many client-server structure.

[0307] And the computing system (1000) can search for and collect documents, metadata, and constraint parameters associated with the target equipment specified by the user from these external resources and obtain them as target-specific information.

[0308] Through this, the computing system (1000) can not only generate a general-purpose vision inspection model using only ideal data from a laboratory environment, but also prevent compatibility issues or performance drops that occur when the model is distributed and fails to reflect the physical characteristics of the actual factory environment.

[0309] Therefore, the computing system (1000) can provide a solid technical foundation for building a customized vision inspection model that perfectly matches the infrastructure constraints and business requirements of the target site.

[0310] Additionally, a computing system (1000) according to one embodiment of the present disclosure can identify missing data based on acquired target-specific information. (S105)

[0311] Here, the missing data according to the embodiment may refer to missing information that is essential for the optimized design, training, and / or stable deployment to the actual target infrastructure of the intended vision inspection model, but is not clearly specified or is absent based only on the user's initial model building request data and collected target-specific information.

[0312] For example, missing data may encompass parameters that need to be supplemented to complete the system's autonomous model learning guide or job planning, such as specific image preprocessing resolution required to match the processing speed of the target equipment, quantified judgment criteria (thresholds) for specific micro-defects, specific labeling guidelines, or specific storage locations where vision data is stored.

[0313] In detail, in an embodiment, the computing system (1000) can identify at least one missing data based on at least one target-specific information and at least one model-building request data obtained as described above.

[0314] More specifically, in the embodiment, the computing system (1000) can cross-validate the intent of the natural language utterance input by the user and target-specific information, such as quality definition documents and development process documents collected through the document search module, by utilizing a planner agent and / or vision language model within the system.

[0315] Through this cross-validation process, the computing system (1000) can precisely detect items where data gaps have occurred among the control variables essential for establishing an overall work plan for model building and identify them as missing data.

[0316] For example, the computing system (1000) can identify the discrepancy and gap in this information as missing data that undermines the technical integrity of the inspection model construction when, in the process of setting a model learning goal based on collected target-specific information (e.g., pilot process quality guidelines), acceptance criteria for a specific type of defect (e.g., fine scratches of 0.5 millimeters or less) are specified in the target-specific information, but the user's model construction request data does not specifically indicate whether the fine scratches are labeled.

[0317] In another example, the computing system (1000) may determine that the physical storage path or location of the data is missing data that needs to be resolved for subsequent work, when the orchestration engine within the computing system (1000) needs to obtain vision data to execute a data preparation workflow, but the user's input text does not contain the physical storage path or location of the data.

[0318] Through this, the computing system (1000) does not remain merely an entity that passively executes user commands, but actively infers physical constraints and quality standards of the target environment and preemptively identifies missing essential information, thereby ensuring technical integrity for planning a defect-free vision inspection pipeline that is perfectly optimized for the actual deployment environment.

[0319] Additionally, a computing system (1000) according to one embodiment of the present disclosure can update a model building work plan based on identified missing data. (S107)

[0320] Here, the model building work plan according to the embodiment may refer to a sequential or complex computational pipeline that autonomously plans a series of entire (i.e., end-to-end) or partial workflows, ranging from the collection and preparation of data for model building to the design, training, evaluation, and final deployment of the model, by identifying user intent and / or target-specific information based on at least one vision inspection agent (or vision language model).

[0321] Such a model building work plan may encompass a work specification that comprehensively defines the call sequence of various functional tools, such as data loaders, data labeling, and model training, data processing methods, and / or integration specifications with target infrastructure.

[0322] Specifically, in an embodiment, the computing system (1000) can generate at least one interactive clarification question to obtain at least one previously identified missing data based on the above-mentioned data and output it through an interactive interface, and can specify and / or update an existing initial model building work plan based on the user's response to the above.

[0323] More specifically, in an embodiment, the computing system (1000) can convert a gap (i.e., identified missing data) derived during the process of verifying the technical integrity of the model building environment into an explicit query in the form of natural language (i.e., conversational query) and transmit it to a user through a conversational interface (e.g., conversational window, etc.) by linking with at least one vision language model, etc.

[0324] And when the computing system (1000) receives response data from a user in response to a transmitted interactive query, it can extract specific parameters, setting values ​​and / or instructions contained in the response through a natural language processing algorithm, etc., and map them to the previously identified missing data area.

[0325] Thus, the computing system (1000) can perform a model building work plan update that supplements the existing incomplete initial model building work plan with a technically completed specific workflow.

[0326] For example, if the computing system (1000) identifies the physical storage path of the vision data as missing data in the user's initial build request, it may output an interactive reverse query asking, "Where is the location where the vision data is stored?" In response, if the user replies with the storage location (e.g., "D: / TireInspection / Data"), the computing system (1000) may update the model build job plan to call a Data Loader tool reflecting the directory path to load the vision data into the system.

[0327] In another example, when the computing system (1000) cross-validates target-specific information (such as quality definition documents) and the directory structure of the currently acquired vision dataset and identifies a gap in the labeling criteria as missing data, it can generate an active, interactive reverse query based on past history, such as, "I have checked the current status. There are 50 defective (NG) vision data and 1,000 normal (OK) vision data. Could you please check if labeling is needed?" or "There was a history of leakage regarding a specific type of defect in the past. Shall we configure the training dataset considering that problem?" If the user responds to the reverse query that labeling is needed, the computing system (1000) can immediately add and update a step to activate an automatic labeling tool or a manual labeling interface within the work plan.

[0328] Through this organic interaction structure, the computing system (1000) in the embodiment can establish a flawless, perfect customized model building work plan by having an agent actively intervene to supplement the insufficient information, even if the user does not fully understand all the constraints or essential parameters of the complex vision inspection system and gives vague instructions.

[0329] Additionally, a computing system (1000) according to one embodiment of the present disclosure can detect ambiguous representation data within an updated model building task plan or model building request data. (S109)

[0330] Here, ambiguous expression data according to the embodiment may refer to linguistic syntax and / or text data, etc., which contain only the user's subjective and qualitative business intentions or requirements without being defined by clear numerical values, mathematical thresholds, and / or quantitative parameters.

[0331] For example, ambiguous expression data can include uncertain expressions that lack specific numerical criteria related to the defect detection sensitivity or production yield of the vision inspection model, such as "inspect strictly," "to the extent that it can detect thoroughly," or "at an appropriate level."

[0332] At this time, according to the embodiment, the computing system (1000) may omit the 'detection' step of explicit ambiguous expression data, and may immediately utilize the 'qualitative requirements' contained in the user's natural language input itself as basic data for multi-agent-based parameter substitution.

[0333] Specifically, in an embodiment, the computing system (1000) can detect at least one ambiguous expression data associated with a vision inspection standard setting by analyzing at least one natural language text included in an updated model building work plan or at least one model building request data described above.

[0334] More specifically, in an embodiment, the computing system (1000) can precisely parse response text accumulated through the user's initial utterance input and / or conversational reverse query, etc., by utilizing a natural language processing (NLP) algorithm of a vision language model provided within the system or through a pre-configured system prompt, in semantic units.

[0335] And the computing system (1000) can actively identify text phrases that have qualitative vocabulary or numerical uncertainty that cannot be 1:1 converted (Mapping) into hyperparameters of the inspection model (e.g., score threshold for defect judgment, defect size pixel standard, etc.) among the text phrases parsed as above, and separate and label them as 'ambiguous expression data' that the system must process subsequently.

[0336] For example, a computing system (1000) can utilize a natural language processing (NLP) algorithm including morphological analysis and semantic inference of a vision language model to calculate that the adverb “meticulously” in the user’s input sentence is logically associated with the model’s ‘defect judgment threshold’ but does not indicate a specific numerical value (e.g., a threshold of 0.95 or higher), and can automatically extract and label this as ambiguous expression data that requires quantification.

[0337] In another example, a computing system (1000) identifies qualitative adjectives or adverbs lacking numerical parameters representing vision inspection criteria in the user's input text, by<Ambiguous_Expression> By pre-injecting a system prompt configured with the content "extract by tag" into the vision language model, the phrase "to an appropriate level" within the phrase "filter to an appropriate level" spoken by the user can be quickly and consistently detected as ambiguous expression data.

[0338] Through such an ambiguous expression data detection configuration, the computing system (1000) can convert instructions into quantitative and concrete control processes within the computing system (1000) even if the user gives instructions using only their subjective language and abstract business goals due to a lack of knowledge regarding complex deep learning internal variables or precise numerical control, thereby intelligently bridging the gap between the user's qualitative requirements and the mathematical parameters of the machine learning model.

[0339] Additionally, a computing system (1000) according to one embodiment of the present disclosure can update a model building work plan based on detected ambiguous representation data. (S111)

[0340] Specifically, in the embodiment, the computing system (1000) can perform a series of simulation processes to replace detected qualitative and uncertain ambiguous representation data with specific and quantitative system parameters, thereby updating the existing unclear work plan into a mathematically and logically optimized model-building work plan.

[0341] FIG. 12 illustrates an example of a flowchart for explaining the operation method of a vision inspection control system based on an interactive artificial intelligence agent according to one embodiment of the present disclosure.

[0342] Specifically, with reference to FIG. 12, a computing system (1000) according to one embodiment of the present disclosure may perform the steps of: generating a plurality of agents having conflicting objective functions based on detected ambiguous representation data (S201); executing a simulation based on the generated plurality of agents (S203); determining an optimal parameter value based on the results of the executed simulation (S205); and updating a model building work plan based on the determined optimal parameter value (S207).

[0343] More specifically, a computing system (1000) according to one embodiment of the present disclosure can generate a plurality of agents having conflicting objective functions based on detected ambiguous representation data. (S201)

[0344] Here, conflicting objective functions according to the embodiment may mean optimization indicators and / or mathematical weighting formulas that are difficult to achieve perfectly at the same time or are mutually exclusive.

[0345] For example, conflicting objective functions may encompass an exclusive optimization relationship between a first optimization metric that aims to maximize 'production yield' by preventing normal products from being discarded as defective through the minimization of false positives, and a second optimization metric that aims to achieve 'zero-defect quality assurance' by completely blocking the leakage of defective products through the zeroing of false negatives.

[0346] In detail, in an embodiment, the computing system (1000) may configure a Multi-Agent Collaboration Structure that dynamically generates multiple artificial intelligence agents assigned different personas and / or optimization goals, instead of relying on the biased judgment of a single model to resolve the abstract requirements of a user (e.g., "check thoroughly") embedded in the ambiguous expression data in a fragmentary manner.

[0347] According to an embodiment, the multi-agent collaboration structure may include a multi-agent debate pattern that discusses mutually exclusive opinions.

[0348] More specifically, in the embodiment, the orchestration engine and / or central supervisory language model (Supervisor LM) within the computing system (1000) can generate a plurality of sub-agents having mutually exclusive logical structures in memory in parallel when it identifies that detected ambiguous expression data is associated with vision inspection parameter tuning.

[0349] For example, a computing system (1000) can interpret the instruction "meticulously," which is ambiguous expression data entered by a user, and generate a first agent based on a first objective function that has a 'quality manager persona' designed to be highly sensitive to defect classification and sets a high false positive (FN) penalty. At the same time, the computing system (1000) can also generate a second agent based on a second objective function that has a 'production yield manager persona' and sets a high false positive (FP) penalty to defend against the phenomenon of yield reduction that may occur as a counter-effect of quality management.

[0350] Through this multi-agent generation configuration based on conflicting objective functions, the computing system (1000) can prevent arbitrary reasoning errors or biases that a single artificial intelligence model may have in advance, and can provide an advanced intelligent decision-making framework in which multiple agents with opposing views on a single ambiguous instruction can logically compete, discuss, and mutually evaluate the same problem.

[0351] In addition, a computing system (1000) according to one embodiment of the present disclosure can execute a simulation based on a plurality of generated agents. (S203)

[0352] In detail, in an embodiment, the computing system (1000) can control each of the multiple agents having mutually exclusive objective functions generated as above to interpret the previously detected ambiguous expression data based on their own perspectives and personas to set different thresholds and / or hyperparameters, and execute a simulation to perform vision inspection prediction operations by individually applying the set parameters to a previously secured validation dataset.

[0353] In this case, according to the embodiment, the simulation may be performed in the form of a parallel simulation, and the parallel simulation may refer to an independent verification process in which a plurality of agents, each assigned different objective functions and parameters, simultaneously execute defect prediction and classification operations on the same bundle of verification vision data in a multi-threaded and / or distributed processing environment.

[0354] More specifically, in the embodiment, the computing system (1000) can perform a simulation as described above based on a Multi-Agent Collaboration Structure in which multiple agents present different opinions or solutions regarding the same problem and compare them to derive an optimal solution.

[0355] For example, the computing system (1000) can run a first simulation logic that induces a first agent, assigned a quality manager persona, to set a defect judgment score threshold conservatively low (e.g., judged as defective (NG) if the defect probability is 0.7 or higher) in response to ambiguous expression data "meticulously" entered by a user, thereby blocking the occurrence of a false positive (FN). At the same time, the computing system (1000) can run a second simulation logic in parallel in the background that induces a second agent, assigned a production yield manager persona, to set the threshold high (e.g., judged as defective (NG) only when the defect probability is 0.95 or higher) thereby minimizing false positives (FP) where a normal product is mistaken for a defect.

[0356] Through this parallel simulation configuration, in the embodiment, the computing system (1000) can not only perfectly map ambiguous natural language instructions from the user to quantitative data processing and inference processes within the computing system (1000), but also pre-verify conflicting results that may occur before field deployment of the actual vision inspection model in a virtual environment, thereby minimizing risk and obtaining multifaceted analysis data to derive optimal model parameters in a short period of time.

[0357] In addition, a computing system (1000) according to one embodiment of the present disclosure can determine optimal parameter values ​​based on the results of an executed simulation. (S205)

[0358] Here, the optimal parameter value according to the embodiment may refer to quantitative and mathematical control variables, including specific neural network weights, loss function penalty weights, defect judgment score thresholds, and / or image preprocessing resolution criteria, verified through simulation to achieve the most ideal balance point among multiple conflicting objective functions (e.g., production yield and defect-free quality).

[0359] Specifically, in the embodiment, the computing system (1000) can synthesize the simulation results based on the Multi-Agent Collaboration Structure executed as described above to derive a quantitative analysis result of the trade-off for mutually conflicting performance evaluation indicators between multiple agents, and determine the final optimal parameter value to be applied to the vision inspection model based on this.

[0360] Here, the trade-off quantitative analysis result according to the embodiment may refer to a comparison indicator that objectifies and visualizes the inverse relationship between multiple mutually exclusive optimization indicators calculated by agents with different personas using independent threshold values, in the form of a multidimensional graph or numerical data.

[0361] For example, the aforementioned conflicting performance evaluation indicators (or optimization indicators) may include production yield, which aims to maximize the normal judgment rate, and defect rate, which aims to minimize the risk of defective product leakage. However, this is merely an example and is not limited thereto, and various performance indicators that are difficult to satisfy simultaneously in a vision inspection system may be applied, such as the relationship between the false positive rate and the false negative rate, and / or the relationship between inspection processing speed and judgment accuracy.

[0362] More specifically, in the embodiment, the computing system (1000) can collect raw prediction results and performance evaluation data calculated by each of the plurality of agents during the simulation process described above, and convert them into a quantitative analysis model that multidimensionally represents the correlation between conflicting indicators to derive a quantitative analysis result of the trade-off.

[0363] For example, the computing system (1000) can derive and visualize multidimensional trade-off quantitative analysis results, such as Pareto Frontiers, by comprehensively calculating detailed raw data such as Confusion Matrix data derived by multiple agents according to their respective unique parameter setting values, inspection processing time logs, and the number of false positives (FP) and false negatives (FN).

[0364] In addition, in the embodiment, the computing system (1000) can determine the optimal parameter value described above based on the quantitative analysis result of the trade-off derived as above.

[0365] In an example, the computing system (1000) can display the results of the quantitative analysis of trade-offs in the form of summary data (e.g., multidimensional graphs, visualization graphs, numerical comparison data and / or comparison text, etc.) on an interactive interface and determine optimal parameter values ​​based on user input based thereon.

[0366] For example, the computing system (1000) can compare yield and emission rate as conflicting indicators and provide comparison text and a visualization graph such as “Yield 98.0% and emission rate 0.00% when Plan A (applying quality manager agent parameters) and Yield 99.5% and emission rate 0.05% when Plan B (applying production manager agent parameters)”.

[0367] In this case, in the embodiment, the computing system (1000) can induce the user to directly select the balance point most suitable for the current process situation based on objective simulation indicators, and determine the parameter corresponding to the user's selection data as the optimal parameter value.

[0368] In another embodiment, the computing system (1000) can autonomously refer to the trade-off quantitative analysis results and pre-set optimal parameter value judgment criteria through an orchestration engine and / or a central supervised language model (Supervisor LM) to determine the optimal parameter value itself through internal calculation without user intervention.

[0369] Here, the criteria for determining the optimal parameter value set according to the embodiment may include the history of past quality complaints regarding a specific target (e.g., whether a claim has occurred due to a specific defect leak within the last 3 months), the daily production target value of the current production line, and / or the upper limit of the allowable maximum emission rate specified in the quality definition document.

[0370] For example, if a recent history of fatal defect leakage is recorded in the context memory, the computing system (1000) may autonomously adopt the parameters of the first agent as the optimal value, which enforces the emission rate to 0.00% even if the yield is slightly reduced (e.g., 98.0%). On the other hand, if the system (e.g., linked ERP system) records that there is no history of quality claims and that achieving the current month's production target is urgent, the computing system (1000) may automatically calculate and determine the parameters of the second agent as the optimal parameter value, which maximizes the yield (e.g., 99.5%) within the range satisfying the condition of the allowable emission rate upper limit (e.g., less than 0.1%).

[0371] As a result, in the embodiment, the computing system (1000) can fully quantify the user's vague and subjective requirements or conflicts of opinion, such as "moderately" or "strictly," into objective indicators through the simulation-based optimal parameter determination configuration as described above.

[0372] Therefore, the computing system (1000) can overcome the limitations of existing systems that rely on the arbitrary judgment or trial and error of non-experts and can build the most optimized and verified vision inspection pipeline that aligns with actual business contexts (e.g., risk avoidance or productivity maximization) and data-based rational decision-making.

[0373] Furthermore, according to an embodiment, the computing system (1000) can create a User Profile by mapping the determined optimal parameter value in a 1:1 manner with the identification information (User ID) of the user and ambiguous expression data (e.g., "meticulously") detected in the initial model building request data, and store and manage it in the context memory, which is the long-term memory area of ​​the system.

[0374] Here, the user profile according to the embodiment may refer to a personalized custom database comprising subjective and qualitative ambiguous expression data frequently used by a specific user, and mapping rules for quantitative optimal parameter values ​​that define how the expression should be specifically replaced by numerical criteria (e.g., defect judgment score threshold) within the system.

[0375] Through this, the computing system (1000) can prevent waste of computational resources by automatically processing and applying the optimal parameter value (e.g., defect threshold 0.95 or higher) matched to the user profile immediately, by bypassing the complex and computationally intensive multi-agent simulation and trade-off analysis process when the same user later utters the same ambiguous expression "strictly" regarding a similar vision inspection process.

[0376] As a result, the computing system (1000) can produce a remarkable effect by providing an advanced customized agent architecture that intelligently evolves by learning personalized context as past conversation content and work history accumulate.

[0377] Additionally, a computing system (1000) according to one embodiment of the present disclosure can update a model building work plan based on a determined optimal parameter value. (S207)

[0378] In detail, in an embodiment, the computing system (1000) can update the initial model building work plan (i.e., work computation pipeline) that was previously established by injecting the optimal parameter value determined as described above into the work specification within the system as a specific control variable.

[0379] More specifically, in the embodiment, the computing system (1000) can specify the entire workflow by collectively updating the setting values ​​of individual function tools leading to preprocessing of vision data, model structure exploration, learning, evaluation, etc., based on optimal parameter values.

[0380] For example, the computing system (1000) can update the work plan to reflect the image resolution determined by optimal parameter values ​​(e.g., scaled up from 512x512 to 1024x1024 for fine defect detection) and the precise crop ratio, etc., in relation to the data loader and preprocessing steps, as execution variables of the data preprocessing tool.

[0381] In another example, the computing system (1000) can update the training plan by directly mapping the penalty weights and defect detection thresholds of the loss function determined as optimal parameters in relation to the model training stage to the search space constraints of the neural architecture search (NAS) or the fine-tuning hyperparameters of the pre-trained inspection foundation model.

[0382] As another example, the computing system (1000) can fix specific target indicators (e.g., emission rate less than 0.05% and yield greater than 99.5%) that the evaluation tool should use as a passing criterion when verifying the performance of the model within the work plan as final verification criteria.

[0383] Through this configuration, the computing system (1000) in the embodiment can perfectly resolve the user's ambiguous initial requirements into clear numbers that the system can understand, and smoothly implement the full-scale model building process with the optimal balance between conflicting business goals mathematically guaranteed.

[0384] Returning to FIG. 11, a computing system (1000) according to one embodiment of the present disclosure can build a model according to an updated model building work plan. (S113)

[0385] In detail, in an embodiment, the computing system (1000) can automatically perform a series of processes to design and train at least one vision inspection model based on an updated model building work plan as described above, and to quantitatively evaluate the performance to derive a final model (hereinafter, optimal target model) to be distributed to a target process.

[0386] More specifically, in the embodiment, the computing system (1000) can automatically design an optimal model architecture that meets the user's requirements and data characteristics by utilizing a pre-trained inspection foundation model or through Neural Architecture Search (NAS) technology, by reflecting the optimal parameters (e.g., preprocessing resolution, threshold value, etc.) determined within the updated model building work plan.

[0387] In addition, in the embodiment, the computing system (1000) can perform model training using at least one training dataset (e.g., good and defective data under the train directory) based on the architecture designed as above.

[0388] During this learning process, the computing system (1000) can continuously run a Noisy Label Detection tool and / or an Out-of-Distribution Data Detection (OOD) tool to detect data uncertainty, thereby strictly monitoring and managing the quality of the learning data.

[0389] For example, the computing system (1000) can ensure the integrity of the data set by automatically excluding the data from training or requesting a review from the user when an outlier (OOD) that is different from the distribution learned by the model is detected or an error (Noisy Label) that occurred during the labeling process is identified.

[0390] In addition, in the embodiment, the computing system (1000) utilizes a communication interface based on an MCP client-server structure to additionally input real-time vision data collected from an actual manufacturing line or inspection equipment and reflect it in the learning even while model learning is in progress.

[0391] Through this, the computing system (1000) can build a robust model that is not limited to a laboratory environment but flexibly adapts to various conditions and changes in an actual process environment.

[0392] In addition, in the embodiment, the computing system (1000) can evaluate the performance of the constructed model using at least one verification dataset.

[0393] At this time, according to the embodiment, the computing system (1000) can generate a training report that calculates standard evaluation indicators such as accuracy, recall, F1 score, and / or defect rate based on these evaluation results.

[0394] In addition, the computing system (1000) can provide the user with the generated training report in various forms, such as multidimensional graphs and / or numerical data, through an interactive interface, so that the user can intuitively grasp the overall detection performance of the trained model and the classification status by defect type, and easily determine the feasibility of field application of the model or the need for additional training based on objective data.

[0395] In addition, the computing system (1000) can autonomously perform a performance trend analysis over a pre-set period based on time-series change data of key performance indicators included in the generated training report.

[0396] Specifically, the computing system (1000) can systematically identify and manage whether the performance of the model is gradually improving through the iterative learning process or is deteriorating due to specific environmental changes or data bias by intensively analyzing the fluctuation trend over a predetermined period (e.g., 7 days) based on at least one indicator (e.g., defect rate indicator) in the training report.

[0397] As such, according to the embodiment, the computing system (1000) can fundamentally prevent the risk of defective product leakage that may occur on the actual production line and preemptively prevent the phenomenon of performance degradation over time after model deployment (model drift) by utilizing the training report described above to thoroughly verify the stability of performance in advance and determine the direction of optimization and whether to operate the retraining pipeline before deploying the model to the target process.

[0398] As a result, the computing system (1000) can stably maintain the reliability and accuracy of the vision inspection model to be built for a long period of time, thereby successfully improving the productivity of the manufacturing process, reducing quality control costs, and implementing an advanced intelligent smart factory environment.

[0399] Additionally, in the embodiment, if the calculated evaluation metric does not reach the target value specified in the work plan (e.g., defect detection accuracy of 99% or higher), the computing system (1000) can analyze the problem itself and establish a new plan for performance improvement, such as adjusting hyperparameters, collecting additional data, augmenting data, or upscaling the input image resolution (e.g., scaling up from 512x512 resolution to 1024x1024 resolution), and can repeatedly perform retraining until the target performance is reached.

[0400] Meanwhile, in the embodiment, if the performance indicator calculated through the evaluation process does not reach the target value specified in the work plan (e.g., defect detection accuracy of 99% or higher), the computing system (1000) can operate a performance improvement loop that analyzes the problem itself, establishes a new plan, and repeatedly performs relearning, rather than simply reporting the failure.

[0401] In an example, the computing system (1000) may establish a new work plan including specific solutions such as fine-tuning of hyperparameters, additional data collection, data augmentation, and / or upscaling of input image resolution (e.g., scaling up from 512x512 resolution to 1024x1024 resolution), and may repeatedly perform retraining until a target performance is reached.

[0402] Furthermore, according to an embodiment, the computing system (1000) may adopt an advanced optimization strategy of configuring multiple candidate models (e.g., model_v1 and model_v2) with different architectures or advantages in parallel to overcome the performance limitations of a single model architecture, evaluating each of them, and then combining them using an ensemble technique.

[0403] Finally, the computing system (1000) can determine an optimal target model that maximizes the overall target performance indicator (e.g., AUROC level 0.99) to the extreme through such autonomous relearning and ensemble fusion processes, and by safely storing this in a Model Registry, it can build a perfect customized model for deployment to target equipment and processes.

[0404] Meanwhile, according to an embodiment, the computing system (1000) can build a model by adopting a 'Bayesian integrated vision inspection framework' that can integrally process various defect and noise labels based on the probability distribution of data and Bayes' Theorem, as the core architecture of the vision inspection model (i.e., optimal target model) described above.

[0405] Specifically, a computing system (1000) (or an agent of the computing system (1000)) can divide each vision data within a training data set into multiple patches and use a pre-trained vision data encoder to extract each patch as a high-dimensional feature, which is a patch embedding.

[0406] Subsequently, the computing system (1000) can group the labeled patch embeddings by class (e.g., good product, scratch defect, etc.) and model a unique data probability distribution for each class using a Gaussian Mixture Model (GMM).

[0407] Furthermore, when test vision data to be inspected is input, the computing system (1000) can perform objective and probabilistic classification by calculating the final posterior probability that each patch belongs to a specific class by performing Bayesian inference on the test patch embeddings extracted from the data, which combines the likelihood and prior probability based on the modeled distribution.

[0408] Through such a Bayesian-based modeling configuration, the computing system (1000) of the present disclosure can mathematically and completely control various constraints and variables that occur during the process of building a conversational model through an artificial intelligence agent.

[0409] Specifically, the computing system (1000) can respond very flexibly by immediately replacing the inspection sensitivity (e.g., “inspect more strictly”) or the detection instruction for a new type of defect that is not present in the training data, which the user vaguely requested through an interactive interface, with the ‘Prior probability parameter for the Anomaly class’ within the Bayes inference process.

[0410] In addition, the computing system (1000) can statistically correct itself through Bayes' theorem, which compares the probability distribution of the entire data, instead of making a misjudgment based on fragmentary distances, even if a noise label (contaminated distribution) occurs in which a normal (OK) patch is mistakenly assigned as a bad (NG) patch during the process of preparing and labeling data through an interactive tool by the user.

[0411] Through this, the computing system (1000) can block error propagation at the source and further maximize the robustness of the model.

[0412] Additionally, a computing system (1000) according to another embodiment of the present disclosure may adopt an 'in-context learning-based vision inspection framework' as another core architecture of the optimal target model, which dynamically controls the operation of the model through a prompt containing a small number of image and label examples without updating model parameters by backpropagation or the like using in-context learning based on a sequence model.

[0413] Specifically, the computing system (1000) can obtain at least one 'Task Prompt' defining an example of an inspection rule and a 'Query Image' which is the actual target of judgment from a user interface or production line system.

[0414] Additionally, the computing system (1000) can generate a prompt embedding sequence by converting image-label pairs of a task prompt into a single embedding vector through an image-label encoder (ILE), and can convert a query image into a query embedding vector through an image encoder (IE).

[0415] Subsequently, the in-context learning sequence model (ICSM) of the computing system (1000) interprets the context according to the generated prompt embedding sequence to dynamically determine the judgment rule of the current task and applies it to the query embedding to infer the final label of the query image immediately without a separate parameter update process.

[0416] By applying such an in-context learning-based model architecture to a conversational agent system, the computing system (1000) of the present disclosure can achieve a remarkable effect of maximizing the operational efficiency and flexibility of the vision inspection pipeline.

[0417] Specifically, when a sudden change in the environment is detected, such as a change in lighting on a production line or the appearance of a new type of defect that does not exist in existing training data, the computing system (1000) can receive several images of the new defect (example) and simple labeling data therefor from the user through an interactive interface.

[0418] And the computing system (1000) can autonomously update the model's judgment criteria in just a few seconds without needing to run a costly model retraining pipeline that takes hours to days by immediately configuring the received data into a new 'task prompt' and injecting it into the model's context.

[0419] As a result, the computing system (1000) can successfully implement an uninterrupted quality management ecosystem that perfectly responds to the multi-product, small-batch production and variable process environment of a smart factory by allowing an intelligent agent to autonomously redefine inspection standards in real time using only a few examples (Few-shot) intuitively provided by the user without complex code modification or parameter tuning.

[0420] In addition, a computing system (1000) according to another embodiment of the present disclosure may autonomously introduce a 'Coarse-to-Fine patch-level classification framework' as the core inference pipeline of the model in order to fundamentally resolve the computational bottleneck that occurs when the constructed optimal target model processes high-resolution vision data in an actual mass production environment in real time.

[0421] Specifically, when target vision data is input, the computing system (1000) can, instead of immediately analyzing the entire area with a heavy deep learning model, first align and divide the target vision data according to the coordinate system of the representative good product data set in advance, and then compare the feature distances for each patch to quickly detect only the area suspected of being defective (patches suspected of being defective) in the first stage (Coarse detection).

[0422] And the computing system (1000) can dynamically search for sample patches with high similarity to the first detected suspected defective patch in the established data pool to obtain a 'Prompt Support Set' and, by referring to it, perform a fine second detection only on the suspected patch.

[0423] By combining such a Coarse-to-Fine architecture with an interactive agent system, the computing system (1000) can flexibly and perfectly satisfy requirements where there is a trade-off between processing speed and accuracy, such as "guarantee a inspection speed of more than 10 frames per second without missing fine defects" when a user directs through an interactive interface.

[0424] That is, the computing system (1000) can compress the target candidate group by boldly omitting computations on the majority of normal background patches that do not require precise analysis, and concentrate the resources for the precise classification described above (e.g., classification based on in-context learning and Bayes inference) only on the core areas where actual defects are possible.

[0425] Through this, the computing system (1000) can overcome the dilemma of speed and accuracy, which was a technical limitation of existing automated inspection, and implement an innovative intelligent quality control network that simultaneously achieves ultra-high speed and high precision vision inspection.

[0426] Furthermore, a computing system (1000) according to another embodiment of the present disclosure can integrate a 'vision inspection camera parameter setting automation framework' into the workflow of an interactive agent as a hardware control step for acquiring raw vision data of the highest quality, beyond the optimization of a software model architecture.

[0427] Specifically, a computing system (1000) (or an agent of the computing system (1000)) can acquire sample vision data by randomly changing combinations of parameter values, such as a camera (e.g., exposure time, focus, sensitivity, etc.) and / or a lighting device (e.g., brightness, color, etc.) through a grid search and / or Bayesian optimization algorithm, for good samples and defective samples placed in actual inspection equipment.

[0428] Additionally, the computing system (1000) can detect a region of interest (ROI) from vision data obtained through a pre-trained Vision Foundation Model and extract high-dimensional features of the region.

[0429] Additionally, the computing system (1000) can measure the distance between the features of the extracted good sample and the features of the defective sample, and determine the combination of parameter values ​​that maximize the distance between the features (i.e., the ability to distinguish between good and defective products is maximized) as the final 'optimal parameter set' and automatically apply it to the inspection equipment.

[0430] By combining this hardware parameter automatic optimization configuration with an interactive agent system, the computing system (1000) can dramatically improve the setup and maintenance efficiency of the vision inspection process.

[0431] Specifically, the computing system (1000) can receive a natural language instruction from a user through an interactive interface, such as “Set up the camera and lighting because a new part has been introduced,” in a situation where a new product line is established or the inspection environment has changed.

[0432] And the computing system (1000) that receives these instructions can autonomously and completely automate the existing analog and manual equipment setting work, which had to rely on numerous repetitive experiments and the subjective intuition of a vision expert, based on quantitative feature distance indicators within the system.

[0433] As a result, the computing system (1000) can implement a true intelligent unmanned smart factory that ensures data consistency by preventing setting errors between equipment due to variations in the skill level of workers, and actively and autonomously corrects hardware conditions even in the event of changes in the physical environment such as lighting aging, thereby permanently maintaining the best inspection accuracy.

[0434] In addition, a computing system (1000) according to one embodiment of the present disclosure can distribute and monitor a constructed model. (S115)

[0435] Here, model deployment and monitoring according to the embodiment may refer to a series of post-management and MLOps (Machine Learning Operations) processes that involve porting the optimal target model built through the aforementioned steps to a target environment, such as inspection equipment on an actual production line or an edge device, and running it, and continuously tracking real-time inference data and equipment status occurring during the actual mass production process to detect performance degradation and abnormal signs.

[0436] In detail, in the embodiment, the computing system (1000) can convert and optimize the above-described optimal target model to suit the physical constraints and / or communication standards of the target environment.

[0437] For example, the computing system (1000) can convert the optimal inspection model into a lightweight format optimized for the target environment, such as TensorRT, ONNX, or OpenVINO, based on the target-specific information initially acquired (e.g., inference speed limit, memory capacity, etc.).

[0438] In addition, in the embodiment, the computing system (1000) can remotely distribute the converted model through a network (e.g., a local area network (LAN) within a specific factory, a closed intranet, and / or a cloud-based wide area network (WAN), etc.) and operate a real-time monitoring system.

[0439] For example, the computing system (1000) can package the model into a container such as Docker and transmit and install it to a vision inspection PC or edge server at the site through an MCP-based communication interface.

[0440] In addition, in the embodiment, when the deployment is completed and real-time vision inspection is initiated on the actual target environment, the computing system (1000) can perform an active monitoring process to evaluate the continuity of inspection performance and system stability by collecting multidimensional real-time log data indicating the operating status and prediction results of the model from the target environment.

[0441] In an example, the computing system (1000) can perform monitoring that comprehensively collects and analyzes software processing indicators such as real-time inference prediction values ​​and inspection time (latency) streamed from a target environment, as well as hardware status data such as system resources of the equipment (e.g., CPU / GPU usage, memory occupancy, etc.).

[0442] As a specific example, the computing system (1000) can continuously calculate and track in the background whether a data drift phenomenon occurs in which the characteristics of the vision data flowing in in real time differ slightly from the distribution of past training data, or whether an abnormal trend (e.g., a sharp increase in false positive rates at a specific time) is detected that deviates from the target emission rate or yield standard set by the user.

[0443] Furthermore, in the embodiment, when an anomaly is detected through the real-time monitoring process in which the performance indicators of a previously distributed model deviate from a preset allowable threshold range, the computing system (1000) can activate an autonomous performance recovery mechanism for model updating along with an immediate administrator notification.

[0444] In an example, the computing system (1000) can immediately send an alarm message and / or a detailed monitoring report to an administrator through an interactive interface to quickly recognize the situation if it detects that the inference performance of the model has deteriorated below a preset threshold due to changes in the physical environment of the actual production site (e.g., fluctuations in lighting conditions, camera lens contamination, etc.) and / or the occurrence of a novel defect pattern that did not exist in the existing training data.

[0445] Additionally, the computing system (1000) can activate a data pipeline that instructs real-time vision data before and after the point in time when performance degradation is detected to be automatically collected and labeled as new training data.

[0446] And the computing system (1000) can dynamically update the existing model building work plan based on the new data obtained in this way, and autonomously trigger a retraining pipeline within the system to optimize the model according to the changed environment.

[0447] Through this autonomous maintenance and relearning configuration, the computing system (1000) can continuously maintain the robustness of the vision inspection model even amidst inevitable physical changes in the production environment, and permanently operate a defect-free quality control network while minimizing manual intervention by a manager.

[0448] In the following, embodiments based on specific usage scenarios will be described to explain how the aforementioned system architecture and methodology can be applied in actual target environments (e.g., specific industrial sites). However, this is merely one example, and the technical concept or scope of rights of this disclosure is not limited to the said specific scenario; it is evident that it can be modified and applied in various forms depending on the characteristics of the applicable industrial field or target environment.

[0449] In an example, if labeled data already exists, a computing system (1000) according to one embodiment of the present disclosure can automatically perform a process of correcting errors and maximizing model performance through interactive interaction when there is training data in which good products and defective products are pre-classified.

[0450] In detail, the computing system (1000) can support the user in creating a project named "assembly_inspect" and uploading a compressed file containing training (train) and verification (val) images classified as good (ok) and defective (ng) products during the project creation and data loading process through an interactive interface.

[0451] At this time, the vision inspection agent of the computing system (1000) can analyze the directory structure of the file to automatically recognize the class label and define it by separating it into data for the prompt and query data for evaluation.

[0452] Additionally, the computing system (1000) can generate and evaluate an initial candidate model (e.g., model_v1) with a preprocessing resolution adjusted based on a preset encoder (e.g., WideResNet-50) for initial model generation and cross-validation.

[0453] Afterward, the computing system (1000) can present a list of false positive and false negative images to the user along with a graph of the AUROC score and NG score distribution, thereby inducing the user to review the integrity of the label.

[0454] Furthermore, during the error analysis and active label correction process, if a user points out that a specific defect (e.g., a crack) among the missing (FN) images was not recognized, the computing system (1000) can actively search within the training dataset to find good (OK) label images that have features similar to the crack and present them to the user.

[0455] When the user checks this and modifies the label to defective (NG), the computing system (1000) can immediately perform a re-evaluation based on the updated dataset to achieve performance improvement.

[0456] Additionally, when a defect that is difficult to distinguish with the existing resolution, such as fine foreign matter, is identified during the resolution adjustment and ensemble optimization process, the computing system (1000) can build and evaluate a new candidate model (e.g., model_v2) with the input resolution increased (e.g., 1024x1024).

[0457] Thus, the computing system (1000) proposes a strategy to combine two models with different advantages (e.g., model_v1, model_v2) using an ensemble technique, thereby generating a final model that is close to perfect at an AUROC level of 0.99, which can be stored in a model registry and distributed to an actual production line.

[0458] In another embodiment, where only unlabeled raw data exists, a computing system (1000) according to another embodiment of the present disclosure can build a complete dataset with minimal user intervention through active interaction simulating an active learning technique, even in a situation where only a large amount of unlabeled vision data exists.

[0459] More specifically, the computing system (1000) may request initial labeling (e.g., passing / failing judgment and indicating defect locations) from the user by randomly sampling a small number (e.g., 20 sheets) of the entire unlabeled dataset to perform iterative sampling and scoring-based active learning.

[0460] When the initial feedback from the user is completed, the computing system (1000) can use this as a prompt to calculate the defect (NG) score of the remaining entire data, sort the images in order of highest defect probability, and request the user to re-label them.

[0461] If the user no longer finds defective images by repeating this process several times, the computing system (1000) can maximize labeling efficiency by automatically processing all remaining large amounts of data as good (OK) in batches.

[0462] Additionally, the computing system (1000) can cross-analyze with the vision data if the user additionally provides a list of images that were found to be defective in another inspection process (e.g., functional inspection) during the process of cross-verifying functional inspection data and handling invisible defects.

[0463] At this time, the computing system (1000) can induce the data to be classified separately by presenting a logical basis that for invisible defects that cannot be visually distinguished as a result of analysis, "they should be treated as OK in vision inspection and detected through other inspections."

[0464] In addition, the computing system (1000) can clearly establish vision inspection standards by sorting images in order of similarity for fine defects that are ambiguous to distinguish with the naked eye and by supporting the user to set an allowable threshold directly.

[0465] In another embodiment, the computing system (1000) can perform immediate response and autonomous updates through interactive agent-based MLOps even if a change in the physical environment occurs after the model is deployed, when performing continuous monitoring and autonomous updates in response to environmental changes.

[0466] In detail, the computing system (1000) can immediately detect, through real-time monitoring, a situation in which a normal product is incorrectly detected as a defect due to a surge in the false detection rate caused by a change in the physical environment, such as a change in lighting angle due to equipment inspection during actual mass production, in the process of real-time anomaly detection and prompt redefinition.

[0467] And the computing system (1000) can display an image of the corresponding time period to the user to report the situation.

[0468] In response to this, when the autonomous relearning pipeline is activated, if the user determines that this is a problem caused by an environmental change rather than a product defect, the computing system (1000) can immediately initiate a workflow to redefine the prompt based on an image of the changed environment (e.g., the day's production, etc.).

[0469] Thus, the computing system (1000) can autonomously perform a series of processes to automatically update the model to a perfectly adapted model by going through a rapid labeling and active learning process through sample images, and immediately reapply it to the production line on the same day without delay.

[0470] Furthermore, a computing system (1000) according to one embodiment of the present disclosure can perform a multi-agent knowledge sharing mechanism that autonomously propagates newly learned knowledge (e.g., task prompts redefined in response to new defect patterns, updated Gaussian mixture model (GMM) parameters, and / or optimized hardware setting values, etc.) in a specific target environment (e.g., a second production line or another factory) that performs a similar process through the autonomous update process described above.

[0471] Specifically, the computing system (1000) (or the agent of the computing system (1000)) can register a knowledge package (e.g., a prompt embedding sequence) for the new defect in the central model registry or knowledge base when a new defect type is detected in the first production line and autonomous relearning and model update are completed.

[0472] Subsequently, the computing system (1000) can analyze the registered knowledge package and / or the edge devices of a second production line that inspect identical or similar products by linking with a higher meta-agent and / or an orchestration engine within the system to preemptively roll out the corresponding updates (e.g., adding a prompt support set for new defect patterns) through a background network.

[0473] Through this knowledge sharing configuration, the computing system (1000) according to the embodiment can immediately transfer learning the ability to respond to unexpected issues or new defects that occur on one production line to multiple lines at the factory level or at the enterprise level.

[0474] Therefore, the computing system (1000) can exponentially accelerate the evolution speed of the enterprise-wide quality control network by preventing the repetition of trial and error or learning costs of a specific line in other lines.

[0475] As described above, according to various embodiments of the present disclosure, the computing system (1000) can automate the entire process of planning, learning, evaluating, and distributing advanced vision inspection models in a one-stop manner using only everyday natural language conversation, even for field workers who do not have coding knowledge or deep learning expertise.

[0476] In particular, the computing system (1000) according to the embodiment can proactively prevent model compatibility issues or risks of on-site performance degradation caused by the discrepancy between the laboratory environment and the actual manufacturing site before deployment by actively collecting target-specific information from external resources, including physical constraints, communication standards, and / or unique quality standards of the target environment where the vision inspection model will finally be operated, and by identifying and supplementing missing data essential for model construction based on this information.

[0477] In addition, the computing system (1000) according to the embodiment can objectively and mathematically guarantee an optimized balance between production yield and defect-free quality by replacing the user's subjective and ambiguous requirements and the physical constraints of the target process with perfect quantitative parameters through multi-agent-based simulation.

[0478] Furthermore, the computing system (1000) according to the embodiment can permanently maintain the inference robustness and reliability of the inspection model even amidst unpredictable changes in the manufacturing environment by operating an intelligent MLOps pipeline that includes real-time data drift detection and autonomous relearning even after deployment.

[0479] Thus, the computing system (1000) according to the embodiment can successfully implement an advanced autonomous intelligent smart factory ecosystem that not only drastically reduces the time and cost required to build a vision inspection pipeline in a target environment, but also fundamentally reduces business risks and quality costs caused by the leakage of defective products, and minimizes manual intervention by workers while simultaneously maximizing productivity and quality control levels.

[0480] Meanwhile, the embodiments according to the present disclosure described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present disclosure or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the present disclosure, and vice versa.

[0481] The specific embodiments described in this disclosure are examples and do not limit the scope of this disclosure in any way. For the sake of brevity of the specification, descriptions of prior electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Additionally, the connections of lines or connecting members between components shown in the drawings are illustrative of functional connections and / or physical or circuit connections, and may be replaced or additionally represented as various functional connections, physical connections, or circuit connections in actual devices. Furthermore, unless specifically stated as “essential,” “importantly,” etc., a component may not be strictly necessary for the application of this disclosure.

[0482] Furthermore, although the detailed description of the present disclosure has been explained with reference to preferred embodiments of the present disclosure, those skilled in the art or those with ordinary knowledge in the art will understand that the present disclosure can be modified and changed in various ways without departing from the spirit and technical scope of the present disclosure as set forth in the claims below. Accordingly, the technical scope of the present disclosure should not be limited to the contents described in the detailed description of the specification but should be determined by the claims.

[0483] The present disclosure relates to a vision inspection control system based on an interactive artificial intelligence agent and a method of operation thereof, and since it is applicable to the artificial intelligence industry, it has industrial applicability.

Claims

1. In a method executed by a computer, At least one processor of the above computer receives at least one input data through at least one interface - wherein the input data is characterized by including at least one natural language input indicating a task purpose; The above at least one processor detects at least one qualitative requirement embedded in the natural language input; The above-mentioned at least one processor generates a plurality of agents having mutually conflicting objective functions based on the detected qualitative requirements—wherein, the plurality of agents are characterized by forming a Multi-Agent Collaboration Structure that interprets the qualitative requirements in a multidimensional manner by being assigned at least one of different personas or optimization goals; The step of the above at least one processor executing a simulation based on the generated plurality of agents; The step of the above at least one processor determining at least one optimal parameter value to substitute the qualitative requirement based on the executed simulation result; The above-mentioned at least one processor takes the received input data as an input (ingest) of at least one artificial intelligence model and reflects the determined optimal parameter value as at least one control variable to generate a task plan required for building at least one target model; wherein the artificial intelligence model is characterized by including an architecture that interprets the intent of the natural language input based on at least one language model and dynamically calls at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model to orchestrate a workflow; The above-mentioned at least one processor constructs the above-mentioned at least one target model based on the above-mentioned generated work plan—wherein the target model is characterized as being an artificial intelligence-based task execution model optimized to perform a task corresponding to the work objective by being trained through a function tool called according to the above-mentioned work plan; and A method comprising the step of the at least one processor applying the constructed target model to at least one target environment, wherein the target environment comprises at least one of at least one infrastructure or edge device in which the constructed target model is deployed to perform at least one target task.

2. In Paragraph 1, The step of detecting the above qualitative requirements is, A step of parsing the natural language input in semantic units by utilizing at least one of a natural language processing (NLP) algorithm or a pre-configured system prompt, and A step of identifying the remaining phrases among the text phrases of the parsed natural language input, excluding the phrases corresponding to the hyperparameters of the target model, and A method comprising the step of extracting and labeling the identified remaining phrases as the qualitative requirements.

3. In Paragraph 1, The step of executing the above simulation is, A method comprising the step of executing a parallel simulation based on the plurality of agents to simultaneously perform at least one of a prediction or classification operation on the same validation dataset based on at least one of a multi-thread or distributed processing environment.

4. In Paragraph 1, The step of executing the above simulation is, A step in which each of the plurality of agents interprets the qualitative requirements based on at least one of the different personas or optimization goals and sets at least one of different thresholds or hyperparameters, and A method comprising the step of each of the plurality of agents applying at least one of the set threshold or hyperparameter to at least one validation dataset to perform at least one of prediction or classification operations.

5. In Paragraph 1, The step of determining the optimal parameter value above is, A step of combining the simulation results above to derive a quantitative analysis result of the trade-off regarding performance evaluation indicators among the plurality of agents, and A method comprising the step of determining the optimal parameter value to be applied to the target model based on the results of the quantitative analysis of the trade-off derived above.

6. In Paragraph 5, The step of determining the optimal parameter value above is, The step of displaying the above trade-off quantitative analysis results on the interface, and A step of receiving user selection data for selecting an optimal balance point in response to the above-mentioned quantitative analysis result of the trade-off, and A method comprising the step of determining the optimal parameter value for a parameter corresponding to the received user selection data.

7. In Paragraph 5, The step of determining the optimal parameter value above is, A step of comparing the above trade-off quantitative analysis result with a pre-set optimal parameter value judgment criterion, and A method comprising the step of determining the optimal parameter value based on the result of the above comparison operation.

8. In Paragraph 1, The above at least one processor maps the determined optimal parameter value to at least one user identification information or at least one of the qualitative requirements and stores it in at least one memory; and A method further comprising the step of applying the optimal parameter value based on the data stored in the memory when the above-mentioned at least one processor receives natural language input identical to the above-mentioned qualitative requirement through the above-mentioned interface.

9. In Paragraph 1, The step of generating the above work plan is, The step of injecting the above-determined optimal parameter value as a control variable for constructing the above-determined target model, and A method comprising the step of specifying the workflow of the work plan by updating the setting value for at least one function tool based on the injected control variable.

10. In Paragraph 1, The above at least one processor obtains target-specific information from at least one external storage system, the data including at least one of constraints or quality standard data for the target environment; and The above-mentioned at least one processor cross-validates the acquired target-specific information and the natural language input to identify at least one missing data, which is missing information required for building the target model but is absent; and A method further comprising the step of at least one processor updating the work plan by supplementing the identified missing data.

11. In Paragraph 10, The step of updating the above work plan is, The step of generating an interactive clarification question to supplement the identified missing data and outputting it through the interface, and A method comprising the step of updating the work plan based on a user response received in response to the outputted interactive reverse query.

12. In Paragraph 11, The step of updating the above work plan is, A method comprising the step of extracting at least one parameter, setting value, or instruction embedded in the above user response and mapping it to the area of ​​the identified missing data.

13. In Paragraph 1, The step of building the above target model is, A step of designing the architecture of the target model by utilizing at least one of an automated model architecture exploration technique or a pre-trained base model based on control variables reflected in the above work plan, and A method comprising the step of performing learning on the designed architecture using at least one functional tool.

14. In Paragraph 10, The step of applying the above target model to the above target environment is, A step of converting the above-established target model into a format optimized for the target environment based on the above-mentioned target-specific information, and A method comprising the step of distributing the converted target model to the target environment through at least one communication network.

15. In Paragraph 1, The step of applying the above target model to the above target environment is, The step of collecting real-time log data from the above target environment, and A method comprising the step of updating the target model by dynamically updating the work plan based on the collected real-time log data.

16. In Paragraph 1, The above-mentioned at least one processor calculates at least one performance evaluation metric for the target model using at least one verification dataset; and A method further comprising the step of the above-mentioned at least one processor manifesting the above-mentioned performance evaluation indicator through the interface.

17. At least one processor; and It includes at least one memory that stores at least one instruction that performs the following when executed by the above at least one processor; and The above at least one instruction is, The above-mentioned at least one processor receives at least one input data through at least one interface—wherein the input data is characterized by including at least one natural language input indicating a task purpose; The above at least one processor detects at least one qualitative requirement embedded in the natural language input; The above-mentioned at least one processor generates a plurality of agents having mutually conflicting objective functions based on the detected qualitative requirements—wherein, the plurality of agents are characterized by forming a Multi-Agent Collaboration Structure that interprets the qualitative requirements in a multidimensional manner by being assigned at least one of different personas or optimization goals; The step of the above at least one processor executing a simulation based on the generated plurality of agents; The step of the above at least one processor determining at least one optimal parameter value to substitute the qualitative requirement based on the executed simulation result; The above-mentioned at least one processor takes the received input data as an input (ingest) of at least one artificial intelligence model and reflects the determined optimal parameter value as at least one control variable to generate a task plan required for building at least one target model; wherein the artificial intelligence model is characterized by including an architecture that interprets the intent of the natural language input based on at least one language model and dynamically calls at least one functional tool for data preparation, training, evaluation, and distribution required for building the target model to orchestrate a workflow; The above-mentioned at least one processor constructs the above-mentioned at least one target model based on the above-mentioned generated work plan—wherein the target model is characterized as being an artificial intelligence-based task execution model optimized to perform a task corresponding to the work objective by being trained through a function tool called according to the above-mentioned work plan; and A system comprising instructions for performing the step of applying the constructed target model to at least one target environment, wherein the target environment comprises at least one of an infrastructure or edge device in which the constructed target model is deployed to perform at least one target task.

18. In Paragraph 17, A plurality of neurons comprising an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; A system further comprising a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.

19. In Paragraph 17, A plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; comprising A system further comprising an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synaptic circuits.

20. In Paragraph 17, A plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; comprising A system further comprising a neuromorphic circuit for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synaptic circuits.