Parseai system and method using intelligent agent crew
The ParseAI system addresses the challenge of converting unstructured data into standardized text by using an intelligent agent crew, enhancing data processing efficiency and accuracy through models like Whisper and GPT-4o.
Patent Information
- Application Number
- PCT/KR2025/008188
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-13
- Publication Date
- 2025-12-26
AI Technical Summary
Existing systems struggle to effectively convert unstructured input data, including voice and image data, into standardized text data for efficient processing and integration with corporate systems.
A ParseAI system utilizing an intelligent agent crew comprising a central receiving agent, voice agent, image agent, text agent, and integrated agent, leveraging models like Whisper, Tesseract, GPT-4o, and Python libraries to convert and process voice and image data into standardized text data.
Improves system operation efficiency by generating accurate, standardized data that can be seamlessly integrated into corporate systems, maintaining context and adapting to new data formats.
Smart Images

Figure KR2025008188_26122025_PF_FP_ABST
Abstract
Description
PARSEAI system and method using intelligent agent crew
[0001] The present invention relates to a ParseAI system and method utilizing an intelligent agent crew, and more particularly, to a system and method utilizing an intelligent agent crew to drive ParseAI, an advanced tool for parsing unstructured data.
[0002]
[0003] Parsing is a step in the computer compiler or translator's process of translating raw code into machine language. It involves analyzing the grammatical structure or syntax of each sentence. In other words, it accepts a sequence of tokens from the source program and constructs a parse tree that conforms to the grammar of the language. Parsing can be broadly divided into top-down and bottom-up parsing.
[0004] Data parsing refers to the process by which a computer program analyzes data to extract or process necessary information. It primarily involves extracting desired information from text or files or converting their format for use.
[0005]
[0006] The technical problem to be solved by the present invention is to provide a ParseAI system and method using an intelligent agent crew that receives unstructured input data and generates standardized data.
[0007] Another technical problem to be solved by the present invention is to provide a ParseAI system and method using an intelligent agent crew that extracts text data from voice data.
[0008] Another technical problem to be solved by the present invention is to provide a ParseAI system and method using an intelligent agent crew that extracts text data from image data.
[0009] The technical problems to be solved by the present invention are not limited to those described above.
[0010]
[0011] According to one embodiment of the present invention, a ParseAI system using an intelligent agent crew may include a server including a central receiving agent that receives unstructured input data; a voice agent that converts voice data included in the unstructured input data into first text data; an image agent that converts image data included in the unstructured input data into second text data; a text agent that performs natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data; and an integrated agent that receives the unstructured input data and the third text data to generate standardized data, and a database that stores a voice-to-text conversion model for driving the voice agent, an image-to-text conversion model for driving the image agent, a natural language processing model for driving the text agent, and a standardization model for driving the integrated agent.
[0012] As one embodiment, the voice agent may convert the voice data into the first text data using at least one of the Whisper model, the Google Cloud Speech-to-Text API, and the Python library, and may improve accuracy by using a confidence score and repeated verification when converting the first text data.
[0013] As one embodiment, the image agent may convert the image data into the second text data using at least one of Tesseract, Google Cloud Vision API, and a Python library, and may perform error correction and accuracy verification using multiple conversion paths during the conversion of the second text data.
[0014] As an embodiment, the text agent may generate the third text data using at least one of a Natural Language Toolkit, a spaCy model, and a regular expression, and may perform a feedback loop and heuristic check when generating the third text data.
[0015] As an example, the integrated agent may generate the standardized data using at least one of GPT-4o and Python libraries, and may verify the standardized data using a reinforcement learning technique when generating the standardized data, and perform continuous improvement of the results.
[0016] According to one embodiment of the present invention, a method of ParseAI using an intelligent agent crew using a server including one or more processors; and one or more memories storing instructions that cause the one or more processors to perform operations when executed by the one or more processors may include: receiving, by the one or more processors, unstructured input data; converting, by the one or more processors, voice data included in the unstructured input data into first text data; converting, by the one or more processors, image data included in the unstructured input data into second text data; performing, by the one or more processors, natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data; and generating, by the one or more processors, standardized data by receiving the unstructured input data and the third text data.
[0017] According to an embodiment of the present invention, a computer program stored in a computer-readable recording medium that can be executed by a computer including one or more processors; and one or more memories storing instructions that cause the one or more processors to perform operations when executed by the one or more processors, may be stored in a computer-readable recording medium so as to be capable of performing the steps of: receiving unstructured input data by the one or more processors; converting voice data included in the unstructured input data into first text data by the one or more processors; converting image data included in the unstructured input data into second text data by the one or more processors; performing natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data; and generating standardized data by receiving the unstructured input data and the third text data by the one or more processors.
[0018]
[0019] The ParseAI system and method using an intelligent agent crew according to one embodiment of the present invention can improve system operation efficiency by generating standardized data from unstructured input data.
[0020] Additionally, the ParseAI system and method using an intelligent agent crew according to one embodiment of the present invention can improve the accuracy of data processing by maintaining context within a single session, improve future data parsing tasks by learning from past interactions, and adapt to new data formats and requirements.
[0021]
[0022] FIG. 1 is an exemplary diagram showing the configuration of a ParseAI system using an intelligent agent crew according to an embodiment of the present invention.
[0023] Figure 2 is an exemplary diagram showing the configuration of a server according to an embodiment of the present invention.
[0024] Figure 3 is an exemplary diagram showing the configuration of a server according to an embodiment of the present invention.
[0025] Figure 4 is a diagram for explaining learning of a neural network according to an embodiment of the present invention.
[0026] FIG. 5 is a flowchart showing the procedure of the ParseAI method using an intelligent agent crew according to an embodiment of the present invention.
[0027]
[0028] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, the embodiments may be modified in various ways, and the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, or alternatives to the embodiments are included within the scope of the patent application.
[0029] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, the embodiments are not limited to the specific disclosed form, and the scope of this specification includes modifications, equivalents, or alternatives that fall within the technical concept.
[0030] Although terms such as "first" or "second" may be used to describe various components, these terms should be interpreted solely to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0031] When it is said that a component is "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.
[0032] The terms used in the examples are for illustrative purposes only and should not be construed as limiting. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood to not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0033] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the embodiments pertain. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0034] In addition, when describing with reference to the attached drawings, identical components will be assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted. When describing embodiments, if a detailed description of a related known technology is judged to unnecessarily obscure the gist of the embodiment, the detailed description will be omitted.
[0035] The embodiments may be implemented in various forms of products, such as personal computers, laptop computers, tablet computers, smart phones, televisions, smart home appliances, intelligent vehicles, kiosks, and wearable devices.
[0036] FIG. 1 is an exemplary diagram showing the configuration of a ParseAI system using an intelligent agent crew according to an embodiment of the present invention.
[0037] As illustrated in FIG. 1, the ParseAI system (100) using an intelligent agent crew may include a plurality of user terminals (110-1,…,110-n), a server (120), and a database (130). According to one embodiment, the database (130) is illustrated as being configured separately from the server (120), but is not limited thereto, and the database (130) may be provided within the server (120). For example, the server (120) may include a plurality of artificial intelligence models for performing a machine learning algorithm. According to one embodiment, the plurality of user terminals (110-1,…,110-n), the server (120), and the database (130) may be connected to each other so as to be able to communicate with each other via a network (N).
[0038] Figure 2 is an exemplary diagram showing the configuration of a server according to an embodiment of the present invention.
[0039] As illustrated in FIG. 2, the server (120) may include a central receiving agent (121), a voice agent (122), an image agent (123), a text agent (124), and an integrated agent (125). According to one embodiment, the central receiving agent (121), the voice agent (122), the image agent (123), the text agent (124), and the integrated agent (125) may be connected to each other so as to be able to communicate with each other via a system bus.
[0040] The central receiving agent (121) can receive unstructured input data. According to one embodiment, the central receiving agent (121) can receive all unstructured data inputs, that is, unstructured input data, and transmit them to each specialized agent (voice agent (122), image agent (123), and text agent (124)) and the integrated agent (125). For example, a speaker included in a user terminal (e.g., 110-1) can receive a user's voice input to form voice data, and the central receiving agent (121) can receive the voice data from the user terminal (110-1). In addition, a camera included in a user terminal (e.g., 110-n) can receive a photo taken by the user to form image data, and the central receiving agent (121) can receive the image data from the user terminal (110-n).
[0041] The voice agent (122) can convert voice data included in the unstructured input data into first text data. According to one embodiment, the voice agent (122) can convert the voice data included in the unstructured input data into first text data using at least one of the Whisper model, the Google Cloud Speech-to-Text API (application programming interface), and the Python library. In addition, the voice agent (122) can improve the accuracy of data conversion by using a confidence score and repeated verification when converting the first text data.
[0042] The Whisper model is OpenAI's speech-to-text model that can be used to convert audio data from audio files into text. The Whisper model is trained on a large-scale English audio and text dataset. The Whisper model is optimized for transcribing the content of audio files containing English speech into text data. It can also be used to transcribe audio files containing speech in other languages. The output of the Whisper model is English text.
[0043] Google Cloud Speech-to-Text API is Google's speech-to-text model that can be used to convert speech data into text data.
[0044] A Python library is a collection of modules and packages available for use in the Python programming language. Libraries typically provide various functions that aid in writing Python code. These functions can be easily performed by calling the library, without having to write code to perform specific tasks. The voice agent (122) can convert speech data into primary text data using the openai and google-cloud-speech Python libraries.
[0045] The image agent (123) can convert image data included in the unstructured input data into second text data using at least one of Tesseract, Google Cloud Vision API, and Python library. In addition, the image agent (123) can perform error correction and accuracy verification using multiple conversion paths when converting the second text data.
[0046] OCR stands for Optical Character Recognition, a technology that recognizes characters using light. It allows a computer to read and copy text from printed materials, document images, and other general images captured by a camera. The image agent (123) can convert image data into secondary text data using various OCR engines.
[0047] Tesseract is an optical character recognition engine for various operating systems. It is free software distributed under the Apache License, Version 2.0, and its development has been sponsored by Google since 2006.
[0048] Google Cloud Vision API provides advanced OCR and image analysis capabilities.
[0049] The image agent (123) can convert image data into second text data using the pytesseract and google-cloud-vision Python libraries.
[0050] The text agent (124) may perform natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data. According to one embodiment, the text agent (124) may generate the third text data using at least one of the Natural Language Toolkit, the spaCy model, and a regular expression. In addition, the text agent (124) may perform a feedback loop and heuristic check when generating the third text data to ensure consistency and accuracy.
[0051] The Natural Language Toolkit is a collection of libraries and programs for symbolic and statistical natural language processing in English, written in the Python programming language. It supports classification, tokenization, morphological analysis, tagging, parsing, and semantic inference.
[0052] Spacey Model is an open-source software library for advanced natural language processing written in the programming languages Python and Cython.
[0053] Regular expressions are a formal language used to represent sets of strings with specific rules. Many text editors and programming languages support regular expressions for string searching and substitution, and Perl and Tcl, in particular, have powerful regular expression implementations built into their languages. Text agents (124) can utilize regular expressions to organize and standardize data.
[0054] The integrated agent (125) can generate standardized data using the unstructured input data received from the central receiving agent (121) and the third text data received from the text agent (124). The integrated agent (125) can generate standardized data by integrating and standardizing the third text data received from the text agent (124). The integrated agent (125) can directly process multi-modal input. According to one embodiment, the integrated agent (125) can generate standardized data using at least one of GPT-4o and a Python library. The integrated agent (125) can verify the standardized data using a reinforcement learning technique when generating the standardized data and perform continuous improvement of the results.
[0055] GPT-4o is OpenAI's multimodal model capable of seamlessly processing text, speech, and image input. GPT-4o can integrate and standardize primary and secondary text data to generate standardized data.
[0056] The integrated agent (125) can generate standardized data using the openai Python library.
[0057] The integrated agent (125) can verify the combined output and promote continuous improvement through repeated refinement and comparison with expected results using reinforcement learning techniques.
[0058] In a data parsing method according to an embodiment of the present invention, a central receiving agent (121) receives unstructured data input, that is, unstructured input data (voice, image, text, etc.), and transmits the received unstructured input data to an appropriate specialized agent (voice agent (122), image agent (123), text agent (124)) and an integrated agent (125).
[0059] Each specialized agent that receives unstructured input data performs preprocessing of the data, and then the voice agent (122) converts the voice data into first text data, the image agent (123) converts the image data into second text data, and the text agent (125) generates third text data. Each specialized agent transmits the first to third text data to the integrated agent (125).
[0060] The integrated agent (125) can generate standardized data using unstructured input data and the first to third texts. In one embodiment, the integrated agent (125) can integrate data using the multi-modal capabilities of GPT-4o and generate standardized data in a standardized format. Furthermore, the integrated agent (125) can verify that the standardized data is accurate and ready for use in external systems.
[0061] Standardized data generated by the server (120) can be seamlessly transferred to external tools, such as a Customer Relationship Management (CRM) system, a database (130), or reporting software, allowing corporate systems to access up-to-date, structured data, thereby enhancing system operational efficiency. Furthermore, the server (120) can maintain context within a single session to enhance the accuracy of data processing, learn from past interactions to improve future data parsing tasks, and adapt to new data formats and requirements.
[0062] A data parsing scenario according to an embodiment of the present invention can receive voice logs of captains from multiple user terminals (110-1,…,110-n) used by captains using a ParseAI system (100) utilizing an intelligent agent crew on a shipping company server and convert them into a structured report. A voice agent (122) can transcribe the logs, and an integration agent (125) can combine the transcribed data with other data inputs to generate a final structured report and integrate it into the CRM system of the shipping company server.
[0063] A data parsing scenario according to an embodiment of the present invention can process receipt photos and handwritten notes using a ParseAI system (100) utilizing an intelligent agent crew on an accounting firm server. An image agent (123) generates second text data from the receipt photo, a text agent (124) standardizes the second text data, and an integration agent (125) integrates the standardized second text data to create structured expense items and seamlessly integrate with the accounting firm server financial system.
[0064] The network (N) can perform wireless or wired communication among a plurality of user terminals (110-1,…,110-n), a server (120), a database (130), etc. For example, the network (N) can perform wireless communication according to a method such as LTE (long-term evolution), LTE-A (LTE Advanced), CDMA (code division multiple access), WCDMA (wideband CDMA), WiBro (Wireless BroadBand), WiFi (wireless fidelity), Bluetooth (Bluetooth), NFC (near field communication), GPS (Global Positioning System), or GNSS (global navigation satellite system). For example, the network (N) can also perform wired communication according to a method such as USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), or POTS (plain old telephone service).
[0065] The database (130) can store various data. The data stored in the database (130) is data acquired, processed, or used by at least one component of a plurality of user terminals (110-1,…,110-n) and the server (120), and may include software (e.g., a program). The database (130) may include volatile and / or non-volatile memory. As an example, the database (130) may store unstructured input data, first to third text data, standardized data, a voice-to-text conversion model for driving a voice agent (122), an image-to-text conversion model for driving an image agent (123), a natural language processing model for driving a text agent (124), and a standardized model for driving an integrated agent (125).
[0066] In the present invention, artificial intelligence (AI) refers to a technology that mimics human learning, reasoning, and perception abilities and implements them on a computer. It may include concepts such as machine learning and symbolic logic. Machine learning (ML) is an algorithmic technology that classifies or learns the characteristics of input data on its own. AI technology analyzes input data using machine learning algorithms, learns the results of that analysis, and makes judgments or predictions based on the results of that learning. Furthermore, technologies that utilize machine learning algorithms to mimic human brain functions such as cognition and judgment can also be understood as falling under the category of AI. For example, this may include technical fields such as linguistic understanding, visual understanding, inference / prediction, knowledge representation, and motion control.
[0067] Machine learning can refer to the process of training a neural network model using data processing experience. Through machine learning, computer software can improve its data processing capabilities. Neural network models are built by modeling correlations between data, and these correlations can be expressed by multiple parameters. Neural network models extract and analyze features from given data to derive correlations between data. This process of iteratively optimizing the parameters of a neural network model can be defined as machine learning. For example, a neural network model can learn the mapping (correlation) between inputs and outputs for data presented as input-output pairs. Alternatively, a neural network model can learn the relationships between inputs and outputs by deriving regularities between the given data, even when presented with only input data.
[0068] An artificial intelligence learning model or neural network model can be designed to implement the structure of the human brain on a computer, and can include multiple network nodes that simulate the neurons of a human neural network and have weights. The multiple network nodes can have connections with each other by simulating the synaptic activity of neurons that exchange signals through synapses. In the artificial intelligence learning model, the multiple network nodes can be located at layers of different depths and exchange data according to convolutional connections. The artificial intelligence learning model can be, for example, an artificial neural network model, a convolutional neural network (CNN), etc. In one embodiment, the artificial intelligence learning model can be machine-learned according to a method such as supervised learning, unsupervised learning, or reinforcement learning. Machine learning algorithms that can be used to perform machine learning include decision trees, Bayesian networks, support vector machines, artificial neural networks, Ada-boost, perceptrons, genetic programming, and clustering.
[0069] CNNs are a type of multilayer perceptron designed to utilize minimal preprocessing. They consist of one or more convolutional layers stacked on top of regular artificial neural network layers, with additional weight and pooling layers. This structure allows CNNs to fully utilize two-dimensional input data. Compared to other deep learning architectures, CNNs demonstrate excellent performance in both image and audio domains. CNNs can also be trained using standard backpropagation. Compared to other feedforward artificial neural network techniques, CNNs are easier to train and have fewer parameters.
[0070] Convolutional networks are neural networks that contain sets of nodes with bounded parameters. The increasing availability of training data and computational power, combined with advances in algorithms such as piecewise linear units and dropout training, have led to significant improvements in many computer vision tasks. With the massive datasets available for many tasks today, overfitting is less of a concern, and increasing network size improves test accuracy. Optimal utilization of computing resources becomes a limiting factor. To address this, distributed, scalable implementations of deep neural networks can be utilized.
[0071] Figure 3 is an exemplary diagram showing the configuration of a server according to an embodiment of the present invention.
[0072] As illustrated in FIG. 3, the server (120) may include one or more processors (126), one or more memories (127), and a transceiver (128). In one embodiment, at least one of these components of the server (120) may be omitted, or another component may be added to the server (120). Additionally or alternatively, some of the components may be implemented in an integrated manner, or may be implemented as a single or multiple entities. At least some of the components inside and outside the server (120) may be connected to each other via a system bus, a general purpose input / output (GPIO), a serial peripheral interface (SPI), or a mobile industry processor interface (MIPI), and may exchange data and / or signals.
[0073] One or more processors (126) may control at least one component of a server (120) connected to the processor (126) by executing software (e.g., commands, programs, etc.). In addition, the processor (126) may perform various operations related to the present invention, such as calculations, processing, data generation, and processing. In addition, the processor (126) may load data, etc. from one or more memories (127), or store data, etc. in one or more memories (127).
[0074] One or more processors (126) may receive non-standard input data. According to one embodiment, the processor (126) may receive non-standard input data from a plurality of user terminals (110-1,…,110-n) via a transceiver (128).
[0075] One or more processors (126) may convert voice data included in the unstructured input data into first text data. According to one embodiment, the processor (126) may convert the voice data included in the unstructured input data into first text data using at least one of the Whisper model, the Google Cloud Speech-to-Text application programming interface (API), and the Python library. In addition, the processor (126) may improve the accuracy of data conversion by using a confidence score and repeated verification when converting the first text data.
[0076] One or more processors (126) may convert image data included in the unstructured input data into second text data. According to one embodiment, the processor (126) may convert the image data included in the unstructured input data into second text data using at least one of Tesseract, Google Cloud Vision API, and Python library. In addition, the processor (126) may verify the accuracy by using error correction and multiple conversion paths when converting the second text data.
[0077] One or more processors (126) may perform natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data. According to one embodiment, the processor (126) may generate the third text data using at least one of the Natural Language Toolkit, the spaCy model, and a regular expression. In addition, the processor (126) may perform a feedback loop and heuristic check when generating the third text data to ensure consistency and accuracy.
[0078] One or more processors (126) can generate standardized data using unstructured input data received from the central receiving agent (121) and third text data received from the text agent (124). According to one embodiment, the processor (126) can integrate and standardize the third text data received from the text agent (124) to generate standardized data. The processor (126) can directly process multi-modal input. For example, the processor (126) can generate standardized data using at least one of GPT-4o and a Python library. The processor (126) can use reinforcement learning techniques to verify the results when generating the standardized data and perform continuous improvement of the results.
[0079] One or more memories (127) can store the above-described unstructured input data, first to third text data, standardized data, a voice-to-text conversion model for driving a voice agent (122), an image-to-text conversion model for driving an image agent (123), a natural language processing model for driving a text agent (124), and a standardized model for driving an integrated agent (125).
[0080] Additionally, one or more memories (127) may store instructions that, when executed by one or more processors (126), cause one or more processors (126) to perform operations.
[0081] According to one embodiment, the server (120) may further include a transceiver (128). The transceiver (128) may perform wireless or wired communication between the server (120) and various external servers (e.g., a Customer Relation Management (CRM) system, reporting software, a shipping company server, an accounting company server, etc.), databases, client devices, and / or other devices. For example, the transceiver (128) may perform wireless communication according to a method such as eMBB (enhanced Mobile Broadband), URLLC (Ultra Reliable Low-Latency Communications), MMTC (Massive Machine Type Communications), LTE (long-term evolution), LTE-A (LTE Advance), UMTS (Universal Mobile Telecommunications System), GSM (Global System for Mobile communications), CDMA (code division multiple access), WCDMA (wideband CDMA), WiBro (Wireless Broadband), WiFi (wireless fidelity), Bluetooth (Bluetooth), NFC (near field communication), GPS (Global Positioning System), or GNSS (global navigation satellite system). For example, the transceiver (128) may also perform wired communication according to a method such as USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), or POTS (plain old telephone service).
[0082] According to one embodiment, one or more processors (126) may control a transceiver (128) to obtain information from various external servers and databases (130). The information obtained from various external servers and databases (130) may be stored in one or more memories (127).
[0083] According to one embodiment, the server (120) may be a device of various forms. For example, the server (120) may be a portable communication device, a computer device, or a device according to a combination of one or more of the devices described above. The server (120) of the present invention is not limited to the devices described above.
[0084] The various embodiments of the server (120) according to the present invention can be combined with each other. Each embodiment can be combined according to a number of cases, and the combined server (120) embodiments also fall within the scope of the present invention. Furthermore, the internal / external components of the server (120) according to the present invention described above can be added, changed, replaced, or deleted depending on the embodiment. Furthermore, the internal / external components of the server (120) described above can be implemented as hardware components.
[0085] Figure 4 is a diagram for explaining learning of a neural network according to an embodiment of the present invention.
[0086] As illustrated in FIG. 4, the learning device can train a neural network (129) to extract text data from unstructured input data. In one embodiment, the learning device may be a separate entity from the server (120), but is not limited thereto.
[0087] The neural network (129) includes an input layer (129-1) into which training samples are input and an output layer (129-2) that outputs training outputs, and can be trained based on the differences between the training outputs and labels. Here, the labels can be defined based on text data corresponding to unstructured input data. The neural network (129) is connected as a group of multiple nodes and is defined by weights between the connected nodes and an activation function that activates the nodes.
[0088] The learning device can train a neural network (129) using the GD (Gradient Descent) technique or the SGD (Stochastic Gradient Descent) technique. The learning device can use a loss function designed based on the outputs and labels of the neural network.
[0089] The learning device can calculate a training error using a predefined loss function. The loss function can be predefined as input variables, including labels, outputs, and parameters, where the parameters can be set by weights within the neural network (129). For example, the loss function can be designed in the form of a Mean Square Error (MSE) or entropy, and various techniques or methods can be employed in the design of the loss function.
[0090] The learning device can use the backpropagation technique to identify weights that influence training errors. Here, the weights represent relationships between nodes within the neural network (129). The learning device can utilize the SGD technique, which utilizes labels and outputs, to optimize the weights identified through the backpropagation technique. For example, the learning device can update the weights of a loss function defined based on the labels, outputs, and weights using the SGD technique.
[0091] According to one embodiment, the learning device can acquire unstructured input data of a training target and extract text data of the training target. The learning device can acquire pre-labeled information (first labels) for each of the training unstructured input data, and can acquire first labels representing predefined text data for the training unstructured input data.
[0092] According to one embodiment, the learning device can generate first training feature vectors based on array features, sequence features, and pattern features of the training unstructured input data. Various methods can be employed to extract features of the training unstructured input data.
[0093] According to one embodiment, the learning device can obtain training outputs by applying the first training feature vectors to the neural network (129). The learning device can train the neural network (129) based on the training outputs and the first labels. The learning device can calculate training errors corresponding to the training outputs and train the neural network (129) by optimizing the connection relationship between nodes in the neural network (129) to minimize the training errors. The server (120) can form first to third text data from unstructured input data using the neural network (129) for which training has been completed.
[0094] FIG. 5 is a flowchart illustrating the procedure of the ParseAI method using an intelligent agent crew according to an embodiment of the present invention. Although the process steps, method steps, and algorithms are described in a sequential order in the flowchart of FIG. 5 , such processes, methods, and algorithms may be configured to operate in any suitable order. In other words, the steps of the processes, methods, and algorithms described in various embodiments of the present invention need not be performed in the order described herein. Furthermore, even if some steps are described as being performed asynchronously, in other embodiments, such some steps may be performed concurrently. Furthermore, the illustration of a process by depiction in the drawings does not imply that the illustrated process excludes other variations and modifications thereof, nor does it imply that the illustrated process or any of its steps is essential to one or more of the various embodiments of the present invention, nor does it imply that the illustrated process is preferred.
[0095] As illustrated in FIG. 5, in step S510, non-standard input data is received. For example, referring to FIGS. 1 to 4, the central receiving agent (121) of the server (120) can receive non-standard input data. According to one embodiment, a user's voice input is received through a speaker included in one of a plurality of user terminals (110-1, ..., 110-n) to form voice data, and the central receiving agent (121) can receive the voice data from the user terminal (110-1) through the network (N). In addition, a camera included in the user terminal (e.g., 110-n) can receive a photo taken by the user to form image data, and the central receiving agent (121) can receive the image data from the user terminal (110-n) through the network (N).
[0096] In step (S520), voice data is converted into first text data. For example, referring to FIGS. 1 to 4, the voice agent (122) of the server (120) can convert voice data included in unstructured input data into first text data. According to one embodiment, the voice agent (122) can convert voice data included in unstructured input data into first text data using at least one of the Whisper model, the Google Cloud Speech-to-Text API (application programming interface), and the Python library. In addition, the voice agent (122) can improve the accuracy of data conversion by using a confidence score and repeated verification when converting the first text data.
[0097] In step (S530), image data is converted into second text data. For example, referring to FIGS. 1 to 4, the image agent (123) of the server (120) can convert image data included in unstructured input data into second text data using at least one of Tesseract, Google Cloud Vision API, and Python library. In addition, the image agent (123) can verify accuracy by using error correction and multiple conversion paths when converting the second text data.
[0098] In step (S540), third text data is generated. For example, referring to FIGS. 1 to 4, the text agent (124) of the server (120) may perform natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate the third text data. According to one embodiment, the text agent (124) may generate the third text data using at least one of the Natural Language Toolkit, the spaCy model, and the regular expression. In addition, the text agent (124) may perform a feedback loop and heuristic check when generating the third text data to ensure consistency and accuracy.
[0099] In step (S550), standardized data is generated. For example, referring to FIGS. 1 to 4, the integrated agent (125) of the server (120) can generate standardized data using unstructured input data received from the central receiving agent (121) and third text data received from the text agent (124). The integrated agent (125) can generate standardized data by integrating and standardizing the third text data received from the text agent (124). The integrated agent (125) can directly process multi-modal input. According to one embodiment, the integrated agent (125) can generate standardized data using at least one of GPT-4o and a Python library. The integrated agent (125) can verify the results using a reinforcement learning technique when generating standardized data and perform continuous improvement of the results.
[0100] While the method has been described through specific embodiments, the method can also be implemented as computer-readable code on a computer-readable recording medium. A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system. Examples of computer-readable recording media include ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. Furthermore, the computer-readable recording medium can be distributed across network-connected computer systems, such that the computer-readable code can be stored and executed in a distributed manner. In addition, functional programs, codes, and code segments for implementing the above embodiments can be readily inferred by programmers skilled in the art to which the present invention pertains.
[0101] While the present invention has been described in detail using preferred embodiments, the scope of the present invention is not limited to the specific embodiments described above, and should be interpreted in accordance with the appended claims. Furthermore, those skilled in the art will appreciate that numerous modifications and variations are possible without departing from the scope of the present invention.
Claims
1. A central receiving agent that receives unstructured input data; A voice agent that converts voice data included in the above unstructured input data into first text data; An image agent that converts image data included in the above unstructured input data into second text data; A text agent that performs natural language processing on at least one of the unstructured input data, the first text data, and the second text data to generate third text data; and A server including an integrated agent that receives the above-mentioned unstructured input data and the above-mentioned third text data and generates standardized data; A database that stores a voice-to-text conversion model for driving the voice agent, an image-to-text conversion model for driving the image agent, a natural language processing model for driving the text agent, and a standardization model for driving the integrated agent. ParseAI system using intelligent agent crew.
2. In paragraph 1, The above voice agent, Converting the speech data into the first text data using at least one of the Whisper model, Google Cloud Speech-to-Text API, and Python library, In the above first text data conversion, accuracy is improved by using confidence scores and repeated verification. ParseAI system using intelligent agent crew.
3. In paragraph 1, The above image agent, Converting the image data into the second text data using at least one of Tesseract, Google Cloud Vision API, and Python library, When converting the second text data, error correction and accuracy verification are performed using multiple conversion paths. ParseAI system using intelligent agent crew.
4. In paragraph 1, The above text agent, Generating the third text data using at least one of the Natural Language Toolkit, the spaCy model, and the regular expression, Performing a feedback loop and heuristic check when generating the third text data, ParseAI system using intelligent agent crew.
5. In paragraph 1, The above integrated agent, Generate the above standardized data using at least one of the GPT-4o, Python libraries, When generating the above standardized data, the reinforcement learning technique is used to verify the above standardized data and to continuously improve the results. ParseAI system using intelligent agent crew.
Citation Information
Patent Citations
Method and system for data gathering, processing andpresentation using computer network
KR1020020061443A
A manufacturing method of Paste containing Ginkgo nuts and Bellflower Root Extract
KR1020210051724A
Apparatus and method for recording and managing counseling contents of counselors and clients
KR102578093B1
SYSTEM AND METHOD FOR ParseAI BY USING INTELLIGENT AGENT CREW
KR102733055B1
Straw for drinking purifier water
KR102882023B1