Ai-based warehouse inventory inquiry and stock inbound / outbound system using voice recognition
Patent Information
- Application Number
- KR1020260101212
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-09
- Estimated Expiration
- 2046-06-04
Smart Images

Figure 112026067674990-PAT00003_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a warehouse inventory inquiry and inbound / outbound system using AI-based voice recognition, and provides a system capable of inquiring about items in a warehouse and processing inbound and outbound shipments using voice recognition. Background Technology
[0002] With the advancement of generative AI and voice interface technologies, the demand for voice-based work support systems is increasing in industrial settings. Voice-based warehouse management systems are designed primarily for large logistics warehouses, resulting in high implementation costs and heavy reliance on cloud services and external networks. Furthermore, they face challenges in adequately reflecting non-standard terminology, pronunciation variations, and differences in the spelling of loanwords used in Korean industrial environments. Additionally, in edge device environments with limited GPU memory, the simultaneous operation of Speech-to-Text (STT), Text-to-Speech (TTS), and Large Language Models (LLM) can lead to memory conflicts and Out-of-Memory (OOM) errors, making stable real-time operation difficult. Consequently, there is a growing need for resource allocation structures and voice-based inventory management technologies that are suitable for actual industrial operating environments and can stably operate multiple AI models simultaneously within resource-constrained conditions.
[0003] At this time, methods for performing general business tasks using a voice recognition-based interface or providing inventory information using voice recognition have been researched and developed. In this regard, prior art Korean Published Patent No. 2023-0040074 (published March 22, 2023) and Korean Published Patent No. 2023-0015137 (published January 31, 2023) disclose, respectively, a configuration for integrated management of production management, sales management, purchasing and sales management, and inventory management data in a smart factory environment, enabling workers to easily perform general business tasks through a voice recognition-based interface, and querying and managing inventory and production information using voice commands; and a configuration for managing restaurant order, inventory, settlement, and delivery person location information using voice recognition, wherein when a store owner inputs a command by voice, the server performs a restaurant management function corresponding to the command and provides order status, inventory information, and settlement information based on a database.
[0004] However, in the case of the former, only configurations for voice recognition-based manufacturing and sales management are disclosed, and configurations for handling non-standard terminology, pronunciation variations, and differences in foreign word spellings found in industrial sites are not disclosed. Similarly, in the case of the latter, only configurations for cafeteria management and inventory inquiry using voice commands are disclosed, and configurations for voice-based inventory inquiry and guidance in limited edge device environments are not disclosed. In consumable rooms and material management environments at industrial sites, there is a problem where workers refer to the same item by different aliases or pronunciations, and multiple AI models must be operated simultaneously in a limited edge device environment. Therefore, research and development of an offline-based voice inventory management system is required that can reliably recognize non-standard field terminology while stably operating voice recognition technology in an environment with limited GPU resources. The problem to be solved
[0005] One embodiment of the present invention provides a warehouse inventory inquiry and inbound / outbound system using AI-based speech recognition, which enables inventory inquiry and inbound / outbound processing solely through the voice input of a worker in an industrial site's consumable room and material management environment. This system reliably recognizes non-standard naming conventions and pronunciation variations in industrial sites by utilizing custom Korean call word detection, a multi-stage STT correction structure, root-suffix separation matching, a category-based alias dictionary, and standard pattern extraction technology. Furthermore, to reliably operate STT, TTS, and LLM simultaneously in a limited edge device environment, it applies a heterogeneous memory allocation separation operation structure that separates model subcomponents into CPU and GPU allocations. By operating on a completely offline basis, it can provide a stable voice-based inventory management service without an external network. However, the technical problem that this embodiment aims to solve is not limited to the technical problem described above, and other technical problems may exist. means of solving the problem
[0006] As a technical means for achieving the aforementioned technical task, one embodiment of the present invention comprises: a voice input device for receiving voice utterances; a receiving unit for receiving voice utterances received from the voice input device; a voice conversion unit for converting voice utterances into text using STT (Speech to Text); an intent classification unit for classifying the intent of voice utterances in text; an extraction unit for extracting item names and specification information within text; a search unit for searching inventory data using item names and specification information; a response generation unit for generating a response based on the searched inventory data; a response conversion unit for converting the generated response into voice using TTS (Text to Speech) and outputting it; and a voice output device for outputting the voice converted by the response conversion unit. Effects of the invention
[0007] According to any one of the means for solving the problem of the present invention described above, non-standard naming conventions, pronunciation variations, and differences in foreign word spelling used in industrial settings can be reliably recognized through a multi-stage STT correction structure and alias dictionary-based matching, thereby enabling accurate inventory inquiry and inbound / outbound processing regardless of the operator's proficiency. Furthermore, even in a limited integrated memory environment, the stability of simultaneous operation of AI models can be ensured while preventing Out Of Memory (OOM) issues through a resource allocation structure that separates STT, TTS, and LLM into CPU and GPU operations. Additionally, since it operates on an offline basis, dependency on external networks and security risks can be reduced, and maintenance and field expansion by non-developers are possible through the structure of the correction dictionary and alias dictionary based on external files. Brief explanation of the drawing
[0008] FIG. 1 is a diagram illustrating a warehouse inventory inquiry and inbound / outbound system using AI-based voice recognition according to an embodiment of the present invention. FIG. 2 is a block diagram illustrating an integrated processing unit included in the system of FIG. 1. FIGS. 3 and 4 are drawings for explaining an embodiment in which a warehouse management solution using voice recognition according to an embodiment of the present invention is implemented. FIG. 5 is an operation flowchart illustrating a method for providing a warehouse management solution using voice recognition according to an embodiment of the present invention. Specific details for implementing the invention
[0009] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0010] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected" but also cases where they are "electrically connected" with other elements interposed between them. Furthermore, when a part is described as "including" a component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components, and it should be understood that this does not preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0011] Terms such as “about,” “substantially,” etc., used throughout the specification, are used to mean at or near the stated value when inherent manufacturing and material tolerances are presented in the stated meaning, and are used to prevent unscrupulous infringers from unfairly exploiting the disclosure in which precise or absolute values are mentioned to aid in understanding the invention. Terms such as “step” or “step of” used throughout the specification of the invention do not mean “step for”.
[0012] In this specification, the term "part" includes a unit realized by hardware, a unit realized by software, and a unit realized using both. Additionally, one unit may be realized using two or more pieces of hardware, and two or more units may be realized by one piece of hardware. Meanwhile, "part" is not limited to software or hardware, and "part" may be configured to reside in an addressable storage medium or configured to run on one or more processors. Accordingly, as an example, "part" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and "parts" may be combined into a smaller number of components and "parts" or further separated into additional components and "parts." In addition, the components and '~parts' may be implemented to play one or more CPUs within the device or secure multimedia card.
[0013] Some of the operations or functions described herein as being performed by a terminal, device, or device may instead be performed by a server connected to said terminal, device, or device. Likewise, some of the operations or functions described as being performed by a server may also be performed by a terminal, device, or device connected to said server.
[0014] In this specification, some of the operations or functions described as mapping or matching with a terminal may be interpreted as meaning mapping or matching the terminal's unique number or personal identification information, which is the terminal's identifying data.
[0015] The present invention will be described in detail below with reference to the attached drawings.
[0016] FIG. 1 is a drawing illustrating a warehouse inventory inquiry and inbound / outbound system using AI-based voice recognition according to an embodiment of the present invention. Referring to FIG. 1, the warehouse inventory inquiry and inbound / outbound system (1) using AI-based voice recognition may include at least one voice input device (100), an integrated processing device (200), a voice output device (300), and a display (400). However, since the warehouse inventory inquiry and inbound / outbound system (1) using AI-based voice recognition of FIG. 1 is merely an embodiment of the present invention, the present invention is not to be interpreted as being limited by FIG. 1.
[0017] In the following, the term "at least one" is defined as a term including both singular and plural forms, and it will be obvious that even if the term "at least one" does not exist, each component may exist in a singular or plural form and may mean singular or plural. Furthermore, whether each component is provided in a singular or plural form may be changed according to the embodiment.
[0018] The voice input device (100) can be configured to receive voice from an industrial site using a USB-based dynamic microphone to collect voice utterances from a worker. The USB dynamic microphone can be, for example, a FIFINE K669B microphone, and by utilizing the directional characteristics of the dynamic microphone, the worker's voice and background noise can be separated even in an industrial site noise environment at a level of about 56 dB. In addition, in the call word detection and voice recording stages, only the actual voice utterance segment can be extracted using voice level analysis techniques based on VAD (Voice Activity Detection), Noise Gate, and RMS (Root Mean Square).
[0019] The integrated processing unit (200) is an Edge AI (Edge Artificial Intelligence) based computing unit that performs integrated speech recognition, natural language processing, and speech synthesis, and can be implemented using a single-board computer such as the NVIDIA Jetson Orin Nano Super. The integrated processing unit (200) uses an integrated memory structure in which the CPU and GPU share the same memory pool, and can be configured to operate Faster-Whisper-based STT, MeloTTS-based TTS, and Ollama-based LLM. In addition, by using a memory management structure based on CUDA (Compute Unified Device Architecture), CTranslate2, PyTorch, Ollama, Python, and Linux, STT or TTS-speech synthesis is processed on the GPU, and LLM and TTS-BERT components are allocated separately to the CPU, thereby preventing memory conflicts and OOM (Out Of Memory) problems.
[0020] The voice output device (300) can be implemented as a USB speaker to output the response generated by the integrated processing unit (200) to the operator. The voice output device (300) can play voice data generated through a MeloTTS-based speech synthesis engine through an ALSA (Advanced Linux Sound Architecture) or PulseAudio-based audio output interface.
[0021] A display (400) can be implemented using a DisplayPort-based monitor to output a web dashboard for administrators. This display (400) can visually display real-time inventory status, incoming and outgoing history, item information, and administrator settings screens using a Flask, SocketIO, and Chromium-based web interface.
[0022] A storage device (not shown) may be implemented using an M.2 NVMe SSD (Solid State Drive) to store the operating system, AI models, and inventory data. The storage device stores the Whisper model, MeloTTS model, proofreading dictionary, alias dictionary, and inventory data, and may be configured to read and modify Excel files and configuration files using Pandas, openpyxl, and JSON-based data processing technologies.
[0023] FIG. 2 is a block diagram for explaining an integrated processing unit included in the system of FIG. 1, and FIG. 3 and FIG. 4 are drawings for explaining an embodiment in which a warehouse management solution using voice recognition according to an embodiment of the present invention is implemented.
[0024] Referring to FIG. 2, the integrated processing unit (200) may include a receiving unit (210), a voice conversion unit (220), an intent classification unit (230), an extraction unit (240), a search unit (250), a response generation unit (260), a response conversion unit (270), a wake-up unit (280), an alias substitution unit (290), a separation assignment unit (291), and a specification automatic extraction unit (293).
[0025] The receiving unit (210) can receive voice utterances received from the voice input device (100). The voice input device (100) can receive voice utterances. In one embodiment of the present invention, the system switches to a mode for receiving voice utterances by receiving a call word before receiving voice utterances, which will be described later in the wake-up unit (280).
[0026] The voice conversion unit (220) can convert voice utterance into text using STT (Speech to Text). In one embodiment of the present invention, a speech recognition engine is used to recognize voice utterance, and this may use a Faster-Whisper-based Korean fine-tuning model (whisper-small-komixv2), but is not limited thereto. At this time, the STT operates on a GPU and has a structure that corrects misrecognitions recognized in industrial settings in multiple stages.
[0027] Multi-stage correction process
[0028] The voice conversion unit (220) can convert the voice utterance into text after undergoing a multi-stage correction process when converting the voice utterance into text using STT. The multi-stage correction process may include a direct correction step that corrects terms within the text using a pre-established correction dictionary, and a separation matching step in which, if a term within the text is not included in the correction dictionary, the term is separated into a root and a suffix, the similarity with the root of a term within the correction dictionary is calculated by comparing only the root, and the term within the text is converted into the term of the root within the correction dictionary that has the highest similarity.
[0029] Direct Correction Stage
[0030] Frequent misrecognitions can be directly replaced using a predefined correction dictionary.
[0031] STT output text-prep After correction reason Pusha Pusher Pronunciation variation Sillimdo cylinder Pronunciation errors sick OPP Abbreviation pronunciation Tactron Teflon Pronunciation errors
[0032] In this case, the correction dictionary is configured to be stored separately as an external JSON file rather than being hardcoded within the code, allowing non-developers to directly add or modify STT misrecognition patterns that occur during operation. For example, by registering correction information such as [pusha → pusher] and [sillimdo → cylinder] in the correction_dict.json file, pronunciation variations and repetitive speech recognition errors found in industrial settings can be continuously corrected during operation. Since this structure allows for the management of correction data without modifying the source code, it enables the protection of core code while enhancing operational independence and maintenance efficiency.
[0033] File path: ~ / l2_config / correction_dict.json Format: {"STT output": "After correction", ...}
[0034] Of course, the file format is not limited to JSON. For example, an Excel sheet is also possible.
[0035] Additionally, the correction dictionary of the present invention can automatically construct a correction dictionary using spoken terms and finally selected terms, without requiring a person to input misrecognition results one by one. This adaptive correction dictionary learning structure is configured to analyze the mapping relationship between the term spoken by the user and the term (item name) finally selected or confirmed, and to automatically add or update the corresponding correction information to the correction dictionary when the same correction result occurs repeatedly. For example, if an operator repeatedly speaks "[Tacfron]" and finally selects "[Teflon Tape]," the relationship between "[Tacfron]" and "[Teflon]" is learned, allowing for automatic correction to be performed upon subsequent input of the same speech. Through this, pronunciation habits, abbreviations, slang, and non-standard names specific to industrial sites can be continuously learned during operation, thereby improving STT recognition accuracy and item matching accuracy over time.
[0036] This adaptive correction dictionary training process can be configured to collect STT result texts and the final selected text, analyze the frequency of occurrence of identical correction patterns, and register them as automatic correction candidates if they are repeated more than a preset threshold. Subsequently, the process can be configured to reflect these changes in the final correction dictionary through cross-validation among multiple workers or administrator approval procedures. To this end, a Python-based log analysis module, Pandas-based frequency analysis, a JSON-based correction dictionary storage structure, a SQLite or PostgreSQL-based training history storage module, a Levenshtein Distance-based similarity calculation algorithm, and rule-based validation logic may be utilized. Additionally, Scikit-learn-based pattern analysis models or lightweight LLM-based semantic similarity analysis models may be used as needed. Of course, the available tools or models are not limited to these.
[0037] <Separation Matching Stage>
[0038] Separation matching according to one embodiment of the present invention means that [roots] and [suffixes] are separated and then matched. That is, it is a method of accurately matching items by separating an item name into a material root and a material suffix, comparing them based on the similarity of the material root portion—which is the core information distinguishing the actual item—and simultaneously verifying whether the suffix matches. In industrial settings, suffixes such as tape, bolt, nut, and spray are used repeatedly, and the core information distinguishing the actual item is often concentrated in the material root portion. However, conventional simple string similarity comparison methods compare the entire item name; consequently, if STT misrecognition occurs, the similarity is calculated as low, leading to a problem where even legitimate items fail to be matched. In particular, since similarity with other items having the same suffix is calculated similarly, lowering the similarity threshold increases the possibility of mismatching.
[0039] Accordingly, one embodiment of the present invention can first separate a pre-set suffix from an item name, perform a similarity calculation based on Hangul consonant-vowel decomposition targeting only the root part, and separately verify whether the suffix matches. Subsequently, by combining the root similarity and the suffix match to calculate a final matching score, accurate item matching can be made possible even when STT misrecognition occurs. For example, in the case of [Tacfron Tape] and [Teflon Tape], the similarity based on consonant-vowel decomposition of the roots [Tacfron] and [Teflon] is high, and since the suffix [Tape] is identical, a high final matching score is calculated, allowing for normal matching. However, items such as [Insulation Tape] or [Box Tape], which have the same suffix but different roots, are judged to have low similarity and are accurately rejected.
[0040] process explanation tool or model 1. Receive item name input Receive STT result or user input item name faster-whisper, Python 2. Loading Suffix Dictionary Loading a list of pre-configured suffixes such as tape, bolt, nut, spray, etc. JSON, Python Dictionary, Pandas 3. Analysis of Item Name Forms Detects the presence of a suffix in the input item name Regex (Regular Expression), String Parsing 4. Separation of Root and Suffix Separating the item name into a material root and a material suffix Python String Processing, Regex 5. Decomposition of Hangul Consonants and Vowels Breaking down the root part into initial, medial, and final consonant units jamo library, KoNLPy, Python 6. Calculation of Root Similarity Similarity is calculated based on decomposed characters. Levenshtein Distance, Fuzzy Matching, RapidFuzz 7. Suffix Match Verification Verify whether the suffixes of the input item and the target item are identical. Rule-based Matching 8. Calculation of Final Matching Score The final score is calculated by combining root similarity and suffix agreement. Weighted Scoring Algorithm, Python 9. Filtering Candidate Items Remove candidates below the threshold and select the optimal item Threshold Filtering, Ranking Logic 10. Final Item Matching The item name with the highest score is confirmed as the final matching result. Python Matching Engine
[0041] Here, the similarity calculation based on Hangul consonant and vowel decomposition takes into account the characteristics of Korean STT misrecognition that occur at the consonant and vowel level, such as changes in final consonants or vowels. Specifically, the input word and the reference word can be configured to calculate similarity by decomposing them into initial, medial, and final consonant units and then performing string distance calculation based on Levenshtein Distance. For example, in the case of [Tactron] and [Teflon], a low similarity is calculated when comparing the original strings, but a high similarity is calculated when comparing after consonant and vowel decomposition, making more accurate STT misrecognition correction possible. If no match is found even after this process, the item name is extracted based on the alias dictionary in the alias replacement unit (290). This will be described later in the alias replacement unit (290).
[0042] In this case, in addition to the method described above, a multiple sum ensemble structure can be used for similarity. The process [Input speech recognition result → Calculate phonological similarity → Calculate word boundary similarity → Calculate syllable N-Gram Jaccard similarity → Calculate weighted sum integrated score → Compare phonological similarity and weighted sum score → Select maximum value → Perform final item matching] can be used. The higher value between the weighted sum of the three measures (phonological similarity, word boundary similarity, and syllable N-Gram Jaccard similarity) and the phonological similarity can be used as the final similarity score. This means that if the pronunciation is sufficiently similar, other measures are ignored and the pronunciation score is used as is.
[0043] Additionally, operator-specific corrections may be performed. For instance, an operator's speech patterns, pronunciation habits, and repetitive STT correction results can be stored as a user profile. Subsequently, a personalized correction structure can be configured to be applied preferentially when processing the same operator's speech. For instance, if a specific operator repeatedly pronounces [Teflon] as [Takfron], the system can be configured to learn that operator's speech pattern and automatically perform corrections thereafter. To achieve this, Speaker Embedding-based speaker identification, Voiceprint-based user differentiation, Personalized Correction Dictionary-based user correction dictionaries, and Python-based learning log analysis technologies may be utilized, but are not limited to these.
[0044] The intent classification unit (230) can classify the intent of a voice utterance in text. Intent classification is a process of analyzing the text of a voice recognition result to determine whether the voice utterance has the purpose of checking inventory, receiving, or sending out. This intent classification can be performed by analyzing the core keywords and context included in inquiry-type expressions such as [Is there stock?], [How many are left?], receiving expressions such as [Please receive], and sending expressions such as [Sending], [Take it out]. To this end, keyword rule analysis based on Rule-based NLP (Natural Language Processing), morphological analysis based on KoNLPy, an intent classification model based on Scikit-learn, and context analysis technology based on LLM (Large Language Model) as needed can be used, and processing can be performed to accurately determine the worker's actual work intent by reflecting multiple expression methods and speech patterns specialized for industrial sites.
[0045] The extraction unit (240) can extract item names and specification information within the text. The extraction unit (240) is configured to automatically extract item names and specification information by analyzing text converted through speech recognition (STT). The extraction unit (240) first performs morphological analysis and string pattern analysis to detect item name candidates within the text, and can extract the official item name when it is derived using a stored item database and an alias dictionary to be described later. Additionally, as will be described later in the specification automatic extraction unit (293), if specification information such as mm, inch, M specification, liter, and size (S / M / L / XL) is extracted by analyzing specification expressions combining numbers and units based on a regular expression (Regex), it can be used. To perform the above-described process, Python-based string processing techniques, Regex pattern matching, KoNLPy-based morphological analysis, Pandas-based item data retrieval, JSON or Excel-based alias dictionaries, and Levenshtein Distance-based similarity analysis techniques may be utilized, and if necessary, Scikit-learn-based Named Entity Recognition (NER) models or LLM-based natural language analysis techniques may be additionally applied.
[0046] The search unit (250) can search for inventory data using item names and specification information. The search unit (250) can search for matching inventory information by querying the inventory database based on the extracted item names and specification information. Inventory data search can be performed using Pandas-based DataFrame Filtering, an Excel-based inventory data management structure, or SQLite-based lightweight database search technology, and can be configured to perform conditional search based on item names, specifications, quantity, and location information. Additionally, the results of alias substitution and specification extraction can be reflected together to accurately distinguish and search for multiple specifications within the same item. After performing the search, the results are branched, and the inventory data search results can be analyzed to determine whether the search result is a single item, multiple items, or a matching failure state. For example, if there is one search result, immediate inquiry or inbound / outbound processing is performed; if multiple results exist, additional specification information is requested from the user; and if no matching result exists, a re-call or administrator verification request can be performed. To this end, Python-based Branch Logic, Rule Engine, and conditional branch processing algorithms may be used, and may be configured to dynamically control the subsequent conversation flow according to the state of the search result.
[0047] The response generation unit (260) can generate a response based on the searched inventory data. The response generation unit (260) is configured to generate a response sentence to be provided to the user based on the searched inventory data. Response generation can be performed by generating a structured response such as "[There are 12 100mm cable ties in stock]" using a Python Template Engine-based string template method when the search result is a single item, and by generating a natural language response using an Ollama-based gemma3:1b LLM (Large Language Model) when multiple candidate guidance, re-question generation, or complex explanation is required. Additionally, it is configured to maintain a response format, item name expression, and work instruction style suitable for an industrial site inventory management environment using Prompt Engineering technology, and can dynamically generate inquiry responses, receiving completion guidance, outbound processing results, and re-question sentences depending on the search result status.
[0048] The response conversion unit (270) can convert the generated response into speech using TTS (Text to Speech) and output it. The speech output device (300) can output the speech converted by the response conversion unit (270). Before converting the generated response into speech using TTS, the response conversion unit (270) can convert the response into speech after undergoing a multi-stage speech synthesis text preprocessing process. At this time, the multi-stage speech synthesis text preprocessing process may include at least one step among the following: a step of converting combinations of English letters and numbers in the response into pre-set combination expressions; a step of converting special characters in the response into pre-set special character expressions or removing them; a step of converting fractions in the response into pre-set fraction expressions; and a step of converting units in the response into Hangul. These four steps are not necessarily all performed, and if each expression exists, the corresponding step may be performed sequentially. Accordingly, it may be structured as a [sequential pipeline + conditional execution structure].
[0049] In addition to this, preprocessing can be performed to facilitate speech synthesis for various expressions through steps such as those shown in Table 4 below.
[0050] step definition explanation Conversion example tool or model English letter and number combination conversion step The step of converting industrial standard expressions combining English letters and numbers into Korean speech forms Converts M specifications, model names, and combination specification expressions used in industrial settings into speech expressions suitable for worker speech patterns. [M8]→[M8], [M5x10]→[M5 multiplied by 10] Python String Processing, Regex (regular expression), Rule-based Text Normalization, TTS Preprocessor Special character conversion / removal step Step of replacing or removing special characters in the response with pronunciation-readable expressions Converts difficult-to-read special characters into natural Korean expressions or removes unnecessary symbols during speech synthesis [×]→[Multiply], [±]→[Plus / Minus], [@]→Remove Regex, Unicode Processing, Python Replace Logic Fractional representation conversion step Step of converting number-based fraction expressions into Korean fraction utterance forms Converts fractions used in industrial standards and inch expressions into natural Korean reading. [1 / 2]→[1 / 2][3 / 8 inch]→[3 / 8 inch] Fraction Parser, Regex, Rule-based NLP Unit Hangul conversion step Step of converting English unit expressions into Korean pronunciation forms Converts industrial units such as mm, kg, and ml into Korean pronunciation that is easy for workers to understand. [mm]→[millimeter][EA]→[piece] Unit Dictionary, Python Mapping Table, Text Normalization Industrial language pronunciation correction step Step of converting industrial site abbreviations and material names into Korean pronunciation Converted English abbreviation-based industrial terms to fit on-site speech patterns. [PVC]→[PVC][OPP]→[OPP] Domain Dictionary, Custom Pronunciation Mapping Naturalization stage of positional representation Step of converting position and sequence expressions into natural Korean speech forms Converts numeric-based location representations into TTS-friendly natural language representations [3rd cell]→[3rd cell] Rule-based NLP, Number-to-Korean Converter Quantity Expression Conversion Step Step of converting number-based quantity expressions into Korean quantity utterances Converted stock quantity to natural Korean reading style
[100] →
[1000] →
[1000] Number Normalization, Korean Number Converter
[0051] Input: "There are 50 M5x10 bolts in the 3rd slot of Shelf A in Supplies Room 1." Output: "There are 50 M5x10 bolts in the 3rd slot of Shelf A in Supplies Room 1." Input: "There are 3 1 / 2-inch torque wrenches on Shelf B in Supplies Room 2." Output: "There are 3 1 / 2-inch torque wrenches on Shelf B in Supplies Room 2."
[0052] The aforementioned multi-stage text preprocessing process for speech synthesis is a preprocessing technique designed to ensure the naturalness of the TTS output and the accurate pronunciation of industrial terminology. Since response sentences for inventory inquiries and inbound / outbound operations in industrial settings contain a mixture of English letters, numbers, special characters, standard expressions, and units, synthesizing them directly can result in unnatural speech or incorrect reading. Accordingly, a multi-stage preprocessing structure is applied prior to speech synthesis to convert the text response generated by the LLM into a format suitable for the Korean pronunciation system.
[0053] This preprocessing is first configured to convert standard expressions combining English letters and numbers to match the speech patterns used in industrial settings. For example, [M8] is converted to [Empal], [M10] to [Emsip], and [M5x10] to [Emo multiply ten] or [Emo esip]. Subsequently, special characters are replaced with phonetically viable forms or unnecessary symbols are removed, converting [×] to [multiply], [±] to [plus minus], and [~] to [from]. Additionally, fractional expressions such as [1 / 2] and [3 / 8] are converted into Korean fractional pronunciations like [one-half] and [three-eighths], while unit expressions such as [mm], [kg], [ml], and [EA] can be converted into Hangul pronunciations like [millimeter], [kilogram], [milliliter], and [gae], respectively. Furthermore, the system can be configured to perform pronunciation correction for abbreviations and material names frequently used in industrial settings. For example, [OPP] is converted to [OPP], [PVC] to [PVC], and [PE] to [PE-E], and positional and quantity expressions can also be corrected into natural Korean speech forms. For example, [the 3rd compartment of shelf A] can be converted to [the third compartment of shelf A], and [100 items] to [one hundred items]. Through this preprocessing, the naturalness of the speech synthesis results and the operator's listening comprehension can be improved, and speech response quality suitable for industrial environments can be provided.
[0054] In addition to processing for natural pronunciation, further configurations can be added to not only respond to user queries but also perform inventory forecasting and provide the results. Inventory forecasting can be configured to predict the likelihood of future stock shortages by analyzing past inbound and outbound history and usage patterns by item. For example, based on the recent increase in usage of a specific bolt, a warning such as "[Stock shortage expected in 3 days]" can be automatically generated. To this end, time-series data analysis, LSTM (Long Short-Term Memory)-based demand forecasting models, Prophet-based consumption pattern analysis, and moving average-based inventory depletion prediction algorithms may be used, but are not limited to these.
[0055] At this stage, since data is insufficient initially, a phased roadmap may be implemented as follows. ① In the inventory depletion prediction based on daily average and moving average, item-specific inbound and outbound history is collected to calculate daily average and moving average usage, and the process can be executed to calculate the estimated depletion time by comparing it with the current inventory level. The calculated results are visualized via a local web UI as estimated depletion dates, shortage warnings, and usage trend information for each item, and can be configured to operate independently of the existing voice pipeline. To this end, Pandas, NumPy, the Moving Average algorithm, SQLite, Flask, or web dashboard technologies based on FastAPI and Chart.js may be used.
[0056] ② In the consumption trend visualization, daily inflow and outflow data by item are aggregated to display usage change trends in a graph. Users can check usage by period, cumulative consumption, and inventory fluctuation status through the web UI, and the system is designed to allow for the intuitive identification of changes in consumption patterns, such as abnormal increases or sharp decreases. To achieve this, technologies such as Pandas-based data aggregation, SQL query processing, time-series graphs based on Plotly or Chart.js, and HTML5-based web dashboards may be utilized. ③ In the advancement of time-series-based inventory forecasting, the system is configured to predict future inventory demand by learning past usage patterns once sufficient inflow and outflow data has been accumulated. In the initial stage, forecast results based on daily averages and moving averages are provided; once data has accumulated for a certain period, Long Short-Term Memory (LSTM), Prophet, or time-series forecasting models are applied to perform more sophisticated demand forecasting. For this purpose, TensorFlow, PyTorch, LSTM, Prophet, Scikit-Learn, and time-series data analysis technologies may be utilized.
[0057] The Wakeup unit (280) can receive a pre-set Wake Word before receiving voice utterance from the receiver (210) and begin preparation for voice utterance. This Korean Wake Word detection technology is a technology that constantly listens to microphone input in an Always-On manner and switches to a command recognition mode when the user utters a pre-registered Korean Wake Word. In one embodiment of the present invention, [L2YA] is used as the Wake Word, and the system may be configured to build a dedicated learning pipeline for Korean Wake Words to overcome the limitation that the existing OpenWakeWord library only supports English-based Wake Words by default.
[0058] To this end, voice samples of wake words with various speakers, speeds, and intonations can be generated using speech synthesis engines based on MeloTTS and Kokoro TTS, and Deep Neural Network (DNN)-based wake word training can be performed in a CoreWorxLab Docker environment. In addition, MIT Room Impulse Responses (RIRs), AudioSet, and FMA datasets can be utilized as negative samples to train the system to reliably distinguish wake words even in noisy industrial environments, and the final training result can be output as a lightweight model in ONNX format to operate in edge device environments.
[0059] In addition, one embodiment of the present invention may apply a learning data adaptation methodology in consideration of the phenomenon where parts of the utterance are removed or distorted even for the same wake word due to the application of a noise gate at the firmware level in a general USB microphone. For example, when [L-tu-ya] is uttered, the [L] part may be removed by the noise gate and recognized as [tu-ya] or [e-tu-ya], and this varies depending on the gate state of the microphone's built-in DSP (Digital Signal Processing). One embodiment of the present invention may be configured to minimize the difference between the domain during learning and the domain during inference by quantitatively analyzing these distortion characteristics through RMS (Root Mean Square) measurement and waveform analysis, and by collecting wake word data using the same microphone under 56dB environmental noise conditions in a consumable room, which is an actual distribution environment.
[0060] In addition, it can be processed to learn various speech patterns by applying speech speed variation (Speed 0.8~1.6). Furthermore, to solve the problem where some commercial wakeword engines require an internet connection for license authentication, one embodiment of the present invention may adopt a completely offline structure using a combination of OpenWakeWord and a custom ONNX model. Through this, Korean wakeword detection is possible without an external network connection, so it can be operated stably even in air-gap environments in industrial sites where security is critical.
[0061] The alias replacement unit (290) may, prior to the extraction unit (240) extracting item name and specification information within the text, query a pre-established alias dictionary if the similarity in the separation matching step is less than a preset threshold, and if the term within the text matches an alias in the alias dictionary, replace the term within the text with an item name that has been mapped and stored with the alias. This may be the final step of the multi-stage correction process described above.
[0062] Alias substitution is a technology designed to automatically replace non-standard names commonly used in industrial settings with official product names. In industrial settings, even for the same product, it is frequently referred to by a wide variety of names depending on individual worker habits, tool usage methods, material names, Japanese loanwords, and abbreviations. Accordingly, one embodiment of the present invention is configured to categorize and manage various types of naming occurrences—such as abbreviated forms, full names, wrench names, pronunciation forms, shape names, manufacturer codes, Japanese expressions, notation forms, STT misrecognition, and half-pronunciation—based on alias data collected from actual industrial sites.
[0063] Category characteristic example Abbreviated form Shorten the long official name Stainless steel bolt → Stainless steel hexagonal bolt Addressing the person A term used to refer to tool head sizes 13mm (=M8 bolt), 17mm (=M10 bolt) Wrench designation Names based on wrench size 5mm wrench (=5mm hex wrench) Pronunciation type Korean pronunciation + number M8, M10 Shape and designation A name given based on shape characteristics Round nut, square washer Manufacturer code Manufacturer Catalog Code PUL-08, PC-04 (Pneumatic Fitting) Japanese Japanese loanwords Vernier calipers Notation Other notation methods M-8 vs M8, 6x20 vs 6x20 STT misrecognition Frequent speech recognition error patterns Pusher, cylinder Half pronunciation Pronounce only part of the word Cable (=cable tie), tape (=box tape)
[0064] The alias dictionary is stored as an Excel file and features a data structure that includes aliases, official item names, specifications, categories, and remarks. For example, non-standard terms that a worker may pronounce are stored in the alias column, along with the corresponding official item name and specification information, while the category information manages the type of alias used. Additionally, operational information such as the registrant, registration date, and memo is stored in the remarks column, allowing for continuous maintenance and expansion during operation.
[0065] The alias substitution process is configured to first receive the STT result text, correct repetitive STT misrecognitions using a correction dictionary, and then attempt primary item matching through root-suffix separation matching. If a matching failure occurs or multiple candidates exist during this process, the alias dictionary is consulted to search for an alias item that matches the input utterance; if a match exists, it is processed to automatically substitute it with the corresponding official item name and specifications. Subsequently, inventory data search is performed using the substituted item name and specification information. Furthermore, unlike conventional technology that uses a simple 1:1 synonym mapping structure, the alias dictionary according to an embodiment of the present invention can categorize and manage the naming generation mechanism itself that occurs in industrial settings. Accordingly, different matching priorities can be applied by category, and new aliases discovered during operation can also be systematically classified and accumulated.
[0066] In particular, while utterances such as "[13mm]" can generally be interpreted as simple length or diameter specifications, in industrial settings they often refer to the "[size of the head side of an M8 hexagonal bolt]." An embodiment of the present invention incorporates this industrial domain-specific knowledge into the "[Side-side Designation]" category, thereby enabling the processing to automatically match an utterance of "[13mm bolt]" to "[M8 hexagonal bolt]."
[0067] Hexagonal bolt (excrement) Wrench bolt (hex wrench) M3 = 5.5mm M3 = 2.5mm M4 = 7mm M4 = 3mm M5 = 8mm M5 = 4mm M6 = 10mm M6 = 5mm M8 = 13mm M8 = 6mm M10 = 17mm M10 = 8mm M12 = 19mm M12 = 10mm M16 = 24mm
[0068] This matching information is registered in the [Large-sized Name] category of the alias dictionary, so when a worker utters natural language-based field terms such as [13mm bolt], it is configured to automatically match them to [M8 hexagonal bolt]. This large-sized name mapping was systematized by the inventor based on the bolt naming system commonly used in Korean industrial sites and the actual terminology used at the Samdasu factory production site, and is characterized by the integration of industrial domain knowledge into a voice interface system.
[0069] User utterance System Analysis [13mm bolt] M8 hex bolt [6mm Wrench Bolt] M8 hex wrench bolt [Hex 13mm] hexagonal bolt series [6mm Wrench] Wrench bolt series
[0070] The separate allocation unit (291) can be configured so that BERT (Bidirectional Encoder Representations from Transformers) within TTS and LLM (Large Language Model) that generates the response are processed on the CPU (Central Processing Unit) rather than the GPU (Graphics Processing Unit). This is because running multiple AI models on limited GPU memory leads to a shortage of memory. For example, if STT, TTS, and LLM all use the GPU, GPU memory usage continues to accumulate, and different CUDA memory allocators may conflict or fail to accurately determine actual usage, eventually resulting in an Out Of Memory (OOM). Therefore, the roles are separated so that real-time performance is important for the GPU, and relatively slower performance is acceptable for the CPU.
[0071] The memory of the present invention is based on an on-device architecture and utilizes 8GB of integrated memory. In order to stably operate speech recognition (STT) and speech synthesis (TTS) models simultaneously in such a limited environment, memory resources must be distributed. To this end, in one embodiment of the present invention, the Faster-Whisper library based on CTranslate2 may be used for STT, and the MeloTTS library based on PyTorch may be used simultaneously for TTS. However, since the two libraries use different CUDA memory allocators, a problem arises where the GPU memory occupied by CTranslate2 is not displayed when a PyTorch-based memory usage check function is called. Consequently, without accurately recognizing the actual GPU memory usage, the memory usage of both sides accumulates, exceeding the physical GPU memory limit. As a result, an Out Of Memory (OOM) error occurs, causing the system to stop functioning. While this problem is not significantly apparent in large-capacity GPU server environments, it acts as a critical issue in edge device environments based on 8GB of integrated memory, such as the Jetson Orin Nano.
[0072] Accordingly, to analyze the cause of this problem, the inventors used the Python gc (Garbage Collection) module to exhaustively examine tensor objects remaining on the GPU and confirmed that approximately 24 768×768 tensors were continuously residing in GPU memory. Subsequently, it was identified that these tensors were weight tensors of the BERT (Bidirectional Encoder Representations from Transformers) model, and it was confirmed that the melo.text.japanese_bert module within MeloTTS maintains the kykim / bert-kor-base model as a GPU global variable. Furthermore, it was analyzed that there is a structural problem in which the GPU memory is not completely released by a standard model.cpu() call alone.
[0073] Accordingly, one embodiment of the present invention resolves memory conflict issues by applying a model subcomponent separation allocation structure. Specifically, ① after the initial warm-up of MeloTTS, the BERT component can be explicitly moved to the CPU to reduce GPU memory occupancy.
[0074] After warming up MeloTTS, access the global variable of the melo.text.japanese_bert module and directly call .cpu(): import melo.text.japanese_bert as jbert jbert.model.cpu()gc.collect()torch.cuda.empty_cache() Effect: GPU Usage 717MB → 269MB (BERT 451MB unlocked)
[0075] ② Apply a source patch that dynamically fits the input tensors inside japanese_bert.py to the model device so that the BERT model can operate normally even when located on the CPU.
[0076] A patch is applied inside get_bert_feature() to dynamically match the input tensor to the model device, ensuring that MeloTTS operates correctly even when BERT is on the CPU. # japanese_bert.py lines 35-38 patch model_device = next(model.parameters()).device inputs[i] = inputs[i].to(model_device)
[0077] In addition, ③ Ollama-based LLM (gemma3:1b) is configured to intensively allocate GPU resources to STT and TTS processing by blocking GPU access and forcing execution to be CPU-only through the setting of the CUDA_VISIBLE_DEVICES environment variable.
[0078] Block CUDA visibility with Ollama system environment variables Environment=CUDA_VISIBLE_DEVICES= Isolate LLM (gemma3:1b) for CPU exclusive use, and save GPU resources for STT / TTS.
[0079] In addition, ④ STT warm-up is performed before TTS initialization to initialize the CUDA context of the STT model first, thereby preventing CUDA context conflicts between Torch and CTranslate2.
[0080] The order in which torch and CTranslate2 share the CUDA context affects memory safety. Avoid context conflicts by performing STT warm-up before MeloTTS initialization. # Order matters stt = WhisperModel('small', device='cuda') stt.transcribe(...) # Warm-up tts = MeloTTS(language='KR', device='cuda') tts.tts_to_file(...) # Warm-up # BERT CPU Separation jbert.model.cpu()
[0081] As a result of applying this structure, the OOM issue was completely resolved, and the stability of simultaneous operation of the [STT(small)+TTS(MeloTTS GPU)+LLM(CPU)] structure was secured. Furthermore, even when BERT was moved to the CPU, there was virtually no degradation in speech synthesis speed; based on actual measurements, the GPU utilization of the Torch series remained at approximately 269MB and the GPU utilization of the CTranslate2 series at approximately 480MB, reducing total GPU memory usage to approximately 749MB. Consequently, more than 3.6GB of free system RAM space was secured, enabling stable real-time voice interface operation even in limited edge device environments.
[0082] GPU CPU STT (Faster-Whisper) TTS-BERT TTS speech synthesis LLM(gemma3:1b)
[0083] The specification automatic extraction unit (293) can automatically extract specification information of the item name within the text using a pre-established specification extraction function when the intent of the text is receiving or shipping. The specification automatic extraction of the present invention is a technology for automatically extracting specification information of an item from text spoken by a user during the receiving and shipping process. In industrial settings, even for the same item name, various specifications exist depending on length, diameter, inches, capacity, and size, so it is often difficult to accurately match inventory using only the item name. Accordingly, it can be configured to enable accurate item matching without additional re-questions by automatically extracting specification information within the speech.
[0084] This automatic specification extraction can be configured to integrate and process various specification expression methods used in industrial settings. For example, length and diameter specifications recognize mm-based expressions such as [100 mm], [5 mm], [50 mm], and [20 millimeters], while screw specifications recognize [M+number] patterns such as [M8], [M10], and [M16]. Additionally, fraction-based inch specifications such as [1 / 2 inch] and [3 / 8 inch], clothing and glove size expressions such as [S], [M], and [XL], and volume expressions such as [1 L] and [500 ml] can also be configured to recognize these as specification patterns.
[0085] To this end, a specification extraction function based on extract_spec_from_text() is implemented, and processing can be performed to detect numbers, units, and specification expressions within an utterance using regular expression (Regex)-based string pattern analysis.
[0086] def extract_spec_from_text(text):# 1).. Number (including decimal) + mm / milli / milli if re.search(r(\d+(?:\.\d+)?)\s*(?:mm|milli|milli)", t): return f"{match.group(1)}mm"# 2). M + Number (bolt spec)if re.search(r"\bM(\d+)\b", t): return f"M{match.group(1)}"# 3). Inch (including fraction)if re.search(r(\d+(?: / \d+)?)\s*inch", t): return f"{match.group(1)}inch"# 4). Single size if re.search(r(?:^|\s)(XS|S|M|L|XL|XXL)(?:\s|$)", t.upper()): return match.group(1)# 5). liter if re.search(r(\d+(?:\.\d+)?)\s*(?:L|liter)", t): return f"{match.group(1)}L"# 6). Diameter×Length(MxN)if re.search(r"\bM(\d+)\s*[xX×*]\s*(\d+)\b", t):return f"M{match.group(1)}x{match.group(2)}"return None
[0087] The extracted specification information is passed to an inbound / outbound processing function based on process_inout(), where it can be used for inventory search along with the item name, quantity, and inbound / outbound intent. Subsequently, if multiple search results exist for the same item, additional filtering is performed using the extracted specification information; if the filtered result is a single item, the system is configured to immediately execute inbound / outbound processing.
[0088] Ignition: "Dispatch 1 pack of 100mm cable ties" Extraction: {"item": "Cable Tie","spec": "100mm","qty": 1,"mode": "out"} Processing: Matching row 1 of 100mm cable tie specifications out of 4 Response: "Dispatch 1 pack of 100mm cable ties completed."
[0089] At this time, the unit is not limited to the aforementioned [bag], and it goes without saying that [piece] is also available. This structure of compatibility and conversion between higher and lower standards is not limited to cable ties but can be commonly applied to items having multiple packaging and management units, such as [box↔piece] for disposable gloves and [bag↔piece] for bolts.
[0090] On the other hand, if specification information is not included in the utterance or if there are multiple filtering results, the process is handled to request the user to input additional specifications.
[0091] Voice: "Dispatching 1 pack of cable ties" Response: "Which size is it: 100mm, 140mm, 270mm, or 370mm?" Voice: "100mm" Response: "Dispatching 1 pack of 100mm cable ties is complete."
[0092] Furthermore, the present invention can comprehensively process various standard utterance patterns actually used in industrial settings and provides a fast standard extraction speed in milliseconds by applying a regular expression-based lightweight processing structure. Accordingly, it can improve user convenience by reducing the frequency of re-questions and can be configured to ensure high recognition accuracy even for unit expressions with a high possibility of confusion, such as [mm] and [ml], by applying explicit processing logic and a unit test-based verification structure.
[0093] In addition, one embodiment of the present invention may further provide a continuous conversation mode and an adaptive noise gate. The continuous conversation mode is a conversation maintenance function that allows subsequent voice commands to be processed continuously for a certain period of time without the user repeating the call word after the user has uttered the initial call word. The continuous conversation mode is applied to improve work convenience in situations such as responding to re-questions regarding specifications, processing the continuous input and output of multiple items, and immediate correction of incorrect utterances. For example, if the system re-questions, "[Which specification is it, 100mm or 140mm?]", the operator is configured to respond immediately with "[100mm]" without the call word. Furthermore, it may be processed to detect negative utterances such as "[No]" or "[Cancel]" to cancel the current processing state, or to analyze the intent of a correction utterance such as "[No, not M8, but M10]". Additionally, it is configured to automatically return to a general call word waiting mode when a designated timeout period has elapsed.
[0094] In addition, an adaptive noise gate structure can be applied to minimize the impact of environmental noise in industrial sites on speech recognition accuracy. Conventional RMS (Root Mean Square) threshold-based noise gates had problems where environmental noise was incorrectly recognized as speech or quiet sounds were blocked depending on the threshold setting, leading to a decrease in recognition rate; furthermore, short silent intervals during speech were cut off, causing speech to be interrupted. Accordingly, one embodiment of the present invention may be configured to measure the RMS value at the unit of every frame (20ms), transmit silent data to VAD (Voice Activity Detection) if the RMS is below the threshold, and process it as normal speech if it is above the threshold.
[0095] In this case, the noise gate threshold can be dynamically loaded via an external configuration file and configured to allow for immediate adjustment in response to environmental changes by performing automatic reloading at regular intervals. Additionally, the administrator UI visually displays real-time RMS values, gate blocking ratios, and recommended threshold information, and allows for real-time adjustment of the threshold through a slider-based interface. In particular, by providing a diagnostic function that automatically calculates the recommended threshold based on the median between environmental noise and the actual ignition RMS, the system can be configured to enable noise gate settings optimized for the field environment.
[0096] As a result of directly applying the above-described adaptive noise gate structure to the applicant's field, VAD malfunctions and recording termination delays caused by environmental noise were reduced, and it was confirmed that the average response waiting time was reduced by approximately 50%. In addition, since there is no interruption during speech and the threshold value can be adjusted immediately according to changes in environmental noise, stable voice interface operation is possible even in industrial environments.
[0097] In addition, in one embodiment of the present invention, during the advancement stage, the system may be configured to continuously collect and learn speech data based on terms actually used in industrial sites, pronunciation variations, and dialects to build a speech recognition model dedicated to industrial sites. For example, the system may be processed to improve recognition accuracy specialized for industrial sites compared to general speech recognition models by additionally training a Whisper-based model focusing on industrial terms related to bolts, piping, equipment, and tools. To this end, Whisper Fine-tuning, LoRA (Low-Rank Adaptation), PEFT (Parameter-Efficient Fine-Tuning), construction of industrial domain speech datasets, and CUDA-based GPU training techniques may be used, but are not limited thereto.
[0098] <Experimental Example>
[0099] A system according to one embodiment of the present invention is currently installed and operating in the No. 2 consumable room of Production Team 1 of the Samdasu Production Headquarters of Jeju Special Self-Governing Province Development Corporation.
[0100] - Installation Location: Entrance to Consumables Room No. 2 - Operation Start Date: March 27, 2026 (First end-to-end operation successful) - Operating Environment: Average ambient noise of 56dB, capable of 24-hour operation - Managed Items: 40 unique items, 336 rows (details by specification)
[0101] Examples for each scenario are as follows.
[0102] Scenario 1: Single Item Inquiry Scenario 2: Lookup using aliases [Operator] "L2" [System] (Wake word detected, recording started) [Operator] "Where are the M8 bolts?" [System] "There are 24 M8 hex bolts on shelf B2 in Supply Room 2." [Worker] "L2, where are the 13mm ones?" "[System] (Alias lookup: 13mm → M8 Bolt Large Size Designation) "[System] "There are 24 M8 hex bolts (large size 13mm) on shelf B2 in Consumables Room 2." ※ For Scenarios 1 and 2, the lookup only goes up to the shelf level. Scenario 3: Inbound / Outbound Processing (Including Standard Ignition) Scenario 4: In / Out Processing (Specification Re-inquiry) [Worker] "L2, I'll ship out 1 pack of 100mm cable ties." [System] (Extract Spec: 100mm) [System] (Check Inventory: Deduct 1 pack of 100mm cable ties) [System] "Shipment of one pack of 100mm cable ties completed." [Worker] "Ltuya, dispatch one pack of cable ties." [System] (Inventory search: 4 cable tie sizes found) [System] "Which size is it: 100mm, 140mm, 270mm, or 370mm?" [Worker] "100mm" (Continuous chat mode) [System] "Dispatch of one pack of 100mm cable ties completed." Scenario 5: Automatic Correction of Pronunciation Errors [Worker] "L2, where is the Tacfron tape?" → Situation involving environmental noise pollution and slurred speech by the worker [System] (STT Output: "Tacfron tape") [System] (Root-suffix separation: Tacfron ↔ Teflon = 0.750) [System] (Automatically corrected to Teflon tape) [System] "There are 3 Teflon tapes on shelf D1 in Consumables Room No. 2."
[0103] In the above-described embodiment, additional information was indicated in parentheses for convenience of explanation and readability, but the actual TTS output was converted into natural Korean pronunciation. In one embodiment of the present invention, operational verification results in an actual industrial environment showed that a success rate of approximately 90% (9 / 10) in detecting a call word was secured under conditions of a distance of 2 to 5 meters, and it was confirmed that the accuracy of speech recognition for industrial terminology improved to over 95% after the application of a correction dictionary. In addition, the average response time from the end of voice input to the start of voice output was measured to be approximately 3 to 5 seconds, and it was verified that stable operation without OOM (Out Of Memory) occurred after the application of a separate memory allocation structure. Furthermore, it was confirmed that it operates stably without interruption even in a 24-hour continuous operation environment, thereby verifying its applicability to industrial sites.
[0104] The quantitative and qualitative effects of the system according to one embodiment of the present invention are summarized as follows.
[0105] item Prior art The present invention Introduction costs 10 million won or more (commercial) Less than 1 million won (1 / 10 or less) Item search time 3~5 minutes About 5 seconds (voice response) Inventory accuracy Occurrence of missing handwritten records Real-time automatic recording New Employee Adaptation Several days to orders Several minutes (learning voice calls only) Operating costs 500,000 to 1,000,000 won per month (license) 0 won (operated offline) GPU memory Server-grade 16GB+ required Edge 8GB integrated memory
[0106] Since the system according to one embodiment of the present invention operates on a completely offline basis, it can perform voice-based inventory management functions without an external network connection, thereby minimizing cyber security risks even in industrial environments where external network connectivity is restricted or security requirements are high, such as food manufacturing processes. Furthermore, inventory inquiry and inbound / outbound processing can be performed solely by voice even when a worker is handling materials with both hands, and user convenience can be enhanced because the system can be used by utilizing various on-site names and aliases even without accurately knowing the official item names.
[0107] Furthermore, by managing calibration dictionaries, alias dictionaries, and item data based on external files, non-developers can directly add and modify item, specification, and alias information without developer intervention, thereby enhancing operational independence and maintenance efficiency. Additionally, based on a modular system structure, horizontal expansion to other consumable rooms or other workplaces is possible, and it provides scalability applicable to various industrial fields, such as pharmaceutical warehouses and maintenance parts warehouses. Moreover, the system according to one embodiment of the present invention is an example of an industrial AI-based voice interface system independently developed by an employee of a domestic public institution, which can provide motivation for the spread of a field-oriented digital innovation culture and subsequent innovation activities.
[0108] Changes
[0109] Post-hoc Filtering - Separating Matching
[0110] Additionally, as the actual implementation code changed while proceeding with the filing of an application according to one embodiment of the present invention, the process of one embodiment of the present invention was changed from a structure in which separate matching is performed first to a structure in which separate matching is performed later when searching for inventory. Accordingly, the changes are summarized as follows. It is clarified that one embodiment of the present invention is based on the modified structure, and the above-described configuration is a configuration according to another embodiment.
[0111] The voice conversion unit (220) can perform a direct correction step by using a pre-established correction dictionary to correct terms within the text when converting voice speech into text using STT. Additionally, the alias replacement unit (290) can replace aliases with a pre-established alias dictionary after the intention classification of the intention classification unit (230) and the extraction of specification information by the extraction unit (240). Then, when the search unit (250) searches for inventory data, if there are multiple search results, it can separate terms within the text into roots and suffixes, and then compare only the roots to calculate the similarity between the roots of the terms within the inventory data and the roots to extract the inventory data. That is, while separation matching was always performed before the change, this process can be used only when two or more pieces of inventory data appear, i.e., as a post-hoc filter.
[0112] <Dictionary of Aliases>
[0113] In an alias dictionary according to one embodiment of the present invention, the category column is removed and a system column is newly established, such as [Item Name, Specification, Alias, Remarks, System]. Here, "system" refers to classification information for identifying interpretation rules or mapping methods applied during the process of mapping an alias to a formal item name, and signifies reference information for processing identical or similar types of alias expressions according to consistent standards. For example, if a worker says, "Ship out a 13mm bolt," the system confirms that the system information stored in the corresponding alias is [Representative System] and can interpret it as "M8 Hex Bolt," which is the formal item name corresponding to the 13mm representative specification. Additionally, if a worker says, "6mm Wrench Bolt," since the system information is set to [Wrench System], it is processed to map to the formal item name corresponding to the hex wrench specification. In this way, the system can be used as reference information to determine which rule to use for interpretation and mapping to the formal item name, even for identical aliases.
[0114] <Dictionary of Variants>
[0115] A variant dictionary is a database that stores the correspondence between standard expressions and various variant expressions arising from a worker's pronunciation habits, abbreviations, misrecognized expressions, spelling variations, or on-site idioms. The variant dictionary can be utilized as dictionary information to convert user utterances or speech recognition results into normalized standard expressions.
[0116] division definition Operational Status example radix It is a core noun part that distinguishes types in item names, and is a major classification unit that a worker can utter independently without modifiers. Operation of 6 root words (bolt, washer, elbow, air fitting, glove, tape) "Bolt", "washer", "glove", "tape" Default It is a representative item that is automatically confirmed without a separate re-question when the operator utters only the root word. Operation of 45 items "Bolt" → "Hex bolt", "Washer" → "Flat washer" Automatic clarification Since the representative item cannot be identified based on the root word alone, the system confirms the item through additional questions. Operation of 4 items "Gloves" → "Which one: sanitary gloves, rubber gloves, or cotton gloves?", "Tape" → Present candidates and select
[0117] Variant dictionary structure hierarchy explanation example radix It is a major product group capable of independent ignition by a worker. bolt sport It is a detailed item group under the root. Hex bolt, wrench bolt, eye bolt Default designation Automatically maps the most frequently used variants to representative items. Bolt → Hex bolt Automatic clarification Perform additional questions if there is no representative item or multiple candidates exist. Gloves → Disposable gloves / Rubber gloves / Cotton gloves
[0118] Input utterance Processing method result Washer Default exists Flat washer automatic confirmation spring washer Direct designation of variants Spring washer confirmed gloves No default Request to select from sanitary gloves, rubber gloves, or cotton gloves rubber gloves Sub-variants exist Request to choose between yellow rubber gloves and red rubber gloves streamer No default Request selection after presenting top-frequency items Air fitting No default Detailed inquiry regarding air fitting types
[0119] division Nickname Dictionary Variant Dictionary purpose Convert aliases to official item names Root-based item hierarchy management Main Column Item name, specification, alias, remarks, system Root, variant, default status, auto-clarification status Target for processing "13mm", "Tacfron", "PVC", etc. "Bolt", "washer", "glove", "tape", etc. result Official Item Name Mapping Automatic confirmation or re-question decision characteristic Number conversion system center Major Category-Detailed Item Hierarchy
[0120] The variant dictionary can be used optionally when term matching fails in the alias dictionary.
[0121] Hereinafter, the operation process according to the configuration of the integrated processing unit of FIG. 2 described above will be explained in detail with reference to FIG. 3 and FIG. 4. However, it is obvious that the embodiment is merely one of the various embodiments of the present invention and is not limited thereto.
[0122] Referring to FIG. 3, the integrated processing unit (200) continuously listens to the USB microphone input while in a constant standby state, and when the operator utters the call word "[L-tu-ya]", it recognizes this and switches to a command recognition mode. Subsequently, in the voice recording stage, only the actual utterance section is recorded using VAD (Voice Activity Detection) and an adaptive noise gate, and voice data is collected when the operator utters "[Ship out 20 13mm bolts]". The collected voice is converted into text using a Faster-Whisper-based speech recognition engine, and, for example, the STT result may be output as "[Ship out 20 3mm bolts]". Subsequently, in the STT correction stage, pronunciation distortion and misrecognition are corrected using a correction dictionary and similarity analysis based on Korean consonant and vowel decomposition.
[0123] Next, in the intent classification stage, the expression "[ship it out]" is analyzed to determine that the user's intent is "[ship]". Then, in the alias substitution stage, the expression "[13mm bolt]" is matched with the alias dictionary's "[correct designation]" category to convert it into the official item name "[M8 hexagonal bolt]". Subsequently, in the specification extraction stage, "[13mm]" specification information is extracted from the "[13mm]" expression, and in the inventory data search stage, the inventory database is queried based on the "[M8 hexagonal bolt]" and "[13mm]" specification conditions. If the search result is confirmed as a single item, the process moves immediately to the shipment processing flow in the result branching stage, and in the response generation stage, a response sentence such as "[20 M8 hexagonal bolts have been shipped. The current inventory is 120]" is generated.
[0124] Subsequently, in the TTS preprocessing stage, numerical, unit, and specification expressions are converted into natural Korean speech forms, and in the speech synthesis stage, voice data is generated using MeloTTS. The generated voice is output to the operator through a USB speaker. Finally, it is maintained in continuous conversation mode for a certain period of time, and is configured to process the next task immediately without a wake word if the operator utters a subsequent command such as "[Please ship 10 M10s as well]". On the other hand, if there is no additional utterance for a specified period, the system returns to a constant standby state and waits for the input of the next wake word.
[0125] Referring to Fig. 4(a), direct correction through a correction dictionary and correction through root-suffix separation matching are performed during the STT correction stage. If correction is still impossible, a correction can be performed by checking whether there is a term matching an alias based on an alias dictionary and replacing it with an alias. If no matching item name is found even after performing alias matching, the system can be configured to determine the term as an unregistered item or an uncertain utterance and branch to a re-question or administrator verification procedure. For example, if the item name uttered by the user does not match in the inventory database, correction dictionary, root-suffix separation matching, or alias dictionary, the system can be processed to generate a re-utterance request response such as "[The item could not be found. Please say it again]." Additionally, if multiple candidates exist but their reliability is low, the system can be configured to re-quest the candidate item to the user, such as "[Is it Teflon tape or insulation tape?]" In addition, if the same unregistered utterance occurs repeatedly, the term corresponding to the utterance is saved as a new alias candidate, and processed to be added as a new entry in the alias dictionary upon administrator approval or when automatic learning conditions are met. Through this, it is configured to adapt to new slang, abbreviations, and non-standard titles that continuously occur in industrial settings.
[0126] (b) is a configuration that sends TTS-BERT and LLM to the CPU because memory may be insufficient when STT, TTS, and LLM are all loaded onto the GPU's memory. Through this, the Out Of Memory (OOM) problem is completely resolved, and normal GPU-based speech synthesis of MeloTTS is possible without fallback to low-quality speech output based on ESpeak-ng. When specifications are automatically extracted as in (c), cases where the item name is the same but the specifications are different can be accurately distinguished to perform inbound / outbound or inquiry, and by performing multi-stage speech synthesis text preprocessing as in (d), there is no awkwardness or difficulty in speech output of special characters, numbers, fractions, etc.
[0127] Software according to one embodiment of the present invention may be implemented with the following libraries. Of course, it is not limited to the following.
[0128] division implementation library role Call word openwakeword Custom Korean wake word 'L2YA' detection Speech Recognition (STT) faster-whisper Korean Fine-tuning Model-Based Speech-to-Text Conversion Natural Language Processing (LLM) Ollama / gemma3:1b Complex response generation, CPU-only operation Text-to-Speech (TTS) MeloTTS Korean speech synthesis, GPU operation Data storage pandas, openpyxl Direct reading / writing of Excel files Admin UI Flask, SocketIO Web dashboard and real-time communication browser Chromium Display administrator screen
[0129] A process according to one embodiment of the present invention is summarized as follows. Each step of the following process may be modified, added, or deleted, but is not limited thereto.
[0130] step explanation tool or model 1. Always on standby The system continuously listens for microphone input and waits for a wake word input. Python, PyAudio, ALSA (Advanced Linux Sound Architecture), Linux Audio Stack 2. Wake word detection Switches to command recognition mode when the user utters a wake word such as "L2ya". openWakeWord, ONNX Runtime, PyTorch, CoreWorxLab Docker 3. Voice recording After detecting the wake word, detect the voice active section to record only the actual command section. VAD (Voice Activity Detection), Noise Gate, WebRTC VAD, RMS analysis 4. Speech Recognition Convert recorded voice to text faster-whisper, whisper-small-komixv2, CUDA, CTranslate2 5. STT Correction Corrects misrecognition of speech recognition results and fixes pronunciation distortions Correction Dictionary, JSON-based correction dictionary, Levenshtein Distance, Korean consonant and vowel decomposition algorithm 6. Intention Classification Determines whether the user's utterance purpose is inquiry, receiving, or shipping. Rule-based NLP, Intent Classification, KoNLPy, Scikit-learn, LLM 7. Alias Substitution Converted site aliases to official item names Alias Dictionary, Excel / openpyxl, Pandas, Category-based Mapping Logic 8. Standard extraction Automatically extracts specification information within the ignition Python Regex, Pattern Matching, NLP Parsing 9. Search Inventory Data Search inventory data based on item name and specifications Pandas, Excel DB, SQLite, DataFrame Filtering 10. Result Branch Determines whether the search result is single, multiple, or a failure Python Branch Logic, Rule Engine 11. Generate response Generates response sentences based on the result Python Template Engine, Ollama, gemma3:1b, Prompt Engineering 12. TTS Preprocessing Converts numbers, units, and special characters into natural Korean speech forms Text Normalization (TN), Regex, Unit Conversion Dictionary 13. Speech Synthesis Converts the generated response sentence into speech MeloTTS, PyTorch, CUDA, BERT-based speech synthesis 14. Voice Output Output the generated voice to the speaker ALSA, PulseAudio, Audio Playback API 15. Return to continuous conversation standby or constant standby Determine whether to issue a follow-up command and return to continuous conversation mode or initial standby state. Session Manager, Context Buffer, Timeout Control
[0131] <Expansion Potential>
[0132] One embodiment of the present invention can be extended in conjunction with automatic weighing. That is, it can be configured to automatically detect the inflow and outflow quantities of small parts, such as bolts, nuts, and washers, using a load cell-based weight sensor installed at the bottom of a small parts box. The load cell is connected via I2C (Inter-Integrated Circuit) communication with an ESP32-based master-slave structure, and the measured weight data is transmitted to an edge device via a USB serial interface. Subsequently, the inflow and outflow quantities are calculated from the changed weight using standard weight data for each item, and if the weight is stably maintained within a range of ±1g for a certain period of time or longer, it can be processed to confirm the actual inflow and outflow. To this end, an HX711 ADC (Analog-to-Digital Converter), ESP32, Python Serial Communication, a Weight Stabilization Algorithm, and rule-based inflow and outflow processing logic may be used.
[0133] One embodiment of the present invention can be extended in conjunction with camera-based monitoring. That is, it can be configured to automatically monitor the status of spare parts, tools, and safety equipment in a consumables room using a MIPI CSI (Camera Serial Interface) camera. In this case, the camera collects images using a periodic shooting method or an event trigger method, and the collected images can be processed to analyze the presence of items, the status of racks, and the storage status of spare parts using a YOLO (You Only Look Once)-based object detection model. Additionally, to prevent increased GPU memory usage and Out-of-Memory (OOM) situations, camera inference tasks and voice AI processing can be configured to operate using a time-slicing method. To this end, MIPI CSI Camera, OpenCV, YOLOv8, PyTorch, CUDA, and event-based scheduling technology may be used.
[0134] One embodiment of the present invention can be extended in conjunction with the advanced integration of LLM. That is, it can be processed to form a self-adaptive learning loop by continuously logging the operator's utterance content and the final item matching result. Accordingly, it can be configured to improve adaptability to industrial site-specific utterance patterns and new aliases by accumulating utterance-matching result data and performing periodic relearning. In addition, to process complex queries such as "[Notify me if the M8 bolt inventory is less than 5]," Natural Language Understanding (NLU)-based intent analysis and conditional expression interpretation structures can be additionally applied. To this end, Ollama-based LLM, Shadow Logging, Vector Database, Prompt Engineering, Fine-tuning, and Rule-based Condition Parser can be utilized.
[0135] One embodiment of the present invention is configured to horizontally extend a voice inventory management system, currently operated based on a single consumables room, to multiple consumables rooms and other business sites during multi-business site expansion. In this expansion structure, item data, alias dictionaries, and inventory data for each business site or consumables room can be managed independently, while enabling integrated monitoring through a central management server. Furthermore, a modular structure is applied to accommodate various item systems, such as pharmaceuticals, maintenance parts, and general office supplies, depending on the industrial characteristics of each business site. To this end, multi-tenant databases, REST APIs, SQLite / PostgreSQL, Docker containers, and central dashboard-based management technologies may be utilized.
[0136] In one embodiment of the present invention, it can be extended into an integrated edge AI platform. That is, a fully offline industrial AI structure based on a single edge device can be configured to extend to various industrial site management areas. This platform can additionally perform predictive maintenance of equipment using vibration sensors, industrial site anomaly detection, and preventive management functions, and can be processed to integrate and operate voice interfaces, vision AI, and sensor-based analysis functions. To this end, vibration sensors, Fast Fourier Transform (FFT)-based frequency analysis, TinyML, Edge AI, Message Queuing Telemetry Transport (MQTT), YOLO-based vision analysis, and time-series anomaly detection models (LSTM, Autoencoder, etc.) may be used, but are not limited thereto.
[0137] As for the details regarding the method of providing a warehouse management solution using voice recognition shown in FIGS. 2 to 4 that are not described, they are identical to or can be easily inferred from the details described above regarding the method of providing a warehouse management solution using voice recognition shown in FIG. 1, so further explanation will be omitted.
[0138] FIG. 5 is a diagram illustrating the process of transmitting and receiving data between each component included in the AI-based voice recognition warehouse inventory inquiry and inbound / outbound system of FIG. 1 according to an embodiment of the present invention. Hereinafter, an example of the process of transmitting and receiving data between each component will be described through FIG. 5, but the present invention is not to be interpreted as being limited to such an embodiment, and it is obvious to those skilled in the art that the process of transmitting and receiving data illustrated in FIG. 5 may be changed according to various embodiments described above.
[0139] Referring to FIG. 5, the integrated processing unit receives a voice utterance received from a voice input device (S5100) and converts the voice utterance into text using STT (Speech to Text) (S5200).
[0140] And. The integrated processing unit classifies the intent of the voice utterance from the text (S5300) and extracts the item name and specification information within the text (S5400).
[0141] Additionally, the integrated processing unit searches for inventory data by item name and specification information (S5500), and generates a response based on the searched inventory data (S5600).
[0142] Accordingly, the integrated processing unit converts the generated response into speech using TTS (Text to Speech) and outputs it (S5700).
[0143] The order of the steps described above (S5100–S5700) is merely an example and is not limited thereto. That is, the order of the steps described above (S5100–S5700) may vary, and some of these steps may be executed simultaneously or deleted.
[0144] As for the details regarding the method of providing a warehouse management solution using voice recognition shown in Fig. 5 that are not described, they are identical to or can be easily inferred from the details described above regarding the method of providing a warehouse management solution using voice recognition shown in Figs. 1 to 4, so further explanation will be omitted.
[0145] A method for providing a warehouse management solution using voice recognition according to an embodiment described through FIG. 5 may also be implemented in the form of a recording medium containing computer-executable instructions, such as an application or program module executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, as well as removable and inseparable media. Additionally, a computer-readable medium may include all computer storage media. Computer storage media include both volatile and non-volatile, removable and inseparable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data.
[0146] The method for providing a warehouse management solution using voice recognition according to one embodiment of the present invention described above may be executed by an application basically installed on a terminal (which may include a program included in a platform or operating system, etc., basically installed on the terminal), or by an application (i.e., a program) directly installed by a user on a master terminal through an application providing server, such as an application store server, an application, or a web server related to the service. In this sense, the method for providing a warehouse management solution using voice recognition according to one embodiment of the present invention described above may be implemented as an application (i.e., a program) that is basically installed on a terminal or directly installed by a user, and may be recorded on a computer-readable recording medium such as a terminal.
[0147] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0148] The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention.
Claims
Claim 1 An integrated processing device comprising: a voice input device for receiving voice utterances; a receiving unit for receiving voice utterances received from the voice input device; a voice conversion unit for converting voice utterances into text using STT (Speech to Text); an intent classification unit for classifying the intent of voice utterances in the text; an extraction unit for extracting item names and specification information within the text; a search unit for searching inventory data using item names and specification information; a response generation unit for generating a response based on the searched inventory data; a response conversion unit for converting the generated response into speech using TTS (Text to Speech) and outputting it; and a wake-up unit that receives a pre-set wake word and starts preparation for voice utterances before receiving voice utterances from the receiving unit. The voice output device that outputs the voice converted by the response conversion unit; wherein the voice conversion unit performs a direct correction step that corrects terms within the text using a pre-established correction dictionary when converting the voice utterance into text using STT, the intent classification unit analyzes the text to determine whether the voice utterance has a purpose among inventory inquiry, receiving, and shipping, the extraction unit performs morphological analysis and string pattern analysis to detect item name candidates within the text, and if an official item name is derived using a pre-stored item database and a pre-established alias dictionary, extracts the derived official item name as the item name, the search unit analyzes the inventory data search results to determine whether the search result is a single item, multiple items, or a matching failure state, and the integrated processing unit includes an alias replacement unit that replaces aliases with the alias dictionary after the intent classification of the intent classification unit and the extraction of specification information by the extraction unit.The search unit further includes, when searching the inventory data, if there are multiple search results, separates the terms within the text into roots and suffixes, compares only the roots to calculate the similarity with the roots of the terms within the inventory data, and extracts the inventory data; the alias substitution unit, if a term cannot be substituted by the alias dictionary, queries using a pre-established variant dictionary with a subdivided term already mapped to the term or receives input from a worker; the voice conversion unit, when converting the voice utterance into text using STT, converts the voice utterance into text after undergoing a multi-stage correction process, and the multi-stage correction process includes a direct correction step of correcting terms within the text using a pre-established correction dictionary; A warehouse inventory inquiry and inbound / outbound system using AI-based speech recognition, comprising: a separation matching step in which, if a term in the text is not included in the proofreading dictionary, the term is separated into a root and a suffix, the similarity with the root of a term in the proofreading dictionary is calculated by comparing only the root, and the term in the text is converted to the term of the root in the proofreading dictionary having the highest similarity; wherein the alias replacement unit, prior to the extraction unit extracting item name and specification information in the text, if the similarity in the separation matching step is less than a preset threshold, a pre-established alias dictionary is queried, and if the term in the text matches an alias in the alias dictionary, the term in the text is replaced with an item name that has been mapped and stored with the alias. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 A warehouse inventory inquiry and inbound / outbound system using AI-based speech recognition, characterized in that, in claim 1, the integrated processing unit further includes a separate allocation unit that sets the BERT (Bidirectional Encoder Representations from Transformers) within the TTS and the LLM (Large Language Model) that generates the response to be processed on a CPU (Central Processing Unit) rather than a GPU (Graphics Processing Unit). Claim 8 A warehouse inventory inquiry and inbound / outbound system using AI-based voice recognition, characterized in that, in claim 1, the integrated processing device further includes a specification automatic extraction unit that automatically extracts specification information of item names within the text using a pre-established specification extraction function when the intent of the text is inbound or outbound. Claim 9 In claim 1, the response conversion unit converts the response into speech after undergoing a multi-stage speech synthesis text preprocessing process before converting the generated response into speech using the TTS, and the multi-stage speech synthesis text preprocessing process comprises at least one step among: a step of converting a combination of English letters and numbers in the response into a preset combination expression; a step of converting or removing special characters in the response into a preset special character expression; a step of converting fractions in the response into a preset fraction expression; and a step of converting units in the response into Hangul; characterized in that the warehouse inventory inquiry and inbound / outbound system using AI-based speech recognition.
Citation Information
Patent Citations
Integrated multimodal AI-based pharmaceutical ordering consultation and distribution quality monitoring system
KR102942086B1