A multimodal manufacturing data collection system based on wearable smart glass fusion to supplement the blind spots of fixed vision sensors, and a multimodal manufacturing data collection method
The multimodal manufacturing data collection system addresses blind spots and inefficiencies in smart factories by integrating fixed vision sensors with smart glasses and real-time collaboration, ensuring data completeness and reliability through multimodal inputs and centralized control.
Patent Information
- Application Number
- KR1020260032131
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-02-20
- Publication Date
- 2026-07-27
- Estimated Expiration
- 2046-02-20
AI Technical Summary
Smart factory environments face challenges due to blind spots in data collection by fixed vision sensors, leading to recognition failures, manual data entry inefficiencies, data redundancy, and lack of real-time collaboration, which compromise data reliability and productivity.
A multimodal manufacturing data collection system integrating fixed vision sensors and smart glasses, with a central control server and AI processing, dynamically switches data collection priorities, utilizes multimodal inputs like voice, gaze, and OCR, and enables real-time collaboration among workers to ensure data completeness and reliability.
The system effectively supplements blind spots with first-person views, maintains productivity by hands-free data collection, enhances data accuracy through multimodal integration, and improves workflow efficiency by reducing errors and data gaps.
Smart Images

Figure R1020260032131_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a multimodal manufacturing data collection system and a multimodal manufacturing data collection method, and more specifically, to a wearable smart glasses fusion-based multimodal manufacturing data collection system and a multimodal manufacturing data collection method for supplementing blind spots of fixed vision sensors. Background Technology
[0002] In smart factory environments, it has become common practice to install multiple fixed vision sensors on ceilings or walls to monitor the entire manufacturing site. These fixed vision sensors are designed to automatically recognize the location, quantity, and movement paths of objects by capturing images of the work area from a fixed height and angle. However, due to their structural characteristics involving fixed viewing angles and shooting ranges, a problem inevitably arises where image data cannot be acquired in blind spots—areas beneath equipment, inside shelves, within boxes or containers, or in areas obscured by workers' bodies or equipment.
[0003] The actual working environment of a manufacturing site is a dynamic one that constantly changes depending on factors such as material stacking, equipment layout, and worker movement. While fixed vision sensors are optimized based on the layout at the time of initial installation, situations frequently arise where the existing camera placement cannot fully cover the entire process when material stacking structures change or new equipment is added. Consequently, some process data is continuously lost, leading to a problem where the reliability of data-driven decision-making is compromised.
[0004] In particular, minute changes in state, label information, and quality markings that occur during processes such as workers lifting materials or performing assembly tasks are often not accurately recognized due to the resolution and distance limitations of fixed cameras. Since the camera's structure is physically unable to perform close-up shots, there is a problem of high recognition error rates when identifying detailed text information or small parts. Such recognition failures can lead to incorrect inventory management, missing process history, and reduced quality traceability.
[0005] In cases of blind spots or recognition failures, conventional technology adopts a method where operators manually input data using separate PC terminals, kiosks, or handheld PDAs. However, since operators typically perform tasks while holding tools or materials, inputting data entails the inconvenience of having to stop work, put down equipment, and operate the input device. This leads to a disruption in the workflow, which in turn results in increased processing time and reduced productivity.
[0006] Furthermore, manual data entry methods are characterized by data quality that varies significantly depending on the operator's skill level and sense of responsibility. Input delays, mis-entries, and omissions occur frequently; in particular, in repetitive processes, the possibility of errors increases due to accumulated fatigue from repeatedly entering the same data. This human-dependent data collection structure presents limitations that conflict with the automation and precision data-driven operations pursued by smart factories.
[0007] Recently, attempts have been made to more intuitively check data at the work site by introducing wearable devices, particularly smart glasses. However, conventional applications of smart glasses are often limited to simple viewer functions, such as displaying work instructions or transmitting remote support video. A structure that dynamically switches data collection priorities by linking in real-time with a fixed vision system has not been sufficiently implemented.
[0008] In environments where fixed sensors and wearable devices operate independently, data regarding the same process is stored in a distributed manner across different systems. This results in data redundancy, requiring additional alignment and refinement operations to perform integrated analysis on a central server. This data integration process increases system complexity and causes delays in real-time analysis.
[0009] Furthermore, mechanisms to quantitatively evaluate the reliability of voice or video data input by operators, or to automatically adjust priorities by comparing them with fixed sensors, are often absent. This leads to situations where it is difficult to determine which data source should serve as the basis for decision-making, resulting in a decline in the consistency of data-driven operations.
[0010] Furthermore, in manufacturing environments where multiple workers operate simultaneously, a collaborative structure is required that allows nearby workers to compensate for situations where a specific worker fails to recognize information located in a blind spot. However, conventional technologies fail to systematically implement real-time collaboration request and verification functions between workers, leading to a problem where data regarding blind spots is lost for extended periods.
[0011] As such, conventional smart factory environments contain complex problems, such as structural limitations of fixed vision sensors, inefficiency of manual input methods, non-interconnection with wearable devices, data redundancy, and the lack of reliability evaluation. Due to these problems, the completeness and accuracy of manufacturing data are degraded, and there are limitations in that the original purpose of smart factories, which is to improve productivity, is not fully achieved. Therefore, the present invention aims to solve the problems of the prior art. Prior art literature
[0012] Korean Patent Publication No. 10-2900896 (Registration Date: December 11, 2025) The problem to be solved
[0013] The present invention aims to solve the problems of blind spots and recognition failures caused by the structural limitations of ceiling-mounted vision sensors, and to implement an integrated environment in which workers can continuously collect manufacturing data without moving to a separate terminal. In addition, the invention aims to minimize data gaps by automatically switching data sources between fixed sensors and wearable smart glasses depending on the situation, and to simultaneously ensure data completeness and reliability without reducing productivity through a hands-free collection method based on multimodal input. means of solving the problem
[0014] A multimodal manufacturing data collection system according to one aspect of the present invention for achieving such objectives may comprise: a fixed vision sensor network that monitors the entire manufacturing site; smart glasses worn by a worker and equipped with a camera, a microphone, and a display; a central control server that receives data from the fixed vision sensors and smart glasses and determines and controls the priority of data sources according to the recognition status; an AI processing engine that converts collected multimodal data into text and numerical data; and a collaboration relay unit that relays collaboration requests through communication between smart glasses worn by a plurality of workers, so that another nearby worker can collect or verify data that one worker failed to collect.
[0015] The present invention may provide a multimodal manufacturing data collection method using the multimodal manufacturing data collection system. A multimodal manufacturing data collection method according to one aspect of the present invention may comprise: an object recognition step of acquiring an image of a work area through a fixed vision sensor fixed to a ceiling and attempting object recognition; an object location detection step of calculating the reliability of the recognition result of the fixed vision sensor and comparing it with a preset threshold value or determining whether the object is located in a blind spot; a data collection priority switching step of automatically switching the priority of data collection from the fixed vision sensor to smart glasses (fail-over) when the reliability is less than the threshold value or is determined to be in a blind spot; a data collection need notification step of displaying a visual notification to an operator through an augmented reality (AR) display of the smart glasses to indicate that data collection is required; and a manufacturing data generation step of generating manufacturing data by receiving at least one multimodal input among the operator's speech (STT), visual image (OCR) based on gaze, and gesture through the smart glasses.
[0016] In one embodiment of the present invention, the visual notification of the data collection needs notification step may be configured to include an augmented reality guide that intuitively indicates the object to be collected by the operator by overlaying a virtual arrow or highlight box on the object to be worked on.
[0017] In one embodiment of the present invention, the voice input of the manufacturing data generation step can be converted into a structured data format by extracting the worker's speech intent and key keywords (item, quantity, status) through natural language processing (NLP) after performing preprocessing to remove noise from the work site.
[0018] In one embodiment of the present invention, the image input of the manufacturing data generation step can automatically detect the text area of a document or label viewed by an operator, perform optical character recognition (OCR), and input the recognized data into a system after verifying whether the recognized data is in a valid format.
[0019] In one embodiment of the present invention, the multimodal manufacturing data collection method may be configured to include: a context-aware step for automatically setting the operating mode (receiving mode, assembly mode, inspection mode) of the smart glasses based on the operator's current location and process step information; and a battery saving step for reducing the battery consumption of the smart glasses by returning the data collection priority from the smart glasses to the fixed vision sensor (fail-back) when the fixed vision sensor recognizes the object normally again. Effects of the invention
[0020] According to the present invention, the omission of manufacturing data can be substantially eliminated by supplementing information in blind spots that fixed vision sensors cannot recognize with the operator's first-person view. In addition, by utilizing multimodal inputs such as voice, gaze, and OCR, data can be naturally generated even during work, thereby preventing interruptions in the workflow and maintaining productivity. Furthermore, by weighting the reliability of data sources and automatically selecting the optimal input path, the accuracy of analysis is improved. Additionally, through AR-based guidance, work errors are reduced, and the proficiency of novice operators is enhanced. Brief explanation of the drawing
[0021] FIG. 1 is an overall configuration diagram of a hybrid data collection system in which a fixed vision sensor and smart glasses are fused according to the present invention. Figure 2 is an example of a multimodal interface combining STT, OCR, and AR provided from the perspective of a worker in smart glasses. Figure 3 is a diagram comparing the global monitoring area of a fixed camera and the first-person view coverage area of smart glasses. Figure 4 is a diagram illustrating the OCR recognition and automatic input process of a transaction statement and barcode using a smart glass camera. FIG. 5 is a block diagram showing a multimodal manufacturing data collection system according to one embodiment of the present invention. FIG. 6 is a flowchart illustrating a multimodal manufacturing data collection method according to one embodiment of the present invention. Specific details for implementing the invention
[0022] Preferred embodiments of the present invention will be described in detail below with reference to the drawings. Prior to this, terms and words used in this specification and claims should not be interpreted as being limited to their ordinary or dictionary meanings, but should be interpreted in a meaning and concept consistent with the technical spirit of the present invention.
[0023] Throughout this specification, when it is stated that one component is located "on" another component, this includes not only cases where one component is in contact with another component, but also cases where another component exists between the two components. Throughout this specification, when it is stated that a part "includes" a component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0024] FIG. 1 shows an overall configuration diagram of a hybrid data collection system in which a fixed vision sensor and smart glasses are fused according to the present invention, FIG. 2 shows an example diagram of a multimodal interface fused with STT, OCR, and AR of smart glasses provided from the operator's perspective, and FIG. 3 shows a diagram comparing the global monitoring area of a fixed camera and the first-person viewpoint coverage area of smart glasses.
[0025] Referring to these drawings, the multimodal manufacturing data collection system (100) and multimodal manufacturing data collection method (S100) according to the present invention can raise the completeness of manufacturing data to 100% by securing detailed work details or obscured material information that cannot be seen by a ceiling camera from the first-person perspective of the worker. In addition, since the worker processes data using only voice or gaze while performing work without needing to move to a separate input terminal, a decrease in productivity caused by data collection can be prevented. Furthermore, through an AR guide, a system and method can be provided that enable even novice workers to perform tasks accurately like skilled workers.
[0026] Hereinafter, with reference to FIGS. 1 to 6, each component constituting the multimodal manufacturing data collection system (100) and the multimodal manufacturing data collection method (S100) according to the present invention will be described in detail.
[0027] Detailed description of the fixed vision sensor network (110)
[0028] A fixed vision sensor network (110) is composed of multiple vision sensor modules placed at regular intervals on the ceiling, walls, or on the top of equipment to continuously monitor the entire manufacturing site. Each vision sensor module includes a high-resolution camera, a wide-angle or variable-focus lens, a light source for illumination correction, and an edge processor for image preprocessing, and is installed according to a coverage map designed in advance considering the movement path of the manufacturing line and the layout structure of the equipment. The fixed vision sensor network (110) operates in a network unit rather than in individual camera units, and is configured so that the image areas between adjacent sensors partially overlap, thereby ensuring the continuity of object tracking.
[0029] The fixed vision sensor network (110) is not limited to simple image acquisition functions but can also incorporate basic analysis functions capable of performing real-time object recognition and status determination. To this end, each sensor module is equipped with a GPU or NPU-based computing unit to perform primary object detection at the edge and transmit the detection results and reliability indicators to a central control server. This structure contributes to ensuring scalability in large-scale manufacturing environments while minimizing network traffic.
[0030] The fixed vision sensor network (110) can form logical clusters in units of process zones. For example, it can be divided into receiving zones, assembly zones, inspection zones, etc., and the recognition model can be optimized for each zone, and different AI models can be applied depending on the process characteristics. Through this, it is configured to enable customized recognition suitable for various work environments even within the same network structure.
[0031] Additionally, the fixed vision sensor network (110) generates metadata such as the position coordinates, movement path, and dwell time of an object and provides it to a central control server. This metadata is utilized not as simple image information but as quantified manufacturing data, and is subsequently converted into structured data and analyzed by an AI processing engine. Through this, it can be utilized for various advanced services such as inventory management, process flow optimization, and worker movement analysis.
[0032] Nevertheless, since the fixed vision sensor network (110) has a structurally fixed viewpoint and shooting range, it has a limitation in recognizing physically obscured areas or fine information requiring close-up shooting. Therefore, the fixed vision sensor network (110) is designed not as a standalone system but under the premise of being linked with a wearable device, and has a structure that is organically connected to a central control server so that the priority of data collection can be automatically switched to another device when the recognition reliability drops below a threshold.
[0033] Detailed description of smart glasses (120)
[0034] Smart glasses (120) are wearable devices worn by a worker and are equipped with a camera, microphone, speaker, micro-display, and wireless communication module. The camera of the smart glasses (120) is positioned to acquire first-person view images that are nearly aligned with the worker's line of sight, and is configured to enable stable image acquisition in various manufacturing environments by including autofocus and low-light correction functions. This allows for the supplementation of close-up information and data on obscured areas that were not captured by fixed vision sensors.
[0035] The smart glasses (120) operate not as a simple video recording device, but as a data collection platform that integrates a multimodal input interface. Voice collected through the microphone undergoes noise removal preprocessing and is then transmitted to an STT and natural language processing module, where the operator's speech intent and key keywords are extracted and converted into structured data. Additionally, text areas of documents or labels captured through the camera undergo automatic detection and OCR processing and are input into the system.
[0036] The display of the smart glasses (120) provides an augmented reality (AR)-based overlay function. Virtual arrows, highlight boxes, or text guides are superimposed on target objects that require data collection, providing intuitive instructions to the operator. These AR guides are dynamically generated based on process steps, the operator's current location, object recognition results, etc., and contribute to providing even novice operators with work accuracy at the level of an expert.
[0037] The smart glasses (120) communicate in real time with a central control server and immediately switch to an active mode upon receiving a data collection priority switching command (Fail-over). Conversely, when the recognition of the fixed vision sensor is normalized, it receives a Fail-back signal and switches to a low-power standby mode. This operational structure minimizes battery consumption of the smart glasses (120) while enabling active data collection only when necessary.
[0038] In addition, the smart glasses (120) are linked with a collaboration relay unit to support collaboration among multiple workers. Information that a specific worker has not recognized can be transmitted to nearby workers in the form of a request message, and other workers can collect or verify the data on their behalf from their own perspective. This collaboration-based structure contributes to enhancing the completeness and reliability of manufacturing data while simultaneously implementing a field-oriented, real-time, mutually complementary data collection system.
[0039] Detailed description of the central control server (130)
[0040] The central control server (130) is a core control node that integrates and manages all video, audio, and metadata collected from the fixed vision sensor network (110) and the smart glasses (120). The central control server (130) includes high-performance computing resources and a large-capacity storage device, aggregates data collected in real-time at the process or line level, and organizes an integrated data stream based on the source and timestamp of each data. Through this, distributed sensor information from the manufacturing site is unified within a single logical system.
[0041] The central control server (130) performs control logic that dynamically determines the priority of data sources, going beyond simple data collection functions. It determines which sensor to prioritize at a specific point in time by comprehensively analyzing the object recognition reliability of fixed vision sensors, blind spot judgment results, and worker location information. In this process, the central control server (130) automatically performs fail-over and fail-back by applying preset threshold comparisons, a context-based policy engine, and a prediction model based on past learning data.
[0042] Additionally, the central control server (130) performs a session control function that manages the smart glasses (120) of multiple workers. It sets the operation mode based on each worker's current process stage, location, work history, etc., and sends a data collection request only to specific workers when necessary. This prevents unnecessary duplicate collection and enables the efficient distribution of system resources.
[0043] The central control server (130) is linked with the collaboration relay unit (150) to relay data supplementation requests between workers. If a failure to recognize a specific object occurs, the central control server (130) determines the worker closest to that location and sends a collaboration request. This structure compensates for data omissions that occurred in a single-sensor dependent structure and serves to implement a field-based distributed collaboration system.
[0044] Furthermore, the central control server (130) is linked with the manufacturing history database to support long-term data analysis. Real-time collected data is structured and stored in a structured database and log storage, and is subsequently used for advanced analysis such as process optimization, quality analysis, and inventory forecasting. In this way, the central control server (130) operates as a core control hub of the system that integrates sensor fusion, priority control, collaboration relay, and data management functions.
[0045] Detailed description of the AI processing engine (140)
[0046] The AI processing engine (140) is an analysis module that analyzes multimodal data transmitted from a fixed vision sensor network (110) and smart glasses (120) and converts it into structured manufacturing data. The AI processing engine (140) integrally includes an image analysis model, a speech recognition model, a natural language processing model, an optical character recognition model, etc., and performs the role of converting data of different input formats into a common data format.
[0047] The AI processing engine (140) can perform the function of re-verifying or correcting the object recognition results transmitted from the fixed vision sensor. It performs additional deep learning-based analysis on the primary recognition results performed at the edge to reduce false detections and missed detections, and precisely determines changes in the state of the object, abnormal signs, etc. Through this, the accuracy of the data is improved and the reliability of process quality control is increased.
[0048] Voice data received from the smart glasses (120) undergoes noise removal, acoustic model application, and language model-based sentence interpretation processes in the voice processing module of the AI processing engine (140). Subsequently, key entities such as item name, quantity, and status are extracted through a natural language processing algorithm and mapped to a predefined manufacturing data schema. This process performs the core function of converting the worker's free speech into structured data.
[0049] In the case of video input, the AI processing engine (140) automatically detects the text area of the document or label viewed by the worker and extracts character information through OCR. The validity of the extracted data is determined through a format verification algorithm, and the consistency of standard codes, date formats, batch numbers, etc., is verified. This prevents the system from reflecting incorrectly recognized data in advance.
[0050] The AI processing engine (140) compares and corrects data collected from different sensors for the same object through a multimodal fusion algorithm. For example, it checks whether the object ID of a fixed sensor matches the voice item name entered from the smart glasses, and if there is a mismatch, it recalculates the reliability weight. This fusion processing structure ensures data consistency and performs the function of substantially improving the decision-making accuracy of the manufacturing site.
[0051] Detailed explanation of the collaborative relay unit (150)
[0052] The collaboration relay unit (150) is a module that relays real-time communication between smart glasses (120) worn by multiple workers and supports nearby workers in collecting or verifying data that could not be collected by a specific worker. The collaboration relay unit (150) is linked with a central control server (130) to continuously receive location information, process stage information, and current work status of each worker, and dynamically selects a collaboration target based on this information. It has a structure that delivers requests to the appropriate worker at the necessary time, while minimizing unnecessary notifications through a context-aware selective relay structure rather than a simple broadcast method.
[0053] The collaboration relay unit (150) automatically initiates a collaboration request process when a blind spot judgment or a decrease in recognition reliability event occurs. For example, if both the fixed vision sensor network (110) and the smart glasses (120) fail to recognize a specific object with sufficient reliability, the collaboration relay unit (150) analyzes the expected location coordinates and work area information of the object to select the nearest worker or a worker belonging to the same process line as a priority candidate. Subsequently, an augmented reality-based collaboration request message is displayed on the smart glasses (120) of the worker to induce additional data collection.
[0054] The collaboration relay unit (150) is not limited to simply transmitting requests but also includes functions for verifying collected collaboration data and managing its history. Data collected by a specific worker is compared with existing data to evaluate its consistency, and if necessary, undergoes a mutual verification process. During this process, the collaboration relay unit (150) records each worker's response time, data accuracy, past collaboration success rate, etc., to create a reliability profile, which is then reflected in the selection of future collaboration targets. This simultaneously improves the efficiency and quality of collaboration.
[0055] Additionally, the collaboration relay unit (150) can perform an extended function of simultaneously transmitting notifications to multiple workers when an emergency situation or quality anomaly is detected. When defects in specific parts are repeatedly detected or abnormal patterns in the process are analyzed, the collaboration relay unit (150) transmits a warning message to all workers in the relevant area, enabling immediate response at the site level. This structure has the effect of integrating the manufacturing site into a single collaboration network, going beyond a single worker-centered data collection system.
[0056] Furthermore, the collaboration relay unit (150) includes a secure communication structure that takes into account network load and personal information protection. Video portions or text data transmitted between workers are limited to the minimum necessary range and are protected through an encryption protocol. In addition, the priority and frequency of collaboration requests are controlled in conjunction with the policy engine of the central control server (130), thereby preventing increased fatigue of workers or notification overload. This configuration of the collaboration relay unit (150) structurally resolves the problem of missing data in blind spots and ensures that the completeness and reliability of manufacturing data are continuously maintained through field-based collaboration.
[0057] The multimodal manufacturing data collection method (S100) begins with an object recognition step (S110) that acquires images of the work area through a fixed vision sensor installed on the ceiling and automatically recognizes objects, and performs an object location detection step (S120) that calculates the reliability of the recognition result or determines whether the object's location corresponds to a blind spot. Subsequently, if the reliability is below a threshold or if it is determined to be a blind spot, the data source is automatically switched to smart glasses through a data collection priority switching step (S130), and an augmented reality-based visual guide is displayed in a data collection needs notification step (S140). Next, structured data is generated by receiving multimodal inputs such as voice, gaze-based images, and gestures in a manufacturing data generation step (S150), and an operation mode is automatically set by reflecting the worker's location and process information in a context-aware step (S160). Finally, when the recognition of the fixed vision sensor is normalized, the smart glasses are switched to standby mode through a battery saving step (S170) to ensure energy efficiency.
[0058] Detailed description of the object recognition step (S110)
[0059] The object recognition step (S110) is a step in which a fixed vision sensor network (110) monitoring the entire manufacturing site continuously acquires images of the work area and detects and classifies work target objects related to the manufacturing process from the acquired images. Since the fixed vision sensor network (110) is installed on the ceiling or above the equipment and captures the work area with a wide field of view, multiple types of objects such as carts, boxes, parts trays, process equipment, and workers can be included in the screen simultaneously. The object recognition step (S110) is configured to include a preprocessing flow capable of stably separating and identifying object candidates based on such complex scenes.
[0060] In the object recognition step (S110), preprocessing is performed immediately after image acquisition to correct image quality degradation factors that frequently occur in the field environment, such as changes in illumination, flicker, and motion blur. For example, since inter-frame shaking caused by LED lighting flickering or equipment vibration interferes with the stable detection of object boundaries, frame stabilization, noise filtering, contrast enhancement, and distortion correction may be applied. In addition, if the background of the work area changes continuously or reflectors are present, corrections are performed to update the background model or suppress reflection highlights, so that the object detection model can focus on practically meaningful features.
[0061] The core of the object recognition step (S110) is to determine the type and state of an object by applying an object recognition model specialized for the manufacturing site. A fixed vision sensor network (110) or a central control server (130) performs deep learning-based object detection to calculate the bounding box, class label, and confidence of the object, and, if necessary, estimates additional attributes such as the object's posture, orientation, and loading status. For example, in the case of a box, state features such as whether the labeled side faces the camera, in the case of a tray, the loading height, and in the case of a trolley, whether it is in motion can be extracted together, and these estimation results are subsequently used directly for data collection policies and process judgments.
[0062] In the object recognition step (S110), continuity regarding the same object is ensured through frame-to-frame tracking, rather than relying solely on one-time detection results. In environments where scene changes are frequent due to the movement of workers, the movement of carts, and the operation of equipment, the same object is often obscured and reappears within a short period of time, so tracking logic is applied to maintain the object ID. At this time, Kalman filter-based prediction, feature point matching, and re-identification models are utilized to reliably calculate the object's movement trajectory and dwell time, and even when the object disappears, the expected location is maintained for a certain period of time so that it is possible to determine whether there is a blind spot in subsequent steps.
[0063] The output of the object recognition step (S110) is not merely a judgment that an object has been found, but is generated as a structured recognition result that enables reliability evaluation and data source switching in subsequent steps. That is, the object recognition step (S110) generates metadata consisting of an object-specific reliability score, the object's location coordinates, the object's size and degree of occlusion, the sensor ID used for recognition, and time information, and transmits this to the object location detection step (S120). Furthermore, by providing the recognition success rate in past frames or recent recognition error patterns for the same object, it provides a stable basis for judgment to prevent accidental errors in a single frame from spreading excessively through priority switching.
[0064] Detailed description of the object location detection step (S120)
[0065] The object location detection step (S120) is a step for determining, based on the recognition result generated in the object recognition step (S110), whether the object exists in a location that can be reliably observed from the perspective of a fixed vision sensor, or whether it is located in a blind spot and there is a high possibility of a recognition gap occurring. The object location detection step (S120) performs reliability threshold comparison and blind spot determination in parallel, and is configured to interpret the cause of unstable recognition from the perspective of location and scene structure, rather than relying solely on whether recognition is successful. This reduces unnecessary switching and ensures that smart glasses-based supplementary collection is initiated only in situations where there is a high risk of actual data loss.
[0066] In the object location detection step (S120), it is first determined whether the reliability score for each object is below a preset threshold. Reliability can be calculated as a comprehensive score combining quality indicators such as tracking stability, inter-frame consistency, resolution relative to object size, and degree of occlusion, in addition to the class probability of the detection model. For example, if the object moves to the edge of the screen and the resolution drops sharply, or if part of the bounding box is obscured by the operator's arm and tool, the reliability score drops rapidly; therefore, the object location detection step (S120) detects this downward pattern to determine the possibility of entering a blind spot early on.
[0067] The object location detection step (S120) can determine whether there is a physical blind spot by utilizing spatial structure information of the work area. Since fixed obstacles such as shelves, equipment, walls, and protective covers exist in the manufacturing site, a blind spot map including the field of view and obscured areas of each vision sensor can be constructed at the time of installation. The object location detection step (S120) compares the coordinates of an object with this blind spot map to determine whether the object is located in an area that is difficult for the fixed sensor to observe directly, such as the bottom of a shelf, the rear of equipment, or the inside of a box. At this time, if the height information of the object can be estimated, the distinction between the bottom and top of the shelf becomes more precise, thereby improving the accuracy of the blind spot determination.
[0068] In the object location detection step (S120), the reliability of determining blind spots can be improved by utilizing observation overlap information between multiple cameras. Even if an object is obscured from view by a specific camera, it may be partially observed by an adjacent camera, so the observation capability of other sensors within the network is evaluated together. If an object is reliably observed by one or more cameras, priority switching can be delayed or the necessity evaluated as low; conversely, if reliability decreases simultaneously across all cameras, it is determined that there is an actual blind spot or a high necessity for close-up shooting, thereby inducing a transition to a subsequent step.
[0069] Since the result of the object location detection step (S120) is used as a trigger for the data collection priority switching step (S130), the judgment result can be generated as state information including causes and grounds, rather than as a binary value. For example, if the cause of reliability degradation is specified as occlusion, distance, illumination, reflection, blind spot map matching, etc., and transmitted to the central control server (130), the central control server (130) can finely adjust the activation level, notification method, and whether to request collaboration of the smart glasses (120) according to the situation awareness policy. Through this, the object location detection step (S120) functions not as a simple detection step, but as a step that provides a judgment basis for sensor fusion control to compensate for blind spots.
[0070] Detailed explanation of the data collection priority switching step (S130)
[0071] The data collection priority switching step (S130) is a step of automatically switching the subject of data collection from the fixed vision sensor network (110) to the smart glasses (120) based on the reliability result or blind spot judgment information calculated in the object location detection step (S120). This step is not a simple device switching step, but includes control logic that selects the optimal data source by comprehensively considering the current process situation, the worker's location, past recognition history, etc. Through this, smart glasses-based collection is activated only when there is a high probability of actual data gaps occurring, while minimizing unnecessary switching.
[0072] In the data collection priority switching step (S130), the central control server (130) determines whether to switch by comprehensively analyzing the recognition reliability of the fixed vision sensor, the location coordinates of the object, the occlusion ratio, and the frequency of recent recognition failures. For example, if a specific object is detected with a reliability below a threshold in consecutive frames, or moves to an area corresponding to the blind spot map, a Fail-over command is immediately generated. This command is selectively transmitted to the smart glasses (120) of the worker performing the task related to the object.
[0073] When a transition is determined, the data collection priority transition step (S130) switches the smart glasses (120) to an active collection mode and temporarily lowers the data reflection weight for the corresponding object in the fixed vision sensor network (110). During this process, object ID mapping is performed so that previously collected data and data subsequently collected from the smart glasses can be linked to the same object. Through this, the continuity and consistency of the process history are maintained even if the data source changes.
[0074] The data collection priority switching step (S130) is characterized by being performed automatically without requiring explicit operation from the operator. The operator naturally performs data collection through the smart glasses based on the system's judgment, without the need for separate button input or menu selection. This automated structure reduces the cognitive burden on the operator and functions as a key step for practically implementing a sensor fusion-based data collection system.
[0075] In addition, the data collection priority switching step (S130) is designed not to be limited to a unidirectional switching, but to return to the original state when the recognition of the fixed vision sensor is normalized in a subsequent step. That is, fail-over and fail-back are operated within a single continuous policy framework, and both the switching point and the return point are determined based on reliability. Through this, balanced operation is possible that optimizes the battery usage of the smart glasses (120) while minimizing data gaps.
[0076] Detailed explanation of the data collection needs notification stage (S140)
[0077] The data collection need notification step (S140) is a step that intuitively conveys the need for data collection to the worker when it is determined that smart glasses (120)-based collection is necessary in the data collection priority switching step (S130). This step aims to provide clear instructions without disrupting the worker's workflow and displays a visual guide using the augmented reality (AR) display of the smart glasses (120).
[0078] In the data collection needs notification stage (S140), a virtual arrow, highlight box, blinking border, or text guidance message is superimposed on the target object. This overlay is dynamically generated based on the object's real-time position coordinates and moves or remains fixed as the operator's gaze moves. Through this, the operator can intuitively recognize which information needs to be collected from which object without separate explanation.
[0079] The data collection needs notification stage (S140) can adjust the intensity and form of the notification according to the situation. For example, in the case of a data collection request related to urgent quality issues or shipment delays, visual emphasis effects may be enhanced or simple voice guidance may be provided in conjunction. Conversely, in the case of general supplementary input, only minimal highlighting is displayed to avoid excessively distracting the operator's attention.
[0080] Additionally, the data collection needs notification step (S140) is linked to the operator's current process stage and work mode. The guidance text displayed or the highlighted targets may vary depending on context-aware information, such as receiving mode, assembly mode, and inspection mode. For example, in inspection mode, detailed text areas such as the label's expiration date or batch number may be highlighted, while in assembly mode, the component assembly status may be highlighted.
[0081] The data collection needs notification step (S140) goes beyond a simple notification function and serves as an interaction interface between the operator and the system. By looking at the displayed object or responding with voice, the operator naturally proceeds to the manufacturing data generation step (S150), and the system immediately receives the response and begins data collection. This structure replaces the operation of a separate input terminal required in conventional technology and has the effect of practically realizing a field-oriented, hands-free data collection environment.
[0082] Detailed description of the manufacturing data generation step (S150)
[0083] The manufacturing data generation step (S150) is a step of generating manufacturing data by receiving at least one input from the worker—speech (STT), eye-based image (OCR), and gesture—via smart glasses (120) after the worker is instructed to collect the data collection target through the data collection needs notification step (S140). The core of this step is to replace the manual input performed by the worker operating a separate terminal in the prior art with an input action naturally performed during work. The data generated in this step is standardized so that it can be immediately utilized as supporting data for manufacturing operations, such as process history, inventory fluctuations, and inspection results.
[0084] In the manufacturing data generation step (S150), voice input includes a sophisticated preprocessing flow based on the noisy environment of the work site. Since voice signals are prone to distortion in an environment where equipment operation noise, pneumatic noise, and mobile equipment noise are mixed, the smart glasses (120) or the central control server (130) perform noise suppression, echo canceling, and voice segment detection to stabilize the STT recognition rate. Subsequently, a natural language processing module extracts key information such as item, quantity, status, and process code from the spoken sentence and maps it to fields of a predefined manufacturing data schema to generate a structured record.
[0085] In the manufacturing data generation step (S150), gaze-based image input is utilized to automatically collect document or label information regarding the object that the worker is actually looking at. The smart glasses (120) estimate the worker's area of interest by linking the direction of gaze with the camera frame, and automatically detect areas containing text within that area of interest. Subsequently, OCR is performed to extract text information such as batch number, LOT, specification code, and expiration date, and the extraction results filter out character combinations with a high probability of error through format verification. Through this, accurate identification information is generated as manufacturing data without additional operation by the worker.
[0086] In the manufacturing data generation step (S150), gesture input can be utilized as an alternative input method in environments where voice usage is restricted or in security zones. The operator can perform basic commands such as confirmation, cancellation, and selecting the next item by performing specific actions with a finger or by simply changing the direction of the head. Gestures are recognized based on camera images of the smart glasses (120) or inertial sensors, and conditions such as holding for a certain period of time and double-checking may be applied to reduce malfunctions. Through this, the operator can confirm or correct data with minimal interaction, even while wearing gloves.
[0087] In addition, the manufacturing data generation step (S150) may include fusion processing that enhances data reliability through the mutual complementarity of multimodal inputs. For example, when a worker speaks the item and quantity, and simultaneously the item code is extracted by label OCR, the consistency between the two results is automatically verified. In case of discrepancy, a simple verification question is displayed on the AR display of the smart glasses (120), allowing the worker to correct it immediately. Through this verification procedure, errors in entry and omissions that frequently occurred in conventional manual input can be structurally reduced.
[0088] Detailed description of the Context-Aware stage (S160)
[0089] The context-aware step (S160) is a step in which the operation mode and data collection policy of the smart glasses (120) are automatically set based on the worker's current location and process step information. In conventional technology, the same input interface is applied uniformly to all areas, resulting in guidance unrelated to the work context or the inconvenience of the worker having to manually change the mode. The context-aware step (S160) enhances field applicability by having the system proactively identify the process context and automatically provide an appropriate collection method and guidance configuration.
[0090] In the context-aware stage (S160), location information can be calculated using indoor positioning technology. For example, Wi-Fi RTT, BLE beacons, UWB tags, or worker location tracking results based on fixed vision sensors can be utilized, and the central control server (130) identifies the current work area by matching the worker's location with a process area map. At this time, if the configuration is set to distinguish detailed areas within the same area, such as the front of the shelf, the front of the inspection table, and the rear of the equipment, the notification target and input expectation value can be set more precisely.
[0091] In the context-aware stage (S160), process stage information can be provided by linking with external systems such as work orders, MES, and WMS. For example, since the required data items vary depending on whether the worker is processing incoming goods, assembling, or inspecting, the central control server (130) automatically switches the mode of the smart glasses (120) to incoming mode, assembly mode, inspection mode, etc., according to the current process stage. Depending on the mode switch, the keyword dictionary of voice commands, priority items for OCR extraction, and the AR guide display format are automatically changed, so that even if the worker uses the same interface, result data is generated according to the purpose of the process.
[0092] The context-aware stage (S160) may include policy functions that reduce worker fatigue by controlling the frequency and intensity of notifications. For example, if repetitive supplementary collection of the same object occurs continuously within a certain period, the system prevents notification overload by grouping them into an integrated request or adjusting the priority. Additionally, the priority of input channels can be automatically adjusted according to field conditions, such as guiding to a voice-centric mode in areas where the worker's hands are busy and to an OCR-centric mode in areas with high noise.
[0093] Furthermore, the context-aware stage (S160) is linked with the collaboration relay unit (150) and is also used as a criterion for selecting the target of the collaboration request. Even if a specific object is located in a blind spot, if the work currently being performed by a worker is urgent or dangerous, the collaboration request can be prioritized and assigned to another worker. At this time, the central control server (130) configures the optimal collaboration path by synthesizing information on the workload, mobility, and process role of each worker, and as a result, the completeness of data collection and the safety of field operations can be simultaneously ensured.
[0094] Detailed explanation of the battery power saving stage (S170)
[0095] The battery power saving step (S170) is a step in which, after the smart glasses (120) are activated and perform supplementary collection in the data collection priority switching step (S130), when it is detected that the fixed vision sensor network (110) has returned to a state where it can recognize objects normally, the data collection priority is returned from the smart glasses (120) to the fixed vision sensor network (110), and the operating mode is switched to a low-power state to minimize energy consumption of the smart glasses (120). In order to mitigate the problem of the battery being consumed quickly while the conventional wearable device is always in operation, this step is characterized by operating a combination of fail-back and power saving control at the system level.
[0096] In the battery saving stage (S170), the normal return of the fixed vision sensor network (110) is not determined by the success of recognition in a single frame, but can be determined based on continuous recognition stability over a certain period. For example, the number of frames in which object recognition reliability is maintained above a threshold, the continuity of object tracking, and the consistency of observation from multiple cameras are comprehensively evaluated to confirm whether the return condition is satisfied. This stability verification structure prevents frequent switching due to temporary changes in illumination or occlusion, thereby reducing situations where the smart glasses (120) are unnecessarily turned on and off and increasing battery efficiency.
[0097] In the battery power saving stage (S170), power saving control of the smart glasses (120) can be performed not simply by turning off the power, but by hierarchically adjusting the power consumption of each function. For example, the camera can lower the frame rate or switch to standby mode, the microphone can be maintained in a low-power detection mode such as keyword wake-up, and the display can stop displaying overlays and switch to minimum brightness or a screen-off state. Additionally, the wireless communication module can maintain only periodic heartbeats or switch to an event-based reception method to respond to urgent requests from the central control server (130) while suppressing unnecessary constant streaming.
[0098] The battery power saving stage (S170) can dynamically adjust the power saving intensity by considering the battery status of the smart glasses (120). When the remaining battery level is low or usage time is long, a policy can be applied to apply a stronger power saving mode to minimize display usage and limit high-load functions such as OCR or high-resolution image processing. Conversely, when the remaining battery level is sufficient, a lightweight standby mode can be applied to maintain the sensor initial state for quick reactivation. This adaptive power saving policy contributes to increasing operational efficiency by adapting to site characteristics such as work shift times, charging infrastructure, and wear patterns by worker.
[0099] Additionally, the battery power saving stage (S170) is maintained in a structure that allows it to be reactivated at any time through a control loop with the central control server (130). When the fixed vision sensor network (110) is switched back to a low-reliability state or a blind spot event recurs, the smart glasses (120) immediately switch from the power saving state to an active state to resume data collection. At this time, thanks to the minimum sensor state and session information maintained during the power saving stage, the reconnection time is shortened, and the data collection needs notification stage (S140) can be provided to the operator without delay. Consequently, the battery power saving stage (S170) ensures the continuity of actual use of wearable-based supplementary collection while achieving a balance between data completeness and energy efficiency through organic sharing with the fixed vision sensors.
[0100] Example 1: Loading work in a narrow space (between racks) of a material warehouse
[0101] In narrow spaces where forklifts have difficulty entering, such as between racks in a materials warehouse, workers often manually transport and organize materials. In such cases, fixed vision sensors installed on the ceiling may fail to recognize material labels or precise storage locations because their view is obstructed by the worker's body or the rack structure. In these environments, the completeness of inventory location registration cannot be ensured relying solely on automatic recognition based on fixed sensors.
[0102] In this embodiment, if the reliability of the fixed vision sensor is determined to be below a threshold value based on the object location detection result, the data collection priority is automatically switched to the smart glasses. The smart glasses take close-up shots of material labels according to the operator's gaze direction and extract item codes and location information through OCR. During this process, the operator can perform inventory registration simply by looking at the label, without operating a separate input device.
[0103] As a result, no data gaps occur even in confined spaces, and workers can perform tasks continuously without putting down or moving materials. By supplementing the blind spots of fixed vision sensors with a first-person view, the accuracy and real-time nature of inventory location information are simultaneously ensured.
[0104] Example 2: Voice input when assembly process barcode is damaged
[0105] In the assembly process, it is common practice to register serial numbers by scanning the barcodes of parts. However, if the barcode is damaged due to oil, dust, scratches, etc., recognition may fail on fixed vision sensors or scanners. Conventionally, there was a problem where the process flow was interrupted because workers had to remove their gloves and manually input data using a keyboard or terminal.
[0106] In this embodiment, when a barcode recognition failure is detected, smart glasses-based voice input is activated. The operator inputs data by speaking, such as "Serial number XYZ-1-2-3 manual registration," and an AI processing engine extracts key keywords and the number system from the voice and converts them into structured data. Noise removal and speech intent analysis are performed in parallel to ensure recognition accuracy.
[0107] This method prevents a decline in productivity by allowing workers to register data without stopping the process, even while wearing gloves. At the same time, it ensures that the process history is maintained without data loss, even in exceptional situations such as damaged barcodes.
[0108] Example 3: Quality Inspection and Defect Reporting
[0109] In quality inspection processes, workers often visually inspect for microcracks or surface defects. However, these defects may be difficult to recognize automatically due to limitations in the resolution or shooting angle of fixed vision sensors. Conventionally, workers had to photograph defects with a separate camera or record them manually, which resulted in potential delays in recording and missing information.
[0110] In this embodiment, when a worker issues a voice command such as "Defect type: Crack, Take photo," the smart glasses immediately capture a high-resolution photo. The captured image is transmitted to a central control server along with the worker's location coordinates and process step information, and an AI processing engine tags the defect type to generate an automatic report. At this time, the location and time information of the defective object are stored together, enhancing traceability.
[0111] This allows operators to immediately record defects without operating separate equipment, and minute defects discovered on-site are reflected in the database in real time. This improves the accuracy of quality history management and enables rapid analysis of defect causes and measures to prevent recurrence.
[0112] As explained above, according to the present invention, the problem of blind spots caused by the structural limitations of a vision sensor fixed to the ceiling can be substantially resolved. In the prior art, data was omitted from areas obscured by the lower part of the equipment, the inside of shelves, or the worker's body, resulting in constant incompleteness of the process history. The present invention has the effect of actively compensating for these physical blind spots and minimizing data gaps by immediately securing a first-person view image through smart glasses worn by the worker.
[0113] In addition, it can resolve the problem of workflow interruption that occurred in conventional manual input methods. Previously, workers had to put down tools or materials and move to a separate terminal to input data, but the present invention enables data generation without interrupting work by providing multimodal input such as speech-to-text (STT), gaze-based OCR, and gestures. As a result, process continuity is maintained, and factors that reduce productivity are eliminated.
[0114] Furthermore, reliance on human input errors can be significantly reduced. In conventional technology, incorrect entries, omissions, and delayed inputs frequently occurred during the repetitive input process. Since the present invention performs voice intent analysis, keyword extraction, automatic document recognition, and format verification through an AI processing engine, it has the effect of improving the consistency and accuracy of input data and strengthening quality traceability.
[0115] In addition, the problem of data redundancy caused by the independent operation of fixed sensors and wearable devices can be resolved. Since the central control server automatically determines the priority of data sources based on recognition reliability and performs failover and failback, data is managed within a single integrated system. Accordingly, real-time analysis delays are reduced, and the complexity of the data integration process is alleviated.
[0116] Furthermore, consistency in decision-making can be ensured through a data reliability evaluation system. Conventionally, there was a possibility of judgment errors because it was unclear which sensor data should be used as the standard. Since the present invention selects the optimal input path by weightedly evaluating the reliability of each data source, it has the effect of improving the reliability of analysis results and increasing operational efficiency.
[0117] Furthermore, mutually complementary data collection among multiple workers is possible through the collaboration relay unit. In conventional technology, information not recognized by a specific worker was often omitted; however, the present invention relays collaboration requests to nearby workers to perform data verification or alternative collection. Consequently, the integrity of manufacturing data is maintained, and the visibility of the entire process is improved.
[0118] Consequently, the present invention resolves the structural problems of conventional technology, such as the physical limitations of fixed vision sensors, the inefficiency of manual input, data redundancy, and the lack of reliability judgment, through an organic sensor fusion structure. This enables the simultaneous assurance of the completeness, accuracy, real-time capabilities, and operational convenience of manufacturing data, and effectively enables the stable implementation of advanced data-based operations required in a smart factory environment.
[0119] The above detailed description of the present invention describes only specific embodiments thereof. However, it should be understood that the present invention is not limited to the specific forms mentioned in the detailed description, but rather should be understood to include all variations, equivalents, and substitutions within the spirit and scope of the invention as defined by the appended claims.
[0120] In other words, the present invention is not limited to the specific embodiments and descriptions described above, and any person skilled in the art to which the present invention pertains can make various modifications without departing from the essence of the invention as claimed in the claims, and such modifications fall within the scope of protection of the present invention. Explanation of the symbols
[0121] 100: Multimodal Manufacturing Data Collection System 110: Fixed Vision Sensor Network 120: Smart Glasses 130: Central Control Server 140: AI processing engine 150: Collaborative Broadcasting Department S100: Multimodal manufacturing data collection method S110: Object recognition stage S120: Object position detection step S130: Data collection priority switching stage S140: Data collection needs notification stage S150: Manufacturing data generation step S160: Context-Aware Step S170: Battery saving stage
Claims
Claim 1 A fixed vision sensor network (110) for monitoring the entire manufacturing site; smart glasses (120) worn by a worker and equipped with a camera, microphone, and display; a central control server (130) for receiving data from the fixed vision sensors and smart glasses (120) and determining and controlling the priority of data sources according to the recognition status; and an AI processing engine (140) for converting collected multimodal data into text and numerical data. The method includes a collaboration relay unit (150) that relays a collaboration request so that another nearby worker can collect or verify data that one worker failed to collect through communication between smart glasses (120) worn by multiple workers; an object recognition step (S110) that acquires an image of a work area through a fixed vision sensor fixed to the ceiling and attempts object recognition; an object location detection step (S120) that calculates the reliability of the recognition result of the fixed vision sensor and compares it with a preset threshold value or determines whether the object is located in a blind spot; a data collection priority switching step (S130) that automatically switches the priority of data collection from the fixed vision sensor to the smart glasses (120) (Fail-over) if the reliability is below the threshold value or if it is determined to be a blind spot; a data collection need notification step (S140) that displays a visual notification to the worker that data collection is necessary through the augmented reality (AR) display of the smart glasses (120); and a method based on the worker's speech (STT) and gaze through the smart glasses (120). A manufacturing data generation step (S150) for generating manufacturing data by receiving at least one multimodal input among an image (OCR) and a gesture; a context-aware step (S160) for automatically setting the operation mode (receiving mode, assembly mode, inspection mode) of the smart glasses (120) based on the operator's current location and process step information;A multimodal manufacturing data collection system characterized by being operated by a multimodal manufacturing data collection method including: a battery saving step (S170) that reduces battery consumption of the smart glasses (120) by returning the data collection priority from the smart glasses (120) to the fixed vision sensor (Fail-back) when the fixed vision sensor recognizes the object normally again. Claim 2 A multimodal manufacturing data collection method using a multimodal manufacturing data collection system according to claim 1, comprising: an object recognition step (S110) of acquiring an image of a work area through a fixed vision sensor fixed to the ceiling and attempting object recognition; an object location detection step (S120) of calculating the reliability of the recognition result of the fixed vision sensor and comparing it with a preset threshold value or determining whether the object is located in a blind spot; a data collection priority switching step (S130) of automatically switching the priority of data collection from the fixed vision sensor to the smart glasses (120) (Fail-over) if the reliability is less than the threshold value or if it is determined to be a blind spot; a data collection need notification step (S140) of displaying a visual notification to an operator through an augmented reality (AR) display of the smart glasses (120) indicating that data collection is necessary; a manufacturing data generation step (S150) of receiving at least one multimodal input among the operator's voice (STT), image based on gaze (OCR), and gesture through the smart glasses (120) to generate manufacturing data; and the operator's current location and process stage A multimodal manufacturing data collection method characterized by including: a context-aware step (S160) for automatically setting the operating mode (receiving mode, assembly mode, inspection mode) of the smart glasses (120) based on information; and a battery saving step (S170) for reducing battery consumption of the smart glasses (120) by returning the data collection priority from the smart glasses (120) to the fixed vision sensor (fail-back) when the fixed vision sensor recognizes the object normally again. Claim 3 A multimodal manufacturing data collection method according to claim 2, wherein the visual notification of the data collection needs notification step (S140) includes an augmented reality guide that intuitively indicates the object to be collected by the operator by overlaying a virtual arrow or highlight box on the work target object. Claim 4 A multimodal manufacturing data collection method according to claim 2, wherein the voice input of the manufacturing data generation step (S150) performs preprocessing to remove noise from the work site, extracts the worker's speech intent and key keywords (item, quantity, status) through natural language processing (NLP), and converts them into a structured data format, and the image input of the manufacturing data generation step (S150) automatically detects the text area of a document or label viewed by the worker, performs optical character recognition (OCR), verifies whether the recognized data is in a valid format, and then inputs it into the system. Claim 5 delete