Method and system for providing training data

The method automates the detection of objects in images and generates learning data to address the resource-intensive challenge of creating sophisticated AI training data, reducing costs and time while ensuring relevance and efficiency.

WO2025116223A1PCT designated stage expired Publication Date: 2025-06-05DEEPFINE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013349
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-09-04
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The resources required to generate sophisticated learning data for artificial intelligence are excessively large, making it time-consuming and costly to create manually for a wide variety of problems.

Method used

A method and system for automatically detecting objects in images and generating learning data by acquiring consecutive images, detecting target objects with the same shape and depth, projecting bounding boxes, labeling information, and providing these images as learning data, while excluding recognizable objects to prevent unnecessary data generation.

Benefits of technology

This approach reduces the time and cost associated with creating learning data by automating the process, focusing on relevant target objects, and minimizing unnecessary data, thereby enhancing the efficiency of artificial intelligence training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013349_05062025_PF_FP_ABST
    Figure KR2024013349_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a method and a system for providing training data. According to an embodiment of the present invention, the method for providing learning data includes the steps of: obtaining consecutive images; detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images; labeling each of the consecutive images with at least one piece of label information about the at least one target object; and providing the consecutive images, projected with the at least one bounding box labeled with the at least one piece of label information, as training data for recognizing the at least one target object.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for providing learning data

[0001] The examples below describe techniques for collecting and providing learning data for objects.

[0002] One of the most crucial elements in advancing artificial intelligence is generating sophisticated training data sets. It's self-evident that AI trained using sophisticated training data can make more sophisticated judgments than those trained without it.

[0003] However, the problem is that the resources required to generate sophisticated training data are prohibitively large. To meet the massive demands of AI, sophisticated training data for a truly diverse range of problems is needed, but manually creating all of this data is prohibitively time-consuming and expensive.

[0004] Accordingly, the embodiments below propose a method and system for automatically detecting objects in an image and generating learning data for recognizing the objects.

[0005] The technical task that one embodiment seeks to achieve is to propose a method and system for detecting an object in an image and generating and providing learning data for recognizing the object.

[0006] At this time, one embodiment proposes a method and system for detecting a target object for learning data while excluding at least one recognizable object to prevent unnecessary learning data generation and provision.

[0007] The technical problems to be achieved by the embodiments are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the embodiments belong from the description below.

[0008]

[0009] In order to achieve the above technical task, a learning data providing method performed by a computer device according to one embodiment may include the steps of: acquiring consecutive images; detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images; labeling at least one label information for the at least one target object onto each of the consecutive images; and providing the consecutive images on which the at least one bounding box is projected and the at least one label information is labeled as learning data for recognizing the at least one target object.

[0010] According to one aspect, the projecting step may be characterized by including a step of detecting at least one target object having the same shape and depth in the consecutive images using point cloud data corresponding to each of the consecutive images.

[0011] According to another aspect, the acquiring step may further include a step of acquiring the point cloud data corresponding to each of the consecutive images.

[0012] According to another aspect, the projecting step may be characterized by including a step of excluding at least one recognizable object from the consecutive images; and a step of detecting at least one unrecognizable target object remaining after excluding the at least one recognizable object from the consecutive images.

[0013] According to another aspect, the at least one label information may be characterized by including information about the shape and information about the depth of the at least one target object.

[0014] According to another aspect, the providing step may further include a step of additionally providing the point cloud data corresponding to each of the consecutive images as the learning data for recognition of the at least one target object.

[0015] According to another aspect, the providing step may be characterized by including: a step of downscaling the resolution of a remaining area excluding the at least one bounding box in each of the consecutive images; and a step of providing the consecutive images, in which the at least one bounding box is projected, the at least one label information is labeled, and the resolution of the remaining area is downscaled, as the learning data for recognizing the at least one target object.

[0016] According to another aspect, the providing step may be characterized as a step of providing only the at least one bounding box and the at least one label information from each of the consecutive images as the learning data for recognition of the at least one target object.

[0017] According to one embodiment, a computer-readable recording medium having recorded thereon a computer program for executing a method for providing learning data on a computer device may include: acquiring consecutive images; detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images; labeling at least one label information for the at least one target object onto each of the consecutive images; and providing the consecutive images on which the at least one bounding box is projected and the at least one label information is labeled as learning data for recognizing the at least one target object.

[0018] According to one embodiment, a computer device for performing a method for providing learning data may include at least one processor configured to execute computer-readable instructions, wherein the at least one processor may include: an acquisition unit for acquiring consecutive images; a projection unit for detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images; a labeling unit for labeling at least one label information for the at least one target object onto each of the consecutive images; and a provision unit for providing the consecutive images on which the at least one bounding box is projected and on which the at least one label information is labeled, as learning data for recognizing the at least one target object.

[0019] According to one aspect, the projection unit may be characterized in that it detects at least one target object having the same shape and depth in the consecutive images by using point cloud data corresponding to each of the consecutive images.

[0020] According to another aspect, the projection unit may be characterized in that it excludes at least one recognizable object from the consecutive images, thereby detecting at least one target object that is unrecognizable and remaining by excluding at least one recognizable object from the consecutive images.

[0021] According to another aspect, the at least one label information may be characterized by including information about the shape and information about the depth of the at least one target object.

[0022] One embodiment may propose a method and system for detecting an object in an image and generating and providing learning data for recognizing the object.

[0023] At this time, one embodiment proposes a method and system for detecting a target object for learning data while excluding at least one recognizable object, thereby preventing unnecessary learning data generation and provision.

[0024] It should be understood that the described technical effects are not limited to the effects described above, but include all effects that can be inferred from the composition of the invention described in the detailed description below or the claims.

[0025] Figure 1 is a diagram illustrating an example of a service environment according to one embodiment.

[0026] FIG. 2 is a block diagram illustrating an example of a computer device according to one embodiment.

[0027] FIG. 3 is a block diagram illustrating an example of components that the processor illustrated in FIG. 2 may include.

[0028] Figure 4 is a flow chart illustrating a learning data provision method that can be performed by the computer device illustrated in Figure 2.

[0029] FIG. 5 is a drawing for explaining projecting at least one bounding box in the learning data providing method illustrated in FIG. 4.

[0030] FIG. 6 is a drawing for explaining labeling at least one piece of label information in the learning data provision method illustrated in FIG. 4.

[0031] Figures 7 to 9 are drawings for explaining providing learning data in the learning data providing method illustrated in Figure 4.

[0032]

[0033] Hereinafter, the present invention will be described with reference to the attached drawings. However, the present invention can be implemented in various different forms and is therefore not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar parts have been designated with similar reference numerals throughout the specification.

[0034] Throughout the specification, when a part is said to be "connected (connected, contacted, or coupled)" to another part, this includes not only cases where it is "directly connected," but also cases where it is "indirectly connected" with another part in between. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather implies that it may include other components, unless otherwise specifically stated.

[0035] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0036]

[0037] In the following embodiments, a method and system for automatically generating and providing learning data for object recognition are described.

[0038] More specifically, the embodiments describe a method and system for providing training data for detecting objects in an image and generating training data for recognition of the objects.

[0039] In addition, in the following embodiments, a method and system for providing learning data that excludes at least one recognizable object and detects a target object for learning data in order to prevent unnecessary generation and provision of learning data are described.

[0040] The learning data provision method can be performed by a learning data provision system implemented in a server or portable terminal comprising at least one computer device including a processor, and the learning data provision system can be operated under the control of a computer program. The above-described computer program can be stored in a computer-readable recording medium so as to be combined with a computer device and execute the learning data provision method on the computer device. The computer program described herein may take the form of an independent program package, or may take the form of an independent program package pre-installed on the computer device and linked with an operating system or other program packages.

[0041] The learning data provision system can input or output necessary information through a service platform running on a portable terminal, and the service platform is implemented in the form of a dedicated program, application, or web page running on a portable terminal, which is a mobile terminal carried by a user, and can be used as an interface in the process of performing the learning data provision method.

[0042]

[0043] FIG. 1 is a diagram illustrating an example of a service environment according to one embodiment. The service environment of FIG. 1 represents an example including a plurality of electronic devices (110, 120, 130, 140), a plurality of servers (150, 160), and a network (170).

[0044] This drawing 1 is an example for explaining the invention, and the number of electronic devices or servers is not limited to that of drawing 1. In addition, the service environment of drawing 1 merely illustrates one example of environments applicable to the present embodiments, and the environments applicable to the present embodiments are not limited to the service environment of drawing 1.

[0045] The plurality of electronic devices (110, 120, 130, 140) may be mobile terminals implemented as computer devices.

[0046] Each of the electronic devices (110, 120, 130, 140) may be a mobile or fixed terminal carried by a user who receives learning data or a user who acquires continuous images used to generate learning data.

[0047] To this end, each of the electronic devices (110, 120, 130, 140) may mean one of various physical computer devices that can communicate with a server (150, 160) via a network (170) using a wireless or wired communication method, assuming that the electronic devices include a camera to capture a target space and obtain continuous images. For example, examples of the plurality of electronic devices (110, 120, 130, 140) include a smart phone, a mobile phone, a navigation system, a laptop, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a tablet PC, etc.

[0048] The communication method between the electronic devices (110, 120, 130, 140) and the servers (150, 160) is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcasting network) that the network (170) may include, but also short-range wireless communication between the devices. For example, the network (170) may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), the Internet, and the like. In addition, the network (170) may include any one or more of a network topology including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree, or a hierarchical network.

[0049] Each of the servers (150, 160) may be implemented as a computer device or multiple computer devices that communicate with multiple electronic devices (110, 120, 130, 140) via a network (170) to provide commands, codes, files, content, services, etc. For example, the server (150) may be a system that implements a learning data provision method based on a service platform installed and operated in multiple electronic devices (110, 120, 130, 140) connected via a network (170).

[0050] FIG. 2 is a block diagram illustrating an example of a computer device according to one embodiment. Each of the plurality of electronic devices (110, 120, 130, 140) or servers (150, 160) described above may be implemented by the computer device (200) illustrated in FIG. 2.

[0051] The computer device (200) may include a memory (210), a processor (220), a communication interface (230), and an input / output interface (240), as illustrated in FIG. 2. The memory (210) may be a computer-readable recording medium, and may include a non-permanent mass storage device such as a random access memory (RAM), a read only memory (ROM), and a disk drive. Here, the non-permanent mass storage device such as a ROM and a disk drive may be included in the computer device (200) as a separate permanent storage device distinct from the memory (210). In addition, the memory (210) may store an operating system and at least one program code. These software components may be loaded into the memory (210) from a computer-readable recording medium separate from the memory (210). The separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. In another embodiment, the software components may be loaded into the memory (210) via a communication interface (230) other than a computer-readable recording medium. For example, the software components may be loaded into the memory (210) of the computer device (200) based on a computer program installed by files received over a network (170).

[0052] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) via the memory (210) or the communication interface (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a storage device such as the memory (210).

[0053] The communication interface (230) may provide a function for the computer device (200) to communicate with other devices (e.g., the storage devices described above) via the network (170). For example, requests, commands, data, files, etc. generated by the processor (220) of the computer device (200) according to program codes stored in a recording device such as the memory (210) may be transmitted to other devices via the network (170) under the control of the communication interface (230). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (200) via the communication interface (230) of the computer device (200) via the network (170). Signals, commands, data, etc. received via the communication interface (230) may be transmitted to the processor (220) or the memory (210), and files, etc. may be stored in a storage medium that the computer device (200) may further include.

[0054] The input / output interface (240) may be a means for interfacing with an input / output device (250). For example, the input device may include a device such as a microphone, keyboard, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (240) may be a means for interfacing with a device that integrates input and output functions, such as a touchscreen. The input / output device (250) may also be configured as a single device with the computer device (200).

[0055] Furthermore, in other embodiments, the computer device (200) may include fewer or more components than those illustrated in FIG. 2. However, it is not necessary to explicitly illustrate most conventional components. For example, the computer device (200) may further include components such as a lidar, a scanner, a camera, and an IMU sensor to scan a target space and obtain scanned space information when implementing each of the plurality of electronic devices (110, 120, 130, and 140).

[0056] Below, specific embodiments of a method and system for providing learning data are described.

[0057]

[0058] FIG. 3 is a block diagram illustrating an example of components that the processor illustrated in FIG. 2 may include, FIG. 4 is a flowchart illustrating a learning data providing method that the computer device illustrated in FIG. 2 may perform, FIG. 5 is a drawing for explaining projecting at least one bounding box in the learning data providing method illustrated in FIG. 4, FIG. 6 is a drawing for explaining labeling at least one piece of label information in the learning data providing method illustrated in FIG. 4, and FIGS. 7 to 9 are drawings for explaining providing learning data in the learning data providing method illustrated in FIG. 4.

[0059] In embodiments of the present invention, the computer device (200) can execute a learning data provision method that detects an object in consecutive images and then generates and provides learning data for recognizing the object, in order to automatically generate and provide learning data for object recognition.

[0060] To this end, a learning data provision system, which is a subject performing a learning data provision method, may be configured in the computer device (200). For example, the learning data provision system may be implemented in the form of an independently operating program, or may be implemented in the form of an in-app of a dedicated application so as to be operable within the dedicated application.

[0061] The processor (220) of the computer device (200) may be implemented as a component for performing the learning data providing method according to FIG. 4. For example, the processor (220) may include an acquisition unit (310), a projection unit (320), a labeling unit (330), and a provision unit (340) as illustrated in FIG. 3 so as to perform steps (S410 to S440) illustrated in FIG. 4. Depending on the embodiment, the components of the processor (220) may be selectively included in or excluded from the processor (220). In addition, depending on the embodiment, the components of the processor (220) may be separated or merged to express the function of the processor (220).

[0062] These processors (220) and components of the processor (220) can control the computer device (200) to perform steps (S410 to S440) included in the learning data provision method of FIG. 4. For example, the processor (220) and components of the processor (220) can be implemented to execute instructions according to the code of the operating system included in the memory (210) and the code of at least one program.

[0063] Here, the components of the processor (220) may be representations of different functions performed by the processor (220) according to commands provided by the program code stored in the computer device (200). For example, the acquisition unit (310) may be used as a functional representation of the processor (220) that controls the computer device (200) to acquire consecutive images.

[0064] The processor (220) can read necessary commands from the memory (210) loaded with commands related to the control of the computer device (200). In this case, the read commands may include commands for controlling the processor (220) to execute steps (S410 to S440) to be described later.

[0065] In step (S410), the processor (220) (more precisely, the acquisition unit (310) included in the processor (220)) can acquire consecutive images (500).

[0066] At this time, before step (S410), the processor (220) (more precisely, the acquisition unit (310) included in the processor (220)) may provide a guide and template for photographing and scanning the target space to enable acquisition of consecutive images (500) using a portable terminal through a service platform on the portable terminal.

[0067] This configuration takes into account the fact that a portable terminal possessed by a non-professional user is different from a scanning terminal for a professional user. A guide and template for photographing and scanning a target space may be provided through a service platform on the portable terminal so that a non-professional user can photograph a target space through the portable terminal and obtain continuous images (500). For example, a guide guiding the user in the direction of photographing the target space may be provided through the service platform on the portable terminal, or a template indicating a photographing area of ​​the target space may be provided through the service platform on the portable terminal.

[0068] Additionally, in step (S410), the processor (220) (more precisely, the acquisition unit (310) included in the processor (220)) can acquire point cloud data corresponding to each of the consecutive images (500).

[0069] In step (S420), the processor (220) (more precisely, the projection unit (320) included in the processor (220)) can detect at least one target object (510, 520) having the same shape and depth in consecutive images and project at least one bounding box (511, 521) including the at least one target object (510, 520) onto each of the consecutive images.

[0070] More specifically, the processor (220) (more precisely, the projection unit (320) included in the processor (220)) can detect at least one target object (510, 520) having the same shape and depth in the consecutive images (500) by using point cloud data corresponding to each of the consecutive images (500).

[0071] At this time, the processor (220) (more precisely, the projection unit (320) included in the processor (220)) can exclude at least one recognizable object (530, 540, 550) from the consecutive images (500) in order to prevent the generation and provision of unnecessary learning data as illustrated in FIG. 5, and then detect at least one unrecognizable target object (510, 520) remaining after excluding at least one recognizable object (530, 540, 550) from the consecutive images (500).

[0072] Distinguishing between at least one recognizable object (530, 540, 550) and at least one unrecognizable target object (510, 520) can be accomplished using an object recognition artificial intelligence model that will provide learning data.

[0073] That is, the processor (220) (more precisely, the projection unit (320) included in the processor (220)) can exclude at least one recognizable object (530, 540, 550) that can be recognized using an object recognition artificial intelligence model among a plurality of objects (510, 520, 530, 540, 550) in the consecutive images (500), and then detect at least one unrecognizable target object (510, 520) that remains after excluding at least one recognizable object (530, 540, 550) from the consecutive images (500).

[0074] However, in another embodiment according to the present invention, the images (500) of at least one object having the movement can be recognized by various types such as adults, children (kids), infants, and / or pets such as dogs, and the distance moved or the movement status or the movement of the joint (which can be linked to a Kinect Sensor, etc.) for a set period of time can be recognized, and the object can be automatically classified by type or the object can be classified and processed (image removed) and stored in the computer device (200) through a computer device (200) connected to an external database according to the average value of the movement distance for a set period of time for each type that has been learned / stored in advance, and stored in the computer device (200), so that it can be used for future learning, etc.

[0075] In addition, according to another embodiment of the present invention, objects such as temporary stands, stalls, or partially-in-progress repair work (e.g., removal and reinstallation of part of floor tiles) installed in a space where an image is being filmed may also be provided to recognize changes or variations over time and remove them based on this, or output them to an input / output interface (240) so that a user can compare and select images before and after installation of the stand or before and after repair work. That is, in the case of an object that is fixedly positioned for a certain period of time or longer, unlike the aforementioned people or pets, the computer device (200) according to the present invention is fixedly positioned within the image for a set period of time, whereas it is not a permanent object like a building, and thus, depending on the user's selection, objects that are not mobile (people, etc.) or fixed (buildings) objects among the images located within the shooting range are separately classified (recognized based on whether they have moved or changed position within a set period of time or time), and in the case of such temporary installations (stands, signboards, objects subject to temporary maintenance, etc.), it is possible to selectively provide for detection, removal, or maintenance by type within the image depending on the user's selection.

[0076] In step (S430), the processor (220) (more precisely, the labeling unit (330) included in the processor (220)) may label at least one piece of label information (610, 620) for at least one target object (510, 520) in each of the consecutive images (500), as illustrated in FIG. 6. Hereinafter, the at least one piece of label information may include information to help an object recognition artificial intelligence model recognize what kind of object at least one target object is when provided as learning data in step (S430) described below. For example, the at least one piece of label information may include information about the shape of at least one target object and information about depth.

[0077] In step (S440), the processor (220) (more precisely, the provision unit (340) included in the processor (220)) may provide continuous images (500) on which at least one bounding box (511, 521) is projected and at least one label information (610, 620) is labeled, as shown in FIG. 7, as learning data (700) for recognition of at least one target object (510, 520).

[0078] When continuous images (500) on which at least one bounding box (511, 521) is projected and at least one label information (610, 620) is labeled are provided as learning data (700), a disadvantage may occur in that learning data (700) with a large data capacity is also provided because unnecessary objects (530, 540, 550) are included in the learning data (700).

[0079] Therefore, in step (S440), the processor (220) (more precisely, the provision unit (340) included in the processor (220)) may provide learning data (700) that includes only information about at least one target object (510, 520) that is the subject of object recognition learning, in order to prevent the above-described disadvantage.

[0080] For example, as illustrated in FIG. 8, in step (S440), the processor (220) (more precisely, the providing unit (340) included in the processor (220)) may downscale the resolution of the remaining area (505) except for at least one bounding box (511, 521) in each of the consecutive images (500), and then provide consecutive images (800) in which at least one bounding box (511, 521) is projected, at least one label information (610, 620) is labeled, and the resolution of the remaining area (505) is downscaled, as learning data (800) for recognition of at least one target object (510, 520).

[0081] For another example, as illustrated in FIG. 9, in step (S440), the processor (220) (more precisely, the providing unit (340) included in the processor (220)) may provide only at least one bounding box (511, 521) and at least one label information (610, 620) from each of the consecutive images (500) as learning data (900) for recognition of at least one target object (510, 520), rather than providing the entire consecutive images (500) as learning data (700) for recognition of at least one target object (510, 520).

[0082] Additionally, in step (S440), the processor (220) (more precisely, the provision unit (340) included in the processor (220)) may additionally provide point cloud data corresponding to each of the consecutive images (500) as learning data for recognition of at least one target object (510, 520).

[0083] As described, by detecting at least one target object (510, 520) in an image and automatically generating and providing learning data (700) for recognizing at least one target object (510, 520), the hassle of having to prepare learning data manually can be prevented.

[0084] In addition, by detecting only at least one unrecognizable target object (510, 520) excluding at least one recognizable object (530, 540, 550) and then generating and providing learning data (700) for at least one target object (510, 520), the disadvantage of the learning data (700) including unnecessary information and thus increasing the data capacity can be prevented.

[0085]

[0086] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0087] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0088] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording media or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0089] In one example for implementing the present invention, the invention may include a function for detecting the location and state of an object that changes in real time in an indoor environment and dynamically generating and updating learning data in accordance with the change.

[0090] When a camera or LiDAR sensor detects a change in an object's position, the system automatically analyzes this change and generates learning data appropriate for the new state. This real-time data generation and update capability maintains the object recognition model's up-to-dateness and enables rapid adaptation to changing environments. This process can be characterized by continuously tracking the object's movement and automatically adjusting data generation rules based on changes in the object's state, as needed.

[0091] In one example for implementing the present invention, a function may be included for automatically generating a label and assigning it to learning data during an object recognition process.

[0092] Labels generated by analyzing the shape, size, location, color, etc. of objects ensure consistent learning data and minimize manual labeling errors that can occur, especially in large-scale datasets.

[0093] Additionally, one embodiment of the present invention may include a function for automatically filtering out unnecessary information when generating learning data. This may include detecting and removing background information or repetitive patterns unnecessary for learning an object recognition model. In an indoor environment, repetitive tile patterns on the floor or wall textures may be considered unnecessary data for learning and thus filtered out.

[0094] In one embodiment of the present invention, a function may be included to simultaneously recognize multiple objects in an indoor environment and reflect interactions or positional relationships between these objects in learning data.

[0095] When recognizing a scene where a person is sitting on a chair, we can help the model learn this relationship by including the spatial relationship between the person and the chair in the training data.

[0096] Additionally, one embodiment of the present invention may include a function for automatically performing data augmentation to increase the diversity of training data. By generating new training data by transforming existing training data into various angles, brightness, colors, and sizes, the generalization performance of the model can be improved.

[0097] In one embodiment of the present invention, a function for updating an object recognition model in real time using new data collected during the learning data generation process may be included. This allows the model to quickly adapt to new objects or environmental changes, continuously improving its accuracy and reliability.

[0098] By learning in real time information about new types of objects or changed object states detected during learning data generation and reflecting this information in the object recognition model, the model can be strengthened in real time.

[0099] Additionally, one embodiment of the present invention may include a function for improving training data by incorporating user feedback. If the user evaluates or inputs corrections to the object recognition results, the system can use this feedback to readjust the training data and further refine the object recognition model.

[0100] In one embodiment of the present invention, a function for generating training data by integrating data collected from various sensors, such as a LiDAR, an infrared sensor, and a depth sensor, in addition to a camera, may be included. This function combines various sensor data to enable more accurate and comprehensive object recognition, and can particularly contribute to improving the performance of 3D object recognition models.

[0101] Additionally, one embodiment of the present invention may include a function for generating learning data by understanding the context of an object. By analyzing the surrounding environment or situational context of an object and generating data useful for object recognition within that context, learning can be achieved that considers not only the object itself but also the relationships between objects.

[0102] In one embodiment of the present invention, a function may be included to continuously monitor the quality of training data and automatically correct or issue an alert if a quality decline is detected. This ensures that training data is always maintained in optimal condition and prevents data quality degradation over time.

[0103] If a specific label in the training data becomes inconsistent or noisy, the system can automatically detect this and take necessary actions.

[0104] Additionally, one embodiment of the present invention may include a function for modularizing and managing learning data. This can be implemented by dividing the dataset into multiple modules and updating or retraining only specific modules as needed. This allows for efficient management of learning data, and when data for a specific object needs to be updated, only that module can be updated instead of retraining the entire dataset.

[0105] In one embodiment for carrying out the present invention, a learning data providing method performed by a computer device may be characterized by including the steps of acquiring consecutive images, detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images, labeling at least one label information for the at least one target object onto each of the consecutive images, and providing the consecutive images on which the at least one bounding box is projected and the at least one label information is labeled as learning data for recognition of the at least one target object.

[0106] In one embodiment for carrying out the present invention, the projecting step may include a learning data providing method characterized in that it includes a step of detecting at least one target object having the same shape and depth in the consecutive images using point cloud data corresponding to each of the consecutive images.

[0107] In one embodiment for carrying out the present invention, the method for providing learning data may be characterized in that the acquiring step further includes a step of acquiring the point cloud data corresponding to each of the consecutive images.

[0108] In one embodiment for carrying out the present invention, the step of projecting may include a learning data providing method characterized in that it includes a step of excluding at least one recognizable object from the consecutive images; and a step of detecting at least one unrecognizable target object remaining after excluding the at least one recognizable object from the consecutive images.

[0109] In one embodiment for carrying out the present invention, the method for providing learning data may be characterized in that the at least one label information includes information on the shape and information on the depth of the at least one target object.

[0110] In one embodiment for carrying out the present invention, the providing step may further include a learning data providing method characterized in that it includes a step of additionally providing the point cloud data corresponding to each of the consecutive images as the learning data for recognition of the at least one target object.

[0111] In one embodiment for carrying out the present invention, the providing step may include a step of downscaling the resolution of the remaining area excluding the at least one bounding box in each of the continuous images; and a step of providing the continuous images, in which the at least one bounding box is projected, the at least one label information is labeled, and the resolution of the remaining area is downscaled, as the learning data for recognizing the at least one target object.

[0112] In one embodiment for carrying out the present invention, the providing step may include a learning data providing method characterized in that the providing step is a step of providing only the at least one bounding box and the at least one label information from each of the consecutive images as the learning data for recognizing the at least one target object.

[0113] In one embodiment for carrying out the present invention, a computer-readable recording medium having recorded thereon a computer program for executing a learning data providing method on a computer device may be characterized by including a step of acquiring consecutive images, a step of detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images, a step of labeling at least one label information for the at least one target object onto each of the consecutive images, and a step of providing the consecutive images on which the at least one bounding box is projected and on which the at least one label information is labeled as learning data for recognizing the at least one target object.

[0114] In one embodiment for carrying out the present invention, a computer device for performing a learning data providing method may be characterized by including at least one processor configured to execute computer-readable instructions, wherein the at least one processor includes an acquiring unit for acquiring consecutive images, a projection unit for detecting at least one target object having the same shape and depth in the consecutive images and projecting at least one bounding box including the at least one target object onto each of the consecutive images, a labeling unit for labeling at least one label information for the at least one target object onto each of the consecutive images, and a providing unit for providing the consecutive images on which the at least one bounding box is projected and on which the at least one label information is labeled as learning data for recognizing the at least one target object.

[0115] In one embodiment for carrying out the present invention, the projection unit may include a computer device characterized in that it detects at least one target object having the same shape and depth in the consecutive images by using point cloud data corresponding to each of the consecutive images.

[0116] In one embodiment for carrying out the present invention, the projection unit may include a computer device characterized in that it excludes at least one recognizable object from the consecutive images, thereby detecting at least one unrecognizable target object remaining after excluding the at least one recognizable object from the consecutive images.

[0117] In one embodiment for carrying out the present invention, the computer device may be characterized in that the at least one label information includes information about the shape of the at least one target object and information about the depth.

[0118] In another embodiment for carrying out the present invention, a method for providing learning data performed by a computer device comprises the steps of: obtaining continuous images by using a portable terminal used by a non-expert, using a guide and template for photographing and scanning a target space so that continuous images can be obtained, through a service platform on the portable terminal; detecting at least one target object having the same shape and depth in the continuous images and projecting at least one bounding box including the at least one target object onto each of the continuous images; labeling at least one label information for the at least one target object onto each of the continuous images; And characterized in that it comprises a step of providing the continuous images on which the at least one bounding box is projected and the at least one label information is labeled as learning data for recognizing the at least one target object, wherein the acquiring step further comprises a step of acquiring point cloud data corresponding to each of the continuous images, and the projecting step comprises a step of detecting the at least one target object having the same shape and depth in the continuous images using the point cloud data corresponding to each of the continuous images, wherein the projecting step further comprises a step of excluding the at least one recognizable object from the continuous images and a step of detecting the at least one unrecognizable target object remaining after excluding the at least one recognizable object from the continuous images, and wherein the projecting step is characterized in that distinguishing the at least one recognizable object from the at least one unrecognizable target object in the continuous images is performed through an object recognition artificial intelligence model providing learning data, and wherein the providing step comprises:The method may further include a step of additionally providing the point cloud data corresponding to each of the consecutive images as the learning data for recognizing the at least one target object, and a step of automatically generating and providing learning data for recognizing the at least one target object, wherein the at least one recognizable object is recognized by a distance moved and a state of movement or a movement of a joint for a set period of time, and is classified, processed (image removed), and stored by the computer device connected to an external database or automatically classified by type according to an average value of the moving distance or a movement pattern for a set period of time for each type that has been learned / stored in advance in the computer device, and wherein the computer device recognizes changes and variations over time and removes them based on the same, or separately types objects that are not moving or fixed objects among images located within the shooting range according to a user's selection, and in the case of temporary installations, selectively provides detection, removal, or maintenance by type within the image according to a user's selection.

[0119] In another embodiment for carrying out the present invention, the method for providing learning data may be characterized in that the at least one label information includes information on the shape and information on the depth of the at least one target object.

[0120] In another embodiment for carrying out the present invention, the providing step may be a learning data providing method characterized in that it includes a step of downscaling the resolution of the remaining area excluding the at least one bounding box in each of the continuous images, and a step of providing the continuous images in which the at least one bounding box is projected, the at least one label information is labeled, and the resolution of the remaining area is downscaled, as the learning data for recognizing the at least one target object.

[0121] In another embodiment for carrying out the present invention, the providing step may be a learning data providing method characterized in that the providing step is a step of providing only the at least one bounding box and the at least one label information from each of the consecutive images as the learning data for recognizing the at least one target object.

[0122] In another embodiment for carrying out the present invention, a computer device for performing a learning data providing method comprises at least one processor configured to execute computer-readable commands, wherein the at least one processor comprises: an acquisition unit for photographing and scanning a target space so that a non-expert can obtain continuous images using a portable terminal, and a guide and template for obtaining the continuous images through a service platform on the portable terminal; a projection unit for detecting at least one target object having the same shape and depth in the continuous images and projecting at least one bounding box including the at least one target object onto each of the continuous images; a labeling unit for labeling at least one label information for the at least one target object onto each of the continuous images; and a providing unit for providing the continuous images on which the at least one bounding box is projected and on which the at least one label information is labeled, as learning data for recognition of the at least one target object, wherein the providing unit additionally provides point cloud data corresponding to each of the continuous images as the learning data for recognition of the at least one target object, and the learning data for recognition of the at least one target object is automatically generated and provided; In each of the consecutive images, the resolution of the remaining area except for the at least one bounding box is downscaled, and the consecutive images in which the at least one bounding box is projected, the at least one label information is labeled, and the resolution of the remaining area is downscaled are provided as the learning data for recognizing the at least one target object, and in each of the consecutive images, only the at least one bounding box and the at least one label information are provided as the learning data for recognizing the at least one target object, wherein the at least one target object is,The computer device may be characterized in that it is recognized as a distance moved and a state of movement or a movement of a joint for a set period of time, and is automatically classified by type or typed, classified, processed (image removed) and stored through the computer device connected to an external database according to an average value or movement pattern of the distance moved for a set period of time for each type that has been learned / stored in advance in the computer device, and the computer device recognizes changes and variations over time and removes them based on the same, or separately types objects that are not moving or fixed objects among images located within the shooting range according to the user's selection, and in the case of temporary installations, it is characterized in that it selectively provides detection, removal or maintenance by type within the image according to the user's selection.

[0123] In another embodiment for carrying out the present invention, the projection unit may be a computer device characterized in that it detects at least one target object having the same shape and depth in the consecutive images by using point cloud data corresponding to each of the consecutive images.

[0124] In another embodiment for carrying out the present invention, the projection unit may be a computer device characterized in that it excludes at least one recognizable object from the consecutive images, thereby detecting the at least one unrecognizable target object remaining after excluding the at least one recognizable object from the consecutive images, and wherein distinguishing between the at least one recognizable object from the at least one unrecognizable target object in the consecutive images is performed through an object recognition artificial intelligence model that provides learning data.

[0125] In another embodiment for carrying out the present invention, the computer device may be characterized in that the at least one label information includes information about the shape of the at least one target object and information about the depth.

[0126] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0127] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

[0128] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single entity may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.

[0129] The scope of the present invention is indicated by the claims set forth below, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present invention.

[0130]

[0131] The mode for carrying out the invention has been described together with the best mode for carrying out the invention above.

[0132] The present invention relates to a technology for collecting and providing learning data for objects, and more specifically, to a method and system for detecting an object in an image and generating and providing learning data for recognizing the object, and to a method and system for detecting a target object for learning data by excluding at least one recognizable object, thereby preventing the generation and provision of unnecessary learning data, thereby having industrial applicability.

Claims

1. A method for providing learning data performed by a computer device, A step of acquiring sequential images; A step of detecting at least one target object having the same shape and depth in the sequential images and projecting at least one bounding box including the at least one target object onto each of the sequential images; A step of labeling at least one label information for at least one target object to each of the above consecutive images; and A step of providing the continuous images on which at least one bounding box is projected and on which at least one label information is labeled as learning data for recognition of at least one target object. A method of providing learning data including:

2. In paragraph 1, The above projecting step is, A step of detecting at least one target object having the same shape and depth in the continuous images by using point cloud data corresponding to each of the continuous images. A method for providing learning data, characterized by including:

3. In paragraph 2, The above obtaining steps are: A step of obtaining the point cloud data corresponding to each of the above consecutive images. A method for providing learning data, characterized by further including:

4. In paragraph 1, The above projecting step is, A step of excluding at least one recognizable object from the above consecutive images; and A step of excluding at least one recognizable object from the above continuous images and detecting at least one unrecognizable target object remaining. A method for providing learning data, characterized by including:

5. In paragraph 1, At least one of the label information above, A method for providing learning data, characterized in that it includes information about the shape and information about the depth of at least one target object.

6. In paragraph 1, The steps provided above are: A step of additionally providing the point cloud data corresponding to each of the above consecutive images as the learning data for recognition of at least one target object. A method for providing learning data, characterized by further including:

7. In paragraph 1, The steps provided above are: A step of downscaling the resolution of the remaining area excluding at least one bounding box in each of the above consecutive images; and A step of providing the continuous images, in which at least one bounding box is projected, at least one label information is labeled, and the resolution of the remaining area is downscaled, as the learning data for recognition of the at least one target object. A method for providing learning data, characterized by including:

8. In paragraph 1, The steps provided above are: A method for providing learning data, characterized in that the step of providing only the at least one bounding box and the at least one label information from each of the above consecutive images as the learning data for recognition of the at least one target object.

9. A computer-readable recording medium having recorded thereon a computer program for executing a method of providing learning data on a computer device, The above learning data provision method is: A step of acquiring sequential images; A step of detecting at least one target object having the same shape and depth in the sequential images and projecting at least one bounding box including the at least one target object onto each of the sequential images; A step of labeling at least one label information for at least one target object to each of the above consecutive images; and A step of providing the continuous images on which at least one bounding box is projected and on which at least one label information is labeled as learning data for recognition of at least one target object. A computer-readable recording medium containing:

10. In a computer device performing a method of providing learning data, At least one processor configured to execute computer-readable instructions Including, At least one processor of the above, An acquisition unit that acquires sequential images; A projection unit that detects at least one target object having the same shape and depth in the consecutive images and projects at least one bounding box including the at least one target object onto each of the consecutive images; A labeling unit that labels at least one label information for at least one target object to each of the above consecutive images; and A providing unit that provides the continuous images on which at least one bounding box is projected and on which at least one label information is labeled as learning data for recognition of at least one target object. A computer device comprising:

11. In paragraph 10, The above projection part, A computer device characterized in that it detects at least one target object having the same shape and depth in the consecutive images by using point cloud data corresponding to each of the consecutive images.

12. In paragraph 10, The above projection part, A computer device characterized in that it excludes at least one recognizable object from the sequential images, thereby detecting the at least one unrecognizable target object remaining after excluding the at least one recognizable object from the sequential images.

13. In paragraph 10, At least one of the label information above, A computer device characterized by including information about the shape and information about the depth of at least one target object.

Citation Information

Patent Citations

  • A method and an apparatus for video decoding, a method and an apparatus for video encoding

    KR1020210025565A

  • Blockchain-base Community comprehensive care system

    KR1020240108730A

  • Chlorine generating electrode and method of manufacturing same

    KR1020240118524A

  • Semiconductor device and method of fabricating the same

    KR102782256B1

  • KR20190124559A