Electronic device for supporting image segmentation
The electronic device enhances image segmentation models by clustering pixels based on entropy and depth information, reducing labeling costs and improving reliability through efficient human intervention and contrastive learning, addressing the limitations of RGB-D data labeling.
Patent Information
- Application Number
- PCT/KR2024/096747
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing image segmentation models using RGB-D data face challenges due to the high cost and time required for labeling large-scale datasets, leading to limited performance improvement and reliability issues, especially when relying on small, labeled real-environment datasets.
An electronic device employs a framework that clusters pixels based on entropy and depth information, allowing for efficient labeling by human experts on high-uncertainty regions and using model inferences for low-uncertainty regions, combined with contrastive learning to enhance model reliability.
This approach reduces labeling costs and time while improving the recognition rate and reliability of image segmentation models by leveraging pixel clusters and entropy-based labeling, enabling more accurate object recognition.
Smart Images

Figure KR2024096747_24072025_PF_FP_ABST
Abstract
Description
Electronic devices to support image segmentation
[0001] The present disclosure relates to an electronic device for supporting image segmentation using an artificial intelligence model.
[0002] Image segmentation, in the field of image recognition, can include techniques for recognizing object boundaries. For example, image segmentation can involve classifying each pixel into a specific class. This technology can be utilized in various fields, including robotics, medicine, and autonomous driving. For example, in industrial robotics applications, image segmentation technology can be used to recognize and classify objects in the environment in which the robot operates. For example, in a manufacturing process, a robot can recognize and classify products, performing appropriate processes for each product. Furthermore, the robot can recognize hazardous objects in its environment and avoid or remove them.
[0003] In applications related to industrial robots, accurate and reliable object segmentation technology is crucial. To achieve this, datasets relevant to the application environment are collected, and these datasets can be used to train an AI model for image segmentation.
[0004] To segment the boundary regions of each object in an image, models can be trained using large-scale RGB datasets or RGB-D (depth) datasets with small labels. However, large-scale RGB datasets primarily contain visual information and relatively little consideration of geometric information. This can limit the model's ability to fully understand the spatial relationships of objects within the environment. Collecting and labeling RGB-D datasets is expensive and time-consuming. Therefore, RGB-D datasets with small labels can be used for model training. However, this strategy struggles to improve model performance due to dataset size limitations. Consequently, to improve the recognition rate of region segmentation models, an automated RGB-D dataset construction pipeline and efficient model training techniques using the resulting dataset are crucial.
[0005] The above information is provided as background information to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0006] RGB semantic segmentation is a technique that uses RGB information in an image to segment each pixel into objects or classes. Electronic devices can use this technique to identify and infer the boundaries and locations of objects within an image. Recently, widely used AI models supporting RGB semantic segmentation can be implemented in various ways, including convolutional neural networks, encoder-decoder networks, and attention-based models.
[0007] RGB-D segmentation is a technique that segments object boundaries using depth information along with RGB information in an image. RGB information can include information about the object's color and brightness. Depth information can also include information about the object's geometric shape and distance. Considering these characteristics, RGB-D segmentation, by utilizing both RGB and depth information, can segment object boundaries more accurately than RGB semantic segmentation. Consequently, utilizing depth information in addition to RGB for region segmentation enables more accurate object segmentation and environmental recognition. These advantages have led to its widespread use in applications such as robotics and autonomous vehicles. Representative research techniques include ShapeConv and Depth-aware CNN, and the core of the technique proposed in this paper is the complementary fusion of RGB and depth information. AI models for RGB-D segmentation typically require more complex structures than RGB segmentation models, resulting in computational complexity and longer training times.
[0008] Although utilizing depth information for segmentation offers advantages in recognition rates, collecting RGB-D image datasets requires separate equipment and is more difficult than RGB image datasets. Most large-scale labeled RGB-D datasets are synthetic, while RGB-D datasets captured in real-world environments are typically small due to the high cost of labeling.
[0009] Because RGB-D data includes color, geometric shape, and distance information, it can achieve more accurate object recognition results than RGB data alone. However, the use of RGB-D data has the following limitations.
[0010] Large-scale labeled RGB-D datasets are scarce. Image segmentation labeling involves assigning labels to pixels corresponding to objects (e.g., sidewalks, vehicles, people, traffic lights, lanes) to identify them. This pixel-by-pixel labeling can be time-consuming and expensive. Furthermore, obtaining objective and accurate labels is challenging. Consequently, labeled RGB-D datasets in real-world environments are rare or small in scale.
[0011] Techniques to reduce the cost of labeling can involve AI models measuring entropy (or uncertainty) for each pixel in an image. Entropy can be a measure of how confident the model is in its inferences. Pixels with high entropy based on the model's inferences can be assigned labels (e.g., true labels) by human experts (annotators), while pixels with low uncertainty can be assigned labels inferred by the model (e.g., pseudo labels). Measuring entropy for each pixel can be sensitive to subtle variations or noise within the image. Furthermore, annotators must make numerous decisions to label numerous small regions distributed across pixels, which can lead to significant labeling time and increased labeling discrepancies among annotators.
[0012] When labeling a dataset using a trained model, the higher the model's discriminative power, the easier it is to automate labeling. During model training, the model can be trained with only a small number of labeled samples secured within the labeling budget. A small labeling budget makes it difficult to improve the model's discriminative power, requiring more human expert intervention.
[0013] Electronic devices can recognize objects using AI models trained on labeled RGB-D data sets. For example, the electronic device can be a device for object recognition (e.g., a robot) or a device connected to it (e.g., a server or personal computer). With a limited number of labels, the reliability of the model's output may be low.
[0014] Various embodiments can improve the object recognition rate of models by providing a framework for labeling RGB-D datasets and model training based on collaboration between humans and models. Large, easily obtainable, unlabeled RGB-D datasets can be used for model training. By labeling the model by pixel clusters rather than pixels, labeling of RGB-D datasets can be completed with less time and effort.
[0015] Various embodiments may provide an electronic device that labels images and retrains a model based on the labeled images, thereby improving the reliability of the output produced by the model.
[0016] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.
[0017] According to one embodiment, an electronic device includes a processor; and a memory that stores a model learned using an RGB-D (depth) dataset for image segmentation. The command, when executed by the processor, may cause the electronic device to obtain entropy information including entropies corresponding to pixels in a first image having RGB information including RGB values corresponding to pixels respectively and depth information including depth values corresponding to the pixels respectively. The command, when executed by the processor, may cause the electronic device to cluster the first image into a plurality of clusters based on the entropy information, the RGB information, and the depth information. The command, when executed by the processor, may cause the electronic device to classify the clusters into a first region and a second region based on the entropy information. The above instructions, when executed by the processor, may cause the electronic device to assign a label to at least one of the clusters classified into the first region based on a user input. The instructions, when executed by the processor, may cause the electronic device to assign a label to a cluster classified into the second region based on an inference of the model. The instructions, when executed by the processor, may cause the electronic device to train the model using a second image having a cluster assigned a label based on the user input and a cluster assigned a label based on an inference of the model.
[0018] The above command, when executed by the processor, may cause the electronic device to input the first image as an input value to the model and obtain the entropy information from an output value output from the model.
[0019] The above command, when executed by the processor, may cause the electronic device to generate a feature map in which the RGB information and the depth information are fused, and to cluster the first image into a plurality of clusters using the feature map and the entropy information.
[0020] The above command, when executed by the processor, may cause the electronic device to classify the third image into a third region that is trustworthy and a fourth region that is not trustworthy, and to perform comparison learning on the model using information about the fourth region.
[0021] The above command, when executed by the processor, may cause the electronic device to input the third image as an input value to the model and obtain information about the third region and the fourth region from an output value output from the model.
[0022] The above command, when executed by the processor, may cause the electronic device to classify the third image into a plurality of classes based on the acquired information and select anchor pixels from the plurality of classes. The above command, when executed by the processor, may cause the electronic device to select positive samples and negative samples for each of the plurality of classes based on the anchor pixels. The above command, when executed by the processor, may cause the electronic device to train the model to position the positive samples closer to the corresponding class and to position the negative samples further from the corresponding class.
[0023] According to one embodiment, a method of operating an electronic device is provided. The method may include obtaining entropy information including entropies corresponding to pixels from a first image having RGB information including RGB values corresponding to pixels, respectively, and depth information including depth values corresponding to the pixels, respectively. The method may include clustering the first image into a plurality of clusters based on the entropy information, the RGB information, and the depth information. The method may include classifying the clusters into a first region and a second region based on the entropy information. The method may assign a label to at least one of the clusters classified into the first region based on a user input. The method may include assigning a label to a cluster classified into the second region based on an inference of the model. The method may include training the model using a second image having a cluster assigned a label based on the user input and a cluster assigned a label based on an inference of an artificial intelligence model.
[0024] The operation of obtaining the entropy information may include an operation of inputting the first image as an input value to the model and obtaining the entropy information from an output value output from the model.
[0025] The clustering operation may include an operation of generating a feature map in which the RGB information and the depth information are fused; and an operation of clustering the first image into a plurality of clusters using the feature map and the entropy information.
[0026] The method may further include an operation of classifying a third image into a trustworthy third region and an untrustworthy fourth region; and an operation of performing contrastive learning on the model using information about the fourth region.
[0027] The above classification operation may include an operation of inputting the third image as an input value to the model and obtaining information about the third region and the fourth region from an output value output from the model.
[0028] According to various embodiments of the present disclosure, labeling costs can be reduced. By utilizing the model to label large-scale RGB-D datasets without labels, the labor and time required for labeling can be reduced.
[0029] According to various embodiments of the present disclosure, while individual pixel-level entropy measurements may be sensitive to subtle variations or noise within an image, clusters of pixels are relatively insensitive to such noise, increasing the human interpretability of each region.
[0030] According to various embodiments of the present disclosure, the use of pixel clusters simplifies the annotation process by allowing annotators to label many pixels in less time.
[0031] According to various embodiments of the present disclosure, pixel clustering considering RGB information, depth information, and entropy can naturally highlight information-rich regions in an image. This approach can help a learning system more efficiently recognize important patterns and prevent resources from being wasted on analyzing regions of relatively low importance.
[0032] According to various embodiments of the present disclosure, the higher the model's discriminative power, the easier automated labeling can be. Consequently, an image segmentation model with a high recognition rate can be provided.
[0033] According to various embodiments of the present disclosure, pixels with high entropy (uncertainty) can be constructed as a negative sample queue. The negative sample queue can be utilized during contrastive learning. This can enhance the reliability of the model's output. A model with higher accuracy can provide higher reliability (lower uncertainty) for multiple unlabeled RGB-D datasets, thereby reducing the labor and time required for labeling.
[0034] In addition, various effects may be provided, either directly or indirectly, through this document.
[0035] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0036] FIG. 2 illustrates a framework for labeling images and training a model based on the labeled images, according to one embodiment.
[0037] FIG. 3 is a flowchart illustrating operations supporting learning of an image segmentation model according to one embodiment.
[0038] FIG. 4 is a block diagram of a fusion module configured to obtain an RGB-D fused feature map according to one embodiment.
[0039] FIG. 5 is a schematic diagram of a spatial-wise cross-modal attention block according to one embodiment.
[0040] FIG. 6 is a flowchart illustrating operations for clustering an image into multiple clusters according to one embodiment.
[0041] FIG. 7 is a flowchart illustrating operations for supporting contrastive learning of an artificial intelligence model according to one embodiment.
[0042] FIG. 8 is a diagram for explaining an operation for managing a voice sample according to one embodiment.
[0043] FIG. 9 is a flowchart illustrating operations for supporting contrastive learning of an artificial intelligence model according to one embodiment.
[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0045] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0046] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0047] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0048] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0049] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0050] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0051] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0052] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0053] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0054] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0055] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0056] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0057] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0058] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0059] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0060] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0061] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0062] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0063] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0064] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0065] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0066] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0067] FIG. 2 illustrates a framework for labeling images and training a model based on the labeled images, according to one embodiment. The framework presented in this document may be configured in an electronic device (e.g., the electronic device (101) or the server (108) of FIG. 1 ). Referring to FIG. 2 , the framework may include an image segmentation model (201), a clustering module (203), a classification module (205), and a labeling module (207). At least one of the image segmentation model (201), the clustering module (203), the classification module (205), and the labeling module (207) may be stored as instructions in a memory (e.g., the memory (130) of FIG. 1) and executed by a processor (e.g., the processor (120) of FIG. 1). At least one of the image segmentation model (201), the clustering module (203), the classification module (205), and the labeling module (207) (e.g., the image segmentation model (201)) may be executed on a processor specialized in processing artificial intelligence models (e.g., a GPU (graphic processing unit)).
[0068] The image segmentation model (201) may include an RGB-based model (e.g., a foundation model) trained using a large-scale RGB dataset (or RGB-D dataset) that is easily acquired. For example, an unlabeled dataset may be used when training the image segmentation model (201). As another example, the dataset used when training the image segmentation model (201) may include labeled images according to an embodiment of the present disclosure.
[0069] The image segmentation model (201) can receive an image (210) as an input value. The image (210) can include RGB information (211) and depth information (213). The RGB information (211) can include RGB values corresponding to each pixel in the image (210). The depth information (213) can include depth values corresponding to each pixel in the image (210).
[0070] The image segmentation model (201) can, in response to receiving an image (210) as input, output a result obtained by the model (201) through inference for each pixel in the image (210). The inference result can include, for example, identification information of objects (e.g., people, roads, vehicles, and sidewalks, etc.) at the pixel level (in other words, pixel-level).
[0071] The image segmentation model (201) can output an inference result including pixel-level entropy information (220). The pixel-level entropy information (220) can include entropies corresponding to each pixel in the image (210). The entropy can include information (e.g., probability or entropy score) indicating how confident the model is about the inference result for the corresponding pixel. For example, if the inference result pixel belongs to a 'person' object, the entropy can be the probability that the pixel is a person.
[0072] The image segmentation model (201) may include a fusion module (203). The fusion module (202) may receive RGB information (211) and depth information (213) as input values. In response to receiving the RGB information (211) and depth information (213), the fusion module (202) may output a feature map (230) as a result obtained by fusing the RGB information (211) and depth information (213).
[0073] The clustering module (203) can cluster an image (210) into multiple clusters using pixel-level entropy information (220) and feature map (230) (or corresponding RGB information (211) and depth information (213)).
[0074] The classification module (205) can classify clusters into a first region (241) with high entropy and a second region (242) with low entropy based on pixel-level entropy information (220). For example, if entropy is a probability or score indicating the reliability of an inference result, the clustering module (203) can obtain the entropy average (probability average or score average) of the pixels of the cluster. If the average is less than a specified threshold, the clustering module (203) can classify the cluster into a first region (241) with high uncertainty (entropy). If the average is greater than or equal to the threshold, the clustering module (203) can classify the cluster into a second region (243) with low uncertainty (entropy).
[0075] The labeling module (207) can assign a true label to a cluster (251) classified as the first region (241) based on user input. The operation of assigning a true label can be performed based on a preset labeling budget. For example, if the set budget is 5%, clusters within 5% of the clusters classified as the first region (241) can be selected, and a user (expert) can input a true label (identification information) for the selected clusters into an electronic device. The labeling module (207) can set the user input (true label) as the label for the corresponding cluster.
[0076] The labeling module (207) can assign a pseudo label, which is considered a correct label, to a cluster (252) classified into the second region (242), based on the inference result (identification information) of the model (201).
[0077] If the ratio of similar labels to correct labels is high in the labeling module (207), the cost for human intervention labeling can be reduced. If the discrimination power (ability to recognize objects in an image) of the image segmentation model (201) is high, the ratio of similar labels can be high. To increase the discrimination power, labeling in the labeling module (207) may be performed in several stages rather than all at once. For example, if the allocated budget is 5%, the ratio of clusters selected for assigning correct labels among the clusters classified into the first region (241) may be gradually increased, such as 1% → 2% → 3% → 4% → 5%. At each stage, pixel-level entropy may be measured for the RGB-D dataset based on the model (201) learned so far. After measurement, similar adjacent pixels are merged by the clustering module (203), and the image can be broadly classified into two regions (a first region with high uncertainty and a second region with low uncertainty) by the classification module (205).
[0078] An image having a cluster labeled by a labeling module (207) can be used as training data for training an image segmentation model (201). The labeling module (207) can provide the image to the image segmentation model (201). The image segmentation model (201) can perform training for image segmentation using the image. In this way, by training the image segmentation model (201) based on the RGB-D dataset having labels, the reliability of the results inferred by the image segmentation model (201) can be gradually improved.
[0079] In the image (210), the first region (241) can be used as negative samples in contrastive learning. The classification module (205) can provide information about the first region (241) to the image segmentation model (201). The image segmentation model (201) can perform contrastive learning using information about the negative samples obtained through the classification module (205). Contrastive learning is a machine learning method that aims to learn similarities and differences between data points, thereby bringing similar data points closer together and pushing different data points away from each other. In contrastive learning, an anchor sample is a reference data point that serves as a basis for comparison. Anchor samples generally refer to instances currently being analyzed. Positive samples are data points with similar characteristics to the anchor sample. Positive samples are data recognized as belonging to the same class as the anchor sample or having similar characteristics or patterns to the anchor sample. Negative samples are data points with characteristics different from the anchor sample. Negative samples are data that are recognized as belonging to a different class from the anchor sample or having different characteristics or patterns than the anchor sample. The goal of contrastive learning is to minimize the distance between anchor-positive pairs and maximize the distance between anchor-negative pairs. Through contrastive learning, the image segmentation model (201) can extract important features from complex image representations. By performing contrastive learning based on negative samples in the image segmentation model (201), the reliability of the results inferred by the image segmentation model (201) can be gradually improved.
[0080] FIG. 3 is a flowchart illustrating operations supporting training of an image segmentation model (201; see FIG. 2 ) according to one embodiment. When instructions stored in a memory (e.g., including modules 202, 203, 205, and 207 of FIG. 2 ) are executed by a processor, the operations of FIG. 3 may be performed by an electronic device (e.g., the electronic device 201 of FIG. 1 ). Before the operations of FIG. 3 begin, a maximum labeling budget ratio that a human expert can label may be set. Thereafter, an artificial intelligence model (e.g., the image segmentation model (201)) may be trained using large-scale RGB data with labels that are easy to obtain. Inference may be performed on large-scale RGB-D data without labels using the trained model.
[0081] In operation 310, the electronic device may obtain entropies corresponding to pixels in a first image (e.g., image (210) of FIG. 2) having RGB information and depth information. According to one embodiment, the electronic device may input the first image as an input value to an artificial intelligence model (e.g., image segmentation model (201) of FIG. 2) and obtain entropy information (e.g., pixel-level entropy information (220)) including entropies corresponding to pixels in an output value output from the artificial intelligence model. According to one embodiment, the electronic device may obtain entropy H of the j-th pixel in image i through the following mathematical equations 1 and 2. i,j can be obtained. In mathematical equations 1 and 2, S is a softmax function, e is a natural constant, and k represents the class number.
[0082]
[0083]
[0084] In operation 320, the electronic device may cluster the first image into a plurality of clusters based on RGB information, depth information, and entropy information. According to one embodiment, the electronic device may generate a feature map by fusing the RGB information and depth information, and perform clustering using the feature map and entropy information.
[0085] In operation 330, the electronic device can classify clusters into a first region with high entropy (e.g., the first region (241) of FIG. 2) and a second region with low entropy (e.g., the second region (242) of FIG. 2) based on entropy information.
[0086] In operation 340, the electronic device may assign a label (e.g., a correct label) to at least one of the clusters classified as a first region based on user input, and may assign a label (e.g., a similar label) to a cluster classified as a second region based on model inference. The user input-based label assignment may be performed within a set budget. For example, if the budget is 5%, label assignment may be performed for no more than 5% of the clusters classified as a first region.
[0087] At operation 350, the electronic device can retrain the model using a second image having labeled clusters based on user input and labeled clusters based on model inference.
[0088] The electronic device may additionally perform contrastive learning on the model. For example, in operation 360, the electronic device may perform contrastive learning on the model using speech samples obtained from the first region.
[0089] Thereafter, the electronic device can perform image segmentation for object recognition using the model learned through the motions 350 and / or 360.
[0090] FIG. 4 is a block diagram of a fusion module (400) configured to obtain an RGB-D fused feature map according to one embodiment. Referring to FIG. 4, the fusion module (400) (e.g., the fusion module (202) of FIG. 2) may include a context feature block (hereinafter, a first block) (410), a spatial-wise cross-modal attention block (hereinafter, a second block) (420), and a channel-wise feature aggregation block (hereinafter, a third block) (430). The fusion module (400) may include a plurality of layers. Each layer may include a first block (410), a second block (420), and a third block (430). A clustering module (e.g., the clustering module (203) of FIG. 2) may utilize a result (feature map) output from the last layer.
[0091] The first block (410) can obtain a plurality of up-sampled RGB feature maps and depth feature maps by performing spatial pyramid pooling (413, 414) on each of the RGB feature map (411) and the depth feature map (412). For example, the first block (410) can encode various spatial scale information on each of the RGB feature map (411) and the depth feature map (412) using spatial pyramid pooling. A pyramid pooling layer can be applied to the feature map of the lth layer to perform a Max pooling operation with a size of 2 to the power of r (pooling ratio). Accordingly, a feature map with a dimension in which the height and width are divided by the power of 2 to the power of r can be generated. In FIG. 2, 'F a l ' means the RGB feature map that is output from the l-1th layer and then input to the first block of the lth layer, and 'F b l' means a depth feature map that is input to the first block of the lth layer after being output from the l-1th layer. The RGB feature map (411) and the depth feature map (412) input to the first block (410) in the first layer may be the original of the feature map, for example, RGB information (e.g., RGB information (211) of FIG. 2) and depth information (e.g., depth information (212) of FIG. 2) of an image input to the model.
[0092] The first block (410) can perform an interpolation method (e.g., neighbor pixel interpolation method) on the RGB feature maps so that the RGB information of the image input to the model has the same scale as the scale of the image. In addition, the first block (410) can perform the same interpolation method on the depth feature maps so that the depth information of the image input to the model has the same scale as the scale of the image.
[0093] The first block (410) can integrate RGB feature maps into one RGB feature map (415) through a concatenation operation (414). In addition, the first block (410) can integrate depth feature maps into one depth feature map (417) through a concatenation operation (416).
[0094] The second block (420) applies a 1*1 convolution (421) to the RGB feature map (415) to create a V(value) feature map (V) of RGB information. a ), K(key) feature map (K a ), and Q(query) feature map (Q a ) can be obtained. In addition, the second block (420) applies a 1*1 convolution (422) to the depth feature map (416) to obtain a V feature map (V) of depth information. b ), K feature map (K b ), and Q feature map (Q b ) can be obtained.
[0095] The second block (420) can obtain RGB modality (423) and depth modality (424) using mathematical expressions 3 and 4.
[0096]
[0097]
[0098] In Equations 3 and 4, softmax() is the softmax function, Z rgb is RGB modality (423), Z d is a depth modulator (424), and C represents the number of channels (e.g., R (red) channel, G (green) channel, B (blue) channel) of the image input to the model. As another example, C may represent a feature channel dimension (FCD) within a model (e.g., a neural network middle layer). For example, if an RGB image of 1280 (width) x 720 (height) x 3 (depth) is input to a neural network middle layer, a feature map with an FCD of 420x180x256 may be generated in the neural network middle layer. Here, C may be '256', which is the number of dimensions in the feature map.
[0099] The third block (430) can integrate the RGB modality (423) and the depth modality (424) into one modality (432) through a concatenation operation (431). The third block (430) can derive the RGB weight (W) from the integrated modality (432). a ) and depth weights (W b ) can be obtained. For example, the third block (430) applies a multi-layer perceptron (MLP) and a softmax layer to the modality (432) to obtain RGB weights (W a ) and depth weights (W b ) can be obtained.
[0100] The third block (430) is the RGB weight (W a) is multiplied by the RGB feature map (411) to obtain the first value and the depth weight (W b ) can be multiplied by the depth feature map (412) to obtain a second value, and by adding the first value and the second value, a feature map (433) (e.g., feature map (230) of FIG. 2) in which RGB information and depth information are fused can be obtained.
[0101] FIG. 5 is a block diagram of a spatial-wise cross-modal attention block according to one embodiment. The spatial-wise cross-modal attention block (hereinafter, referred to as the second block) (500) illustrated in FIG. 5 may be a block corresponding to the second block (420) in the fusion module (400).
[0102] In order to reduce the complexity of the calculation in the second block (500), the H (height) and W (width) of the image can be reduced during the query-key-value operation. For this purpose, the second block (500) The feature map is G in the spatial dimension as l Divide into groups and G at the channel level l The ship can be expanded. The attention map is "(HW / G l )*(HW / G l )" is reduced to O(CH) and the computational complexity is O(CH) 2 W 2 ) in O(CH 2 W 2 / G l ) will be reduced. To address the loss of some pixel location information during query-key-value operations, spatial split and splice operators may be used. Spatial splitting divides the feature map into non-overlapping flat patches. can be decomposed into . Here, P is the patch size, and the relationship between patches can be encoded using cross-attention. During the computation, the number of channels C can be changed (increased or decreased) to C'.
[0103] FIG. 6 is a flowchart illustrating operations for clustering an image into multiple clusters according to one embodiment. The operations of FIG. 6 (e.g., operation 320 in FIG. 3) may be performed by an electronic device (e.g., the electronic device (201) in FIG. 1) when instructions stored in a memory (e.g., the clustering module (203) in FIG. 2) are executed by a processor. The operations of FIG. 6 may be initiated after an operation for measuring entropy per pixel in an image (e.g., operation 310) is performed.
[0104] At action 610, the electronic device is P i,j , RGBD i,j , and U i,j Using the feature space Xi,j =[P i,j , RGBD i,j , U i,j ] can be defined. For example, P i, j represents the position value of the (i (row), j (column))th pixel in the image (210) of Fig. 2, and RGBD i,j represents the feature value of the (i, j)th pixel in the feature map in which the RGB information (211) and depth information (213) output from the fusion module (202) are fused, and U i,j can represent the entropy of the (i, j)th pixel in the pixel-level entropy information (220).
[0105] In operation 620, the electronic device displays an image (e.g., image (210) of FIG. 2) with a side length of (where N is the total number of pixels in the image) can be classified into K clusters.
[0106] At action 630, the electronic device is P i,j and the center point P of the cluster u,v The distance D between them can be determined. According to one embodiment, the electronic device can calculate D using the following mathematical expression 3. In the mathematical expression 5, α, β, and γ represent weights that are arbitrarily assigned.
[0107]
[0108] At operation 640, the electronic device, based on the distance D calculated at operation 630, i,j can be assigned to the nearest cluster.
[0109] At operation 650, when cluster assignment for all pixels in the image is completed, the electronic device can update the average of the position values of the pixels belonging to each cluster as the new center point of the cluster.
[0110] The electronic device can repeatedly perform operations 630, 640, and 650 until a difference between a center point set before operation 650 is performed and a center point updated by performing operation 650 is less than or equal to a specified threshold.
[0111] FIG. 7 is a flowchart illustrating operations for supporting contrastive learning of an artificial intelligence model according to one embodiment. The operations of FIG. 7 (e.g., operation 360) may be performed by an electronic device (e.g., the electronic device (201) of FIG. 1) when instructions stored in a memory (e.g., including the classification module (205) of FIG. 2) are executed by a processor.
[0112] In operation 710, the electronic device may select an anchor pixel from an image (e.g., image (210)). In one embodiment, the electronic device may select an anchor pixel Ac of a c-th class from an unlabeled image i of a mini-batch using the following mathematical expression (6).
[0113]
[0114] In the above mathematical equation 6, Z ij is the jth pixel in image i. P ij δ is the softmax probability for the jth pixel of image i. p is the specified threshold value (positive number). yij is a pseudo label with high reliability. According to Equation 6, P ij Ga δ p If it is greater than, pixel j is the anchor pixel A in the cth class. c is selected as
[0115] In operation 720, the electronic device may select a pixel in the image that has a high similarity to an anchor pixel as a positive sample (e.g., an average of all anchor pixels). In one embodiment, the electronic device may select a positive sample in the image using the following mathematical expression 7.
[0116]
[0117] In the above mathematical formula 7, Z + denotes a positive sample. |A| denotes the size of all anchor pixels (e.g., total count), and Z C represents an anchor pixel in class c.
[0118] In operation 730, the electronic device determines from the image a threshold value specified by δ P ) can be determined as a negative sample. Here, the entropy may be, for example, a value obtained from pixel-level entropy information (220) output by an image segmentation model (201). Mathematical expression 8 below represents the entropy H of the probability distribution for all pixels in the image.
[0119]
[0120] In the above mathematical equation 8, P ij is the softmax probability for pixel j in unlabeled image i, and c represents the number assigned to each class.
[0121] According to one embodiment, according to one embodiment, the electronic device can select a voice sample using the following mathematical expression 9.
[0122]
[0123] In the above mathematical equation 9, N C Z represents a set of speech samples of class c. - ci,j δ denotes the jth pixel in image i selected as a voice sample. P is a specified threshold value (positive). According to Equation 9, P ij Ga δ P If it is smaller than , pixel j is selected as a negative sample in the cth class.
[0124] Additionally, in operation 740, the electronic device may determine pixels in the image that do not belong to any class as negative samples. For example, the electronic device may identify a class for each pixel in the image based on the model's inference results. In this case, pixels that do not belong to any class may be determined as negative samples.
[0125] At operation 750, the electronic device can perform contrastive learning on a model using speech samples.
[0126] FIG. 8 is a diagram illustrating an operation for managing voice samples according to one embodiment. The voice sample management operation (e.g., included in operation 360) may be performed by an electronic device (e.g., the electronic device (201) of FIG. 1) when instructions stored in a memory (e.g., including the classification module (205) of FIG. 2) are executed by the processor.
[0127] An electronic device may store a voice sample selected from an image (e.g., a first region (241)) in a memory (e.g., a memory (130) of FIG. 1). For example, the electronic device may set a portion of the memory as a storage for storing voice samples (e.g., a voice sample queue or a memory bank). The electronic device may manage the voice sample storage by category (e.g., a class). In addition, the electronic device may manage (e.g., store and delete) the voice samples stored in the voice sample storage in a first in first out (FIFO) manner. For example, referring to FIG. 8, the electronic device may store a new N in the voice sample storage for class c. c7 can be stored. The electronic device stores the first N stored voice samples among the voice samples stored in the storage. c1 can be extracted and used for contrastive learning.
[0128] The electronic device performs a loss L comparison between positive and negative samples among unlabeled samples in the RGB-D dataset. c can calculate the loss L c can be obtained through the following mathematical expression 10. Here, the loss may refer to, for example, the error between the inference value (predicted value) provided by the model and the actual correct answer. The electronic device can improve the discriminative ability (e.g., class discrimination ability) of the model by contrastively training the model to position positive samples closer to the anchor pixels of the corresponding class and to position negative samples belonging to the same class as the positive samples further from the anchor pixels.
[0129]
[0130] In the above mathematical expression 10, M is the total number of anchor pixels, C is the number of classes, and Z ciis the i-th anchor pixel of class c, N is the total number of pixels, τ is the temperature parameter used in the artificial intelligence model, and sim(,) represents the cosine similarity function.
[0131] FIG. 9 is a flowchart illustrating operations for supporting contrastive learning of an artificial intelligence model according to one embodiment. The operations of FIG. 9 (e.g., operation 360) may be performed by an electronic device (e.g., the electronic device (201) of FIG. 1) when instructions stored in a memory (e.g., including the classification module (205) of FIG. 2) are executed by the processor.
[0132] An image (910) to be input to an image segmentation model (901) may be an unlabeled image that is divided into multiple classes (e.g., three classes (911, 912, 913)). The image segmentation model (901) may divide the image (910) into three classes (911, 912, 913) and classify each class into a reliable region and an unreliable region. The electronic device may determine an anchor pixel in the reliable region and, based on this, select a positive sample and a negative sample in the corresponding class. For example, if an anchor pixel (911a) is determined in a first class (911), based on this, the electronic device may select a positive sample (911b) similar to the anchor pixel (911a) and a negative sample (911c) having a low similarity to the anchor pixel (911a) in the first class (911). The electronic device can train the model to position the positive sample (911b) closer to the first class (911) and the negative sample (911c) farther from the first class (911). For example, the electronic device can store the negative sample (912a) from the second class (912) and the negative samples (913a, 913b) from the third class (913) in the storage (902) for the first class (911). The electronic device can extract the negative samples (912a, 913a, 913b) from the storage (902) and cause the model to perform contrastive learning for the first class (911) using the extracted negative samples.
[0133] The following mathematical expression 11 is the overall loss function L. In mathematical expression 10, Ls is the supervised learning loss function from the correct labels labeled by human experts, Lu is the supervised learning loss function from the reliable virtual labels from the previous model inference results, and Lc represents contrastive learning using the negative sample cue λ. u is the weight multiplied by the virtual label loss function, and λ crefers to the weight multiplied by the contrastive learning loss function.
[0134]
[0135] According to one embodiment, an electronic device includes a processor; and a memory that stores a model learned using an RGB-D (depth) dataset for image segmentation. The command, when executed by the processor, may cause the electronic device to obtain entropy information including entropies corresponding to pixels in a first image having RGB information including RGB values corresponding to pixels respectively and depth information including depth values corresponding to the pixels respectively. The command, when executed by the processor, may cause the electronic device to cluster the first image into a plurality of clusters based on the entropy information, the RGB information, and the depth information. The command, when executed by the processor, may cause the electronic device to classify the clusters into a first region and a second region based on the entropy information. The above instructions, when executed by the processor, may cause the electronic device to assign a label to at least one of the clusters classified into the first region based on a user input. The instructions, when executed by the processor, may cause the electronic device to assign a label to a cluster classified into the second region based on an inference of the model. The instructions, when executed by the processor, may cause the electronic device to train the model using a second image having a cluster assigned a label based on the user input and a cluster assigned a label based on an inference of the model.
[0136] The above command, when executed by the processor, may cause the electronic device to input the first image as an input value to the model and obtain the entropy information from an output value output from the model.
[0137] The above command, when executed by the processor, may cause the electronic device to generate a feature map in which the RGB information and the depth information are fused, and to cluster the first image into a plurality of clusters using the feature map and the entropy information.
[0138] The above feature map can be obtained by the electronic device performing the first operation, the second operation, and the third operation.
[0139] The first operation comprises: performing spatial pyramid pooling on the RGB information to obtain a plurality of up-sampled RGB feature maps, and performing spatial pyramid pooling on the depth information to obtain a plurality of up-sampled depth feature maps; performing interpolation (e.g., neighboring pixel interpolation) on the RGB feature maps to have the same scale as the scale of the RGB information, and performing interpolation (e.g., neighboring pixel interpolation) on the depth feature maps to have the same scale as the scale of the depth information; combining the RGB feature maps into one RGB feature map through a concatenation operation, and combining the depth feature maps into one depth feature map through a concatenation operation; and applying a 1*1 convolution to the RGB feature map combined into one to obtain V rgb Feature map, K rgb Feature maps, and Q rgbObtain the feature map and apply 1*1 convolution to the above depth feature map that is combined into one V d Feature map, K d Feature maps, and Q d It may include an operation to obtain a feature map.
[0140] The second operation may include an operation of obtaining RGB modality and depth modality using mathematical expressions 3 and 4.
[0141] The third operation may include an operation of combining the RGB modality and the depth modality into a single modality through a concatenation operation; an operation of obtaining an RGB weight and a depth weight from the combined modality; and an operation of obtaining a feature map in which the RGB information and the depth information are fused by adding a first value obtained by multiplying the RGB weight by the RGB information and a second value obtained by multiplying the depth weight by the depth information.
[0142] The above instruction, when executed by the processor, causes the electronic device to: i,j , RGBD i,j , U i,j ](Here, Pi, j represent the position value of the i (row), j (column)th pixel, and RGBD i,j represents the feature value of the i, jth pixel in the above feature map, and U i,j An operation that defines (where the i, j-th pixel represents the entropy); the first image has a side length of (where N is the total number of pixels in the first image) is an operation of classifying into K clusters; P i,j and the center point P of the cluster u,v An action to check the distance between pixels; based on the checked distance, pixel P i,jAn operation of assigning pixels to the closest cluster; and when the assignment for all pixels in the first image is completed, an operation of updating the average of the position values of the pixels belonging to each cluster as a new center point of the cluster may be performed. The instruction, when executed by the processor, may cause the electronic device to repeat the calculation operation, the assignment operation, and the update operation until a difference between the center point set before the update operation is performed and the center point obtained through the update operation becomes less than or equal to a threshold value.
[0143] The above command, when executed by the processor, may cause the electronic device to calculate the distance using mathematical expression 5.
[0144] The above instruction, when executed by the processor, causes the electronic device to select an anchor pixel of the cth class in an unlabeled image i of a mini batch. (Here, Z ij is the jth pixel in image i, and p ij is the softmax probability for the jth pixel of image i, and δ p is a specified threshold (positive), and y ij is a pseudo label with high confidence, and p ij Ga δ p If pixel j is greater than the anchor pixel A in the cth class, c ) is selected, and a pixel with high class similarity to the anchor pixel in the c-th class is selected as a positive sample. (where |A| is all anchor pixels and Z cis an anchor representation of class c), and pixels having an entropy obtained as a result of inference by the model in the c-th class less than the threshold value and pixels not belonging to any class are determined as negative pixels, and the model can be used to perform contrastive learning using the negative pixels.
[0145] The above command, when executed by the processor, may cause the electronic device to classify the third image into a third region that is trustworthy and a fourth region that is not trustworthy, and to perform comparison learning on the model using information about the fourth region.
[0146] The above command, when executed by the processor, may cause the electronic device to input the third image as an input value to the model and obtain information about the third region and the fourth region from an output value output from the model.
[0147] The above command, when executed by the processor, may cause the electronic device to classify the third image into a plurality of classes based on the acquired information, select anchor pixels from the plurality of classes, select positive samples and negative samples for each of the plurality of classes based on the anchor pixels, and train the model to position the positive samples closer to the corresponding classes and position the negative samples farther from the corresponding classes.
[0148] According to one embodiment, a method of operating an electronic device is provided. The method may include obtaining entropy information including entropies corresponding to pixels from a first image having RGB information including RGB values corresponding to pixels and depth information including depth values corresponding to the pixels. The method may include clustering the first image into a plurality of clusters based on the entropy information, the RGB information, and the depth information. The method may include classifying the clusters into a first region and a second region based on the entropy information. The method may include assigning a label to at least one of the clusters classified into the first region based on a user input, and assigning a label to a cluster classified into the second region based on an inference of the model. The method may include training the model using a second image having a cluster assigned a label based on the user input and a cluster assigned a label based on an inference of an artificial intelligence model.
[0149] The operation of obtaining the entropy information may include an operation of inputting the first image as an input value to the model and obtaining the entropy information from an output value output from the model.
[0150] The clustering operation may include an operation of generating a feature map in which the RGB information and the depth information are fused; and an operation of clustering the first image into a plurality of clusters using the feature map and the entropy information.
[0151] The method may further include an operation of classifying a third image into a trustworthy third region and an untrustworthy fourth region; and an operation of performing contrastive learning on the model using information about the fourth region.
[0152] The above classification operation may include an operation of inputting the third image as an input value to the model and obtaining information about the third region and the fourth region from an output value output from the model.
[0153] In the above explanation, the prefixes “first,” “second,” and “third” are only used to distinguish between the same names and do not have any special meaning in themselves, such as importance or order.
[0154] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0155] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0156] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. In one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0157] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0158] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0159] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, processor; and A memory for storing a model trained using an RGB-D (depth) dataset for instructions and image segmentation, wherein the instructions, when executed by the processor, cause the electronic device to: Obtain entropy information including entropies corresponding to each pixel in a first image having RGB information including RGB values corresponding to each pixel and depth information including depth values corresponding to each pixel, Based on the entropy information, the RGB information, and the depth information, the first image is clustered into a plurality of clusters, Based on the above entropy information, the clusters are classified into the first region and the second region, At least one of the clusters classified into the first region is assigned a label based on user input, and a cluster classified into the second region is assigned a label based on inference of the model. An electronic device that trains the model using a second image having labeled clusters based on the user input and labeled clusters based on the model's inference.
2. In the first paragraph, when the command is executed by the processor, the electronic device, An electronic device that inputs the first image as an input value to the model and obtains the entropy information from an output value output from the model.
3. In the first paragraph, when the command is executed by the processor, the electronic device, Generate a feature map in which the RGB information and the depth information are fused, An electronic device that clusters the first image into a plurality of clusters using the feature map and the entropy information.
4. In the third paragraph, the feature map is obtained by the electronic device performing the first operation, the second operation, and the third operation, The above first operation is, An operation of obtaining a plurality of up-sampled RGB feature maps by performing spatial pyramid pooling on the RGB information, and obtaining a plurality of up-sampled depth feature maps by performing spatial pyramid pooling on the depth information; An operation of performing an interpolation method (e.g., neighboring pixel interpolation method) on the RGB feature maps so that they have the same scale as the scale of the RGB information, and performing an interpolation method (e.g., neighboring pixel interpolation method) on the depth feature maps so that they have the same scale as the scale of the depth information; An operation of combining the RGB feature maps into one RGB feature map through a concatenation operation and combining the depth feature maps into one depth feature map through a concatenation operation; and Applying 1*1 convolution to the above RGB feature maps that are collated into one V rgb Feature map, K rgb Feature maps, and Q rgb Obtain the feature map and apply 1*1 convolution to the depth feature map that is collated into one V d Feature map, K d Feature maps, and Q d Contains an operation for obtaining a feature map, The second operation above is, It includes an operation of obtaining RGB modality and depth modality using the following mathematical formula, , In the above mathematical formula, softmax() is the softmax function, Z rgb is the RGB modality, and Z d represents the above depth modalitator, The third operation above is, An operation of combining the RGB modality and the depth modality into one modality through a concatenation operation; An operation of obtaining RGB weights and depth weights from the modalities collated as above; and An electronic device including an operation of obtaining a feature map in which the RGB information and the depth information are fused by adding a first value obtained by multiplying the RGB weight by the RGB information and a second value obtained by multiplying the depth weight by the depth information.
5. In the third paragraph, when the command is executed by the processor, the electronic device, Feature space X i,j = [P i,j , RGBD i,j , U i,j ] (Here, P i, j represents the position value of the i (row), j (column)th pixel, and RGBD i,j represents the feature value of the i, jth pixel in the above feature map, and U i,j An operation that defines (where i and j represent the entropy of the i, jth pixel); The length of one side of the first image above is An operation of classifying into K clusters (where N is the total number of pixels in the first image); P i,j and the center point P of the cluster u,v Action to check the distance between the two; Based on the above confirmed distance, pixel P i,j The action of assigning to the nearest cluster; and When the above allocation for all pixels in the above first image is completed, an operation is performed to update the average of the position values of the pixels belonging to each cluster as the new center point of the cluster, An electronic device that repeats the calculation operation, the allocation operation, and the update operation until the difference between the center point set before the update operation is performed and the center point obtained through the update operation becomes less than or equal to a threshold value.
6. In the fifth paragraph, the command, when executed by the processor, causes the electronic device to calculate the distance using the following mathematical formula. (Here, α, β, and γ represent randomly assigned weights) 7. In the first paragraph, when the command is executed by the processor, the electronic device, Anchor pixel of cth class in unlabeled image i of a mini batch (Here, Z ij is the jth pixel in image i, and p ij is the softmax probability for the jth pixel of image i, and δ p is a specified threshold (positive), and y ij is a pseudo label with high confidence, and p ij Go δ p If pixel j is greater than the anchor pixel A in the cth class, c Select ) and In the cth class, the pixel with high class similarity to the anchor pixel is called a positive sample. (Here, |A| is any anchor pixel and Z c is determined as an anchor expression of class c), In the cth class, pixels whose entropy obtained as a result of inference by the above model is less than the threshold value and pixels that do not belong to any class are determined as negative pixels, An electronic device for performing contrastive learning on the model using the above-described speech pixels.
8. In the first paragraph, when the command is executed by the processor, the electronic device, The third image is classified into a trustworthy third area and an untrustworthy fourth area. An electronic device that performs comparative learning on the model using information about the fourth area.
9. In the 8th paragraph, when the command is executed by the processor, the electronic device, An electronic device that inputs the third image as an input value to the model and obtains information about the third area and the fourth area from the output value output from the model.
10. In the 9th paragraph, when executed by the command, the electronic device, Based on the information obtained above, the third image is divided into multiple classes, and anchor pixels are selected from the multiple classes. Based on the above anchor pixels, positive samples and negative samples are selected for each of the multiple classes, An electronic device that trains the model to position the positive sample closer to the corresponding class and to position the negative sample farther from the corresponding class.
11. A method of operating an electronic device, An operation of obtaining entropy information including entropies corresponding to each pixel in a first image having RGB information including RGB values corresponding to each pixel and depth information including depth values corresponding to each pixel; An operation of clustering the first image into a plurality of clusters based on the entropy information, the RGB information, and the depth information; An operation of classifying the clusters into a first region and a second region based on the above entropy information; An operation of assigning a label to at least one of the clusters classified into the first region based on a user input, and assigning a label to a cluster classified into the second region based on an inference of the model; A method comprising the action of training the model using a second image having labeled clusters based on the user input and labeled clusters based on inference of the artificial intelligence model.
12. In the 11th paragraph, the operation of obtaining the entropy information is as follows: A method including an action of inputting the first image as an input value to the model and obtaining the entropy information from an output value output from the model.
13. In the 11th paragraph, the clustering operation is: An operation of generating a feature map in which the RGB information and the depth information are fused; and A method comprising an operation of clustering the first image into a plurality of clusters using the feature map and the entropy information.
14. In paragraph 11, The action of classifying the third image into a trustworthy third region and an untrustworthy fourth region; and A method further comprising an operation of performing comparative learning on the model using information about the fourth region.
15. In the 14th paragraph, the classifying operation is: A method including an action of inputting the third image as an input value to the model and obtaining information about the third region and the fourth region from an output value output from the model.
Citation Information
Patent Citations
RGB video frame hand-held object detection method based on deep learning
CN117095339A
Method and apparatus for assigning ratings to training images
KR1020130126916A
Method and apparatus of object recognition, Method and apparatus of learning for object recognition
KR1020170034226A
Electronic device
KR1020230110192A
System and Method for Serving Artificial intelligence model to utilize spatial data
KR1020250064005A