Learning method and information processing device

The learning method and device enhance self-supervised learning by assigning image region classifications and using mask-induced attention bias to reduce training data needs for object attribute estimation in images.

JP2026041400APending Publication Date: 2026-03-10KYOCERA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Pre-training using self-supervised learning is not effective in reducing training data requirements for object attribute estimation in images.

Method used

A learning method and device that pre-train a classification-based image recognition model using self-supervised learning, assigning classifications to image regions through segmentation, and utilizing a mask-induced attention bias to enhance feature extraction in self-attention mechanisms.

Benefits of technology

Reduces the need for supervised data while maintaining learning effectiveness by leveraging weakly supervised information in self-supervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041400000001_ABST
    Figure 2026041400000001_ABST
Patent Text Reader

Abstract

Improve the effectiveness of reducing training data. In this learning method, each region of an acquired image is assigned a classification, and a learning model is pre-trained using self-supervised learning using the classification. The learning model estimates the attributes of objects in the target image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a learning method and an information processing device. [Background technology]

[0002] It has been proposed to perform pre-training in the training of a learning model that estimates the attributes of objects, such as the position and name of objects included as subjects in acquired images. Because pre-training also requires a large dataset, technological development is underway to reduce annotation costs. It has been proposed to perform pre-training using self-supervised learning, which does not require labeling (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin, “Emerging properties in self-supervised vision transformers”, In ICCV, 2021 Summary of the Invention [Problem to be solved by the invention]

[0004] However, pre-training using self-supervised learning was not very effective in reducing training data.

[0005] Therefore, an object of the present disclosure is to provide a learning method and an information processing device that improve the effect of reducing training data. [Means for solving the problem]

[0006] The first perspective of learning is Classification of each area of ​​the acquired image is assigned, A learning model for estimating the attributes of objects in a target image is pre-trained by self-supervised learning using the classification.

[0007] An information processing device according to a second aspect comprises: an acquisition unit that acquires an image; The system also includes a control unit that assigns a classification to each region of the image acquired by the acquisition unit and performs pre-learning of a learning model that estimates the attributes of objects in the target image through self-supervised learning using the classification. [Effects of the Invention]

[0008] According to the learning method and information processing device configured as described above according to the present disclosure, the training data is reduced. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a schematic configuration of an information processing apparatus according to this embodiment. [Figure 2] 10A and 10B are diagrams illustrating a mask image generated from an image obtained by segmentation processing. [Figure 3] 2 is a flowchart for explaining a calculation process executed by the control unit of FIG. 1. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of an information processing device that executes a learning method to which the present disclosure is applied will be described with reference to the drawings.

[0011] 1 shows an example of the configuration of an information processing device 10. The information processing device 10 includes an acquisition unit 11 and a control unit 12.

[0012] The acquisition unit 11 acquires images as information from various devices, such as other information processing devices, storage media, and imaging devices.

[0013] The acquisition unit 11 may employ, for example, a physical connector and a wireless communication device. The physical connectors include an electrical connector compatible with transmission by electrical signals, an optical connector compatible with transmission by optical signals, and an electromagnetic connector compatible with transmission by electromagnetic waves. The electrical connectors include a connector conforming to IEC 60603, a connector conforming to the USB standard, a connector compatible with an RCA terminal, a connector compatible with an S terminal specified in EIAJ CP-1211A, a connector compatible with a D terminal specified in EIAJ RC-5237, a connector conforming to the HDMI (registered trademark) standard, and a connector compatible with a coaxial cable including a BNC. The optical connectors include various connectors conforming to IEC 61754. The wireless communication device includes a wireless communication device conforming to various standards including Bluetooth (registered trademark) and IEEE 802.11.

[0014] The control unit 12 is configured to include at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for a specific process. The dedicated circuit may be, for example, an FPGA (Field-Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or the like. The control unit 12 may control the operation of the information processing device 10.

[0015] The control unit 12 may further include a storage unit. The storage unit may include any storage device, such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The storage unit may store various programs that cause the control unit 12 to function and various information used by the control unit 12.

[0016] The control unit 12 may train an image recognition model (learning model). The image recognition model estimates the attributes of an object included as a partial image in the target image. The learning model is, for example, a Vision Transformer. The image recognition model is not limited to a Vision Transformer, and may be any model that uses a self-attention mechanism. The self-attention mechanism is a mechanism that performs feature transformation processing based on the self-attention of the data itself. The attributes of an object include the name of the object, the drawing area of ​​the object in the target image, etc.

[0017] More specifically, the self-attention mechanism considers n image token vectors x = (x1, ..., x n ), (x i ∈R dx ) output y=(y1,...,y n ), (y i ∈R dy ) is calculated.

[0018]

number

[0019] In equation (1), W V is the learnable weight. Weight coefficient α ij is calculated using the following formula (2).

[0020]

number

[0021] In equation (2), e ij is the scaled dot product attention for the input x. ij is calculated using equation (4) described later.

[0022] The control unit 12 performs at least pre-training of the image recognition model. Pre-training is a method of initially training a network using a large-scale data set. In the present disclosure, the control unit 12 performs pre-training by self-supervised learning using unlabeled images. The self-supervised learning may be performed by a known learning method such as contrastive learning or masked image modeling (MIM).

[0023] The control unit 12 may further perform main training of the pre-trained image recognition model. Main training is a means for adjusting the image recognition model learned by the pre-training so that it is specialized for a specific task.

[0024] In the pre-learning, the control unit 12 uses the image acquired by the acquisition unit 11 to assign a classification to each of the constituent regions. The size of the regions may be the same as each region that characterizes the input image to the image recognition model. The control unit 12 may, for example, perform a segmentation process on the image to classify each of the regions that constitute the image by object.

[0025] For example, as shown in Figure 2, a mask image 16 may be generated by segmentation processing from an image 15 captured from above of a bagged product 14 placed on a stand 13, in which a first classification and a second classification are assigned to the areas where the stand 13 and the product 14 are depicted, respectively.

[0026] The control unit 12 uses the assigned classification to perform self-supervised learning using the images acquired by the acquisition unit 11. Specifically, the control unit 12 may use, for each region, a parameter indicating the consistency of classification with other regions in learning based on a self-attention mechanism. Learning based on a self-teaching mechanism using a parameter indicating the consistency of classification will be described below using a specific example.

[0027] The control unit 12 performs a resizing operation on the mask image 16 to obtain a mask token vector m=(m 1, …,m n)∈{x∈Z|0≦x≦c} n The mask token vector may have the same number of elements as the image token vector of the image to be input to the image recognition model. The control unit 12 calculates a mask-induced attention bias b=(b 1,1, …,b n,n )∈R n×n may be calculated using the following formula (3).

[0028]

number

[0029] In equation (3), u pos is a learnable scalar used for the classification matching relationship between any region i in the mask image 16 and other regions j. neg is a learnable scalar used for the classification discrepancy relationship between any region i in the mask image 16 and other regions j.

[0030] The control unit 12 uses the mask-induced attention bias to calculate the scaled dot-product attention e ij may be calculated.

[0031]

number

[0032] In equation (4), W Q , W K are the learnable weights.

[0033] As described above, the mask-induced attention bias b may be used to calculate the output y in the self-attention mechanism as a parameter indicating the consistency of classification for each region with other regions.

[0034] Next, the calculation process of the loss function executed by the control unit 12 in this embodiment will be described with reference to the flowchart in Fig. 3. The calculation process of the loss function starts when the acquisition unit 11 acquires an image set of multiple frames.

[0035] In step S100, the control unit 12 performs segmentation processing on each of the acquired images 15 to generate a mask image 16. After the segmentation processing, the process proceeds to step S101.

[0036] In step S101, the control unit 12 calculates a mask token vector by performing a resizing operation on the mask image 16 generated in step S100. After the resizing operation, the process proceeds to step S102.

[0037] In step S102, the control unit 12 calculates the mask-induced attention bias using the mask token vector calculated by the resize operation in step S101. After calculating the mask-induced attention bias, the process proceeds to step S103.

[0038] In step S103, the control unit 12 reflects the mask-induced attention bias calculated in step S102 in the feature extraction process in self-supervised learning for the acquired image 15. Specifically, the control unit 12 reflects the mask-induced attention bias in feature extraction by applying a self-attention mechanism. After reflection, the process proceeds to step S104.

[0039] In step S104, the control unit 12 calculates a loss function for the feature extraction process reflected in step S103. After the calculation, the calculation process ends.

[0040] The information processing device 10 of this embodiment configured as described above includes an acquisition unit 11 that acquires an image 15, and a control unit 12 that assigns a classification to each region of the image 15 acquired by the acquisition unit 11 and performs pre-learning of a learning model that estimates attributes of objects in a target image through self-supervised learning using the classification. With this configuration, the information processing device 10 uses weakly supervised information in self-supervised learning, thereby improving the effectiveness and efficiency of pre-learning. Therefore, the information processing device 10 can reduce artificially added supervised data while maintaining the effectiveness of learning, thereby improving the effectiveness of reducing supervised data.

[0041] The above has described an embodiment of the information processing device 10, but the embodiment of the present disclosure can also be implemented as a method or program for implementing the device, or as a storage medium on which a program is recorded (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.).

[0042] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module incorporated into an operating system. Furthermore, the program may or may not be configured so that all processing is performed solely by the CPU on the control board. The program may also be configured so that part or all of it is executed by another processing unit mounted on an expansion board or expansion unit added to the board as needed.

[0043] The drawings illustrating the embodiments of the present disclosure are schematic, and the dimensional ratios and the like in the drawings do not necessarily correspond to the actual ones.

[0044] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art could make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies, and multiple components can be combined or divided into one.

[0045] All of the features described in this disclosure and / or all steps of all of the disclosed methods or processes may be combined in any combination except combinations in which these features are mutually exclusive. Furthermore, each feature described in this disclosure may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly denied. Thus, unless expressly denied, each disclosed feature is only one example of a generic series of identical or equivalent features.

[0046] Furthermore, embodiments of the present disclosure are not limited to the specific configurations of any of the above-described embodiments, but rather extend to any novel feature or combination thereof described herein, or any novel method or process step or combination thereof described herein. [Explanation of symbols]

[0047] 10. Information processing equipment 11 Acquisition Department 12 Control Unit 13 units 14 items 15 images 16 Mask Images

Claims

1. Classification of each area of ​​the acquired image is assigned, Using the classification, a learning model is pre-trained to estimate the attributes of objects in the target image through self-supervised learning. How to learn.

2. 2. The learning method according to claim 1, The learning model is a model that uses a self-attention mechanism. How to learn.

3. 3. The learning method according to claim 1 or 2, For each of the regions, a parameter indicating the consistency of classification with other regions is used for learning based on the self-attention mechanism. How to learn.

4. an acquisition unit that acquires an image; a control unit that assigns a classification to each region of the image acquired by the acquisition unit and performs pre-learning of a learning model that estimates attributes of an object in a target image through self-supervised learning using the classification. Information processing device.