Information processing device and information processing method

By extracting features from image data only and grouping prediction items in a loss function, the method addresses noise issues in heterogeneous data sets, enhancing prediction accuracy and resource efficiency.

JP2025112394APending Publication Date: 2025-08-01TOPPAN HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024006587
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Machine learning methods applied to heterogeneous data sets combining image and non-image data face issues with noise inclusion, leading to reduced accuracy and unclear basis for prediction results.

Method used

Perform feature extraction only on image data, combine extracted image features with non-image data, and use a loss function to group prediction items based on relevance criteria, calculating loss values and gradients to generate accurate basis information.

Benefits of technology

Outputs highly accurate basis information for prediction results with reduced noise, saving computing resources and shortening processing time while improving model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025112394000001_ABST
    Figure 2025112394000001_ABST
Patent Text Reader

Abstract

To provide information processing means which can output highly accurate basis information.SOLUTION: An information processing device includes: an acquisition unit which acquires an input data set including image data and non-image data; a feature extraction unit which extracts an image feature indicating a characteristic of the image data, from the image data; a coupling unit which couples the non-image data and the image feature extracted from the image data to generate coupled data; a prediction unit which generates a prediction result indicating prediction items related to the input data set, on the basis of the coupled data; a loss management unit which groups the prediction items which satisfy predetermined relevancy criteria, among the prediction items in the prediction result, and calculates a loss value about each group; and a basis management unit which calculates the gradient for minimizing the loss value, and generates and outputs basis information indicating an influence degree of an element which has affected the prediction item, about each prediction item in the prediction result, on the basis of at least the gradient.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus and an information processing method.

Background Art

[0002] In recent years, machine learning has attracted attention as a means for making predictions and judgments on target data. Among them, according to machine learning using a neural network, especially deep learning, high-precision results can be obtained even in predictions with complex solutions. However, since the calculation method for the resulting prediction value is black-boxed, there is a problem that the elements affecting the prediction are unknown and the basis for the prediction result cannot be grasped.

[0003] Regarding this problem, several proposals have been made. For example, Japanese Patent Application Laid-Open No. 2023-514282 (Patent Document 1) discloses that "an automatic data analysis technique for non-tabular data sets may include: (1) automatically developing a model that executes tasks in the fields of computer vision, acoustic processing, speech processing, text processing, or natural language processing; (2) automatically developing a model that analyzes a heterogeneous data set including image data and non-image data, and / or a heterogeneous data set including tabular data and non-tabular data; (3) determining the importance of image features with respect to a modeling task; (4) explaining the value of a modeling target based at least in part on the image features; and (5) detecting drift in the image data. In some cases, a multi-stage model may be developed, a pre-trained feature extraction model extracts low-level, medium-level, high-level, and / or top-level features of non-tabular data, and a data analysis model uses those features (or features obtained therefrom) to execute a data analysis task."

Prior Art Documents

Patent Documents

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2023-514282 [Summary of the Invention] [Problems to be Solved by the Invention]

[0005] Currently, machine learning methods are used in various applications, and there is a need for a method that can handle various forms of input and output information and provide highly accurate prediction and basis output.

[0006] Patent Document 1 describes a method for determining the importance of feature quantities that affect the analysis results for a heterogeneous data set including image data and non-image data. However, when extracting feature quantities for a heterogeneous data set that combines image data and non-image data as in the method described in Patent Document 1, noise may be included in the feature quantities, and the prediction results and the accuracy of the basis for the prediction results may be limited. However, Patent Document 1 does not consider the reduction in accuracy due to noise included in the feature quantities.

[0007] Therefore, an object of the present disclosure is to provide information processing means that can output highly accurate basis information for prediction results with reduced noise by using a loss function that groups prediction items after inputting non-image data after image feature extraction when applying a machine learning method to a heterogeneous data set including image data and non-image data. [Means for Solving the Problems]

[0008] To solve the above problems, a representative information processing apparatus of the present invention includes a processor and a memory. The memory includes an acquisition unit that acquires an input data set including image data and non-image data, a feature extraction unit that extracts image features indicating characteristics of the image data from the image data, a combination unit that generates combined data by combining the image features extracted from the image data and the non-image data, a prediction unit that generates a prediction result indicating a prediction item related to the input data set based on the combined data, a loss management unit that groups prediction items satisfying a predetermined relevance criterion among the prediction items in the prediction result and calculates a loss value for each group, and a basis management unit that calculates a gradient for minimizing the loss value and generates and outputs basis information indicating the degree of influence of elements that have influenced each prediction item in the prediction result based at least on the gradient, and includes processing instructions for causing the processor to function as the above units.

Advantages of the Invention

[0009] According to the present disclosure, when a machine learning method is applied to a heterogeneous data set including image data and non-image data, after inputting non-image data after image feature extraction and using a loss function in which prediction items are grouped, it is possible to provide information processing means capable of outputting highly accurate basis information for a prediction result with reduced noise. Problems, configurations, and effects other than the above will be clarified by the description in the following embodiments for carrying out the invention.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the present invention is not limited by this embodiment. Also, in the description of the drawings, the same parts are denoted by the same reference numerals. Also, terms such as "first", "second", "third", etc. may be used in the present disclosure to describe various elements or components, but it will be understood that these elements or components should not be limited by these terms. These terms are only used to distinguish one element or component from another. Therefore, the first element or component discussed below can also be referred to as the second element or component without departing from the teachings of the present inventive concept.

[0012] (Overview of the Present Disclosure) In principle, in a method of performing basis output, the feature map extracted by the feature amount extraction means is important, and how to reduce noise from this feature map affects the prediction accuracy and the accuracy of the basis output. However, as described above, when extracting feature amounts for a heterogeneous data set that combines image data and non-image data, there is a problem that noise is included in the feature amounts, and the prediction result and the accuracy of the basis for the prediction result are limited.

[0013] In the present disclosure, in order to reduce noise, feature extraction is performed only on image data, and prediction is made based on combined data obtained by combining the features of the image thus extracted and non-image data. In this way, by performing feature extraction only on the image data, noise can be reduced compared to the case where feature extraction is performed on the image data and non-image data together.

[0014] Further, in the present disclosure, among each prediction item, prediction items having a certain correlation, such as similar numerical characteristics, are grouped, and a loss function that sums the loss values obtained for each group is used. As a result, without significantly increasing the amount of calculation, the accurate loss value of each prediction item can be grasped, and the prediction accuracy of the entire machine learning model can be improved. Further, by also using this loss function when performing the basis output, the loss for each prediction item can be accurately grasped.

[0015] By these contrivances, it becomes possible to output basis information with reduced noise, and the degree of influence on the respective prediction results of the elements that have affected the prediction results (for example, a specific region in the image data, setting parameters included in the non-image data, etc.) can be easily grasped. Further, by omitting the feature extraction for the non-image data, computing resources can be saved and the processing time can be shortened.

[0016] Next, with reference to FIG. 1, a computer system 100 for implementing an embodiment of the present disclosure will be described. The mechanisms and apparatuses of the various embodiments disclosed herein may be applied to any suitable computing system. The main components of the computer system 100 include one or more processors 102, a memory 104, a terminal interface 112, a storage interface 113, an I / O (input / output) device interface 114, and a network interface 115. These components may be interconnected via a memory bus 106, an I / O bus 108, a bus interface unit 109, and an I / O bus interface unit 110.

[0017] The computer system 100 may include one or more general-purpose programmable central processing units (CPUs) 102A and 102B collectively referred to as the processor 102. In certain embodiments, the computer system 100 may comprise multiple processors, and in other embodiments, the computer system 100 may be a single CPU system. Each processor 102 executes instructions stored in the memory 104 and may include an on-board cache. Also, in certain embodiments, the computer system 100 may include a GPU (Graphics Processing Unit) in addition to the processor 102. By using the GPU, processing such as a machine learning model used in the information processing application 150 described later can be accelerated.

[0018] In certain embodiments, the memory 104 may include a random access semiconductor memory, a storage device, or a storage medium (either volatile or non-volatile) for storing data and programs. The memory 104 may store all or part of the programs, modules, and data structures for implementing the functions described herein. For example, the memory 104 may store the information processing application 150. In certain embodiments, the information processing application 150 may include instructions or descriptions for executing the functions described later on the processor 102.

[0019] In some embodiments, information processing application 150 may be implemented in hardware via a semiconductor device, chip, logic gate, circuit, circuit card, and / or other physical hardware device, instead of or in addition to a processor-based system. In some embodiments, information processing application 150 may include data other than instructions or descriptions. In some embodiments, a camera, sensor, or other data input device (not shown) may be provided to communicate directly with bus interface unit 109, processor 102, or other hardware of computer system 100.

[0020] Computer system 100 may include a bus interface unit 109 that facilitates communication between processor 102, memory 104, display system 124, and I / O bus interface unit 110. I / O bus interface unit 110 may be coupled to an I / O bus 108 for transferring data between various I / O units. I / O bus interface unit 110 may communicate with a plurality of I / O interface units 112, 113, 114, and 115, also known as I / O processors (IOPs) or I / O adapters (IOAs), via I / O bus 108.

[0021] Display system 124 may include a display controller, display memory, or both. The display controller may be capable of providing video, audio, or both data to display device 126. Also, computer system 100 may include devices such as one or more sensors configured to collect data and provide the data to processor 102.

[0022] For example, the computer system 100 may include a biometric sensor that collects heart rate data, stress level data, etc., an environmental sensor that collects humidity data, temperature data, pressure data, etc., and a motion sensor that collects acceleration data, motion data, etc. Other types of sensors can also be used. The display system 124 may be connected to a display device 126 such as a single display screen, a television, a tablet, or a portable device.

[0023] The I / O interface unit has a function of communicating with various storage or I / O devices. For example, the terminal interface unit 112 can be attached with a user I / O device 116 such as a user output device such as a video display device, a speaker television, or a user input device such as a keyboard, a mouse, a keypad, a touch pad, a trackball, a button, a light pen, or other pointing devices. The user can use the user interface to operate the user input device to input input data and instructions to the user I / O device 116 and the computer system 100, and receive output data from the computer system 100. The user interface may be displayed on the display device, reproduced by the speaker, or printed via the printer, for example, via the user I / O device 116.

[0024] The storage interface 113 can be attached to one or more disk drives and direct access storage devices 117 (usually magnetic disk drive storage devices, but can also be an array of disk drives configured to appear as a single disk drive or other storage devices). In certain embodiments, the storage device 117 may be implemented as any secondary storage device. The contents of the memory 104 may be stored in the storage device 117 and read from the storage device 117 as needed. The I / O device interface 114 may provide an interface to other I / O devices such as printers, fax machines, etc. The network interface 115 may provide a communication path for the computer system 100 to communicate with other devices mutually. This communication path may be, for example, the network 130.

[0025] In certain embodiments, the computer system 100 may be a device that receives requests from other computer systems (clients) without a direct user interface, such as a multi-user mainframe computer system, a single-user system, or a server computer. In other embodiments, the computer system 100 may be a desktop computer, a portable computer, a laptop computer, a tablet computer, a pocket computer, a phone, a smartphone, or any other suitable electronic device.

[0026] FIG. 2 is a diagram showing an example of the configuration of an information processing system 200 according to an embodiment of the present disclosure. The information processing system 200 is a system for generating and outputting basis information regarding prediction results by a machine learning method. As shown in FIG. 2, the information processing system 200 mainly includes an information processing device 210, a communication network 250, and a user terminal 260. The information processing device 210 and the user terminal 260 may be connected to each other via the communication network 250.

[0027] The information processing device 210 is a device for generating and outputting basis information regarding prediction results by a machine learning method. As shown in FIG. 2, it mainly includes a memory 220, a storage unit 230, a processor 244, and an input / output unit 246. In a certain embodiment, the information processing device 210 may be implemented by the computer system 100 shown in FIG. 1.

[0028] The memory 220 may be a memory for storing an information processing application 150 for implementing the functions of the information processing means according to the embodiments of the present disclosure. This information processing application 150 may include processing instructions for implementing the functions of software modules such as a feature extraction unit 222, a combination unit 224, a prediction unit 225, a loss management unit 226, and a basis management unit 228, as shown in FIG. 2.

[0029] The feature extraction unit 222 is a functional unit for extracting feature quantities (also referred to as "image features" in the present disclosure) from the image data input from the input / output unit 246. The feature quantity here is information representing numerical values, symbols, or categories indicating the characteristics and attributes of the data, and can be represented by a multi-dimensional vector. As will be described later, the feature quantity extracted by the feature extraction unit 222 is used by the prediction unit 225 to generate a prediction result regarding the input data. The feature extraction unit 222 here may be composed of a multi-layer neural network that extracts feature quantities from images, such as a convolutional neural network or a Vision Transformer. By performing feature extraction in each layer of the feature extraction unit 222, it is possible to extract feature quantities for grasping more complex features as the layer becomes deeper. Note that since the data input to the feature extraction unit 222 is only image data, it should be noted that noise is less likely to enter the feature quantities extracted in each layer compared to the case where feature quantity extraction is performed collectively on image data and non-image data. The noise here means the error from the information that is originally desired to be obtained in the calculation of extracting feature quantities in each layer.

[0030] The combining unit 224 is a functional unit for generating combined data by combining the image features extracted from the image data by the feature extraction unit 222 and the non-image data. The combining unit 224 may combine the image data and the non-image data, for example, by concatenating a numerical value indicating the characteristics of the non-image data to a multi-dimensional vector indicating the image features. In an embodiment, the combining unit 224 may be composed of a network having a multi-layer structure called a fully connected layer or the like.

[0031] The prediction unit 225 is a functional unit for generating a prediction result regarding the combined data obtained by combining the image data and the non-image data included in the input data. The prediction unit 225 may be composed of, for example, a network having a multi-layer structure or the like. In the present disclosure, the "prediction result" refers to information indicating a prediction generated by analyzing a set of input data composed of image data and non-image data by a predetermined machine learning method. Further, this prediction result may include a plurality of prediction items that are the subjects to be predicted.

[0032] In the present disclosure, the "feature extraction unit 222" and the "prediction unit 225" may be collectively referred to as a "learning unit". Each layer of this learning unit is embedded with "weight parameters", which are adjustable coefficients that define the behavior of the machine learning model. In the learning stage of the machine learning model, by adjusting and updating these weight parameters, the learning unit performs learning and the prediction accuracy is improved. Further, as will be described later, based on the weight parameters of the learning unit, it is possible to generate basis information regarding the prediction result generated by the prediction unit 225. In the present disclosure, the "basis information" refers to information indicating the degree of influence of each element that has influenced the prediction result when generating the prediction result.

[0033] The loss management unit 226 is a functional unit that calculates the loss value for the prediction generated by the prediction unit 225. Further, the loss management unit 226 according to the embodiment of the present disclosure can group prediction items that satisfy a predetermined relevance criterion among the prediction items included in the prediction result, and calculate the loss value for each group. Here, the loss value for each group of prediction items can be calculated by applying a predetermined loss function to the prediction items generated by the prediction unit 225 and the ground truth indicating the correct result of the prediction items. Note that the loss here is information indicating the difference between the prediction items included in the prediction result generated by the prediction unit 225 and the ground truth indicating the correct numerical value of the prediction items.

[0034] The relevance criterion here is a criterion for evaluating the similarity between prediction items, and may be defined based on, for example, the measurement unit, scale, numerical characteristics, semantic information, etc. of the prediction items. By grouping prediction items that satisfy the relevance criterion into the same group and applying the loss function, the loss value can be obtained on an appropriate scale. Further, by performing error backpropagation to the learning unit using the loss value obtained here, it is possible to update the weight parameters of the learning model and obtain a gradient for calculating the basis information.

[0035] The basis management unit 228 is a functional unit that generates basis information for the prediction result by the prediction unit 225. The basis management unit 228 can generate basis information by using the weight parameters embedded in each layer of the learning unit, the image features extracted by the feature extraction unit 222, and the gradient obtained from the loss value calculated by the loss management unit 226. In this way, by suppressing the noise included in the three types of data of the weight parameters, feature amounts, and gradients, high-precision basis information becomes possible.

[0036] The storage unit 230 is a storage area that houses a database (hereinafter, "DB") for storing various information according to the embodiment of the present disclosure, and may include an information processing DB 236 as shown in FIG. 2.

[0037] The information processing DB236 is a database for storing input data (image data and non-image data) used in the present disclosure, image features extracted from the image data by the feature extraction unit 222, weight parameters and gradients of the learning unit, and various information of the machine learning model obtained by learning.

[0038] The processor 244 is a processing unit for executing processing instructions that define the functions of the respective functional units of the information processing application 150 stored by the memory 220.

[0039] The input / output unit 246 is a functional unit for receiving information input to the information processing apparatus 210 (for example, input data including image data and non-image data) and outputting information (such as basis information) generated by the information processing apparatus 210. In one embodiment, the input / output unit 246 may include, for example, a keyboard, a mouse, a display for displaying a GUI (Graphical User Interface), and the like. In one embodiment, the input / output unit 246 may provide a GUI for inputting and outputting various information to the user terminal 260.

[0040] The communication network 250 may include, for example, a local area network (LAN), a wide area network (WAN), a satellite network, a cable network, a WiFi network, or any combination thereof.

[0041] The user terminal 260 is a terminal device that can be used by the user of the information processing apparatus 210. The user can input image data and non-image data to the information processing apparatus 210 and check the basis information output from the information processing apparatus 210 by using the user terminal 260. As an example, the user terminal 260 may include, for example, a smartphone, a smartwatch, a tablet, a personal computer, etc., of a user who subscribes to an information processing service provided by the information processing system 200, and is not particularly limited. In FIG. 2, for the sake of convenience of explanation, a configuration including one user terminal 260 is described as an example. However, the number of user terminals 260 is not limited, and a configuration including a plurality of user terminals 260 is also possible.

[0042] According to the information processing system 200 of the present disclosure described above, when a machine learning method is applied to a heterogeneous data set including image data and non-image data, after the non-image data is input after image feature extraction, by using a loss function in which prediction items are grouped, it is possible to provide information processing means capable of outputting highly accurate basis information for the prediction result with reduced noise.

[0043] Next, with reference to FIG. 3, the data flow in the information processing apparatus according to the embodiment of the present disclosure will be described.

[0044] FIG. 3 is a diagram showing the data flow in the information processing apparatus 210 according to the embodiment of the present disclosure.

[0045] First, the input / output unit 246 receives input data including image data 302 and non-image data 304 from, for example, the user terminal 260 shown in FIG. 2, transfers the image data 302 to the feature extraction unit 222, and transfers the non-image data 304 to the combining unit 224.

[0046] The feature extraction unit 222 extracts image features 306 by performing extraction of feature amounts on the image data 302. After that, the feature extraction unit 222 transfers the extracted image features 306 to the combining unit 224.

[0047] The combining unit 224 generates combined data 308 by combining the image features 306 received from the feature extraction unit 222 and the non-image data received from the input / output unit 246. After that, the combining unit 224 transfers the combined data 308 to the prediction unit 225.

[0048] The prediction unit 225 generates a prediction result 310 regarding the input data based on the combined data 308 received from the combining unit 224. After that, the prediction unit 225 transfers the prediction result 310 to the loss management unit 226.

[0049] The loss management unit 226 groups prediction items that meet a predetermined relevance criterion among the prediction items included in the prediction result 310 received from the prediction unit 225, and calculates a loss value 312 for each group. Then, the loss management unit 226 transfers the loss value 312 to the basis management unit 228. Also, the loss management unit 226 can update the weight parameters of each layer so as to minimize the loss by performing error backpropagation on the feature extraction unit 222 and the prediction unit 225 using the calculated loss value 312.

[0050] The basis management unit 228 uses the gradient obtained from the loss value 312 received from the loss management unit 226, the image features 306 extracted by the feature extraction unit 222, and the weight parameters embedded in each layer of the feature extraction unit 222 and the prediction unit 225 to generate basis information 315 indicating the influence degree of each element that influenced the prediction result 310. Then, the basis management unit 228 outputs the basis information 315 to the user terminal 260 shown in FIG. 2, for example, via the input / output unit 246.

[0051] According to the information processing apparatus 210 of the present disclosure described above, when a machine learning method is applied to a heterogeneous data set including image data and non-image data, after inputting the non-image data after image feature extraction and using a loss function in which prediction items are grouped, it is possible to provide information processing means capable of outputting highly accurate basis information for the prediction result with reduced noise.

[0052] Next, with reference to FIG. 4, an example of the flow of the information processing method according to the embodiment of the present disclosure will be described.

[0053] FIG. 4 is a flowchart showing an example of the flow of an information processing method 400 according to an embodiment of the present disclosure. The information processing method 400 shown in FIG. 4 is a method for generating basis information with suppressed noise for a prediction result by a machine learning method, and may be implemented by each functional unit of the information processing apparatus 210 according to the embodiment of the present disclosure shown in FIGS. 2 and 3, for example.

[0054] First, in step S405, the input / output unit 246 inputs input data including image data 302 and non-image data 304 from, for example, the user terminal 260 shown in FIG. 2. The image data here may be unstructured information visually showing figures, pictures, images, etc. Also, the non-image data here may be structured information indicating conditions or parameters related to the image data in text or numerical values. As an example, the image data may be an image showing, for example, the cross-sectional shape of an optical element, and the non-image data may be numerical information indicating experimental conditions (characteristics such as the wavelength and pitch of light irradiated on the optical element) of an experiment performed on the optical element.

[0055] Next, in step S410, the feature extraction unit 222 extracts image features indicating the characteristics of the image data by performing an existing feature quantity extraction method such as a Convolutional Neural Network (CNN), SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), Vision Transformer, etc. on the image data. The image features here may be represented, for example, as a 256-dimensional multi-dimensional vector. In a certain embodiment, the feature extraction unit 222 may be a neural network with a multi-layer structure having different network architectures such as a convolutional layer and a Vision Transformer layer. In this way, by extracting local features in the convolutional layer and overall features of the image in the Vision Transformer, it is possible to obtain highly robust image features considering the relationship between fine features and overall features.

[0056] Next, in step S415, the combining unit 224 generates combined data by combining the image features extracted from the image data by the feature extraction unit 222 in step S410 and the non-image data received in step S405. In an embodiment, the combining unit 224 may combine the image data and the non-image data by concatenating a numerical value indicating the characteristics of the non-image data to a multi-dimensional vector indicating the image features. As an example, when the image features are, for example, vectors of 256 dimensions and the non-image data are vectors of 4 dimensions, the combining unit 224 may generate a vector of 260 dimensions as the combined data by concatenating the respective vectors of the image features and the non-image data. In this way, by performing feature amount extraction only on the image data and combining the obtained image features with the non-image data, it is possible to suppress the noise included in the feature amount as compared with the case of collectively extracting the feature amounts of the image data and the non-image data.

[0057] Next, in step S420, the prediction unit 225 generates a prediction result regarding the combined data generated in step S415. Here, when configured by a neural network having a multi-layer structure, the prediction unit 225 can generate a prediction result regarding the combined data by treating the calculations in each layer as one-dimensional numerical data instead of two-dimensional data and reducing the dimensions. As described above, this prediction result may include a plurality of prediction items that are the subjects to be predicted. The prediction items to be predicted may be set in advance by the user, and may include, for example, a plurality of items having different units and semantic contents. In this way, by setting a plurality of prediction items with high degrees of freedom, a large number of predictions can be accurately made with one learning model.

[0058] As an example, when performing so-called optical simulation by analyzing input data including image data showing the cross-sectional shape of an optical element and non-image data showing the characteristics (such as wavelength) of light irradiated to the optical element by a machine learning method, the prediction unit 225 may generate a prediction result showing, as a prediction item, a predicted value of the optical characteristics (such as transmittance, reflectance, diffraction angle, etc.) of the optical element based on the combined data obtained by combining these image data and non-image data.

[0059] Note that in the case of a neural network having a multi-layer structure including an input layer, an intermediate layer, and an output layer that are sequentially connected, for example, the combined data may be input to the input layer of the prediction unit 225 or may be input to the intermediate layer (that is, the layer immediately before the output layer) of the prediction unit 225. In this way, by inputting the combined data to a layer as close as possible to the output layer, the computational amount of the prediction unit 225 can be suppressed.

[0060] Next, in step S425, the loss management unit 226 groups prediction items that satisfy a predetermined relevance criterion among the prediction items in the prediction result generated in step S420. As described above, this relevance criterion is a criterion for evaluating the similarity between prediction items, and may be defined based on, for example, the measurement unit, scale, numerical characteristics, semantic information, etc. of the prediction items. As an example, the loss management unit 226 may group, as prediction items that satisfy the relevance criterion, prediction items in the prediction result that have the same measurement unit (nanometer, degree, rate), prediction items having the same numerical characteristics, or prediction items having semantic information obtained by analyzing the prediction result by a predetermined natural language processing and satisfying a predetermined similarity criterion. The number of groups of prediction items is not particularly limited here, but in principle, the smaller the number, the smaller the computational amount. However, if the number of groups is small and prediction items with different units and numerical characteristics are included in the same group, the prediction accuracy may decrease. Therefore, it is desirable to set the number of groups according to the characteristics of the input data, the number, and the types of prediction items. In the present disclosure, numerical characteristics are information that characterizes prediction items represented by numerical values, and may include, for example, numerical ranges (lower limit and upper limit), number of digits, increase patterns (linear, exponential, logarithmic, etc.), and the like. As an example, prediction items of percentages represented by numerical values within the range of "0 to 100" can be grouped into the same group, and prediction items of angles represented by numerical values within the range of "0 to 360" can be grouped into the same group.

[0061] Next, in step S430, the loss management unit 226 calculates a loss value based on the prediction items grouped in step S425. Here, the loss management unit 226 can calculate the loss value for each group of prediction items by applying a predetermined loss function to the prediction items grouped in step S425 and the ground truth indicating the correct result of the prediction items. It should be noted that here, a loss function configured to receive a plurality of grouped prediction items is used to obtain the loss value for each group of prediction items. In this way, by obtaining the loss value for each group of prediction items, it is possible to calculate the loss between similar items without mixing different characteristics (numerical characteristics, units, semantic information) of the prediction items, so that a more accurate loss value can be obtained and the accuracy of the learning model can be improved. Note that when performing the loss function, the loss management unit 226 may apply a mask to any number of prediction items among the prediction items. This mask is for excluding predetermined prediction items from loss calculation and subsequent generation of basis information. By applying a mask to prediction items other than the prediction items for which generation of basis information is desired, basis information is generated only for the target prediction items. This enables fine adjustment of the basis information and saves computing resources.

[0062] Next, in step S435, the loss management unit 226 uses the loss values for each group of prediction items calculated in step S430 to obtain a gradient for minimizing the loss value by performing a gradient descent method such as error backpropagation on each layer of the feature extraction unit 222 and the prediction unit 225. The gradient here is, for example, a vector of partial derivatives with respect to the weight parameters embedded in each layer of the feature extraction unit 222 and the prediction unit 225 configured by a neural network, and for each weight parameter, it indicates the rate of change of the loss with respect to the parameter. In this way, when a gradient descent method such as error backpropagation is performed on each layer of the feature extraction unit 222 and the prediction unit 225, the weight parameters of each layer advance in the gradient direction and are updated to minimize the loss value, and the learning model is trained.

[0063] Next, in step S440, the basis management unit 228 calculates the attention degree for the combined data based on the gradient obtained in step S435 and the weight parameters embedded in each layer of the feature extraction unit 222 and the prediction unit 225. The attention degree here is a measure indicating the importance of each piece of information that has influenced the prediction result when generating the prediction result. The attention degree for this combined data may include the attention degree for the image data and the attention degree for the non-image data. In a certain embodiment, the attention degree for the combined data can be obtained, for example, by the following mathematical formula 1.

Equation

[0064] In step S445A and step S450A, the basis management unit 228 generates a saliency map as basis information for the prediction result. This saliency map is information indicating the degree of influence of the area that has affected the prediction item of the prediction result in the image data.

[0065] In step S445A, the basis management unit 228 calculates a value obtained by multiplying the degree of attention to the image data by the image features of the final layer of the neural network constituting the feature extraction unit. In step S450A, the basis management unit 228 sums up the value obtained by multiplying the degree of attention and the image features calculated in step S445A, and generates and outputs a saliency map by imaging. In this way, it becomes possible to generate a saliency map considering the influence of the mask described above.

[0066] In step S445B and step S450B, the basis management unit 228 generates parameter attention information as basis information for the prediction result. This parameter attention information is information indicating the degree of influence of parameters (experimental conditions, etc.) that have affected the prediction item of the prediction result in non-image data.

[0067] In step S445B, the basis management unit 228 calculates the average value of the degree of attention to the image data to calculate the degree of attention of the entire image. In step S450B, the basis management unit 228 compares the degree of attention of the entire image calculated in step S445B with the degree of attention to the non-image data, and generates and outputs parameter attention information by quantifying their relative relationships. In this way, parameter attention information considering the influence of the mask described above is generated, and it is possible to compare different forms of data such as image data and non-image data.

[0068] According to the information processing method 400 described above, when applying a machine learning method to a heterogeneous data set including image data and non-image data, after inputting the non-image data after image feature extraction, by using a loss function in which prediction items are grouped, it is possible to provide information processing means capable of outputting highly accurate basis information for the prediction result with reduced noise.

[0069] Next, with reference to FIGS. 5 to 6, a specific example of applying the information processing means according to the embodiment of the present disclosure to optical simulation will be described.

[0070] FIG. 5 is a diagram showing an example of the configuration of an information processing apparatus 510 when the information processing means according to the embodiment of the present disclosure is applied to optical simulation. The information processing apparatus 510 shown in FIG. 5 can generate a learning model for predicting a prediction result 610 indicating the optical characteristics of an optical element based on image data 602 showing the cross-sectional shape of the optical element and experimental conditions which are non-image data 604 showing characteristics of light irradiated on the optical element. Here, the non-image data 604 may include, for example, the pitch, polarization, wavelength, etc. of the light irradiated on the optical element.

[0071] Note that, for convenience of explanation, in the configuration of the information processing apparatus 510 shown in FIG. 5, illustration of a part of the functional units is omitted, but it should be noted that the information processing apparatus 510 is configured substantially the same as the information processing apparatus 210 described above. Also, hereinafter, for convenience of explanation, descriptions overlapping with the above description will be omitted.

[0072] As shown in FIG. 5, the information processing apparatus 510 includes a feature extraction unit 522 and a prediction unit 525 having a multi-layer structure, and combined data obtained by combining image features extracted from the image data 602 and the non-image data 604 is input to an intermediate layer of the neural network by a combining unit 524 ("Concat" in FIG. 5). The prediction unit 525 predicts a predicted value of the optical characteristics of the optical element shown in the input data as a prediction item of the prediction result 610 based on this combined data.

[0073] After that, the loss management department (not shown in FIG. 6) groups the prediction items that meet the relevance criteria among these prediction items into the same group, and calculates the loss value for each group. As an example, when there are eight types of predicted values as prediction items, the loss management department may group these eight types of predicted values into four groups of transmittance, transmission diffraction angle, reflectance, and reflection diffraction angle, and obtain the loss value.

[0074] FIG. 6 is a diagram showing an example of input / output information when the information processing means according to the embodiment of the present disclosure is applied to optical simulation. In FIG. 6, the input / output information of the information processing apparatus 510 according to the embodiment of the present disclosure may include image data 602, non-image data 604, saliency map 606, parameter attention information 608, prediction result 610, and ground truth 612.

[0075] As described above, the image data 602 is image data showing the cross-sectional shape of the optical element that is the object of optical simulation. The non-image data 604 is a numerical value indicating the experimental conditions of the optical simulation, and may include, for example, the pitch, polarization, wavelength, etc. of the light irradiated to the optical element.

[0076] The saliency map 606 is a type of basis information indicating the influence degree of the region of the image data 602 that has affected the prediction items in the prediction result 610 in the image data 602. As shown in FIG. 6, the saliency map 606 may indicate the influence degree of the region of the image data 602 that has affected the prediction items in different colors (for example, the blue region has a low influence degree, the yellow region has a medium influence degree, and the red region has a high influence degree). As an example, in the saliency map 606 shown in FIG. 6, the region on the right side of the cross-sectional shape of the optical element shown in the image data 602 has a high influence degree on the prediction items. This saliency map 606 is obtained by summing up all the values obtained by multiplying the feature amounts in each layer of the feature extraction unit by the gradient obtained from the loss value.

[0077] The parameter attention degree information 608 is a kind of basis information indicating the influence degree of the experimental conditions (set parameters) that have affected the prediction items of the prediction result 610. As shown in FIG. 6, the parameter attention degree information 608 may be expressed, for example, as a percentage. As an example, in the parameter attention degree information 608 in the first column shown in FIG. 6, the influence degree on the prediction result 610 of the image data 602 ("Img" in FIG. 6) is "40.573%", and the influence degree on the prediction result 610 of the pitch ("Pitch" in FIG. 6) in the non-image data 604 is "21.189%". This parameter attention degree information 608 is obtained by comparing the weight parameters of the prediction unit with the total of the weight parameters applied to each input information.

[0078] The prediction result 610 is information including a plurality of prediction items that are the subjects to be predicted. As shown in FIG. 6, the prediction result 610 may include, for example, transmittance, reflectance, etc. as the predicted optical characteristics of the optical element. Note that the prediction items included in the prediction result 610 may be freely set by the user and are not particularly limited here.

[0079] The ground truth 612 is information indicating the correct optical characteristics of the optical element. By using this ground truth 612, the loss value of the prediction items included in the prediction result 610 can be calculated.

[0080] With reference to the input / output information in FIG. 6 described above, an example of the operation of the information processing apparatus according to the embodiment of the present disclosure will be described. First, when the image data 602 is input to the feature extraction unit of the information processing apparatus according to the embodiment of the present disclosure, feature amounts are calculated in each layer within the feature extraction unit. The image features finally obtained by the feature extraction unit and the information on the experimental conditions of the non-image data 604 are combined and input to the prediction unit. Similar to the feature extraction unit, feature amounts are also calculated in each layer constituting the prediction unit. Then, the set of prediction items finally obtained by the prediction unit becomes the prediction result 610.

[0081] Next, the loss management unit groups the prediction items of the prediction result 610 based on the above-described predetermined relevance, and obtains the loss value of each group by comparing with the ground truth 612 using a predetermined loss function. Here, the loss management unit may group the prediction items of the prediction result 610 into four groups: "-1st transmittance, 0th transmittance, 1st transmittance" (Transmittance), "transmittance diffraction angle" (Transmittance angle), "-1st reflectance, 0th reflectance, 1st reflectance" (Reflectance), and "transmittance reflection angle" (Reflectance angle).

[0082] Thereafter, by performing error backpropagation on the feature extraction unit and the prediction unit, the gradients used to generate the basis information can be obtained for each layer. By multiplying and calculating the gradients, image features, and weight parameters of each layer obtained in this way, the saliency map 606 and parameter attention information 608 shown in FIG. 6 can be generated as the basis information.

[0083] According to the research of the inventors of the present disclosure, when the information processing according to the embodiment of the present disclosure is applied to optical simulation, the difference between the ground truth 612 and the prediction items in the prediction result 610 is 1% or less, and the calculation speed is about 900 times that of the conventional method. This is an improvement in accuracy obtained by inputting combined data into an intermediate layer close to the output layer of the neural network constituting the prediction unit and calculating the loss value for each group in which similar prediction items are grouped.

[0084] <Application Example of the Present Disclosure> The information processing means according to the embodiment of the present disclosure described above can be suitably applied to various fields. For example, when the information processing means according to the embodiment of the present disclosure is applied to the above-described optical simulation, fluid simulation, etc., design data (2D or 3D) and experimental conditions are set as inputs, and the simulation result is predicted, so that basis information that cannot be obtained by simulation software can be obtained, and the parts to be emphasized in the design can be easily grasped.

[0085] Also, in a program that makes important determinations such as a medical diagnosis system, whether to trust the determination by artificial intelligence is sometimes decided by humans. Therefore, the transparency of the basis that affects the determination by artificial intelligence becomes important. Thus, for example, when an X-ray image and other clinical data are input, by determining the presence or absence of a fracture using the information processing means according to the embodiment of the present disclosure described above, it is possible to obtain basis information indicating which bones shown in the X-ray image have affected the determination of the presence or absence of a fracture. Therefore, the reliability of the determination by artificial intelligence can be more accurately evaluated.

[0086] Also, in an in-vehicle system or a security system, there is a scenario where real-time segmentation is performed on an image obtained from a camera, and feature amounts in a learning model are used for the segmentation. Thus, for example, when an in-vehicle camera image and the speed of the vehicle are input, by determining whether there is a person within a dangerous range using the information processing means according to the embodiment of the present disclosure, noise during feature extraction can be reduced, and the accuracy and calculation speed can be improved. Therefore, a high-precision prediction result can be obtained in real time, and driving control considering safety becomes possible.

[0087] As described above, in the conventional method, when performing machine learning analysis on heterogeneous data that combines image data and non-image data with other complex structures, noise is included in the feature amounts extracted from the data, so the accuracy of the basis information may deteriorate. Therefore, according to the present disclosure, when applying a machine learning method to a heterogeneous data set including image data and non-image data, after inputting the non-image data after image feature extraction and using a loss function that groups prediction items, it is possible to provide information processing means capable of outputting high-precision basis information for the prediction result with reduced noise.

[0088] As described above, conventionally, when extracting feature amounts from a heterogeneous data set that combines image data and non-image data, there is a problem that noise is included in the feature amounts, and the accuracy of the prediction result and the basis for the prediction result is limited. Therefore, in the present disclosure, in order to reduce noise, feature extraction is performed only on the image data, and prediction is performed based on the combined data obtained by combining the features of the image thus extracted and the non-image data. By performing feature extraction only on the image data in this way, noise can be reduced compared to the case where feature extraction is performed on the image data and the non-image data together.

[0089] In addition, in the present disclosure, among the respective prediction items, for example, prediction items having relevance such as similar numerical characteristics are grouped, and a loss function that sums the loss values obtained for each group is used. As a result, without significantly increasing the amount of calculation, it is possible to grasp the accurate loss value of each prediction item, and the prediction accuracy of the entire machine learning model can be improved. Further, by also using this loss function when performing the basis output, it is possible to accurately grasp the loss for each prediction item.

[0090] As described above, according to the information processing means according to the embodiment of the present disclosure, it is possible to output basis information with reduced noise, and it is possible to easily grasp the degree of influence of each element (for example, a specific region in the image data, setting parameters included in the non-image data, etc.) that affects the prediction result on the prediction result. In addition, by omitting feature extraction for non-image data, computing resources can be saved and the processing time can be shortened.

[0091] As described above, the information processing means according to the embodiment of the present disclosure includes the following aspects.

[0092] (Aspect 1) An information processing apparatus, comprising a processor and a memory, wherein the memory has an acquisition unit that acquires an input data set including image data and non-image data, a feature extraction unit that extracts image features indicating the characteristics of the image data from the image data, A combining unit that generates combined data by combining the image features extracted from the image data and the non-image data; A prediction unit that generates a prediction result indicating a prediction item regarding the input data set based on the combined data; A loss management unit that groups prediction items satisfying a predetermined relevance criterion among the prediction items in the prediction result and calculates a loss value for each group; A basis management unit that calculates a gradient for minimizing the loss value and generates and outputs basis information indicating the degree of influence of elements that have influenced the prediction item for each prediction item in the prediction result based at least on the gradient; An information processing apparatus comprising processing instructions for causing the processor to function as described above.

[0093] (Aspect 2) The feature extraction unit is composed of a neural network including at least a first layer and a second layer, wherein the first layer and the second layer have different network architectures, The information processing apparatus according to aspect 1, characterized in that.

[0094] (Aspect 3) In the feature extraction unit, the first layer is a convolutional layer, and the second layer is a Vision Transformer, The information processing apparatus according to aspect 2, characterized in that.

[0095] (Aspect 4) The prediction unit includes at least an input layer, an intermediate layer, and an output layer that are sequentially connected, The combining unit inputs the combined data to the intermediate layer, The information processing apparatus according to any one of aspects 1 to 3, characterized in that.

[0096] (Aspect 5) The loss management unit Among the prediction items in the prediction result, Group prediction items with the same measurement unit as prediction items that satisfy the relevance criterion, and calculate the loss value. The information processing apparatus according to Aspects 1 to 4, characterized in that.

[0097] (Aspect 6) The loss management unit Among the prediction items in the prediction result, Group prediction items having the same numerical characteristics as prediction items that satisfy the relevance criterion, and calculate the loss value. The information processing apparatus according to Aspects 1 to 5, characterized in that.

[0098] (Aspect 7) The loss management unit By analyzing the prediction result by a predetermined natural language processing, determine semantic information regarding each prediction item. Among the prediction items in the prediction result, Group prediction items having semantic information that satisfies a predetermined similarity criterion as prediction items that satisfy the relevance criterion, and calculate the loss value. The information processing apparatus according to Aspects 1 to 6, characterized in that.

[0099] (Aspect 8) The basis management unit Among each prediction item in the prediction result, Specify a first prediction item, Generate the basis information only for the first prediction item by applying a mask to prediction items other than the first prediction item included in the prediction result. The information processing apparatus according to Aspects 1 to 7, characterized in that.

[0100] (Aspect 9) The basis information The image data is A saliency map indicating the degree of influence of the region in the image data that affected the prediction item of the prediction result, and In the non-image data, parameter attention information indicating the degree of influence of the setting parameter that has influenced the prediction item of the prediction result is included. The information processing apparatus according to Aspects 1 to 8, characterized in that.

[0101] (Aspect 10) The basis management unit By multiplying the gradient and the weight parameter in the layer of the prediction unit, the attention degree for the image data and the attention degree for the non-image data are calculated. The saliency map is generated based on the value obtained by summing the values obtained by multiplying the attention degree for the image data and the image features of the final layer of the feature extraction unit. The information processing apparatus according to Aspect 9, characterized in that.

[0102] (Aspect 11) The basis management unit By multiplying the gradient and the weight parameter in the layer of the prediction unit, the attention degree for the image data and the attention degree for the non-image data are calculated. The average value of the attention degree for the image data is calculated. Based on the relative relationship between the average value and the attention degree for the non-image data. The parameter attention information is generated. The information processing apparatus according to Aspect 9, characterized in that.

[0103] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist of the present invention.

Explanation of Reference Numerals

[0104] 150 Information processing application 200 Information processing system 210 Information processing apparatus 220 Memory 222 Feature extraction unit 224 Combining unit 225 Prediction Unit 226 Loss Management Unit 228 Basis Management Unit 230 Memory Unit 236 Information Processing DB 244 Processor 246 Input / Output Unit 250 Communication Network 260 User Terminal

Claims

1. An information processing apparatus comprising: a processor and a memory, wherein the memory has an acquisition unit that acquires an input data set including image data and non-image data, a feature extraction unit that extracts image features indicating characteristics of the image data from the image data, a combination unit that generates combined data by combining the image features extracted from the image data and the non-image data, a prediction unit that generates a prediction result indicating a prediction item regarding the input data set based on the combined data, a loss management unit that groups prediction items in the prediction result that satisfy a predetermined relevance criterion, and calculates a loss value for each group, a basis management unit that calculates a gradient for minimizing the loss value, and generates and outputs basis information indicating the degree of influence of elements that have influenced each prediction item in the prediction result based at least on the gradient; An information processing apparatus characterized by including processing instructions for causing the processor to function as described above.

2. The feature extraction unit is composed of a neural network including at least a first layer and a second layer, wherein the first layer and the second layer have different network architectures, The information processing apparatus according to claim 1, characterized in that.

3. In the feature extraction unit, the first layer is a convolutional layer, the second layer is a Vision Transformer, The information processing apparatus according to claim 2, characterized in that.

4. The prediction unit includes at least an input layer, an intermediate layer, and an output layer that are sequentially connected, The combination unit inputs the combined data to the intermediate layer of the prediction unit, The information processing apparatus according to claim 1, characterized in that.

5. The loss management unit groups prediction items in the prediction result that have the same measurement unit as prediction items that satisfy the relevance criterion, and calculates the loss value, The information processing apparatus according to claim 1, characterized in that.

6. The loss management unit groups prediction items in the prediction result that have the same numerical characteristics as prediction items that satisfy the relevance criterion, and calculates the loss value, The information processing apparatus according to claim 1, characterized in that.

7. The loss management unit analyzes the prediction result by a predetermined natural language processing to determine semantic information regarding each prediction item, Among the prediction items in the prediction result, group the prediction items having semantic information satisfying a predetermined similarity criterion as the prediction items satisfying the relevance criterion, and calculate the loss value. The information processing apparatus according to claim 1, characterized in that.

8. The basis management unit Among each prediction item in the prediction result, specify a first prediction item, generate the basis information only for the first prediction item by applying a mask to the prediction items other than the first prediction item included in the prediction result. The information processing apparatus according to claim 1, characterized in that.

9. The basis information The image data is a saliency map indicating the degree of influence of the region in the image data that affected the prediction item of the prediction result, and in the non-image data, including parameter attention information indicating the degree of influence of the set parameters that affected the prediction item of the prediction result. The information processing apparatus according to claim 1, characterized in that.

10. The basis management unit calculate the attention degree for the image data and the attention degree for the non-image data by multiplying the gradient and the weight parameters in the layer of the prediction unit, generate the saliency map based on the value obtained by summing the values obtained by multiplying the attention degree for the image data and the image features of the final layer of the feature extraction unit. The information processing apparatus according to claim 9, characterized in that.

11. The basis management unit calculate the attention degree for the image data and the attention degree for the non-image data by multiplying the gradient and the weight parameters in the layer of the prediction unit, calculate the average value of the attention degree for the image data, generate the parameter attention information based on the average value and the relative relationship of the attention degree for the non-image data. The information processing apparatus according to claim 9, characterized in that.

12. An information processing method implemented in an information processing apparatus, The information processing apparatus includes a processor and a memory, The information processing method by the processing instructions stored in the memory, a step of acquiring an input data set including image data and non-image data, a step of extracting image features indicating the characteristics of the image data from the image data, a step of generating combined data by combining the image features extracted from the image data and the non-image data. generating a prediction result indicating prediction items regarding the input data set based on the combined data; grouping prediction items that satisfy a predetermined relevance criterion among the prediction items in the prediction result, and calculating a loss value for each group; calculating a gradient for minimizing the loss value, and generating and outputting basis information indicating the degree of influence of elements that have influenced each prediction item in the prediction result based at least on the gradient; An information processing method, characterized in that the processor is caused to execute the method.

13. An information processing method implemented in an information processing apparatus, wherein the information processing apparatus comprises a processor and a memory, and the information processing method obtains an input data set including image data indicating a cross-sectional shape of an optical element and non-image data indicating experimental conditions of an optical simulation for the optical element by a processing instruction stored in the memory; extracting image features indicating characteristics of the image data from the image data; generating combined data by combining the image features extracted from the image data and the non-image data; generating a prediction result indicating prediction items regarding optical characteristics of the optical element based on the combined data; grouping each prediction item in the prediction result into a group of transmittance, transmission diffraction angle, reflectance, and reflection diffraction angle, and calculating a loss value for each group; calculating a gradient for minimizing the loss value, and generating and outputting basis information including a saliency map indicating the degree of influence of a region in the image data that has influenced the prediction item of the prediction result and parameter attention information indicating the degree of influence of experimental conditions that have influenced the prediction item of the prediction result in the non-image data for each prediction item in the prediction result based at least on the gradient; An information processing method, characterized in that the processor is caused to execute the method. ​

Citation Information

Patent Citations

  • Automated data analysis method for non-tabular data, related system and apparatus

    JP2023514282A