Information processing device, information processing method, and program

By generating virtual spectra through dimensionality compression and restoration of virtual data, the apparatus addresses spectral measurement inconsistencies, enhancing the accuracy and cost-effectiveness of spectral data for machine learning models.

JP2026074764AActive Publication Date: 2026-05-07NAT AGRI & FOOD RES ORG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NAT AGRI & FOOD RES ORG
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing technologies face challenges in reducing differences in spectral measurement conditions, which affect the accuracy and consistency of spectral data across different devices and measurement setups, leading to inconsistencies in machine learning models trained on such data.

Method used

An information processing apparatus and method that involves acquiring multiple spectra, generating correspondence data, compressing dimensions using algorithms like PCA, adding virtual data to a feature space, and restoring it to original dimensions to generate virtual spectra, thereby reducing measurement condition differences.

Benefits of technology

This approach allows for the generation of virtual spectra that reflect measurement conditions at a lower cost, improving the accuracy and consistency of spectral data, thus enhancing the performance of machine learning models trained on these data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074764000001_ABST
    Figure 2026074764000001_ABST
Patent Text Reader

Abstract

The present invention provides an information processing device and method that reduces differences in spectral measurement conditions at a lower cost. [Solution] In an information processing device, the control unit includes an acquisition unit that acquires multiple spectra, a first correspondence generation unit that generates correspondence data representing the correspondence between spectra for the acquired multiple spectra, a space generation unit that compresses the dimensions in the correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data, a data addition unit that adds one or more virtual data to the feature space, a second correspondence generation unit that restores the added virtual data to the same dimensions as the correspondence data generated by the first correspondence generation unit and generates further correspondence data, an additional acquisition unit that acquires one or more additional spectra, and a virtual spectrum generation unit that generates one or more virtual spectra using one or more additional spectra and the correspondence data generated by the first correspondence generation unit and the second correspondence generation unit, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, learning data, an information processing method, and a program. <00000​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

[0006] To solve the above problems, an information processing apparatus according to one aspect of the present invention includes: an acquisition unit that acquires a plurality of spectra; a first correspondence generation unit that generates correspondence data representing the correspondence between spectra for the plurality of spectra acquired by the acquisition unit; a space generation unit that compresses the dimensions of the correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data; a data addition unit that adds one or more virtual data to the feature space; a second correspondence generation unit that restores the added virtual data to the same dimensions as the correspondence data generated by the first correspondence generation unit and further generates the correspondence data; an additional acquisition unit that acquires one or more additional spectra; and a virtual spectrum generation unit that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated by the first correspondence generation unit and the second correspondence generation unit, respectively.

[0007] To solve the above problems, an information processing method according to one aspect of the present invention includes: an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing the correspondence between spectra for the plurality of spectra acquired in the acquisition process; a space generation process for compressing the dimensions of the correspondence data using a predetermined algorithm and generating a feature space which is a space containing the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimensions as the correspondence data generated in the first correspondence generation process and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; and a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, respectively.

[0008] To solve the above problems, a program according to one aspect of the present invention causes a computer to execute: an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing the correspondence between spectra for the plurality of spectra acquired in the acquisition process; a space generation process for compressing the dimensions of the correspondence data using a predetermined algorithm and generating a feature space which is a space containing the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimensions as the correspondence data generated in the first correspondence generation process and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; and a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, respectively. [Effects of the Invention]

[0009] According to one aspect of the present invention, differences in spectral measurement conditions can be reduced at a lower cost. [Brief explanation of the drawing]

[0010] [Figure 1] This is a block diagram showing an example configuration of an information processing device according to the present invention. [Figure 2] This figure shows an example of spectral division according to the present invention. [Figure 3] This figure shows an example of a difference spectrum according to the present invention. [Figure 4] This figure shows an example of principal component analysis in a difference spectrum according to the present invention. [Figure 5] This figure shows an example of adding virtual data according to the present invention. [Figure 6] This figure shows an example of a difference spectrum after adding virtual data according to the present invention. [Figure 7] This figure shows an example of generating a virtual spectrum according to the present invention. [Figure 8] This is a flowchart showing an example of the flow of the information processing method according to the present invention. [Figure 9] This is a diagram showing an example of a regression coefficient according to the present invention. [Figure 10] This is a diagram showing an example of a spectral conversion coefficient according to the present invention. [Figure 11] This is a diagram showing an example of principal component analysis in the spectral conversion coefficient according to the present invention. [Figure 12] This is a diagram showing an example of adding virtual data according to the present invention. [Figure 13] This is a diagram showing an example of generating a virtual spectrum according to the present invention. [Figure 14] This is a diagram showing an example of an estimation result by an estimation model according to the present invention. [Figure 15] This is a diagram showing an example of adding virtual data according to the present invention. [Figure 16] This is a diagram showing an example of an estimation result by an estimation model according to the present invention.

Mode for Carrying Out the Invention

[0011] 〔Embodiment 1〕 Hereinafter, an embodiment of the present invention will be described in detail.

[0012] (Configuration of Information Processing Apparatus 1) The configuration of the information processing apparatus 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing a configuration example of the information processing apparatus 1. The information processing apparatus 1 includes, for example, as shown in FIG. 1, a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.

[0013] (Control Unit 10) The control unit 10 comprehensively controls each unit of the information processing apparatus 1. The control unit 10 includes, for example, as shown in FIG. 1, an acquisition unit 11, a first correspondence generation unit 12, a space generation unit 13, a data addition unit 14, a second correspondence generation unit 15, an additional acquisition unit 16, and a virtual spectrum generation unit 17.

[0014] (Acquisition Unit 11) The acquisition unit 11 acquires multiple spectra. A spectrum is obtained by decomposing information or a signal into its (continuous or discontinuous) components and arranging each component by its "value (magnitude, intensity, frequency, etc.)". Specific examples of spectra acquired by the acquisition unit 11 include spectral spectra, frequency spectra, electromagnetic spectra, and mass spectra. For example, the spectra acquired by the acquisition unit 11 may be those measured using a measuring device.

[0015] For example, the information processing device 1 and the measuring device may be connected via a network N, as described later. Furthermore, for example, the information processing device 1 may further include a configuration for measuring the spectrum provided by the aforementioned measuring device. The multiple spectra acquired by the acquisition unit 11 may be measured using, for example, one or more measuring devices. That is, the acquisition unit 11 may acquire spectra measured by each of the multiple measuring devices, or it may acquire multiple spectra measured by a single measuring device.

[0016] Figure 2 shows an example of segmentation in a spectrum measured using a measuring device. For example, the measuring device may be configured with one master unit and multiple slave units, as shown in Figure 2.

[0017] The spectra measured using the measuring device may be divided into, for example, a conversion set, a training set (parent unit only), and a validation set, as shown in Figure 2. The conversion set consists of spectra used to convert spectra between the parent unit and the child unit. The conversion set also consists of spectra of a sample common to both the parent unit and the child unit, as shown in Figure 2. The training set consists of data in which the spectrum of a sample is linked to a target variable whose explanatory variable is that spectrum. Specific examples of the target variable include sugar content and protein content. The validation set consists of spectra used to validate a machine learning model trained using the processing results from the information processing device 1, both in the parent unit and the child unit. The validation set also consists of spectra of a sample common to both the parent unit and the child unit, as shown in Figure 2.

[0018] Here, the conversion set is an example of multiple spectra acquired by the acquisition unit 11. The learning set is an example of additional spectra acquired by the additional acquisition unit 16, which will be described later. In the embodiments described later, an embodiment using a verification set will be explained.

[0019] In the example described with reference to Figures 2 to 4, as shown in Figure 2, for convenience, three slave units are configured, with the master unit assigned the number #1 (Machine #1) and the three slave units assigned the numbers #2 to #4 (Machine #2 to #4).

[0020] (First correspondence generation unit 12) The first correspondence generation unit 12 generates correspondence data representing the correspondence between spectra for the multiple spectra acquired by the acquisition unit 11. The correspondence data may include, for example, data representing the correspondence between any two spectra among the multiple spectra acquired by the acquisition unit 11. For example, if the acquisition unit 11 acquires spectra measured in each of multiple measuring devices, the correspondence data may include data representing the correspondence between any two of those multiple devices. Also, for example, if the acquisition unit 11 acquires spectra of multiple samples measured in one measuring device, the correspondence data may include data representing the correspondence between any two spectra among the spectra of those multiple samples.

[0021] For example, the correspondence data may cover all combinations of two spectra from among the multiple spectra acquired by the acquisition unit 11.

[0022] Furthermore, for example, the correspondence data may be a difference spectrum representing the difference between spectra. A difference spectrum is calculated by determining the difference for each wavelength between the measured values ​​in two spectra. Alternatively, for example, the correspondence data may be a conversion coefficient that converts spectra between spectra. Examples of conversion coefficients will be discussed later with reference to Figures 8 to 12.

[0023] (When the corresponding data is a difference spectrum) Figure 3 shows a difference spectrum, which is an example of correspondence data. For example, when the acquisition unit 11 acquires the transformation sets corresponding to each of the Machines #1 to #4 shown in Figure 2 (hereinafter also referred to as transformation sets #1 to #4), the first correspondence generation unit 12 may generate a difference spectrum representing the difference between two of the transformation sets #1 to #4, as shown in Figure 3.

[0024] Figure 3 shows, as an example, the difference spectrum obtained when the acquisition unit 11 acquires a spectrum showing the wavelength dependence of absorbance. In the difference spectrum graph illustrated in Figure 3, the horizontal axis represents wavelength, and the vertical axis represents the difference in absorbance.

[0025] In the following explanation, we will refer to the difference spectrum "2->1" as an example in Figure 3, but the same applies to other difference spectra. The difference spectrum "2->1" in Figure 3 represents the conversion from the source conversion set #2 to the destination conversion set #1. In other words, the difference spectrum "2->1" represents the result of subtracting the absorbance in conversion set #2 from the absorbance in conversion set #1 for each wavelength.

[0026] The first correspondence generation unit 12 may, for example, as shown in Figure 3, generate a difference spectrum that covers all combinations of two conversion sets from conversion sets #1 to #4 acquired by the acquisition unit 11.

[0027] The above configuration allows for the generation of data representing the differences in spectral measurement conditions.

[0028] (Space generation unit 13) The spatial generation unit 13 compresses the dimensions of the correspondence data using a predetermined algorithm and generates a feature space which is the space containing the compressed data. Specific examples of the predetermined algorithm include principal component analysis (PCA), kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or autoencoder.

[0029] Figure 4 shows an example of the result of dimensionality compression performed by the spatial generation unit 13 using principal component analysis in the difference spectrum exemplified in Figure 3. Here, the result of dimensionality compression performed by the spatial generation unit 13 using principal component analysis may represent, for example, a feature space which is the space containing the data after dimensionality compression. That is, the spatial generation unit 13 may, for example, compress the dimensions of the correspondence data using principal component analysis, as shown in Figure 4, and generate a feature space which is the space containing the compressed data. In the example in Figure 4, as a result of dimensionality compression of the difference spectrum by the spatial generation unit 13, mutually orthogonal first principal component axis PC1 and second principal component axis PC2 are set. Also, in the example in Figure 4, points corresponding to each difference spectrum (difference spectrum "2->1", etc.) exemplified in Figure 3 are shown in the feature space where the first principal component axis PC1 and second principal component axis PC2 are set.

[0030] The above configuration makes it possible to generate a feature space, which is a low-dimensional space characterized by differences in the measurement conditions of the spectrum.

[0031] (Data addition section 14) The data addition unit 14 adds one or more virtual data points to the feature space. Virtual data points are virtual points newly added to arbitrary coordinates (hereinafter also referred to as "grids") on the feature space generated by the space generation unit 13 where no data exists.

[0032] Figure 5 shows an example of adding virtual data to the feature space illustrated in Figure 4 by the data addition unit 14. The feature space illustrated in Figure 5 is a space generated by the space generation unit 13 using principal component analysis, and, similar to the feature space illustrated in Figure 4, a first principal component axis PC1 and a second principal component axis PC2 are set. In the example in Figure 5, "Data" represents data generated by compressing the dimensions of the correspondence data. Here, the data addition unit 14 adds virtual data (Sampled) to the feature space generated by the space generation unit 13, for example, as shown in Figure 5.

[0033] With the above configuration, it is possible to add any quantity of virtual data that reflects the differences in the spectral measurement conditions.

[0034] (Second correspondence generation unit 15) The second correspondence generation unit 15 restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit 12, and generates further correspondence data. In other words, the processing in the second correspondence generation unit 15 corresponds to the inverse transformation of the processing in the spatial generation unit 13.

[0035] Figure 6 shows an example of a difference spectrum including the difference spectrum generated by the first correspondence generation unit 12 and the difference spectrum generated by the second correspondence generation unit 15. Here, the difference spectrum generated by the second correspondence generation unit 15 is the difference spectrum corresponding to the virtual data added by the data addition unit 14.

[0036] (Additional acquisition section 16) The additional acquisition unit 16 acquires one or more additional spectra. The additional spectra acquired by the additional acquisition unit 16 are similar to the spectra acquired by the acquisition unit 11, but are acquired in addition to the spectra acquired by the acquisition unit 11. For example, each spectrum in the one or more additional spectra may be associated with a predetermined target variable. The learning set described above in the explanation with reference to Figure 2 is an example of additional spectra acquired by the additional acquisition unit 16.

[0037] (Virtual spectrum generation unit 17) The virtual spectrum generation unit 17 generates one or more virtual spectra using one or more additional spectra and the correspondence data generated by the first correspondence generation unit 12 and the second correspondence generation unit 15, respectively. For example, if the correspondence data generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 are difference spectra, the virtual spectrum generation unit 17 may generate a virtual spectrum by adding the absorbance of one or more additional spectra and the absorbance of the difference spectrum for each wavelength. Alternatively, for example, when the virtual spectrum generation unit 17 generates one or more virtual spectra from one spectrum, it may associate a predetermined target variable with the one or more virtual spectra.

[0038] Figure 7 shows an example of virtual spectrum generation using difference spectra by the virtual spectrum generation unit 17. The virtual spectrum generation unit 17 may, for example, generate virtual spectra (composite 1 to composite 100) by adding, wavelength by wavelength, the absorbance in the learning set of the master unit acquired by the additional acquisition unit 16 and the absorbance in the difference spectra generated by the first corresponding generation unit 12 and the second corresponding generation unit 15, respectively, as shown in Figure 7.

[0039] With the above configuration, it is possible to generate virtual spectra in any quantity that reflect the differences in the measurement conditions of the spectra.

[0040] Furthermore, the virtual spectrum in the above configuration can be generated at a lower cost than measuring the spectrum using a measuring device. Therefore, the above configuration can reduce the differences in spectrum measurement conditions at a lower cost.

[0041] (Training data generated using information processing device 1) The data containing one or more virtual spectra generated by the information processing device 1 may, for example, be training data for training a machine learning model that takes spectra as input and a target variable as output. For example, PLS (Partial Least Squares) regression, LASSO (Least Absolute Shrinkage and Selection Operator) regression, SVM (Support Vector Machine) regression, ensemble learning, neural networks, etc., may be applied to the training data and the corresponding target variable.

[0042] With the above configuration, the machine learning model can be trained to be less affected by differences in spectral measurement conditions.

[0043] (Storage unit 20) The storage unit 20 stores various data referenced by the control unit 10, as well as various data generated by the control unit 10. Specific examples of data stored in the storage unit 20 include: • Spectral data SPE • Mapping data REL • Algorithm Data ALG • Feature Spatial Data SPA • Virtual spectral data IMA • Training data TD These are some examples.

[0044] Spectral data (SPE) is obtained by decomposing information or signals into their (continuous or discontinuous) components and arranging each component by its "value (magnitude, intensity, frequency, etc.)". Specific examples of spectral data (SPE) include spectral spectra, frequency spectra, electromagnetic spectra, and mass spectra. For example, spectral data (SPE) may be measured using a measuring device. Also, for example, spectral data (SPE) may be acquired by the acquisition unit 11, or additionally acquired by the additional acquisition unit 16. Furthermore, for example, each spectrum in the spectral data (SPE) acquired by the additional acquisition unit 16 may be associated with a predetermined target variable.

[0045] The correspondence data REL is data that represents the correspondence between multiple spectral data SPEs. The correspondence data REL may include, for example, data representing the correspondence between any two spectral data SPEs among multiple spectral data SPEs. For example, if the acquisition unit 11 acquires spectral data SPEs measured in each of multiple measuring devices, the correspondence data REL may include data representing the correspondence between any two of those multiple devices. Also, for example, if the acquisition unit 11 acquires spectral data SPEs of multiple samples measured in one measuring device, the correspondence data REL may include data representing the correspondence between any two spectral data SPEs among those multiple samples.

[0046] For example, the first correspondence generation unit 12 may generate correspondence data REL for the spectral data SPE acquired by the acquisition unit 11. Alternatively, for example, the second correspondence generation unit 15 may restore the virtual data added by the data addition unit 14 to the feature space data SPA to the same dimension as the correspondence data REL generated by the first correspondence generation unit 12, and generate correspondence data REL.

[0047] Furthermore, for example, the correspondence data REL may cover all combinations of two spectral data SPEs from among multiple spectral data SPEs.

[0048] Furthermore, for example, the correspondence data REL may be a difference spectrum representing the difference between multiple spectra. Here, the difference spectrum is calculated by determining the difference for each wavelength between the measured values ​​in two spectra.

[0049] Furthermore, for example, the correspondence data REL may be a conversion coefficient that transforms spectra between multiple spectra. For example, the conversion coefficient may be a regression coefficient included in a regression equation that approximately transforms one spectrum into another. The regression coefficient may be calculated, for example, by wavelength-specific linear regression. Specific examples of analytical methods used in wavelength-specific linear regression include PLS regression and LASSO (Least Absolute Shrinkage and Selection Operator) regression. Also, for example, when determining the regression coefficient, the entire set of regression coefficients may be calculated at once using the proximal method.

[0050] Algorithmic data ALG is data containing algorithms used by the space generation unit 13 to compress the dimensions of the correspondence data REL and generate feature space data SPA, which is the space containing the compressed data. Specific examples of algorithms used in Algorithmic data ALG include Principal Component Analysis (PCA), Kernel Principal Component Analysis, Independent Component Analysis, Factor Analysis, Multidimensional Scaling, Non-negative Matrix Factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or autoencoders.

[0051] Feature space data (SPA) is a space containing data after compressing the dimensions of the correspondence data (REL) using algorithmic data (ALG). In other words, feature space data (SPA) is a low-dimensional space characterized by differences in the measurement conditions of the spectrum.

[0052] The virtual spectrum data IMA is a virtual spectrum generated by the virtual spectrum generation unit 17 using the spectrum data SPE acquired by the additional acquisition unit 16 and the correspondence data REL generated by the first correspondence generation unit 12 and the second correspondence generation unit 15, respectively.

[0053] For example, if the correspondence data REL generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 are difference spectra, the virtual spectrum generation unit 17 may generate a virtual spectrum by adding the absorbance in the spectrum data SPE additionally acquired by the additional acquisition unit 16 and the absorbance in the correspondence data REL for each wavelength.

[0054] Furthermore, for example, if the correspondence data REL generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 are regression coefficients, which are an example of conversion coefficients, the virtual spectrum generation unit 17 may generate virtual spectrum data IMA by multiplying the matrix representing each of the additionally acquired spectrum data SPEs by the additional acquisition unit 16 by the matrix representing the regression coefficients.

[0055] Furthermore, for example, when the virtual spectrum generation unit 17 generates one or more virtual spectra included in the virtual spectrum data IMA from one spectrum, it may associate a predetermined target variable with the one or more virtual spectra.

[0056] The training data TD is data containing one or more virtual spectra generated by the information processing device 1, and is used to train a machine learning model that takes spectra as input and an objective variable as output. For example, PLS (Partial Least Squares) regression, LASSO (Least Absolute Shrinkage and Selection Operator) regression, SVM (Support Vector Machine) regression, ensemble learning, neural networks, etc., may be applied to the training data TD and the corresponding objective variable.

[0057] (Communications Section 30) The communication unit 30 communicates with devices outside the information processing device 1. For example, the communication unit 30 communicates with various external devices connected to the information processing device 1 via a network N (not shown). The communication unit 30 transmits data supplied from the control unit 10 to the outside and supplies data received from the external devices to the control unit 10. The specific configuration of the network N is not limited to this exemplary embodiment, but as an example, a wireless LAN (Local Area Network), wired LAN, WAN (Wide Area Network), public telephone network, mobile data communication network, or a combination of these networks can be used.

[0058] (Input / output section 40) The input / output unit 40 is configured to include at least one of the following input / output devices: a keyboard, mouse, display, printer, touch panel, etc. Alternatively, the input / output unit 40 may be configured to have input / output devices such as a keyboard, mouse, display, printer, touch panel, etc. connected to it. In this configuration, the input / output unit 40 receives various types of information from the connected input device to the information processing device 1. The input / output unit 40 also outputs various types of information to the connected output device under the control of the control unit 10. An interface such as USB (Universal Serial Bus) can be used as the input / output unit 40.

[0059] (Information processing method S1 flow) The flow of the information processing method S1 executed by the information processing device 1 will be explained with reference to Figure 8. Figure 8 is a flowchart showing an example of the flow of the information processing method S1. The information processing method S1 includes, for example, an acquisition process (step) S11, a first correspondence generation process (step) S12, a space generation process (step) S13, a data addition process (step) S14, a second correspondence generation process (step) S15, an additional acquisition process (step) S16, and a virtual spectrum generation process (step) S17, as shown in Figure 8.

[0060] (Step S11) In step S11, the acquisition unit 11 acquires multiple spectra.

[0061] (Step S12) In step S12, the first correspondence generation unit 12 generates correspondence data representing the correspondence between spectra for the multiple spectra acquired by the acquisition unit 11.

[0062] (Step S13) In step S13, the space generation unit 13 compresses the dimensions of the correspondence data using a predetermined algorithm and generates a feature space which is the space containing the compressed data.

[0063] (Step S14) In step S14, the data addition unit 14 adds one or more virtual data to the feature space.

[0064] (Step S15) In step S15, the second correspondence generation unit 15 restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit 12, and generates further correspondence data.

[0065] (Step S16) In step S16, the additional acquisition unit 16 acquires one or more additional spectra.

[0066] (Step S17) In step S17, the virtual spectrum generation unit 17 generates one or more virtual spectra using one or more additional spectra and the correspondence data generated by the first correspondence generation unit 12 and the second correspondence generation unit 15, respectively.

[0067] [Embodiment 2] Other embodiments of the present invention are described below. For the sake of clarity, components having the same function as those described in the above embodiments will be denoted by the same reference numerals, and their descriptions will not be repeated.

[0068] (Summary of Embodiment 2) Embodiment 2 describes the first correspondence generation unit 12 and the second correspondence generation unit 15 provided in the information processing device 1 according to Embodiment 1. More specifically, the case where the correspondence data generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 are conversion coefficients will be described with reference to Figures 9 to 13. The other configurations of the information processing device 1 according to Embodiment 2 are the same as those of the information processing device 1 according to Embodiment 1.

[0069] (When the corresponding data is a conversion coefficient) Correspondence data may, for example, be conversion coefficients that convert spectra between multiple spectra. For example, the conversion coefficients may be regression coefficients included in a regression equation that approximately converts one spectrum to another. Regression coefficients may be calculated, for example, by wavelength-specific linear regression. Specific examples of analytical methods used in wavelength-specific linear regression include PLS regression and LASSO regression. Alternatively, for example, when determining regression coefficients, the entire set of regression coefficients may be calculated at once using the proximal method.

[0070] Figure 9 shows an example of regression coefficients for converting the spectrum from the master unit to the spectrum from the slave unit of a measuring device. In the example in Figure 9, both the master and slave units show the wavelength dependence of the measured values. As illustrated in Figure 9, the spectra from both the master and slave units may be represented by matrices where the number of samples is the number of rows and the number of wavelengths is the number of columns. In this case, for example, the regression coefficient B for converting the spectrum from the master unit to the spectrum from the slave unit may be a matrix where both the number of rows and columns are the number of wavelengths, as shown in Figure 9.

[0071] Figure 10 shows a transformation coefficient, which is an example of correspondence data. In the following explanation, the transformation coefficient exemplified in Figure 10 will be explained using the example of the transformation coefficient "1->2," but the same applies to other transformation coefficients. The transformation coefficient "1->2" in Figure 10 is a transformation coefficient that represents the transformation from the source transformation set #1 to the destination transformation set #2. For example, the transformation coefficient "1->2" may also be a regression coefficient included in a regression equation for approximately transforming transformation set #1 to transformation set #2.

[0072] Figure 10 illustrates the conversion coefficients "1->2" and "3->1". In each conversion coefficient illustrated in Figure 10, the vertical axis represents the wavelength in the source spectrum, and the horizontal axis represents the wavelength in the destination spectrum. The numerical values ​​included in the conversion coefficients in Figure 10 are indicated by the color scale to the right of each conversion coefficient.

[0073] Figure 11 shows an example of the result of dimensionality reduction performed by the space generation unit 13 using principal component analysis with the transformation coefficients exemplified in Figure 10. In the example in Figure 11, points corresponding to the respective transformation coefficients are shown in the feature space where the first principal component axis PC1 and the second principal component axis PC2 are set. Furthermore, the example in Figure 11 shows an example of the result of dimensionality reduction including the transformation coefficients exemplified in Figure 10.

[0074] Figure 12 shows an example of adding virtual data to the feature space illustrated in Figure 11 by the data addition unit 14. The feature space illustrated in Figure 12 is a space generated by the space generation unit 13 using principal component analysis, and, similar to the feature space illustrated in Figure 11, a first principal component axis PC1 and a second principal component axis PC2 are set. In the example in Figure 12, the black dots represent data generated by compressing the dimensionality in the transformation coefficients. Here, the data addition unit 14 adds virtual data (× dots) to the feature space generated by the space generation unit 13, for example, as shown in Figure 12.

[0075] The ellipse in the feature space illustrated in Figure 12 is an example of a confidence ellipse set for any range and interval between 0% and 100% in the principal component scores. For example, the range in the feature space to which the data addition unit 14 adds virtual data may be determined by the confidence ellipse in the principal component scores.

[0076] Figure 13 shows an example of virtual spectrum generation by the virtual spectrum generation unit 17 using regression coefficients, which are examples of conversion coefficients generated by the first correspondence generation unit 12 and the second correspondence generation unit 15, respectively. The virtual spectrum generation unit 17 may, for example, generate virtual spectra (composite 1 to composite 50) by multiplying a matrix representing the learning set of the master unit acquired by the additional acquisition unit 16 with a matrix representing the regression coefficients (B1 to B50) generated by the first correspondence generation unit 12 and the second correspondence generation unit 15, respectively, as shown in Figure 13.

[0077] With the above configuration, a virtual spectrum can be generated using conversion coefficients between multiple spectra.

[0078] [Examples of implementation using software] The functions of the information processing device 1 (hereinafter referred to as "the device") are programs that cause the device to function as a computer, and these programs can be realized by programs that cause each control block of the device (especially each part included in the control unit 10) to function as a computer.

[0079] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.

[0080] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.

[0081] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.

[0082] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI ​​may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).

[0083] 〔summary〕 An information processing device according to embodiment 1 of the present invention comprises: an acquisition unit that acquires a plurality of spectra; a first correspondence generation unit that generates correspondence data representing the correspondence between spectra for the plurality of spectra acquired by the acquisition unit; a space generation unit that compresses the dimensions in the correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data; a data addition unit that adds one or more virtual data to the feature space; a second correspondence generation unit that restores the added virtual data to the same dimensions as the correspondence data generated by the first correspondence generation unit and further generates the correspondence data; an additional acquisition unit that acquires one or more additional spectra; and a virtual spectrum generation unit that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated by the first correspondence generation unit and the second correspondence generation unit, respectively.

[0084] In the information processing apparatus according to aspect 2 of the present invention, in aspect 1 above, a predetermined target variable is associated with each spectrum in the one or more additional spectra, and when the virtual spectrum generation unit generates one or more virtual spectra from one spectrum, it associates the predetermined target variable with the one or more virtual spectra.

[0085] In the information processing device according to aspect 3 of the present invention, in aspect 1 or 2 above, the correspondence data is a difference spectrum representing the difference between spectra, or a conversion coefficient that converts spectra between spectra.

[0086] In any of the above embodiments 1 to 3, the information processing device according to embodiment 4 of the present invention is characterized in that the predetermined algorithm is principal component analysis, kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or autoencoder.

[0087] The training data according to aspect 5 of the present invention is data including one or more virtual spectra generated by the information processing device described in aspect 2 above, and is data for training a machine learning model that takes spectra as input and an objective variable as output.

[0088] An information processing method according to aspect 6 of the present invention includes: an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing the correspondence between spectra for the plurality of spectra acquired in the acquisition process; a space generation process for compressing the dimensions of the correspondence data using a predetermined algorithm and generating a feature space which is a space containing the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimensions as the correspondence data generated in the first correspondence generation process and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; and a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, respectively.

[0089] A program according to aspect 7 of the present invention causes a computer to execute: an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing the correspondence between spectra for the plurality of spectra acquired in the acquisition process; a space generation process for compressing the dimensions of the correspondence data using a predetermined algorithm and generating a feature space which is a space containing the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimensions as the correspondence data generated in the first correspondence generation process and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; and a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, respectively.

[0090] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Examples]

[0091] [Example 1] One embodiment of the present invention is described below.

[0092] Four Kubota K-BA800 spectrometers were used as measuring devices to measure the spectra of 740 mini tomatoes, and the Brix sugar content was measured after squeezing the juice. Of the four spectrometers, unit 1 was designated as the master unit, and the rest as slave units. Ten of the obtained spectra were designated as the "conversion set," 472 as the "learning set (master unit only)," and 150 as the "verification set." The acquisition unit 11 acquired the conversion set, and the first correspondence generation unit 12 calculated the difference spectrum between each master-slave pair. In the spatial generation unit 13, principal component analysis was performed on the difference spectra obtained in the first correspondence generation unit 12 to calculate a feature space consisting of the first and second principal components. In the data addition unit 14, sampling was performed to add 100 virtual data points to the grid on the feature space calculated by the spatial generation unit 13. In Example 1, the maximum and minimum values ​​of the first and second principal component scores of the difference spectrum obtained in the first correspondence generation unit 12 were set as the sampling range, and 100 virtual data points were uniformly sampled within the sampling range. In the second correspondence generation unit 15, the inverse transform of the processing in the spatial generation unit 13 was applied to generate 100 new corresponding difference spectra. In the virtual spectrum generation unit 17, the difference spectra generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 were added to the learning set acquired by the additional acquisition unit 16 via the master unit to generate 47,200 virtual spectra. Machine learning was applied to the spectra generated by the virtual spectrum generation unit 17 and the corresponding measured Brix sugar content values ​​to construct a model for estimating Brix sugar content from spectra.

[0093] Figure 14 shows the results of applying the estimation model based on the spectrum generated by the information processing device 1 to the validation set in Example 1. In Figure 14, the horizontal axis represents the measured Brix value, and the vertical axis represents the estimated value. The circles represent the estimation results for the training set, and the triangles represent the estimation results for the validation set. Figure 14 also shows the results for machine #1, machine #2, machine #3, and machine #4, respectively. As shown in Figure 14, the absolute value of the bias for each of machines #1 to #4 in Example 1 was at most about 0.1.

[0094] [Example 2] Other embodiments of the present invention are described below.

[0095] In Example 2, the sampling range of the virtual data was changed in the data addition unit 14 of Example 1. The other conditions in Example 2 were the same as in Example 1.

[0096] Figure 15 shows the sampling range of virtual data by the data addition unit 14 in Example 2. In Example 2, the sampling range was set to "1x", the same as in Example 1. Furthermore, in Example 2, even when the sampling range was changed between 0.2x and 10x, 100 virtual data points were uniformly sampled within each sampling range, and a new difference spectrum was generated based on these. Figure 15 illustrates the cases of sampling ranges of 0.5x, 1x (same as in Example 1), and 2x.

[0097] Figure 16 shows the results of applying the estimation model based on the spectrum generated by the information processing device 1 to the validation set in Example 2. Figure 16 shows the results for slave unit 2. Also, in Figure 16, similar to Figure 15, examples are shown for sampling ranges of 0.5x, 1x (same as slave unit 2 in Example 1), and 2x. As shown in Figure 16, the results of Example 2 showed that the absolute value of the bias for sampling ranges of 0.5x and 2x was at most about 0.1.

[0098] (Results regarding estimation errors in Examples 1 and 2) Table 1 below shows the results for estimation error (RMSE (Root Mean Squared Error)) in Examples 1 and 2. The value in the "1" column for the sampling range multiplier in Table 1 is the result for Example 1. Table 1 also shows the results when the information processing device 1 is not used ("No spectral expansion" column).

[0099] Table 1 Results for Estimation Error in Examples 1 and 2

[0100] [Table 1] As shown in Table 1, when the information processing device 1 was not used, there were slave units with estimation errors exceeding 1, whereas in Example 1, the estimation error was below 0.4 for all slave units. In other words, in Example 1, it was found that, similar to the master unit, the measured values ​​and estimated values ​​tended to match with high accuracy in the slave units as well.

[0101] Furthermore, as shown in Table 1, in Example 2, the estimation error was less than 1 for all sub-units, regardless of whether the sampling range was changed to 0.2x, 0.5x, 2x, 5x, or 10x. In other words, in Example 2, it was found that even when the sampling range was changed between 0.2x and 10x, the measured values ​​and estimated values ​​tended to match accurately in the sub-units. [Explanation of symbols]

[0102] 1. Information Processing Device 10 Control Unit 11 Acquisition Department 12 First correspondence generation unit 13 Space generation part 14. Data Addition Section 15 Second correspondence generation unit 16 Additional acquisition section 17 Virtual Spectrum Generation Unit 20 Memory Department 30 Ministry of Communications 40 Input and Output Force Section

Claims

1. An acquisition unit that acquires multiple spectra, A first correspondence generation unit generates correspondence data representing the correspondence between the spectra for the plurality of spectra acquired by the acquisition unit, A space generation unit compresses the dimensions of the aforementioned correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data, A data addition unit that adds one or more virtual data to the feature space, A second correspondence generation unit restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit, and further generates the correspondence data. An additional acquisition unit that acquires one or more additional spectra, A virtual spectrum generation unit generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated by the first correspondence generation unit and the second correspondence generation unit, respectively. An information processing device equipped with the following features.

2. Each of the one or more additional spectra is associated with a predetermined target variable. The virtual spectrum generation unit, when generating one or more virtual spectra from one spectrum, associates the predetermined target variable with the one or more virtual spectra. The information processing apparatus according to claim 1.

3. The aforementioned correspondence data is a difference spectrum representing the difference between spectra, or a conversion coefficient that converts spectra between spectra. The information processing apparatus according to claim 1.

4. The aforementioned predetermined algorithm is principal component analysis, kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or an autoencoder. The information processing apparatus according to claim 1.

5. Training data comprising one or more virtual spectra generated by the information processing device described in claim 2, wherein the data is used to train a machine learning model that takes spectra as input and an objective variable as output.

6. The acquisition process for obtaining multiple spectra, A first correspondence generation process generates correspondence data representing the correspondence between the spectra for the plurality of spectra acquired in the acquisition process, A spatial generation process that compresses the dimensions of the aforementioned correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data, A data addition process that adds one or more virtual data to the feature space, A second correspondence generation process that restores the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process, and further generates the correspondence data. An additional acquisition process to acquire one or more additional spectra, A virtual spectrum generation process that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, Information processing methods, including those mentioned above.

7. On the computer, The acquisition process for obtaining multiple spectra, A first correspondence generation process generates correspondence data representing the correspondence between the spectra for the plurality of spectra acquired in the acquisition process, A spatial generation process that compresses the dimensions of the aforementioned correspondence data using a predetermined algorithm and generates a feature space which is a space containing the compressed data, A data addition process that adds one or more virtual data to the feature space, A second correspondence generation process that restores the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process, and further generates the correspondence data. An additional acquisition process to acquire one or more additional spectra, A virtual spectrum generation process that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated in the first correspondence generation process and the second correspondence generation process, A program that executes something.