Information Processing Apparatus, Information Processing Method, and Program
The information processing apparatus addresses the high cost of reducing spectral measurement condition differences by generating virtual spectra through data compression and virtual data addition, achieving efficient and cost-effective data processing.
Patent Information
- Application Number
- JP2024185293
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Existing techniques for reducing differences in spectral measurement conditions require a large number of spectral measurement data, leading to increased costs.
An information processing apparatus and method that acquire spectra, generate correspondence data, compress dimensions using algorithms like PCA, add virtual data, restore it to original dimensions, and generate virtual spectra to reduce measurement condition differences at a lower cost.
The solution effectively reduces differences in spectrum measurement conditions while minimizing costs, enabling more efficient data processing and analysis.
Smart Images

Figure 0007687748000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, learning data, an information processing method, and a program.
Background Art
[0002] Techniques for reducing differences in spectra that can occur due to differences in measurement conditions are known. Examples of differences in measurement conditions include individual differences in measurement devices used for measurement, differences in the production areas of agricultural crops, differences in the types of agricultural crops, and annual differences in agricultural crops. For example, Non-Patent Document 1 discloses a technique for creating an estimation model that can be applied regardless of the measurement device.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the technique disclosed in Non-Patent Document 1, a large number of spectral measurement data are required when learning the estimation model. Therefore, there is a problem that the cost for reducing the difference in measurement conditions tends to increase.
[0005] One aspect of the present invention aims to realize a technique for reducing differences in spectral measurement conditions at a lower cost.
Means for Solving the Problems
[0006] In order to solve the above problems, an information processing apparatus according to an aspect of the present invention includes an acquisition unit that acquires a plurality of spectra, a first correspondence generation unit that generates correspondence data representing a correspondence between the spectra for the plurality of spectra acquired by the acquisition unit, a space generation unit that compresses the dimension in the correspondence data using a predetermined algorithm and generates a feature space that is a space including the compressed data, a data addition unit that adds one or more virtual data to the feature space, a second correspondence generation unit that restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit and further generates the correspondence data, an additional acquisition unit that acquires one or more additional spectra, and a virtual spectrum generation unit that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated by each of the first correspondence generation unit and the second correspondence generation unit.
[0007] In order to solve the above problems, an information processing method according to an aspect of the present invention includes an acquisition process of acquiring a plurality of spectra, a first correspondence generation process of generating correspondence data representing a correspondence between the spectra for the plurality of spectra acquired in the acquisition process, a space generation process of compressing the dimension in the correspondence data using a predetermined algorithm and generating a feature space that is a space including the compressed data, a data addition process of adding one or more virtual data to the feature space, a second correspondence generation process of restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process and further generating the correspondence data, an additional acquisition process of acquiring one or more additional spectra, and a virtual spectrum generation process of generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process.
[0008] To solve the above problems, a program according to one aspect of the present invention causes a computer to perform an acquisition process of acquiring a plurality of spectra, a first correspondence generation process of generating correspondence data representing a correspondence between the plurality of spectra acquired in the acquisition process, a space generation process of compressing dimensions in the correspondence data using a predetermined algorithm and generating a feature space which is a space including the compressed data, a data addition process of adding one or more virtual data to the feature space, a second correspondence generation process of restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process and further generating the correspondence data, an additional acquisition process of acquiring one or more additional spectra, and a virtual spectrum generation process of generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process.
Advantages of the Invention
[0009] According to one aspect of the present invention, differences in spectrum measurement conditions can be reduced at a lower cost.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Mode for Carrying Out the Invention
[0011] 〔Embodiment 1〕 Hereinafter, an embodiment of the present invention will be described in detail.
[0012] (Configuration of Information Processing Apparatus 1) The configuration of the information processing apparatus 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing a configuration example of the information processing apparatus 1. The information processing apparatus 1 includes, for example, as shown in FIG. 1, a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.
[0013] (Control Unit 10) The control unit 10 comprehensively controls each unit of the information processing apparatus 1. The control unit 10 includes, for example, as shown in FIG. 1, an acquisition unit 11, a first correspondence generation unit 12, a space generation unit 13, a data addition unit 14, a second correspondence generation unit 15, an additional acquisition unit 16, and a virtual spectrum generation unit 17.
[0014] (Acquisition Unit 11) The acquisition unit 11 acquires a plurality of spectra. A spectrum is obtained by decomposing information or a signal into its (continuous or discontinuous) components and arranging them in terms of "values (such as magnitude, intensity, frequency, etc.)" for each component. Specific examples of the spectra acquired by the acquisition unit 11 include a spectroscopic spectrum, a frequency spectrum, an electromagnetic spectrum, a mass spectrum, and the like. For example, the spectrum acquired by the acquisition unit 11 may be one measured using a measuring device.
[0015] For example, the information processing device 1 and the measuring device may be connected via a network N or the like described later. Also, for example, the information processing device 1 may further include a configuration for measuring the spectrum provided in the aforementioned measuring device. The plurality of spectra acquired by the acquisition unit 11 may be, for example, those measured using one or a plurality of measuring devices. That is, the acquisition unit 11 may, for example, acquire spectra measured by each of the plurality of measuring devices, or may acquire a plurality of spectra measured by one measuring device.
[0016] FIG. 2 is a diagram showing an example of division in a spectrum measured using a measuring device. For example, in the measuring device, as shown in FIG. 2, one master device and a plurality of slave devices may be set.
[0017] The spectrum measured using the measuring device may be divided into, for example, a conversion set, a learning set (parent machine only), and a verification set, as shown in FIG. 2. The conversion set is a spectrum used to convert the spectrum between the parent machine and the child machine. Also, as shown in FIG. 2, the conversion set is obtained by measuring the spectra of samples common to the parent machine and the child machine. The learning set is data in which a target variable having the spectrum as an explanatory variable is associated with the spectrum of the sample. Specific examples of the target variable include sugar content, protein content, and the like. The verification set is a spectrum for verifying the machine learning model learned using the processing result by the information processing device 1 in the parent machine and the child machine. Also, as shown in FIG. 2, the verification set is obtained by measuring the spectra of samples common to the parent machine and the child machine.
[0018] Here, the conversion set is an example of a plurality of spectra acquired by the acquisition unit 11. Also, the learning set is an example of additional spectra acquired by the additional acquisition unit 16 described later. In the embodiments described later, embodiments using the verification set will be described.
[0019] In the example in the description with reference to FIGS. 2 to 4, for convenience, as shown in FIG. 2, three child machines are set, the parent machine is assigned the number #1 (Machine #1), and the three child machines are assigned the numbers #2 to #4 (Machine #2 to #4), respectively.
[0020] (First correspondence generation unit 12) The first correspondence generation unit 12 generates association data representing the association between the plurality of spectra acquired by the acquisition unit 11. The association data may include, for example, data representing the association between any two spectra among the plurality of spectra acquired by the acquisition unit 11. For example, when the acquisition unit 11 acquires the spectra measured by each of a plurality of measuring devices, the association data may include data representing the association between any two of the plurality of devices. Also, for example, when the acquisition unit 11 acquires the spectra of a plurality of samples measured by one measuring device, the association data may include data representing the association between any two spectra among the spectra of the plurality of samples.
[0021] For example, the association data may cover all combinations of pairs of two spectra among the plurality of spectra acquired by the acquisition unit 11.
[0022] Also, for example, the association data may be a difference spectrum representing the difference between spectra. The difference spectrum is obtained by calculating the difference for each wavelength of the measured values at two spectra. Also, for example, the association data may be a conversion coefficient for mutually converting spectra between spectra. Examples of the conversion coefficient will be described later with reference to FIGS. 8 to 12.
[0023] (When the association data is a difference spectrum) FIG. 3 is a diagram showing a difference spectrum which is an example of the association data. For example, when the acquisition unit 11 acquires the conversion sets (hereinafter also referred to as conversion sets #1 to #4) corresponding to each of Machine#1 to #4 shown in FIG. 2, the first correspondence generation unit 12 may generate, as shown in FIG. 3, a difference spectrum representing the difference between two of the conversion sets #1 to #4.
[0024] In FIG. 3, as an example, it shows a difference spectrum when the acquisition unit 11 acquires a spectrum showing the wavelength dependence of absorbance. In the graph of the difference spectrum illustrated in FIG. 3, the horizontal axis represents the wavelength, and the vertical axis represents the difference in absorbance, respectively.
[0025] Hereinafter, regarding the difference spectrum illustrated in FIG. 3, an example of the difference spectrum "2->1" will be described with reference to an example, but the same applies to other difference spectra. The difference spectrum "2->1" in FIG. 3 is a difference spectrum representing the conversion from the conversion set #2, which is the source of conversion, to the same #1, which is the destination of conversion. That is, the difference spectrum "2->1" represents the result of subtracting the absorbance in the conversion set #2 from the absorbance in the conversion set #1 for each wavelength.
[0026] For example, as shown in FIG. 3, the first correspondence generation unit 12 may generate difference spectra covering all combinations of two conversion sets out of the conversion sets #1 to #4 acquired by the acquisition unit 11.
[0027] With the above configuration, data representing the difference in the measurement conditions of the spectrum can be generated.
[0028] (Spatial generation unit 13) The spatial generation unit 13 compresses the dimensions in the association data using a predetermined algorithm and generates a feature space, which is a space including the compressed data. Specific examples of the predetermined algorithm include principal component analysis (PCA: Principal Component Analysis), kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or autoencoder.
[0029] FIG. 4 is a diagram showing an example of the result of dimension compression using principal component analysis in the difference spectrum illustrated in FIG. 3. Here, the result of dimension compression by the space generation unit 13 using principal component analysis may represent, for example, a feature space that is a space including the data after dimension compression. That is, the space generation unit 13 may compress, for example, the dimensions in the association data using principal component analysis as shown in FIG. 4, and generate a feature space that is a space including the compressed data. In the example of FIG. 4, as a result of the space generation unit 13 compressing the dimensions of the difference spectrum, a first principal component axis PC1 and a second principal component axis PC2 that are orthogonal to each other are set. Also, in the example of FIG. 4, points corresponding to the respective difference spectra (difference spectrum "2->1", etc.) illustrated in FIG. 3 are shown in the feature space in which the first principal component axis PC1 and the second principal component axis PC2 are set.
[0030] With the above configuration, it is possible to generate a feature space that is a low-dimensional space characterized by differences in the measurement conditions of the spectrum.
[0031] (Data addition unit 14) The data addition unit 14 adds one or more virtual data to the feature space. The virtual data is a virtual point newly added to an arbitrary coordinate (hereinafter also referred to as "grid") where no data exists on the feature space generated by the space generation unit 13.
[0032] FIG. 5 is a diagram showing an example of adding virtual data to the feature space illustrated in FIG. 4 by the data addition unit 14. The feature space illustrated in FIG. 5 is a space generated by the space generation unit 13 using principal component analysis, and the first principal component axis PC1 and the second principal component axis PC2 are set in the same manner as the feature space illustrated in FIG. 4. "Data" in the example of FIG. 5 represents data generated by compressing the dimensions in the association data. Here, the data addition unit 14 adds virtual data (Sampled) to the feature space generated by the space generation unit 13 as shown in FIG. 5, for example.
[0033] With the above configuration, virtual data reflecting differences in the measurement conditions of the spectrum can be added in any quantity.
[0034] (Second correspondence generation unit 15) The second correspondence generation unit 15 restores the added virtual data to the same dimension as the association data generated by the first correspondence generation unit 12, and further generates association data. That is, the process in the second correspondence generation unit 15 corresponds to the inverse transformation of the process in the space generation unit 13.
[0035] FIG. 6 is a diagram showing an example of a difference spectrum including the difference spectrum generated by the first correspondence generation unit 12 and the difference spectrum generated by the second correspondence generation unit 15. Here, the difference spectrum generated by the second correspondence generation unit 15 is a difference spectrum corresponding to the virtual data added by the data addition unit 14.
[0036] (Additional acquisition unit 16) The additional acquisition unit 16 acquires one or more additional spectra. The additional spectra acquired by the additional acquisition unit 16 are spectra similar to the spectra acquired by the acquisition unit 11, but are spectra acquired in addition to the spectra acquired by the acquisition unit 11. For example, a predetermined target variable may be associated with each spectrum in the one or more additional spectra. The learning set described above in the description with reference to FIG. 2 is an example of the additional spectra acquired by the additional acquisition unit 16.
[0037] (Virtual spectrum generation unit 17) The virtual spectrum generation unit 17 generates one or more virtual spectra using one or more additional spectra and the association data generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15. For example, when the association data generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 is a difference spectrum, the virtual spectrum generation unit 17 may add the absorbance in one or more additional spectra and the absorbance in the difference spectrum for each wavelength to generate a virtual spectrum. Further, for example, when the virtual spectrum generation unit 17 generates one or more virtual spectra from one spectrum, a predetermined target variable may be associated with the one or more virtual spectra.
[0038] FIG. 7 is a diagram showing an example of generating a virtual spectrum using a difference spectrum by the virtual spectrum generation unit 17. The virtual spectrum generation unit 17 may, for example, as shown in FIG. 7, add the absorbance in the learning set of the master unit acquired by the additional acquisition unit 16 and the absorbance in the difference spectrum generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 for each wavelength to generate virtual spectra (Synthesis 1 to Synthesis 100).
[0039] With the above configuration, an arbitrary number of virtual spectra reflecting the differences in the spectrum measurement conditions can be generated.
[0040] Also, the virtual spectrum in the above configuration can be generated at a lower cost than measuring the spectrum using a measuring device. Therefore, with the above configuration, the differences in the spectrum measurement conditions can be reduced at a lower cost.
[0041] (Learning data generated using the information processing apparatus 1) Data including one or more virtual spectra generated by the information processing apparatus 1 may be, for example, learning data for training a machine learning model that takes a spectrum as an input and outputs a target variable. For example, for the target variable corresponding to the learning data, PLS (Partial Least Squares) regression, LASSO (Least Absolute Shrinkage and Selection Operator) regression, SVM (Support Vector Machine) regression, ensemble learning, neural network, etc. may be applied.
[0042] With the above configuration, the machine learning model can be trained so as not to be easily affected by differences in the measurement conditions of the spectrum.
[0043] (Storage unit 20) The storage unit 20 stores various data referred to by the control unit 10 and various data generated by the control unit 10. Specific examples of the data stored in the storage unit 20 include · Spectrum data SPE · Association data REL · Algorithm data ALG · Feature space data SPA · Virtual spectrum data IMA · Learning data TD and the like.
[0044] The spectrum data SPE decomposes information or signals into their (continuous or discontinuous) components and arranges them by "value (magnitude, intensity, frequency, etc.)" for each component. Specific examples of the spectrum data SPE include spectral spectrum, frequency spectrum, electromagnetic spectrum, mass spectrum, etc. For example, the spectrum data SPE may be measured using a measuring device. Also, for example, the spectrum data SPE may be acquired by the acquisition unit 11, or may be additionally acquired by the additional acquisition unit 16. Also, for example, a predetermined target variable may be associated with each spectrum in the spectrum data SPE acquired by the additional acquisition unit 16.
[0045] The association data REL is data representing the association between a plurality of spectral data SPEs. The association data REL may include, for example, data representing the association between any two of the plurality of spectral data SPEs. For example, when the acquisition unit 11 acquires the spectral data SPEs measured by each of a plurality of measuring devices, the association data REL may include data representing the association between any two of the plurality of devices. Also, for example, when the acquisition unit 11 acquires the spectral data SPEs of a plurality of samples measured by one measuring device, the association data REL may include data representing the association between any two of the spectral data SPEs of the plurality of samples.
[0046] For example, the first association generation unit 12 may generate the association data REL for the spectral data SPE acquired by the acquisition unit 11. Also, for example, the second association generation unit 15 may restore the virtual data added to the feature space data SPA by the data addition unit 14 to the same dimension as the association data REL generated by the first association generation unit 12 and generate the association data REL.
[0047] Also, for example, the association data REL may cover all combinations of pairs of two of the plurality of spectral data SPEs.
[0048] Also, for example, the association data REL may be a difference spectrum representing the difference between a plurality of spectra. Here, the difference spectrum is obtained by calculating the difference for each wavelength of the measured values in two spectra.
[0049] Further, for example, the association data REL may be a conversion coefficient that converts spectra mutually between a plurality of spectra. For example, the conversion coefficient may be a regression coefficient included in a regression equation for approximately converting one spectrum into another spectrum. The regression coefficient may be calculated, for example, by linear regression for each wavelength. Specific examples of the analysis method used in the linear regression for each wavelength include PLS regression, LASSO (Least Absolute Shrinkage and Selection Operator) regression, and the like. Further, for example, when obtaining the regression coefficient, the proximal method may be used to calculate the entire regression coefficient at once.
[0050] The algorithm data ALG is data including an algorithm used for the space generation unit 13 to compress the dimension in the association data REL and generate a feature space data SPA which is a space including the compressed data. Specific examples of the algorithm data ALG include principal component analysis (PCA: Principal Component Analysis), kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling method, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or an autoencoder.
[0051] The feature space data SPA is a space including the data after compressing the dimension in the association data REL using the algorithm data ALG. In other words, the feature space data SPA is a low-dimensional space characterized by the difference in the measurement conditions of the spectra.
[0052] The virtual spectrum data IMA is a virtual spectrum generated by the virtual spectrum generation unit 17 using the spectrum data SPE additionally acquired by the additional acquisition unit 16 and the association data REL generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15.
[0053] For example, when the association data REL generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 is a difference spectrum, the virtual spectrum generation unit 17 may generate a virtual spectrum by adding the absorbance in the spectrum data SPE additionally acquired by the additional acquisition unit 16 and the absorbance in the association data REL for each wavelength.
[0054] Also, for example, when the association data REL generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 is a regression coefficient which is an example of a conversion coefficient, the virtual spectrum generation unit 17 may generate virtual spectrum data IMA by multiplying the matrix representing each of the spectrum data SPE additionally acquired by the additional acquisition unit 16 and the matrix representing the regression coefficient.
[0055] Also, for example, when the virtual spectrum generation unit 17 generates one or more virtual spectra included in the virtual spectrum data IMA from one spectrum, a predetermined target variable may be associated with the one or more virtual spectra.
[0056] The training data TD is data including one or more virtual spectra generated by the information processing apparatus 1, and is data for training a machine learning model that takes a spectrum as an input and outputs a target variable. For example, PLS (Partial Least Squares) regression, LASSO (Least Absolute Shrinkage and Selection Operator) regression, SVM (Support Vector Machine) regression, ensemble learning, neural network, etc. may be applied to the target variable corresponding to the training data TD.
[0057] (Communication unit 30) The communication unit 30 communicates with devices external to the information processing apparatus 1. As an example, the communication unit 30 communicates with each external device connected to the information processing apparatus 1 via a network N (not shown). The communication unit 30 transmits the data supplied from the control unit 10 to the outside, or supplies the data received from each external device to the control unit 10. Note that the specific configuration of the network N does not limit this exemplary embodiment, but as an example, a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public switched telephone network, a mobile data communication network, or a combination of these networks can be used.
[0058] (Input / output unit 40) The input / output unit 40 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. Alternatively, the input / output unit 40 may be configured such that input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel are connected thereto. In the case of such a configuration, the input / output unit 40 receives the input of various types of information to the information processing apparatus 1 from the connected input devices. Further, the input / output unit 40 outputs various types of information to the connected output devices under the control of the control unit 10. Examples of the input / output unit 40 include an interface such as USB (Universal Serial Bus).
[0059] (Flow of the information processing method S1) The flow of the information processing method S1 executed by the information processing apparatus 1 will be described with reference to FIG. 8. FIG. 8 is a flowchart showing an example of the flow of the information processing method S1. The information processing method S1 includes, for example, as shown in FIG. 8, an acquisition process (step) S11, a first correspondence generation process (step) S12, a space generation process (step) S13, a data addition process (step) S14, a second correspondence generation process (step) S15, an additional acquisition process (step) S16, and a virtual spectrum generation process (step) S17.
[0060] (Step S11) In step S11, the acquisition unit 11 acquires a plurality of spectra.
[0061] (Step S12) In step S12, the first correspondence generation unit 12 generates association data representing the association between the plurality of spectra acquired by the acquisition unit 11.
[0062] (Step S13) In step S13, the space generation unit 13 compresses the dimensions in the association data using a predetermined algorithm, and generates a feature space which is a space including the compressed data.
[0063] (Step S14) In step S14, the data addition unit 14 adds one or more virtual data to the feature space.
[0064] (Step S15) In step S15, the second correspondence generation unit 15 restores the added virtual data to the same dimension as the association data generated by the first correspondence generation unit 12, and further generates association data.
[0065] (Step S16) In step S16, the additional acquisition unit 16 acquires one or more additional spectra.
[0066] (Step S17) In step S17, the virtual spectrum generation unit 17 generates one or more virtual spectra using the one or more additional spectra and the association data generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15.
[0067] 〔Embodiment 2〕 Other embodiments of the present invention will be described below. For the sake of convenience of explanation, members having the same functions as those described in the above embodiments are denoted by the same reference numerals, and the description thereof will not be repeated.
[0068] (Overview of Embodiment 2) In Embodiment 2, the first correspondence generation unit 12 and the second correspondence generation unit 15 included in the information processing apparatus 1 according to Embodiment 1 will be described. More specifically, a case where the association data generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 is a conversion coefficient will be described with reference to FIGS. 9 to 13. Note that other configurations of the information processing apparatus 1 according to Embodiment 2 are the same as those of the information processing apparatus 1 according to Embodiment 1.
[0069] (Case where the association data is a conversion coefficient) The association data may be, for example, a conversion coefficient for mutually converting spectra between a plurality of spectra. For example, the conversion coefficient may be a regression coefficient included in a regression equation for approximately converting one spectrum into another spectrum. The regression coefficient may be calculated, for example, by linear regression for each wavelength. Specific examples of the analysis method used in the linear regression for each wavelength include PLS regression, LASSO regression, and the like. Further, for example, when obtaining the regression coefficient, the proximal method may be used to calculate the entire regression coefficient at once.
[0070] FIG. 9 is a diagram showing an example of a regression coefficient for converting the spectrum in the master unit of the measuring device into the spectrum in the slave unit. In the example shown in FIG. 9, the spectra in the master unit and the slave unit are both spectra showing the wavelength dependence of the measured values. Further, as illustrated in FIG. 9, the spectra in the master unit and the slave unit may each be represented by a matrix having the number of samples as the number of rows and the number of wavelengths as the number of columns. At this time, for example, the regression coefficient B for converting the spectrum in the master unit into the spectrum in the slave unit may be a matrix in which both the number of rows and the number of columns are the number of wavelengths, as shown in FIG. 9.
[0071] FIG. 10 is a diagram showing a conversion coefficient which is an example of the association data. Hereinafter, the conversion coefficient illustrated in FIG. 10 will be described with reference to the example of the conversion coefficient "1->2", but the same applies to other conversion coefficients. The conversion coefficient "1->2" in FIG. 10 is a conversion coefficient representing the conversion from conversion set #1 which is the source of conversion to the same #2 which is the destination of conversion. For example, the conversion coefficient "1->2" may be a regression coefficient included in a regression equation for approximately converting conversion set #1 to the same #2.
[0072] In FIG. 10, the conversion coefficients "1->2" and "3->1" are illustrated. In each of the conversion coefficients illustrated in FIG. 10, the vertical axis indicates the wavelength in the spectrum of the source of conversion, and the horizontal axis indicates the wavelength in the spectrum of the destination of conversion. Further, the numerical values included in the conversion coefficients in FIG. 10 are indicated by the color scales attached to the right side of the respective conversion coefficients.
[0073] FIG. 11 is a diagram showing an example of the result of dimension compression by the space generation unit 13 using principal component analysis in the conversion coefficients illustrated in FIG. 10. In the example in FIG. 11, points corresponding to the respective conversion coefficients are shown in the feature space in which the first principal component axis PC1 and the second principal component axis PC2 are set. Further, in the example in FIG. 11, an example of the result of dimension compression including the conversion coefficients illustrated in FIG. 10 is shown.
[0074] FIG. 12 is a diagram showing an example of addition of virtual data to the feature space illustrated in FIG. 11 by the data addition unit 14. The feature space illustrated in FIG. 12 is a space generated by the space generation unit 13 using principal component analysis, and the first principal component axis PC1 and the second principal component axis PC2 are set in the same manner as the feature space illustrated in FIG. 11. The black dots in the example of FIG. 12 represent the data generated by compressing the dimensions in the conversion coefficients. Here, the data addition unit 14 adds virtual data (× points), for example, to the feature space generated by the space generation unit 13 as shown in FIG. 12.
[0075] The ellipse in the feature space illustrated in FIG. 12 is an example of a confidence ellipse set at an arbitrary range and interval between 0% and 100% in the principal component scores. For example, the range in which the data addition unit 14 adds virtual data on the feature space may be determined by the confidence ellipse in the principal component scores.
[0076] FIG. 13 is an example of the generation of a virtual spectrum using regression coefficients, which are an example of the conversion coefficients generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 by the virtual spectrum generation unit 17. For example, as shown in FIG. 13, the virtual spectrum generation unit 17 may multiply the matrix representing the learning set of the master device acquired by the additional acquisition unit 16 by the matrix representing the regression coefficients (B1 to B50) generated by each of the first correspondence generation unit 12 and the second correspondence generation unit 15 to generate virtual spectra (synthesis 1 to synthesis 50).
[0077] With the above configuration, a virtual spectrum can be generated using the conversion coefficients between a plurality of spectra.
[0078] 〔Example of Realization by Software〕 The functions of the information processing apparatus 1 (hereinafter referred to as “apparatus”) are programs for causing a computer to function as the apparatus, and can be realized by programs for causing a computer to function as each control block of the apparatus (particularly each unit included in the control unit 10).
[0079] In this case, the above apparatus includes a computer having at least one control device (for example, a processor) and at least one storage device (for example, a memory) as hardware for executing the above program. By executing the above program by this control device and storage device, each function described in the above embodiments is realized.
[0080] The above program may be recorded on one or more computer-readable recording media, rather than being temporary. This recording medium may or may not be provided in the above device. In the latter case, the above program may be supplied to the above device via any wired or wireless transmission medium.
[0081] Also, part or all of the functions of each of the above control blocks can also be realized by a logic circuit. For example, an integrated circuit in which a logic circuit functioning as each of the above control blocks is formed is also included in the scope of the present invention. In addition to this, for example, it is also possible to realize the functions of each of the above control blocks by a quantum computer.
[0082] Also, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may operate in the above control device, or may operate in another device (for example, an edge computer or a cloud server, etc.).
[0083] 〔Summary〕 The information processing apparatus according to aspect 1 of the present invention includes an acquisition unit that acquires a plurality of spectra, a first correspondence generation unit that generates correspondence data representing the correspondence between the plurality of spectra acquired by the acquisition unit, a space generation unit that compresses the dimension in the correspondence data using a predetermined algorithm and generates a feature space that is a space including the compressed data, a data addition unit that adds one or more virtual data to the feature space, a second correspondence generation unit that restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit and further generates the correspondence data, an additional acquisition unit that acquires one or more additional spectra, and a virtual spectrum generation unit that generates one or more virtual spectra using the one or more additional spectra and the correspondence data generated by each of the first correspondence generation unit and the second correspondence generation unit.
[0084] In the information processing apparatus according to aspect 2 of the present invention, in the above aspect 1, a predetermined target variable is associated with each spectrum in the one or more additional spectra, and when the virtual spectrum generation unit generates one or more virtual spectra from one spectrum, the predetermined target variable is associated with the one or more virtual spectra.
[0085] In the information processing apparatus according to aspect 3 of the present invention, in the above aspect 1 or 2, the association data is a difference spectrum representing a difference between spectra, or a conversion coefficient for mutually converting spectra between spectra.
[0086] In the information processing apparatus according to aspect 4 of the present invention, in any of the above aspects 1 to 3, the predetermined algorithm is principal component analysis, kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or an autoencoder.
[0087] The learning data according to aspect 5 of the present invention is data including the one or more virtual spectra generated by the information processing apparatus described in the above aspect 2, and is data for training a machine learning model that takes a spectrum as an input and outputs a target variable.
[0088] The information processing method according to aspect 6 of the present invention includes an acquisition process of acquiring a plurality of spectra, a first correspondence generation process of generating correspondence data representing a correspondence between the spectra for the plurality of spectra acquired in the acquisition process, a space generation process of compressing the dimension in the correspondence data using a predetermined algorithm and generating a feature space which is a space including the compressed data, a data addition process of adding one or more virtual data to the feature space, a second correspondence generation process of restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process and further generating the correspondence data, an additional acquisition process of acquiring one or more additional spectra, and a virtual spectrum generation process of generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process.
[0089] The program according to aspect 7 of the present invention causes a computer to execute an acquisition process of acquiring a plurality of spectra, a first correspondence generation process of generating correspondence data representing a correspondence between the spectra for the plurality of spectra acquired in the acquisition process, a space generation process of compressing the dimension in the correspondence data using a predetermined algorithm and generating a feature space which is a space including the compressed data, a data addition process of adding one or more virtual data to the feature space, a second correspondence generation process of restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process and further generating the correspondence data, an additional acquisition process of acquiring one or more additional spectra, and a virtual spectrum generation process of generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process.
[0090] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
Example
[0091] 〔Example 1〕 An embodiment of the present invention will be described below.
[0092] Four Kubota K-BA800 spectrometers were used as measuring devices to measure the spectra of 740 mini-tomatoes, and the juice was squeezed to measure the Brix sugar content. Among the four devices, No. 1 was the master device, and the others were slave devices. Of the obtained spectra, 10 points were used as the "conversion set", 472 points were used as the "learning set (master device only)", and 150 points were used as the "verification set". The acquisition unit 11 acquired the conversion set, and the first correspondence generation unit 12 calculated the difference spectra between each master-slave device pair. The space generation unit 13 performed principal component analysis on the difference spectra obtained by the first correspondence generation unit 12 to calculate a feature space composed of the first and second principal components. The data addition unit 14 sampled 100 virtual data points to be added on the grid on the feature space calculated by the space generation unit 13. In Example 1, the maximum and minimum values of the first and second principal component scores of the difference spectra obtained by the first correspondence generation unit 12 were set as the sampling range, and 100 virtual data points were uniformly sampled within the sampling range. The second correspondence generation unit 15 applied the inverse transformation of the process in the space generation unit 13 to generate 100 corresponding new difference spectra. The virtual spectrum generation unit 17 added the difference spectra generated by the first correspondence generation unit 12 and the second correspondence generation unit 15 to the learning set acquired by the additional acquisition unit 16 via the master device to generate 47,200 virtual spectra. Machine learning was applied to the measured Brix sugar content values corresponding to the spectra generated by the virtual spectrum generation unit 17 to construct a model for estimating the Brix sugar content from the spectra.
[0093] FIG. 14 is a diagram showing the result of applying the estimation model based on the spectrum generated by the information processing apparatus 1 in Example 1 to the verification set. In FIG. 14, the horizontal axis represents the actually measured Brix value, and the vertical axis represents the estimated value. Also, the round dots represent the estimation results of the learning set, and the triangular dots represent the estimation results of the verification set. Further, in FIG. 14, the results of machine #1, machine #2, machine #3, and machine #4 are shown respectively. As shown in the result of Example 1 in FIG. 14, the absolute value of the bias in each of machines 1 to 4 was at most about 0.1.
[0094] [Example 2] Another embodiment of the present invention will be described below.
[0095] In Example 2, the sampling range of the virtual data was changed in the data addition unit 14 of Example 1. Other conditions in Example 2 were the same as those in Example 1.
[0096] FIG. 15 is a diagram showing the sampling range of the virtual data by the data addition unit 14 in Example 2. In Example 2, the same sampling range as in Example 1 was set to "1 time". Then, in Example 2, even when the sampling range was changed between 0.2 times and 10 times, 100 points of virtual data were uniformly sampled within each sampling range, and a new difference spectrum was generated based on this. In FIG. 15, the cases of sampling ranges of 0.5 times, 1 time (the same as in Example 1), and 2 times are illustrated.
[0097] FIG. 16 is a diagram showing the result of applying the estimation model based on the spectrum generated by the information processing apparatus 1 in Example 2 to the verification set. In FIG. 16, the result of slave machine 2 is shown. Also, in FIG. 16, the cases of sampling ranges of 0.5 times, 1 time (the same as slave machine 2 in Example 1), and 2 times are illustrated as in FIG. 15. As shown in the result of Example 2 in FIG. 16, the absolute value of the bias in each of the sampling ranges of 0.5 times and 2 times was at most about 0.1.
[0098] (Results on Estimation Errors in Examples 1 and 2) Table 1 shown below is a diagram showing the results in Examples 1 and 2 regarding the estimation error (RMSE (Root Mean Squared Error)). The value of the item with the magnification of the sampling range "1" in Table 1 is the result in Example 1. Also, in Table 1, the results when the information processing apparatus 1 is not used (the item of "without spectrum expansion") are also shown together.
[0099] Table 1 Results in Examples 1 and 2 Regarding Estimation Error
[0100] [Table 1] As shown in Table 1, when the information processing apparatus 1 was not used, there were slave devices with an estimation error exceeding 1, while in Example 1, the estimation error was less than 0.4 for all slave devices. That is, in Example 1, it was found that the measured values and the estimated values tended to match accurately in the slave devices as well as in the master device.
[0101] Also, as shown in Table 1, in Example 2, even when the sampling range was changed to any of 0.2 times, 0.5 times, 2 times, 5 times, and 10 times, the estimation error was less than 1 for all slave devices. That is, in Example 2, it was found that the measured values and the estimated values tended to match accurately in the slave devices even when the sampling range was changed between 0.2 times and 10 times. [Description of Signs]
[0102] 1 Information processing apparatus 10 Control unit 11 Acquisition unit 12 First correspondence generation unit 13 Space generation unit 14 Data addition unit 15 Second correspondence generation unit 16 Additional acquisition unit 17 Virtual spectrum generation unit 20 Memory unit 30 Communication unit 40 Input / output unit
Claims
1. An acquisition unit for acquiring a plurality of spectra; a first correspondence generating unit configured to generate correspondence data representing correspondence between the spectra for the plurality of spectra acquired by the acquiring unit; a space generating unit that compresses dimensions in the association data using a predetermined algorithm and generates a feature space that is a space including the compressed data; a data adding unit that adds one or more virtual data to the feature space; a second correspondence generation unit that restores the added virtual data to the same dimension as the correspondence data generated by the first correspondence generation unit, and further generates the correspondence data; An additional acquisition unit for acquiring one or more additional spectra; a virtual spectrum generating unit that generates one or more virtual spectra by using the one or more additional spectra and the correspondence data generated by each of the first correspondence generating unit and the second correspondence generating unit; Equipped with the correspondence data is a difference spectrum representing a difference between the spectra, or a conversion coefficient for converting the spectra into each other, When the correspondence data is the difference spectrum, the virtual spectrum generation unit adds, for each wavelength, a measurement value in the one or more additional spectra and a measurement value in the correspondence data generated by the second correspondence generation unit to generate the one or more virtual spectra; When the correspondence data is the conversion coefficient, the virtual spectrum generation unit multiplies a matrix representing the one or more additional spectra by a matrix representing the correspondence data generated by the second correspondence generation unit to generate the one or more virtual spectra. Information processing device.
2. Each spectrum in the one or more additional spectra is associated with a predetermined response variable; the virtual spectrum generation unit, when generating one or more virtual spectra from one spectrum, associates the predetermined response variable with the one or more virtual spectra; The information processing device according to claim 1 .
3. The predetermined algorithm is principal component analysis, kernel principal component analysis, independent component analysis, factor analysis, multidimensional scaling, non-negative matrix factorization, t-SNE (t-distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), or an autoencoder. The information processing device according to claim 1 .
4. an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing correspondence between the spectra for the plurality of spectra acquired in the acquisition process; a space generation process of compressing dimensions in the association data using a predetermined algorithm and generating a feature space that is a space including the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process, and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process; Including, the correspondence data is a difference spectrum representing a difference between the spectra, or a conversion coefficient for converting the spectra into each other, When the correspondence data is the difference spectrum, in the virtual spectrum generation process, a measurement value in the one or more additional spectra is added to a measurement value in the correspondence data generated in the second correspondence generation process for each wavelength to generate the one or more virtual spectra; When the correspondence data is the conversion coefficient, in the virtual spectrum generation process, a matrix representing the one or more additional spectra is multiplied by a matrix representing the correspondence data generated in the second correspondence generation process to generate the one or more virtual spectra. Information processing methods.
5. On the computer, an acquisition process for acquiring a plurality of spectra; a first correspondence generation process for generating correspondence data representing correspondence between the spectra for the plurality of spectra acquired in the acquisition process; a space generation process of compressing dimensions in the association data using a predetermined algorithm and generating a feature space that is a space including the compressed data; a data addition process for adding one or more virtual data to the feature space; a second correspondence generation process for restoring the added virtual data to the same dimension as the correspondence data generated in the first correspondence generation process, and further generating the correspondence data; an additional acquisition process for acquiring one or more additional spectra; a virtual spectrum generation process for generating one or more virtual spectra using the one or more additional spectra and the correspondence data generated in each of the first correspondence generation process and the second correspondence generation process; Run the command, the correspondence data is a difference spectrum representing a difference between the spectra, or a conversion coefficient for converting the spectra into each other, When the correspondence data is the difference spectrum, in the virtual spectrum generation process, a measurement value in the one or more additional spectra is added to a measurement value in the correspondence data generated in the second correspondence generation process for each wavelength to generate the one or more virtual spectra; When the correspondence data is the conversion coefficient, in the virtual spectrum generation process, a matrix representing the one or more additional spectra is multiplied by a matrix representing the correspondence data generated in the second correspondence generation process to generate the one or more virtual spectra. program.
Citation Information
Patent Citations
Data generation method, learning model generation method, computer program, information processor, and analysis device
JP2023074746A
Signal processing method and signal processing device
WO2024043009A1