METHOD FOR CLASSIFIING SPECTRA OF OBJECTS WITH COMPLEX INFORMATION CONTENT
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- TECHNISCHE UNIVERSITAT DRESDEN
- Filing Date
- 2017-09-15
- Publication Date
- 2026-05-13
AI Technical Summary
Existing classification methods for optical molecular spectra, particularly in-ovo spectroscopy of chicken eggs, face challenges in achieving a balance between accuracy and robustness due to high variability and external influences, often leading to overtraining and reduced classification accuracy.
A method involving multiple classification procedures with different data pretreatments and classifiers, iteratively calculated and validated, to emphasize certain features while suppressing others, using techniques like linear discriminant analysis, neural networks, and cluster analysis to determine the probability of class membership.
This approach enhances the accuracy and robustness of sex determination in chicken eggs by minimizing interference and variability, ensuring stable classification results despite high natural variability and external disturbances.
Description
[0001] The invention relates to a method for classifying spectra of objects with complex information content with at least two different object information types, in particular optical molecular spectra for assigning the object information.
[0002] Publications on methods for the aided classification of optical spectra are known. These include, among others, the publication by AE Nikulin, B. Dolenko, T. Bezabeh, RL Somorjai: Near-optimal region selection for feature space reduction: novel preprocessing methods for classifying MR spectra, NMR Biomed. 11 (4-5), 1998, pp. 209-216, the publication by BK Lavine, CE Davidson, AJ Moores: Genetic algorithms for spectral pattern recognition, Vibrational Spectroscopy, Volume 28, Issue 1, 2002, Pages 83-95, in which the algorithm is based on the principal components and a weighting of spectral regions is used for classification, and the publication by J. Jacques, C. Bouveyron, S. Girard, O. Devos, L. Duponchel, C. Ruckebusch: Gaussian mixture models for the classification of high-dimensional vibrational spectroscopy data, Journal of Chemometrics, Volume 24, Issue 11-12, pp. 719-727 (Dec. 2010).
[0003] It describes a method in which particularly high-dimensional spectral data are decomposed into so-called subspaces, which are subsequently classified using discriminant analysis.
[0004] Also known are Eladio Rodriguez-Diaz, Satish K. Singh, Irving J. Bigio, David A. Castañón, "Spectral classifier design with ensemble classifiers and misclassification rejection: application to elastic-scattering spectroscopy for detection of colonic neoplasia", J. Biomed. Opt. 16(6) 067009 (1 June 2011), which classifies different spectral ranges separately and combines their results using an ensemble system; as well as Lu Xu, Yan-Ping Zhou, Li-Juan Tang, Hai-Long Wu, Jian-Hui Jiang, Guo-Li Shen, Ru-Qin Yu, Ensemble preprocessing of near-infrared (NIR) spectra for multivariate calibration, Analytica Chimica Acta, Volume 616, Issue 2, 2008, Pages 138-143, ISSN 0003-2670, which applies several preprocessing methods of NIR spectra in parallel and combines the resulting calibration models in an ensemble to obtain more robust and accurate multivariate calibration models and thus reduce the dependence on a single "best" preprocessing method.
[0005] Optical molecular spectra contain a wealth of information about the molecular properties of the object under investigation. Vibrational spectra, in particular, are considered a fingerprint of molecules due to their high information density about molecular structure. In the spectroscopic analysis of complex biological objects, the relevant information, as specified in the parameters, must be separated from less or insignificant information, as well as from interference. This is typically achieved using chemometric methods, and for higher-dimensional data, multivariate methods are also employed.
[0006] If important spectral features of the sought-after molecular information of objects are known, supervised classification methods can be used, as described, for example, in the publication G. Steiner, S. Kuchler, A. Herrmann, E. Koch, R. Salzer, G. Schackert, M. Kirsch: Cytometry, Part A 2008, 73A, 1158-1164.
[0007] Supported classification methods are distinguished from other methods by their higher accuracy in the identification and quantitative evaluation of the sought-after information. In the well-known supported classification method according to... Fig. 4a A classifier 50 is calculated using a training set 19 consisting of representative spectra with assigned properties. The classifier 50 is then tested against an independent test set 29; that is, the spectra are not used to construct the classifier 50, as in Fig. 4a demonstrated, checked and evaluated for validation, and a classified test set of 24 was achieved.
[0008] The construction of the classifier 50 using the training set 19 is verified by a test set 29 with, for example, a maximum of 30% of the spectra (dashed line to the created classifier) according to Fig. 4b verified, so that as a result a classified test set 24 (dashed line from classifier 50) is achieved for the classified test set 24.
[0009] A general problem with supported classification methods is the trade-off between the accuracy of the resulting classifications and the robustness of the classification. Very high accuracies can often only be achieved through so-called overtraining of the classifier. This means that the classifier can only correctly classify certain spectra with very high accuracy. Conversely, even the slightest deviations or disturbances lead to a dramatically reduced accuracy of the classification. Therefore, a balance between the best possible classification and high robustness of the classification is sought.
[0010] When dealing with spectra exhibiting very high variability, such as in-ovo spectra used for sex determination of chicken eggs, compromises in accuracy must inevitably be made to maintain sufficient robustness of the classification. This inherent contradiction is fundamentally unsolvable. To nevertheless achieve good stability with sufficient accuracy, various classification methods have been developed in recent years. The basic approach involves parallelizing the classification process using different decision trees. The Random Forest method is based on a network of uncorrelated decision trees, which are randomly generated and linked during the training process. Each tree makes a decision.The group of trees with the most identical decisions determines the classification result, i.e., the spectrum assignment. However, the Random Forest method cannot compensate for differently occurring disturbances or variations in the spectra. Here, too, simply setting up too many trees can lead to overtraining.
[0011] Document US20120321174 A1 describes a classification method based on the Random Forest method for image analysis. This supported classification method is designed to take into account particularly small but relevant features for classification.
[0012] These relevant features of the general classification procedure can also be defined and play a role, for example, in in-ovo spectroscopy of chicken eggs in the form of small spectra-related signals of sex information.
[0013] In in-ovo spectroscopy of chicken eggs, a supported classification method is used to identify the sex.
[0014] However, optical in-ovo spectra are often characterized by very high natural variability, which significantly overshadows the comparatively small signals of sex information. In addition, there are unavoidable external influences from the measurement environment itself.
[0015] Currently, the following different methods for classifying the spectra of objects, in particular for determining the sex of fertilized and / or incubated eggs, are described in the publications listed below: Publication WO 2010 / 150265 A describes a method based on coloration, specifically of the feathers of the developing embryo. This method is based on the fact that, in the advanced stage of development (day 12 of incubation), the color of the feathers in certain chicken breeds allows conclusions to be drawn about the sex. The evaluation is carried out using a classification algorithm.
[0016] Furthermore, publication WO 2014 / 021715 A2 describes a procedure in which the sex of the embryo is determined by means of endocrinological analysis.
[0017] The publication DE 10 2007 013 107 A1 describes the application of Raman spectroscopy for the sex determination of birds, generally examining cell-containing material. However, no method for in-ovo sex determination is described.
[0018] The molecular spectra are recorded using methods and devices according to the following publications: A method and devices for determining the sex of chicken eggs, based on optical, preferably fiber-coupled, spectroscopy, are described in publication DE 10 2010 006 161 B3. However, no methods for analyzing the spectra and for classification are described.
[0019] Documents DE 10 2014 010 150 A1 and WO 2016 / 000678 A1 describe methods and devices for Raman spectroscopic in ovo sex determination. The evaluation of the spectra can advantageously be carried out using chemometric methods.
[0020] The publication EP 2 336 751 A1 describes a method for determining the sex of bird eggs. In this method, the germinal disc of an egg is illuminated with light, and the emitted fluorescence is recorded with time resolution. Sex identification is achieved using a supported classification method, whereby a classifier is calculated using the fractal dimension method.
[0021] Document US 6 029 080 B describes a method for in ovo sex determination. From a certain stage of embryonic development, the sex organs can be identified by analyzing MRI images of the egg and used for sex determination.
[0022] The disadvantage of evaluating these methods is that each of these methods ultimately uses a procedure for classifying spectra of objects with only one classifier.
[0023] In summary, the disadvantages include the fact that, in order to reliably extract the desired sex information from the recorded spectra, the use of only a single classifier is insufficient to guarantee adequate detection accuracy. Instead, the variable influences and variations in the biochemical composition of the egg, as well as in the different developmental stages, must be taken into account. To incorporate this wide range of variation into a reliable classification procedure, it is estimated that the calculation using only one classifier will therefore be inadequate.
[0024] The invention therefore aims to provide a method for classifying the spectra of objects with complex information content, designed to achieve maximum accuracy in determining the assigned selected features of the objects, while simultaneously maintaining at least the stability of the classification. The goal is to achieve a balance between optimal classification accuracy and high robustness. At the same time, overtraining the classifier should be avoided.
[0025] The problem is solved using the features of claim 1.
[0026] In the method for classifying spectra of objects with complex information content with at least two different object information, using a method for registering and pretreating spectral data and a method of classification associated with the data pretreatment with the calculation of a classifier, according to the characterizing part of claim 1, after the registration of the spectra and the pretreatment of spectral data, a multiple classification method is carried out with at least two different methods of data pretreatment of spectral data and a method of classification associated with the respective data pretreatment.
[0027] Within the scope of the invention, the registration of spectra is understood to mean the recording, determination and storage of spectra and the generation of digitized signals for storage, which are available for further data preprocessing of the spectral data.
[0028] Depending on the pretreatment algorithm used, different corrected, pretreated spectra with many data points are generated during data pretreatment, and these are assigned to at least one classification procedure.
[0029] The following steps are carried out in the evaluation process of registered spectra of objects after registration and data preprocessing: a calculation of multiple classifiers per type of data pretreatment, calculation of multiple classifiers per type of data pretreatment by iterative calculation and validation starting from the training set, classification of the pretreated spectra of the test set with all classifiers calculated from the training data set, - classification of the spectra into a class of object information with an expression of a probability for classification membership, calculation of a classification result by calculating a median or by performing a cluster analysis to represent the probability result of the object information belonging to a class.
[0030] When defining and determining the number of calculated classifiers NG in the series per classification group, both the scope of the spectral data points are taken into account. vS and the twice half-width of the spectral regions w S as well as the number of selected spectral regions RS for the classification in equation (I) are taken into account: N G = v S 2 w S ⋅ R S where equation (I) ensures that each data point v S can be selected with equal probability.
[0031] The data points belonging to a set of spectral data points can also be weighted.
[0032] At least one of the spectral pretreatments is structured in such a way that certain features are emphasized and other specific features are suppressed, so that differently defined features are used for classification.
[0033] At least one spectral pretreatment can be designed with identically defined features, and at least one of the aforementioned spectral pretreatments with differently defined features can be used for classification.
[0034] The pre-treated spectra can be designed as variable training sets, and several classifiers of the series or classifiers are iteratively determined and validated.
[0035] Within the scope of the invention, classification is understood to mean the assignment of pretreated spectra to a respective class according to a predetermined algorithm. The classification process is carried out using predetermined parameters, and the result of the classification is expressed by a calculated classifier.
[0036] At least one supported and / or unsupported classification method can be used to select spectral regions or individual wavelength ranges and subsequently analyze them. Linear or nonlinear discriminant analysis can be employed.
[0037] Neural network methods and / or linear time-frequency transformation (wavelet) methods can also be used for classification.
[0038] The spectra of optical molecular spectroscopy, such as absorption, emission, scattering or UV / vis, NIR, IR absorption, fluorescence or Raman, can be classified.
[0039] Data pretreatments of the recorded spectral data or raw spectra can include baseline corrections, normalizations, derivations, covariance and / or principal component analysis.
[0040] To evaluate the classifiers for a classification result, a calculation of a median or a cluster analysis can be performed.
[0041] The median, or central value, is given as a mean value for distributions in statistics. The median of a list of numerical values is the value that occupies the middle (central) position when the values are sorted in order of magnitude. The value of the magnitude here represents either the score of a classifier or the probability of belonging to a particular class as determined by the classifier.
[0042] The inventive method is carried out in detail with the following steps: Acquisition and recording of spectra using at least one optical device with at least one spectrometer and / or further detectors, generation of digitized signals in the form of data points and storage of the recorded spectra in storage units of classification units of an evaluation unit, performance of spectral pretreatment by individually pretreating the recorded and stored spectra in the individual storage units and making the associated digitized evaluated signals available for further processing, separation of the pretreated spectra as a training set and as a test set, design and use of the pretreated spectra as a training set and of a test set separate from the training set, calculation of several classifiers per type of data pretreatment by iterative calculation and validation starting from the training set,by selecting spectral regions RS from a coordinate of relative wavenumbers and subsequently classifying the intensity values of the selected regions RS using discriminant analysis, wherein in a repeated step a further selection of spectral regions RS and the classification of the intensity values is performed, and wherein the cycle is repeated iteratively until an accuracy that can no longer be improved is reached, specifying a termination criterion, classifying the pre-treated spectra of the test set with all outgoing classifiers calculated from the training data set, - assigning the spectra to a class of object information with an expression of a probability for classification membership, calculation of a classification result by calculating the median or by performing a cluster analysis to represent the probability result of the object information belonging to a class.
[0043] Any bird eggs, or optionally chicken eggs, can be used as objects with object information, and in a specific use case, the dual information about the female or male sex can be used as object information.
[0044] The procedure thus comprises steps for performing multiple classification, based on various conventional evaluation methods following spectral pretreatment after spectral detection and registration, and subsequent multiple calculations using different classifiers. At least one spectral pretreatment is involved, structured in such a way that, while considering the equivalence of features, certain features are emphasized more strongly, while others are more strongly suppressed. The pretreated spectra are then used as a training set, with several series of classifiers being calculated. The classifiers are calculated and validated iteratively. In this way, multiple classifiers can be determined. The spectra of the test set are then classified using all classifiers.The classification of spectra into a specific class of object information / characteristics (in chickens: male, female) is preferably done as a point value (score) or as a probability expression for class membership. To derive a meaningful conclusion from the classifiers, the relationships within the classifiers are determined. Simple methods for expressing these relationships include, for example, calculating the median or performing cluster analysis.
[0045] A device for classifying spectra of objects with complex information content, preferably with objects in the form of chicken eggs for determining dual egg information - female or male -, wherein the aforementioned method is implemented in the device, can comprise at least the following units At least one detecting optical device with at least one spectrometer and / or further detectors for recording and registering the spectra; a unit for generating digitized signals in the form of data points that represent the spectra; storage units for storing the registered spectra in the classification units / groups of an evaluation unit comprising the classification units; spectral preprocessing units in which the registered spectra are individually preprocessed and the associated digitized evaluated signals – the determined data points – are made available for further processing; training sets for designing and applying the preprocessed spectra; at least one classification unit for calculating multiple classifiers per type of data preprocessing iterative procedure and validation based on the training set in the classification units.Test sets for classifying the pre-treated spectra with all classifiers per type of data pre-treatment; a unit for classifying the pre-treated spectra into at least one dual class – in the case of chickens: male or female – of object (egg) information with an expression of a probability for class membership; an evaluation unit for calculating the classification result, e.g., in the form of the median or by performing a cluster analysis to determine the probability result of at least one of the object (egg) information belonging to the dual class – in the case of chickens: female or male.
[0046] Further developments and embodiments of the invention are specified in further dependent claims.
[0047] The invention is explained by means of exemplary embodiments with reference to drawings.
[0048] They show: Fig. 1 a schematic block representation of a method according to the invention for classifying spectra of objects with complex information content, in particular optical molecular spectra for assigning and determining dual object information, wherein the method is implemented as a multiple classification method, Fig. 2a a schematic representation of a single classification method with spectral pretreatment: raw spectra, wherein all features are considered equivalent (circles of equal size), Fig. 2b a schematic representation of a single classification method with spectral pretreatment: linear baseline correction, with a favored large feature circle of fluorescence intensity and several less relevant small feature circles, Fig.Fig. 2a A schematic representation of a single classification procedure with spectral pretreatment: normalization, with a favored large feature set of the fluorescence profile and several less relevant small feature sets, Fig. 2a A schematic representation of a single classification procedure with spectral pretreatment: Raman spectra, with a favored large feature set of the molecular composition and several less relevant small feature sets, Fig. 3a A schematic representation of the spectra assigned to the raw spectrum classification procedure according to . Fig. 2a , where the dotted spectrum is assigned to the female hen egg spectrum, Fig. 3 schematic representation of the spectra assigned to the linear baseline correction classification method according to Fig. 2b , where the dotted spectrum is assigned to the female hen egg spectrum, Fig. 3 is a schematic representation of the spectra assigned to the nomination classification procedure according to Fig. 2c , where the dotted spectrum is assigned to the female hen egg spectrum, Fig. 3 a schematic representation of the spectra assigned to the Raman spectrum classification method according to Fig. 2d , wherein the dotted spectrum is assigned to the female hen egg spectrum, Fig. 4a a schematic representation of the process of a classification with determination of a classifier according to the prior art, Fig. 4b a flowchart for the multiple classification method according to the invention with training set and test set in algorithmic connection with classification result design from a large number of classifiers, Fig. 5 a schematic representation of a probability / number of classifier bar chart for twenty classifiers according to Fig. 1 For an example of a chicken egg, where the columns above the dashed line – dividing line – are assigned to the female sex, Fig. 6 shows a top view of the column representation of an egg according to the Fig. 5 and a possible view on a display, Fig. 7 a probability (female) / number of classifiers representation showing the calculated median for an egg with the 20 classifiers according to Fig. 6 , where a bold dashed line represents the cutoff point at 0.5 of the probability and a thin dashed median value represents the probability of being "female" at approximately 0.72 of the probability, so that the sex of the egg can be identified as female, Fig. 8 a schematic representation of a probability / number of classifiers bar chart for an egg for optionally 120 classifiers, where the end regions of the bars above the bold dashed line – cutoff point – are assigned to the female sex of an egg and the end regions of the bars below the bold dashed line are assigned to the male sex of an egg, Fig. 9 a representation of the calculated, dashed median for an egg identified as female with the 120 classifiers according to Fig. 8 In a probability / number of classifier representation, Fig. 10a shows a schematic representation of a single classification procedure with spectral pretreatment: raw spectra, where all features are considered equivalent, with eight selected spectral regions RS1 to RS8, with a range of spectral data points across the entire spectral range between wavenumbers 570 cm⁻¹ to 2750 cm⁻¹ to determine the number of classifiers per data pretreatment. Fig. 10 shows an enlarged section (II) of the data point representation in the specified spectral region RS8 according to Fig. 10a , Fig. 10 shows an enlarged section (II) of the data point representation in the spectral region R S8 with an indication of the weighting of data points in the area of the region R S8 according to Fig. 10a und Fig. 10b , Fig. 11a a representation of the calculated, dashed median for an egg identified as male with the 120 classifiers similar to the Fig. 8 in a probability / classifier count representation, Fig. 11 shows a first histogram representation of the dependence between the number of elements in the cluster and the centroid of the cluster, and Fig. 11 shows a second histogram representation of the dependence between the number of elements in the cluster and the centroid of the cluster.
[0049] The following will be the Fig. 1 and the Fig. 2a, 2b, 2c, 2d considered together. In Fig. 1 A schematic block diagram shows a method 1 according to the invention for classifying spectra 4 of an object 2 with complex information content with at least two / dual and different object information / features, in particular optical molecular spectra 4 for assigning object information / features 3 for a probable determination of an example dual object information 31, 32.
[0050] Bird eggs, for example chicken eggs, can be used as objects 2 to be examined, and the dual object information 3 can be, for example, the characteristic 31 about the female sex and the characteristic 32 about the male sex.
[0051] The following is a description of the inventive method 1 for carrying out the classification.
[0052] This is in the Fig. 1 A block-wise sequence of the inventive method 1 is shown.
[0053] In the procedure 1 for classifying spectra 4 of objects 2 with complex information content with at least two different object information, a classifier is calculated after registration using a data pretreatment procedure and a classification procedure associated with data pretreatment.
[0054] According to the invention, after registration and data pretreatment of spectra 4, a multiple classification method is carried out with at least two different data pretreatment methods 5, 6, 7, 8 of the spectra 4 and the method assigned to the respective data pretreatment 5, 6, 7, 8 for classification in the groups 9, 10, 11, 12 to determine several, e.g. five, classifiers per group 9, 10, 11, 12, i.e. a total of many, e.g. twenty (five classifier / group x four groups) classifiers 131, 132, 133, 134, 135, etc. for the series 14, 15, 16.
[0055] The following steps are carried out after registration and data preprocessing of the spectra, whereby the steps relate to the Fig. 1 relate: a calculation of five classifiers from series 13, 14, 15, 16 for each type of data pretreatment 5, 6, 7, 8, so that finally twenty classifiers 131, 132, 133, 134, 135, etc. are determined; a determination of the five classifiers from series 13, 14, 15, 16, iteratively adapted and validated; a calculation of probabilities of class membership; an equal inclusion of all five classifiers from series 13, 14, 15, 16 or classifiers 131, 132, 133, 134, 135, etc. in the determination of a classification result 18, e.g. in the form of a median 30.
[0056] When determining the number of classifiers NG to be calculated in series 13, 14, 15, 16 with respect to groups 9, 10, 11, 12, a range of spectral data points is used. vS and a double half-width w S of spectral regions RS as well as a number of the selected spectral regions RS in the following equation (I) are taken into account: N G = v S 2 w S ⋅ R S where equation (I) ensures that each data point v S can be selected with equal probability.
[0057] For a total of twenty classifiers of the four series 13, 14, 15, 16 with NG (13), NG (14), NG (15) and NG (16) according to Fig. 1 as well as according to the Fig. 5, Fig. 6 and Fig. 7 For example, the following parameters are specified for the entire spectral range from 500 cm⁻¹ to 2750 cm⁻¹: Scope of spectral data points v S in a given entire spectral range 500 cm⁻¹ to 2750 cm⁻¹ with v S = 800, number of selected spectral regions RS with RS = 8, width B of the spectral regions RS with B = 2 · w S = 5, i.e., twenty data points can be found in one region RS vS is located. The half-width w S is therefore w S = 2.5.
[0058] The spectral pretreatments 6, 7, 8 can be performed according to Fig. 2b, 2c, 2d are structured in such a way that certain features are favored and emphasized, while other features are suppressed.
[0059] During the spectral pretreatment 5 of the raw spectra 25 according to Fig. 2a All considered characteristics can be treated equally and given priority.
[0060] The pre-treated spectra 4 are presented as training set 24 according to Fig. 4b designed and several classifiers of series 13, 14, 15, 16 or, for example, in detail for series 13, classifiers 131, 132, 133, 134, 135 etc. were iteratively determined and validated.
[0061] After passing at least two classification procedures 9, 10, 11, 12, each with a preceding spectrum-different data pretreatment 5, 6, 7, 8, with at least one determined classifier 131, 141, 151, 161, according to Fig. 1 Overall, at least two of the determined classifiers 131, 141, 151, 161 are achieved and used for the evaluation and subsequent determination of a probability result 18 with regard to the given different object information 31, 32, whereby the probability result 18 is output, so that a conclusion is made at least on the object information 31 or 32 determined with the highest value.
[0062] At least one supported classification and / or unsupported classification method can be used to select spectral regions RS or individual wavelength ranges / wavenumber ranges and subsequent analysis.
[0063] The subsequent analysis can be a linear discriminant analysis or a nonlinear discriminant analysis.
[0064] However, it can also be used as a method for classification in groups 9, 10, 11, 12, a method of neural networks and / or a method of linear time-frequency transformations (wavelets).
[0065] The spectra 4 of optical molecular spectroscopy, such as absorption, emission, scattering or UV / vis, NIR, IR absorption, fluorescence, Raman, can be classified using the method according to the invention.
[0066] Among the in Fig. 1 The data pretreatments shown 5, 6, 7, 8 allow for the definition and application of raw spectra 25, baseline corrections 26, normalizations 27, derivatives, covariance, and / or principal component analysis / Raman spectra 28. Data pretreatment consists of generating digital signals, i.e., data points, which, when concatenated, yield the respective calculated spectral curves 25, 26, 27, 28 and which can thus be assigned to the individual classification procedures used.
[0067] To evaluate the classifiers of series 13, 14, 15, 16 for a classification result of 18, a median of 30 can be calculated ( Fig. 7 . Fig. 9 , Fig. 11a ) for an object 2 or a run ( Fig. 11a , Fig. 11b ) a cluster analysis is planned.
[0068] The well-known k-means cluster analysis can be used as an example of an evaluation. In this analysis, Fig. 11a At least two clusters are specified for "male" and "female". The cluster to which the most elements, i.e., probabilities, are assigned defines the gender ( Fig. 11b - male). A cluster is a group of combined elements with similar properties (here: probabilities - classification results).
[0069] The Fig. 11a The graph shows a curve of the relationship between the probability of an egg being classified as "male" and a classifier count of 120. The plot of the sorted probabilities shows that the median (30) lies just above the cutoff value (42) of 0.5 of the probability coordinate. Therefore, egg 2 will just barely be classified as male. The well-known k-means cluster analysis yields a clearer result.
[0070] This includes in the Fig. 11b a first histogram representation 43 of the dependence between the number of elements of the cluster and the centroid of the cluster and in the Fig. 11c A second histogram representation 44 of the dependence between the number of elements of the cluster and the centroid of the cluster is shown, where the elements are the classification results achieved.
[0071] This includes in Fig. 11b Two clusters were formed, with their centers at 0.84 and 0.17. Since the cluster with the center at 0.84 has more elements, i.e., classification results, assigned to it, egg 2 can be definitively assessed as male.
[0072] When selecting five clusters in the histogram representation 44 according to Fig. 11c The result is also confirmed. Here, 65 calculated probability values are classified as male, with the cluster (number 4 = #4) having the strongest value of 0.96 for "male" also containing the most elements according to Fig. 11a It contains [something]. Therefore, the sex of egg 2 is clearly "male".
[0073] The same and similar results apply to cluster analysis if the sex of egg 2 is determined to be female.
[0074] The inventive method 1 can be implemented by means of the following steps using hardware components of an associated device: Recording and registration of the spectra 4 by means of at least one optical device with at least one spectrometer and / or further detectors, generation of digitized signals, the data points, and storage of the detected spectra 4 in storage units of the classification units of an evaluation unit, spectral pretreatment 5, 6, 7, 8, by extracting the data points vThe existing stored spectra 4 in the individual storage units are evaluated individually and the associated digitized evaluated signals are provided for further processing, design or development of the pretreated spectra 25, 26, 27, 28 as training set 19 and as a separate test set 29, calculation of the classifiers of the series 13, 14, 15, 16 in the form of individual classifiers 131, 132, 133, 134, 135 of a series 13 etc.The individual classification procedures 9, 10, 11, 12 considered, including iterative procedures and validation in the classification units / groups, classification of the evaluated spectra 25, 26, 27, 28 of test set 24 with all classifiers of series 13, 14, 15, 16, classification of the spectra 25, 26, 27, 28 into a class of object information with an expression of a probability for class membership, calculation of the median 30 or performance of a previously mentioned cluster analysis to represent the probability result in the form of a classification result 18 of the object information belonging to a class.
[0075] The construction of classifiers with respect to series 13, 14, 15, 16 using a training set 19 is performed by a test set 29 with, for example, a maximum of 30% of the spectra (dashed line to classifiers 13, 14, 15, 16) according to Fig. 4b verified, so that the result is a classified test set 24 (dashed line from the classifiers of series 13, 14, 15, 16 to the classified test set 24).
[0076] It should be noted that, in general, the recorded in-ovo spectra are inherently highly variable. This is due, firstly, to the inherent variability of biological systems and, secondly, to the sensitivity of Raman spectroscopic measurements.
[0077] External disturbances of a systematic and random nature lead to a high variability of the spectral characteristics and thus obscure the gender-relevant information.
[0078] Furthermore, the Raman spectroscopy method also produces fluorescence light, which also contains molecular information, but simultaneously overlays the usually much weaker Raman spectroscopic molecular information about the composition of the object 2 under investigation.
[0079] In Fig. 2a is a schematic representation of the raw spectra 25 as one of all the individual classification methods considered in Fig. 1 the four individual classification procedures are specified.
[0080] Generally speaking, according to Fig. 2a, Fig. 2b, Fig, 2c und Fig. 2d form at least four classes of signals or specific characteristics: molecular composition 20, fluorescence intensity 21, fluorescence profile 22 and variation of physical parameters 23, where these specific features 20, 21, 22, 23 are used for visible representation in Fig. 3a, Fig. 3b, Fig. 3c and Fig. 3d are formed as circles of equal size with borders and / or differently dashed borders.
[0081] Behind each classifier is a mathematical expression for separating the signals according to the object information 3 (31 female, 32 male).
[0082] Three of the four classes / features 20, 21, 22 contain gender-relevant information. However, it is not possible to eliminate the variation 23 of the physical parameters from the spectra 4 in such a way that no or only a minimal loss of information occurs in the other three classes 20, 21, 22. The raw spectra 25 thus have the highest information content due to the equal weighting of all specific features, but also the highest level of interference. By adding at least one of the aforementioned data pretreatments, e.g., 26 from the data pretreatments 26, 27, 28 with differently weighted features, the interference is reduced. By using further data pretreatments 27, 28, the original interference is minimized or even eliminated.
[0083] In Fig. 2b It has been shown that the in-ovo spectra 4, recorded as digital signals, are subjected to a linear baseline correction 26, which highlights the fluorescence intensity signal 21 (large circle). Simultaneously, signals 23 of physical parameters in the spectra are suppressed (small circle). Due to the typically large intensity differences between the fluorescence signal and the Raman signal, information on the molecular composition 20 (small circle) also recedes into the background. However, the fluorescence intensity 21 (large circle) itself is a potential marker for sex determination, since male embryos frequently, but not always, exhibit a biochemical composition in their blood that shows an increased fluorescence intensity 21 compared to female embryos or the blood of female embryos.
[0084] In Fig. 2c It has been shown that variations in fluorescence intensity 21 (small circle) can be compensated for and the random influences of physical parameters 23 (small circle) minimized using spectral normalization methods 27, for example, vector or area normalization. This allows the fluorescence profile 22 (large circle) to be preferentially emphasized. At the same time, only a small amount of information about the molecular composition 20 (small circle), based on the Raman signals, is minimized. Since the fluorescence profile 22, i.e., the spectral characteristics of the fluorescence, is determined by the molecular composition 20, gender-relevant information can be highlighted.
[0085] In Fig. 2d It has been shown that a correction of the so-called background of Raman spectra 28 as complete as possible leads to the sole highlighting of the Raman bands, i.e. the information on the molecular structure and composition 20 (large circle) of the object 2 under investigation.
[0086] In the Fig. 3a, 3b, 3c und 3d Each of these is a schematic representation of the spectra assigned to the individual classification procedures (based on the relative wavenumber) in relation to the Fig. 2a, 2b, 2c, 2d specified.
[0087] Dabei wird at least the spectral pretreatment 5 with identically defined characteristics and at least one of the spectral pretreatments 6, 7, 8 with differently defined characteristics were added for evaluation.
[0088] In Fig. 4b Figure 1 shows a flowchart for the multiple classification method according to the invention, with a training set and a test set in spatial separation but algorithmically linked, and a classification result design based on a large number of classifiers. The respective spectral pretreatment is structured such that certain features become more pronounced and other features are more strongly suppressed. The pretreated spectra 4 are then processed according to Figure 2. Fig. 4b now designed as a variable training set 19, in which several series 13, 14, 15, 16 of classifiers are calculated, e.g., in detail for a series 131, 132, 133, 134, 135, etc. Typically, all classifiers of series 13, 14, 15, 16 are calculated and validated iteratively. Variability means that, without further preconditions, any spectra can be selected for each classifier to be calculated. In this way, for example, several classifiers 131, 132, 133, 134, 135 can be determined for the first series 13, etc. The same applies to the other series 14, 15, 16. The selected spectra of test set 29, e.g., 30%, are calculated according to... Fig. 4b The following section classifies all the samples into a classified test set 24. The classification of spectra 4, 25, 26, 27, and 28 into a specific class of characteristics (male, female) is preferably done as a score or as a probability expression for class membership. To derive an independent statement from the classifiers of series 13, 14, 15, 16, 131, 132, 133, 134, 135, etc., the ratios within the classifiers 13, 14, 15, 16, 131, 132, 133, 134, and 135 are determined. A simple method for this is to calculate the median 30 or, as previously mentioned, to perform a cluster analysis.
[0089] In the Fig. 4b In the flowchart shown, at the point marked with reference symbol 45 = ①, each classified spectrum of each form of pretreatment is compared with the characteristic. The result is simply "correct" or "incorrect".
[0090] Example: Training set 19 comprises 100 spectra. Of these, 60 are selected for calculating the classifiers. If four data preprocessing methods 5, 6, 7, and 8 are used, 60 x 4 = 240 classified spectra are available. The comparison with the feature list thus yields 240 statements that are either "true" or "false." This result is achieved, for example, in the defined 129th iteration stage.
[0091] At point 46 = ②, the classified spectra are evaluated against one or more defined criteria. These criteria could include, for example, a limit on accuracy or a maximum number of iterative steps. The criteria can be logically ANDed or logically ORed.
[0092] Example: Of the 240 possible statements, 205 are "true" and 35 are "false". This results in an accuracy rate of 85% for the training set.
[0093] The following criteria are defined before the classification process begins: 1. Accuracy > 80% and 2. a maximum number of iterations: 1000. i.e.: after the 129th iteration stage is completed, < 1000 and with an achieved accuracy of 85% > specified accuracy.
[0094] With logical AND, "schlecht" (whereby the classifiers are stored as the best intermediate result) and with logical OR, "gut" can be achieved.
[0095] Wenn an der When the number of predefined classifiers to be determined is reached at the point indicated by reference 47 = ⑤, all classifiers (specifically those that led to the best result at point ② = 45) are passed to the validation of the entire training set 19. For example: 30 classifiers are specified to be calculated for each data preprocessing stage 5, 6, 7, 8, and a multiple classification is performed. This results in 30 x 4 = 120 classifiers being passed to the validation.
[0096] At the point / reference point indicated by reference 48 = ③, a joint evaluation of the classification of all spectra of training set 19 is carried out using the leave-one-out or cross-validation method.
[0097] If the test is "passed", the classifiers for classifying the "unknown" spectra of test set 29 are passed.
[0098] If the test is failed, classification according to the specified criteria is not possible.
[0099] At the point / reference junction indicated by reference 49 = ④, a final evaluation of the classification of the spectra of test set 24 is carried out based on their known characteristics. Example:
[0100] Test set 29 comprises 50 spectra. Each of these was classified using 120 classifiers, meaning that each spectrum is assigned 120 probabilities for class membership. The class membership is then determined according to the median or cluster analysis. This is the result of the multiple classification for each individual spectrum. For example, if 41 of the 50 spectra are correctly classified, this results in an accuracy of 82% for the entire test set 24.
[0101] The resulting method 1 of multiple classification is then evaluated based on a comparison with the feature list. The method is thus complete and can now be used for spectra without knowledge of the features.
[0102] In Fig. 5 is a schematic representation of a probability / number of classifiers column chart 38 for 20 classifiers according to Fig. 1 For the display of an egg identified as female, 2 is shown. The unhatched end regions / faces 33 of columns 34 can be located above a specific line – the dividing line 42 – and the columns 35 with the hatched end regions / faces 36 can be located below the dividing line 42 and thus belong to the female sex. The dividing line value is in Fig. 5 At 0.5, the median 30 has a value of 0.72. Therefore, egg 2 is clearly classified as "female".
[0103] In Fig. 6 The top view corresponding to the perspective column representation is shown, which is displayed on a color display as classification result image 37 with the majority of the unhatched end areas / faces 33 indicating the object information 31 "female". On the color display, the unhatched facets 33 can be red and the hatched facets 36 blue, so that a color-coded visual representation of the gender assessment is also possible.
[0104] The hatched squares can be shown in blue and the unhatched squares in red. The few blue squares indicate male object information 32. The predominantly red squares indicate female object information 31. Since the red squares predominate, the sex of the incubated chicken egg 2 can be identified as female trait 31.
[0105] In Fig. 7 is a representation of the calculated median 30 with 10 classifiers in relation to the number of twenty classifiers for an egg 2 according to Fig. 1 with a series of five classifiers each 131, 132, 133, 134, 135 per group for four groups 9, 10, 11, 12. In the bar and median representation, 17 classifiers result in the indication of a female egg 2. The overall classification result 18 can be expressed with the calculated median 30.
[0106] In Fig. 8 Figure 39 is a schematic representation of another exemplary probability / classifier count bar chart for 120 classifiers, shown as bars, for the display of an egg 2 identified as female. The cutoff line 42 again shows the boundary between the characteristic "male" and the characteristic "female". Here too, the bars 34 (31) ending above the cutoff line 42 are shown unhatched on their end faces, and the bars 35 (32) ending below the cutoff line 42 are shown hatched.
[0107] In the case of an egg 2 with male sex characteristic 32, a different column representation may be formed, whereby the front surfaces of the formed columns 35 lying above the dividing boundary 42 are hatched in their majority compared to the unhatched front surfaces of the columns 34 (not drawn).
[0108] In Fig. 9 is a representation of the calculated median 30 in relation to the total number of 120 classifiers according to the bar chart in Fig. 8 In a probability / classifier count representation, the number of classifiers for an egg 2 with female sex characteristics is given as 31, sorted by increasing points. The median 30 is half of the 120 determined classifiers and has a probability value of 0.95.
[0109] The classification units / groups 9, 10, 11, 12 contained in an evaluation unit for determining the object information in the form of dual sex characteristics 31, 32 - female or male - of fertilized and unincubated and incubated eggs 2 function as follows: The functionality is explained.
[0110] For each class 25, 26, 27, 28, several classifiers from series 13, 14, 15, 16 are calculated after spectral pretreatment 5, 6, 7, 8. The determination of the classifier series 13, 14, 15, 16 is carried out according to an algorithm that, in a kind of tandem procedure, first selects spectral regions RS from the coordinate of the relative wavenumbers and subsequently classifies the intensity values of the selected regions RS using discriminant analysis.
[0111] In a subsequent step, a further selection of spectral classes and the classification of intensity values are performed in comparison to the training data for class affiliation. This cycle is repeated iteratively until a level of accuracy that can no longer be improved is reached, whereby the termination criterion can be specified.
[0112] The risk of overtraining, and thus the achievement of high instabilities, increases with the number of spectral classes 25, 26, 27, 28 used for classification. Therefore, it is desirable to use only a few (3 to a maximum of 20) spectral classes for creating the classifier series 13, 14, 15, 16. However, since the sex-relevant information is distributed across the entire spectral range, albeit differently, creating only one classifier would leave significant spectral information unused. For this reason, it is advisable to calculate several (10 to 30) classifiers in series 13, 14, 15, 16 for each group of data preprocessing 5, 6, 7, 8.
[0113] This has the advantage that, firstly, the accuracy of the classification is improved, based solely on the inclusion of as much spectral information as possible, and secondly, the robustness, i.e., the stability, is increased, since several classifiers of the 13, 14, 15, 16 series support the assignment and individual misassignments are compensated.
[0114] The hardware units assigned to the classifications operate identically for all four groups 9, 10, 11, and 12. Therefore, instead of four parallel-controlled units, only one unit can be used, which serially generates the series 13, 14, 15, and 16 of classifiers in a predefined sequence. When determining the number of calculated classifiers NG in series 13, 14, 15, and 16 for each group 9, 10, 11, and 12, the scope of the spectral data points is... vS and the double half-width of the spectral regions w S as well as the number of selected spectral regions RS to be taken into account: N G = v S 2 w S ⋅ R S
[0115] Equation (I) ensures that each data point v S can be selected with equal probability.
[0116] In the Fig. 10a und Fig. 10b Using the raw spectra (intensity / wavenumber curves) as an example, twenty classifiers per data preprocessing step 25 are specified. This is further detailed in... Fig. 10a A section of the corresponding male intensity / wavenumber curves is shown enlarged. The range of spectral data points is... v S across the entire spectral range between 500 cm⁻¹ and 2750 cm⁻¹ data points. The number of selected spectral regions RS is in the Fig. 10a and Fig. 1b RS = 8 with RS1, RS2, RS3, RS4, RS5, RS6, RS7, and Rss.
[0117] According to equation (I), twenty classifiers NG can be calculated for the raw spectrum 25. With four data preprocessing operations (25, 26, 27, 28), this results in a total of 80 generated classifiers (20 classifiers / group x 4 groups).
[0118] After the enlarged section in Fig. 10c can the data points v S can also be weighted additionally. A weighting diagram 40 is provided, which shows that the highest weighting value is assigned to the middle data point 41.
[0119] This can be done with both the male and female spectrum.
[0120] The evaluation 17 and the classification of the results assigned to the classifiers of series 13, 14, 15, 16 are carried out in an evaluation unit and lead to a classification result 18 (30).
[0121] Finally, a classification result 18 is output in the form of the median 30, which, in the determination of the sex of chicken eggs, represents the dual sex information 31, 32 (male or female) with the highest probability.
[0122] The method according to the invention can generally be carried out in detail in the following steps: Acquisition and recording of the spectra by means of at least one optical device with at least one spectrometer and / or further detectors, generation of digitized signals in the form of data points and storage of the detected spectra in storage units of classification units of an evaluation unit, spectral pretreatment by evaluating the stored spectra individually in the individual storage units and making the associated digitized evaluated signals available for further processing, separation of the pretreated spectra as a training set and as a test set, design of the pretreated spectra as a training set and of a test set separate from the training set, wherein according to the invention at least a calculation of the classifiers of the series of the individual classification methods considered, including iterative methods and a validation in the classification units,Classification of the evaluated spectra of the training set with all classifiers of the series, assignment of the spectra of the training set to a class of object information with an expression of a probability for class membership, calculation of the median or performance of a cluster analysis to represent the probability result of the object information of the training set belonging to a class, classification of the evaluated spectra of the test set with all classifiers of the series, assignment of the spectra of the test set to a class of object information with an expression of a probability for class membership and calculation of the median or performance of a cluster analysis to represent the probability result / classification result of the object information of the test set belonging to a class. be performed.
[0123] Additionally, it should be noted that, in general, the recorded spectra are inherently highly variable. This is due, firstly, to the inherent variability of biological systems and, secondly, to the sensitivity of Raman spectroscopic measurements. External disturbances of a systematic and random nature lead to a high variability of the spectral features and thus obscure the feature-relevant information. Furthermore, the Raman spectroscopy method also produces fluorescence light, which, while also containing molecular information, simultaneously obscures the generally much weaker Raman spectroscopic molecular information about the composition of the object under investigation.
[0124] Based on this preliminary remark and the Fig. 1 Another example of an implementation with more than two features of object information will be explained. Fig. 1 The schematic block diagram shows a method 1 according to the invention for classifying spectra 4 of an object 2 with complex information content, in particular optical molecular spectra 4 for assigning object information / features 3 for a probable determination of, for example, a dual object information 31, 32 or of four object information 51, 52, 53, 54.
[0125] Tissue samples, for example brain tumors, can also be used as objects 2 to be examined, and instead of the dual feature information 31, 32, four different features 3 with 51, 52, 53, 54 can, for example, be selected and determined, for example > Feature 51: healthy tissue, > Feature 52: tumor tissue with tumor grade I and II according to the histological grading scheme of the World Health Organization (WHO), > Feature 53: tumor tissue with tumor grade III and IV according to WHO and > Feature 54: necrotic tissue.
[0126] The recording and registration of backscattered radiation from the tissue sample is carried out using at least one optical device, for example as described in German patent application DE 10 2014 010 150 A1. The recorded backscattered spectra are digitized and stored in an evaluation unit. Data preprocessing is performed using, for example, three different methods (5, 6, 7). The resulting data sets can contain, for example, raw spectra, normalized spectra, and spectra with nonlinear baseline correction. The stored spectra in the individual storage units are evaluated separately, and the corresponding digitized signals are made available for further processing.The pre-treated spectra are designed as a training set, wherein, according to the invention, the classifiers of the series of the individual classification methods considered are calculated using iterative methods and validation in the classification units. Furthermore, the evaluated spectra of the test set are classified using all classifiers of the series, and the tissue spectra are assigned to a class of object information according to features 51 to 54, with a probability expression for class membership. The classification is evaluated by calculating the median or by means of cluster analysis, and the probability result / classification result of the object information belonging to a class in the test set is presented.This means that for each registered spectrum of a tissue sample, a point value is calculated using multiple classification, which, according to defined cut-off limits, lies in one of the 4 probability ranges that correspond to the histological findings of the characteristics: 51 - Healthy / 52 - WHO I,II / 53 - WHO III,IV / 54 - Necrosis.
[0127] A device for classifying spectra 4 of objects 2 with complex information content, preferably with objects in the form of chicken eggs 2 for determining dual egg information 31, 32 - female or male - in which the aforementioned method is implemented and which largely corresponds to the block (box) representation in Fig. 1 If appropriately trained, it can include at least the following units at least one detecting optical device with at least one spectrometer and / or further detectors for recording and registering the spectra 4, a unit for generating digitized signals in the form of data points that realize the spectra 4, storage units for storing the spectra 4 in the classification units / groups 9, 10, 11, 12 of an evaluation unit comprising the classification units, units for spectral pretreatment 5, 6, 7, 8 in which the stored spectra 4 in the individual storage units are evaluated individually and the associated digitized evaluated signals are made available for further processing, training sets 19 for the design and application of the pretreated spectra 4;25, 26, 27, 28, at least one classification unit for groups 9, 10, 11, 12 for calculating the classifiers of series 13, 14, 15, 16 of the considered individual conventional classification procedures 25, 26, 27, 28, including iterative procedures and validation in the classification units, test sets 29 for classifying the evaluated spectra 4 with all classifiers of series 13, 14, 15, 16, a unit for classifying the spectra 4 into, for example, a dual class - male 32 or female 31 - of object (egg) information with an expression of a probability for class membership, an evaluation unit for calculating the classification result 18, e.g., in the form of the median 30, or after performing a cluster analysis to represent the probability result of, for example, one of the dual classes - female or male - associated object (egg) information 31, 32. ;
[0128] An identical device can be set up for the multiple classification procedure with the four features 51, 52, 53, 54 or with other predefined features. Reference symbol list
[0129] 1 Procedure in a box representation 2 Object / Egg 3 Object information / Characteristics 4 Recorded spectra 5, 6, 7, 8 Pretreatment 9, 10, 11, 12 Classification 13, 14, 15, 16 Series of classifiers 131, 132, 133, 134, 135, 136 Classifier 17 Evaluation 18 Classification result 19 Training set 20 Molecular composition 21 Fluorescence intensity 22 Fluorescence profile 23 Variation of physical parameters 24 Classified test set 25, 26, 27, 28 Pretreated spectra 29 Test set with preferably selected 30% of the spectra 30 Median 31 Object information / Characteristic female 32 Object information / Characteristic male 33 Unhatched frontal area, assigned to the female sex 34 Column of the female sex 35 Column of the male sex 36 Hatched frontal area, assigned to the male sex 37 Classification result image 3839 Column plot 40 Weighting diagram 41 Mean data point in the region of a spectrum curve 42 Defined cutoff point 43 First histogram of the cluster analysis 44 Second histogram of the cluster analysis 45 = 1 Comparison point 46 = 2 Comparison point 47 = 2 Comparison point 48 = 3 Comparison point 49 = 5 Comparison point 50 State-of-the-art classifier 51, 52, 53, 54 Feature,
Claims
1. Method for classifying spectra of objects having complex information content with at least two different pieces of object information, using a method for recording and pretreating spectral data and a classification method associated with the data pretreatment, which classification method includes the calculation of a classifier, wherein a multiple classification method with at least two different methods for the data pretreatment of the spectra and for the classification method assigned to the corresponding data pretreatment is carried out, and the following steps are carried out: - capturing and recording the spectra by means of at least one optical device comprising at least one spectrometer and / or further detectors, - carrying out a spectra pretreatment, in which the recorded spectra are individually pretreated and the associated digitized evaluated signals are provided for further processing, - separating the pretreated spectra as a training set and as a test set - calculating a plurality of classifiers per type of data pretreatment through iterative calculation and validation on the basis of the training set by spectral regions RS being selected from a coordinate of relative wave numbers and the intensity values of the selected regions RS being subsequently classified by means of discriminant analysis, wherein a renewed selection of spectral regions RS and the classification of the intensity values take place in a repeated step, and the cycle is repeated iteratively until an accuracy that can no longer be improved is achieved, subject to an abortion criterion, - classifying the pretreated spectra of the test set with all classifiers that are calculated on the basis of the training data set, - sorting the spectra into a class of object information with an expression of a probability of belonging to the classification, - calculating a classification result by calculating a median or by carrying out a cluster analysis in order to represent the probability result of the piece of object information belonging to a class.
2. Method according to claim 1, in which the scope of the spectral data points (vS) and twice the half-width of the spectral regions wS as well as the quantity of selected spectral regions (RS) are taken into account when determining the quantity of calculated classifiers NG in the series per classification group: N G = v S 2 w S ⋅ R S wherein equation (I) ensures that each data point can be selected with equal probability.
3. Method according to claim 2, in which the data points belonging to the scope of the spectral data points vs are weighted.
4. Method according to claim 1, in which raw spectra, baseline corrections, normalizations, derivations, covariance and / or principal-component analysis are used for at least one of the spectra pretreatments, such that in each case defined features are emphasized and other defined features are suppressed.
5. Method according to claim 4, in which at least one spectra pretreatment having equally defined and equally weighted features is added to at least one of the spectra pretreatments having differently defined features of the evaluation.
6. Method according to claim 1, in which the pretreated spectra are designed as a training set and a plurality of classifiers of the series or classifiers are defined and validated.
7. Method according to claim 1, in which a method relating to neural networks and / or a method relating to linear time-frequency transformations (wavelets) is used as a classification method in the classification groups.
8. Method according to claim 1, in which pretreated spectra of optical molecular spectroscopy, preferably absorption, emission, scattering or UV / vis, NIR, IR absorption, fluorescence, Raman, are classified.
9. Method according to claims 1 to 8, in which optionally bird eggs, preferably chicken eggs, are used as objects and the dual information items about the female egg sex and the male egg sex are used as pieces of object information.
10. Method according to claims 1 to 9, in which optionally tissue samples, for example brain tumors, are used as objects and four features are used, selected and determined as pieces of object information, namely: - healthy tissue, - tumor tissue having tumor grade I and II according to the histological grading schema of the World Health Organization, - tumor tissue having tumor grade III and IV according to WHO and - necrotic tissue.
11. Method according to claim 1, in which, after at least two classification methods have been performed with at least one determined classifier, each classification method having a preceding spectrum-different data pretreatment, overall at least two determined classifiers for evaluating and subsequently determining a probability result with respect to the predefined different pieces of object information are obtained and used, wherein the probability result is output, so that it is possible to make a conclusion at least about the piece of object information determined to have the highest value.
12. Device for classifying spectra of objects having complex information content, in which the aforementioned method according to any of claims 1 to 11 is implemented, comprising at least the following units - at least one detecting optical device comprising at least one spectrometer and / or further detectors for capturing and recording the spectra, - a unit for generating digitized signals in the form of data points that realize the spectra, - storage units for storing the recorded spectra in the classification units / groups of an evaluation unit that comprises the classification units, - spectra pretreatment units in which the recorded spectra are individually pretreated and the associated digitized evaluated signals are provided for further processing, - training sets for the interpretation and use of the pretreated spectra, - at least one classification unit for calculating a plurality of classifiers per type of data pretreatment through an iterative calculation and a validation based on the training set in the classification units, in which unit spectral regions Rs are selectable from a coordinate of relative wave numbers and the intensity values of the selected regions RS are subsequently classifiable by means of discriminant analysis, wherein a renewed selection of spectral regions Rs and the classification of the intensity values are implementable in a repeated step, and the cycle is iteratively repeatable until an accuracy that can no longer be improved is achieved, subject to an abortion criterion, - test sets for classifying the pretreated spectra with all classifiers per type of data pretreatment, - a unit for sorting the pretreated spectra into at least one dual class of pieces of object information with an expression of a probability for belonging to the class, - a unit of evaluation for calculating the classification result in the form of the median or carrying out a cluster analysis in order to determine the probability result of at least one of the predefined pieces of object information.