Method and medium for predicting the category of data samples using a deep neural network

By integrating batch normalization and sample normalization in the normalization layer of deep neural networks and dynamically adjusting the weights according to the category distinction index, the problem of low classification accuracy caused by a single normalization processing method in the existing technology is solved, and higher classification accuracy and wider scope of application are achieved.

CN119719991BActive Publication Date: 2025-06-03GENERAL HOSPITAL OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510214140.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-03
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

When the existing technology uses deep neural networks to classify data samples, the normalization layer has a single processing method, resulting in low classification accuracy, and heavy reliance on domain knowledge and insufficient generalization ability.

Method used

By fusing batch normalization and sample normalization in the normalization layer, and dynamically adjusting their respective weight coefficients according to the category distinction index of the statistical feature matrix of the training sample to improve the classification accuracy.

Benefits of technology

It effectively avoids the problem of poor classification effect due to the inapplicable normalization processing method, improves the prediction accuracy of the category to which the data samples belong, and does not rely on feature recognition and network structure changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719991B_ABST
    Figure CN119719991B_ABST
Patent Text Reader

Abstract

This application relates to a method and medium for predicting the category of data samples using a deep neural network. The method includes inputting training samples in a training sample set into the deep neural network; obtaining a statistical feature matrix based on the output feature matrix of the convolutional layer; calculating a category discrimination index characterizing the classification ability of the statistical feature matrix; determining the weight coefficient of batch normalization and the weight coefficient of sample normalization based on the category discrimination index, and weighted summing the feature matrix after batch normalization and the feature matrix after sample normalization in the normalization layer to obtain the output feature matrix of the training samples passing through the normalization layer, and finally using the trained deep neural network to predict the category to which the data samples belong. The method according to this application can avoid the problem of poor classification effect caused by single-mode normalization processing, automatically make the normalization layer have a processing method more matching the data set, thereby improving the accuracy of classifying data samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of deep learning and time series data prediction, and more particularly, to a method and medium for predicting the category of data samples using a deep neural network. Background Art

[0002] Deep neural networks have powerful non-linear feature expression capabilities, so they can be used to predict and classify data samples in complex datasets with non-linear features such as time series data.

[0003] Traditional techniques mostly adopt a two-stage method, that is, first, feature representation is performed on the data samples to be classified, and then in the downstream task, a certain similarity metric and classifier are combined to operate on the extracted features to achieve the prediction of the category to which the data samples belong. It can be seen that the quality of the features will have an important impact on the execution results of the downstream task. However, in the existing two-stage methods, once the feature representation is determined, it cannot be adjusted according to the downstream task, and poor feature quality may even directly lead to the failure of the downstream task. Therefore, such methods rely heavily on domain knowledge and need to manually construct high-quality features with high category discrimination ability through, for example, expert knowledge. In the case of a lack of domain knowledge, the generalization ability of this method is insufficient. In another type of data-driven model, the features are regarded as variables to be solved and embedded in the loss function, and the features and the downstream task can be jointly learned, that is, the features are associated with the downstream task. In this type of method, the feature variables and model parameters are solved by minimizing the loss function. However, for the convenience of calculation, usually only simple transformations are performed on the data, resulting in the lack of non-linear feature expression ability of the model. Therefore, for the data samples in complex datasets, the classification effect of this type of method is often poor.

[0004] It can be seen that there is no method in the prior art that can fully consider the impact of the feature representation of data samples on the classification performance and can accurately predict the categories of data samples in complex datasets. Summary of the Invention

[0005] This application is provided to solve the above problems existing in the prior art.

[0006] When using a deep neural network to classify data samples, if a single batch normalization process is adopted in the normalization layer, or a single sample normalization process is adopted, the problem of low classification accuracy will occur in some cases. This application intends to provide a method and medium for predicting the category of data samples using a deep neural network, which can improve the accuracy rate when the deep neural network predicts the category of data samples by improving the processing method of the normalization layer of the deep neural network.

[0007] According to the first aspect of the present application, a method for predicting the category of data samples using a deep neural network is provided. The deep neural network is used to predict the category to which the data samples belong and includes a convolutional layer and a normalization layer. The method includes, during the process of training the deep neural network based on a training sample set: inputting the training samples in the training sample set into the deep neural network; obtaining a statistical feature matrix of the training samples based on the output feature matrix of the convolutional layer; calculating a category discrimination index of the statistical feature matrix, where the larger the category discrimination index, the stronger the category discrimination ability of the statistical feature matrix for the training samples in the training sample set; determining a first coefficient of the weight of batch normalization and a second coefficient of the weight of sample normalization based on the category discrimination index, such that when the category discrimination index is larger, the first coefficient is larger and the second coefficient is smaller; in the normalization layer, performing batch normalization on the output features of the convolutional layer to obtain a first normalized feature matrix of the training samples, and performing sample normalization on the output feature matrix of the convolutional layer to obtain a second normalized feature matrix of the training samples; weighting the first normalized feature matrix of the training samples and the second normalized feature matrix of the training samples based on the first coefficient and the second coefficient to obtain an output feature matrix of the normalization layer of the training samples; and predicting the category to which the data samples belong using the trained deep neural network.

[0008] According to the second aspect of the present application, a deep neural network for predicting the category of data samples is provided. The deep neural network includes a convolutional layer and a normalization layer. During the process of training the deep neural network based on a training sample set, the method for predicting the category of data samples using a deep neural network as described in various embodiments of the present application is executed, and the category to which the data samples belong is predicted using the trained deep neural network.

[0009] According to the third aspect of the present application, a non-transitory computer-readable storage medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the method for predicting the category of data samples using a deep neural network as described in various embodiments of the present application is executed.

[0010] Using the method and medium for class prediction of data samples by means of a deep neural network according to various embodiments of the present application, during the process of training the deep neural network based on a training sample set, instead of pre-selecting a single normalization processing method in the normalization layer, batch normalization and instance normalization are fused and applied. Specifically, by obtaining a class discrimination index that characterizes the class discrimination ability of the statistical quantity of the feature matrix encoded by the convolutional layer for the training samples, the applicability of batch normalization processing and instance normalization processing is judged, and the respective weight coefficients of batch normalization and instance normalization are dynamically adjusted according to the class discrimination index. In this way, the deep neural network trained can effectively avoid the problem of poor classification effect caused by using an inappropriate normalization processing method in the normalization layer, and can improve the accuracy of predicting the category to which the data sample to be measured belongs without having to pre-identify the features of the data sample set in advance and without changing the basic network structure of the deep neural network.

[0011] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specific embodiments of the present application are specifically exemplified.

[0012] It should be understood that the foregoing general description and subsequent detailed description are exemplary and explanatory only, and are not restrictive of the claimed invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In the drawings which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The drawings generally illustrate, by way of example and not limitation, various embodiments, and are used in conjunction with the specification and the claims to explain the disclosed embodiments. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be an exhaustive or exclusive embodiment of the apparatus or method.

[0014] Figure 1 An exemplary structural block diagram of a deep neural network according to an embodiment of the present application is shown.

[0015] Figure 2 A schematic flowchart of a method for class prediction of data samples by means of a deep neural network according to an embodiment of the present application is shown.

[0016] Figure 3 A schematic diagram of a method for obtaining a statistical quantity feature matrix of training samples based on the output feature matrix of a convolutional layer according to an embodiment of the present application is shown.

[0017] Figure 4Schematic diagram showing the calculation method of the class discrimination index of the statistical feature matrix according to an embodiment of the present application.

[0018] Figure 5 Schematic diagram showing the process of predicting the category to which a data sample to be measured belongs by using a trained deep neural network according to an embodiment of the present application.

[0019] FIG. 6(a) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix of the convolutional layer before batch normalization processing according to an embodiment of the present application.

[0020] FIG. 6(b) shows a visualization schematic diagram of the statistical feature matrix of the first normalized feature matrix after batch normalization processing according to an embodiment of the present application.

[0021] FIG. 6(c) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix of the convolutional layer before sample normalization processing according to an embodiment of the present application.

[0022] FIG. 6(d) shows a visualization schematic diagram of the statistical feature matrix of the second normalized feature matrix after sample normalization processing according to an embodiment of the present application.

[0023] FIG. 7(a) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix of the convolutional layer before batch normalization processing according to another embodiment of the present application.

[0024] FIG. 7(b) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix of the convolutional layer before sample normalization processing according to another embodiment of the present application.

[0025] Figure 8 Performance analysis of the method for predicting the category of data samples by using a deep neural network according to an embodiment of the present application. Detailed implementation manners

[0026] To enable those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation manners. The embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, but it is not a limitation to the present application.

[0027] The terms "first", "second" and similar terms used in this application do not denote any order, quantity or importance, but are merely used for distinction. Terms such as "comprising" or "including" mean that the elements before the term cover the elements listed after the term, and do not exclude the possibility of also covering other elements. The execution order of each step in the method described in this application in combination with the drawings is not limited. As long as the logical relationship between each step is not affected, several steps can be integrated into a single step, a single step can be decomposed into multiple steps, or the execution order of each step can be adjusted according to specific requirements.

[0028] It should also be understood that the term "and / or" in this application is merely a relational description of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application generally represents an "or" relationship between the associated objects before and after.

[0029] Figure 1 An exemplary structural block diagram of a deep neural network according to an embodiment of the present application is shown. The deep neural network according to an embodiment of the present application at least includes a convolutional layer and a normalization layer. Figure 1 Taking a fully convolutional neural network (FCN) as an example, a typical network structure is shown. As Figure 1 shown, the fully convolutional neural network 10 includes three serially connected convolutional modules, namely convolutional module 11, convolutional module 12 and convolutional module 13. Among them, each convolutional module internally includes a serially connected convolutional layer and a normalization layer. In some embodiments, only as an example, an activation layer (i.e., a ReLU activation function layer) and the like can be serially connected after the normalization layer in each convolutional module as needed. Other network components included in the fully convolutional neural network 10 are not specifically limited in this application and can be adjusted as needed according to their functional design.

[0030] Figure 2 A schematic flowchart of a method for class prediction of data samples using a deep neural network according to an embodiment of the present application is shown, where the deep neural network can be a fully convolutional neural network similar to the example in Figure 1 , or an MLP (Multi-Layer Perceptron) and a ResNet (Residual Network), etc. This application does not make specific limitations on this, as long as it can be used for predicting the category to which the data sample belongs and includes a convolutional layer and a normalization layer. Different types of deep neural networks may have certain differences in network performance, but it does not affect the applicability of this application therein.

[0031] As Figure 2 shown, the method according to an embodiment of the present application includes steps 201 - 207, wherein steps 201 - 206 are executed during the process of training the deep neural network based on a training sample set.

[0032] First, in step 201, the training samples in the training sample set can be input into the deep neural network, where the training samples are usually a batch of data containing multiple samples.

[0033] The convolutional layer in the deep neural network will perform feature encoding on the input training samples, thereby generating a feature matrix corresponding to the training samples. Thus, in the next step 202, a statistical feature matrix of the training samples can be obtained based on the output feature matrix of the convolutional layer. In some embodiments, the statistics may include, for example, mean and variance, or other applicable variables that can represent the statistical features of the feature matrix of the training samples obtained by performing corresponding linear transformations based on the mean and variance, etc. The present application does not limit this, but for the sake of description, in the following embodiments, the mean and variance are used as typical statistics for illustration. In some embodiments, the calculation method of the statistical feature matrix of the training samples can be correspondingly selected according to the type of the deep neural network and / or the specific meaning of each dimension of the output feature matrix of the convolutional layer, as long as each element in the calculated statistical feature matrix can correspond to the statistics of each training sample relative to its category.

[0034] Then, in step 203, further in combination with the true label of the category to which the training sample belongs, a category discrimination index of the statistical feature matrix is determined, where the larger the category discrimination index, the stronger the classification ability of the statistical feature matrix for the training samples in the training sample set. It can be understood that the training sample set for training is with true labels. Therefore, in some embodiments, the category discrimination index of the statistical feature matrix as described above can be calculated by combining the true labels of the categories to which each training sample belongs, so as to measure the ability of the statistics of the feature matrix output by the convolutional layer in the deep neural network to distinguish the categories of the training samples after feature encoding. The category discrimination index defined in the present application is positively correlated with the classification ability, and the specific numerical relationship can be set as required. The present application does not limit this.

[0035] On the basis of determining the class discrimination index of the statistic feature matrix, in step 204, the first coefficient of the weight of batch normalization (BN) and the second coefficient of the weight of instance normalization (IN) can be further determined based on the class discrimination index, so that when the class discrimination index is larger, the first coefficient is larger and the second coefficient is smaller. That is to say, the first coefficient used to weight batch normalization has a positive correlation with the class discrimination index, while the second coefficient used to weight instance normalization has an inverse correlation with the class discrimination index. Only as an example, for instance, it can be set that the sum of the first coefficient and the second coefficient is 1. In other embodiments, other mapping relationships can also be adopted as long as the above-mentioned correlation rules are not violated.

[0036] In step 205, in the normalization layer, the output feature matrix of the convolutional layer is subjected to batch normalization to obtain the first normalized feature matrix of the training sample, and the output feature matrix of the convolutional layer is subjected to instance normalization to obtain the second normalized feature matrix of the training sample. That is to say, in the normalization layer according to the embodiment of the present application, two kinds of normalization processing need to be performed. In addition, the calculation process of the first coefficient and the second coefficient executed in the above steps 202-204 and step 205 can be executed in any order or simultaneously, and the present application does not make any limitation in this regard.

[0037] In step 206, based on the first coefficient and the second coefficient determined by using steps 202-204, the first normalized feature matrix of the training sample and the second normalized feature matrix of the training sample obtained in step 205 are weighted to obtain the output feature matrix of the normalization layer of the training sample.

[0038] The following is further illustrated with reference to FIGS. 6(a)-6(d). FIG. 6(a) shows a visualization diagram of the statistic feature matrix of the output feature matrix of the convolutional layer before batch normalization according to the embodiment of the present application; FIG. 6(b) shows a visualization diagram of the statistic feature matrix of the first normalized feature matrix after batch normalization according to the embodiment of the present application; FIG. 6(c) shows a visualization diagram of the statistic feature matrix of the output feature matrix of the convolutional layer before instance normalization according to the embodiment of the present application; FIG. 6(d) shows a visualization diagram of the statistic feature matrix of the second normalized feature matrix after instance normalization according to the embodiment of the present application. Among them, the training sample sets adopted in FIGS. 6(a)-6(d) are all ECG200, and the type of the deep neural network is similar Figure 1The FCN in [description] and a single batch normalization process or sample normalization process is adopted in the normalization layer. Among them, ECG200 contains 200 electrocardiogram data of 2 categories, the length of each sample is 96, and there are 100 samples in each of the training and test sets. Figures 6(a)-6(d) only show the visualization diagrams of the statistical feature matrices of each convolution module in the FCN before and after the normalization process. However, it should be noted that the normalization layers in other convolution modules also perform the same batch normalization process or sample normalization process as the third convolution module, and the corresponding statistical feature matrices are not shown. In some embodiments, a method similar to that for calculating the statistical feature matrix of the output feature matrix of the convolution layer can be used to calculate the statistical feature matrix of the first normalized feature matrix after the batch normalization process and the statistical feature matrix of the second normalized feature matrix after the sample normalization process. Figures 6(a)-6(d) respectively show the schematic diagrams of the results of visualizing the statistical feature matrices using the t-SNE method when all data samples of the ECG200 training sample set are input into the FCN deep neural network and the statistics are the mean and variance. Among them, t-SNE is a commonly used non-linear dimensionality reduction algorithm in the field of machine learning, which can usually be used to reduce high-dimensional data to 2D or 3D and visualize it (in each embodiment of this application, the high-dimensional data is reduced to 2D and then visualized, where the abscissa is dimension 1 and the ordinate is dimension 2), and make similar data (that is, data of the same category) close when displayed, and dissimilar data (that is, data of different categories) far when displayed. The specific algorithm of t-SNE is not elaborated in this application.

[0039] As can be seen from Figures 6(a) and 6(c), before batch normalization or sample normalization, samples of the same category are closer and samples of different categories are farther apart. That is, the statistical feature matrices before batch normalization and sample normalization both have good category discrimination. The experimental data of the embodiments of this application also show that, in this case, the category discrimination index of the statistical feature matrix is also higher. As can be seen from Figure 6(b), after the batch normalization process is performed, the statistical feature matrix of the first normalized feature matrix can still distinguish samples of different categories, while as can be seen from Figure 6(d), after the sample normalization process is performed, the statistical feature matrix of the second normalized feature matrix can no longer distinguish samples of different categories. From this, it can be inferred that for the ECG200 data set, performing batch normalization in the normalization layer is more applicable than performing sample normalization, and it will also be more beneficial to the execution of subsequent steps and the classification ability of the deep neural network after training. That is to say, the higher the category discrimination index of the statistical feature matrix of the output feature matrix of the convolution layer, the higher the weight of performing batch normalization in the normalization layer should be, and correspondingly, the weight of performing sample normalization will be reduced accordingly.

[0040] Figures 7(a) and 7(b) show situations different from Figures 6(a) and 6(c). Other conditions in Figures 7(a) and 7(b) are similar to those in Figures 6(a)-6(d), but the training sample set used is the MedicalImages training sample set. Figure 7(a) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix before batch normalization processing in the third convolutional module, and Figure 7(b) shows a visualization schematic diagram of the statistical feature matrix of the output feature matrix before sample normalization processing in the third convolutional module. The MedicalImages training sample set contains 10 categories from class0-class9, with 1,141 samples. The length of each sample is 99, and there are 381 and 760 samples in the training and test sets respectively.

[0041] As shown in the legends of Figures 7(a) and 7(b). It can be seen from Figures 7(a) and 7(b) that the class discrimination ability of the statistical feature matrix of the output feature matrix of the convolutional layer is not high, and the corresponding class discrimination index is also low. In this case, it can be predicted that no matter whether batch normalization or sample normalization is adopted in the normalization layer, its class discrimination ability will not be very high. Therefore, it is not practical to retain the original statistics of batch normalization. Therefore, the weight of sample normalization processing should be adjusted upward accordingly, while the weight of batch normalization processing should be adjusted downward accordingly.

[0042] The above steps 201-step 206 are the process of training a deep neural network using training samples. By reasonably integrating and applying the two normalization methods in the normalization layer, when the training is completed and the deep neural network converges, the application stage can be entered. That is, in step 207, the trained deep neural network can be used to predict the category of the data sample to be measured. Among them, the data sample to be measured can be data belonging to the same sample set as the training sample, or a data sample with the same characteristics as the training sample. Only as an example, when the training sample is the electrocardiogram data of a specific patient within a specific time period, the data sample to be measured should be the electrocardiogram data to be classified of the same patient within the same or similar time period. In this case, the data sample to be measured can be considered to have the same characteristics as the training sample.

[0043] Different from the single normalization method in the prior art that either adopts batch normalization or instance normalization in the normalization layer, the method for predicting the category of data samples using a deep neural network according to various embodiments of the present application dynamically fuses the above two normalization methods (Dynamic Batch Instance Normalization, DBIN, that is, dynamic batch instance normalization) in the normalization layer of the deep neural network, which can effectively avoid the problem of poor classification effect caused by using an inappropriate normalization method in the normalization layer. Specifically, during the process of training the deep neural network based on each batch of data samples in the training sample set, combined with the true annotation of the category to which the training samples belong, a category discrimination index is obtained to characterize the category discrimination ability of the feature matrix after the training samples are encoded by the convolutional layer, so as to judge the applicability of batch normalization and instance normalization to the current batch of data samples, and accordingly dynamically adjust the respective weight coefficients of batch normalization and instance normalization. As the training process progresses, the matching degree between the processing method of the normalization layer and the data samples in the training sample set will gradually increase. Therefore, when using the trained deep neural network to predict the category of the data sample to be tested, it will also have higher accuracy. The method according to the embodiments of the present application does not depend on the specific characteristics of the data samples to be classified, so it has a very wide range of applications. In addition, the method according to the embodiments of the present application does not change the basic network structure of the deep neural network, and the network complexity and the time complexity of training and prediction do not increase significantly.

[0044] Figure 3 Illustrates a method for obtaining the statistical feature matrix of training samples based on the output feature matrix of the convolutional layer according to an embodiment of the present application when the statistics adopt the mean and variance.

[0045] In Figure 3 taking FCN as an example, assuming that the output feature matrix of a batch of training samples after passing through the one-dimensional convolutional layer of FCN is , where N is the number of samples included in a batch of training samples, Figure 3 in which N = 3, that is, the number of samples is 3; H is the sample length of each sample; C is the number of channels, Figure 3 in which C = 5, that is, the number of channels is 5.

[0046] In Figure 3In the example, the deep neural network 300 includes a convolutional layer 301 and a normalization layer 302. After inputting the training sample t into the convolutional layer 301 of the deep neural network 300, the output feature matrix z of the convolutional layer can be obtained. Then, the following processing is performed in the normalization layer 302: Based on the output feature matrix z of the convolutional layer, for each channel, the mean and variance of each channel of the training sample are calculated respectively to obtain the mean matrix and the variance matrix , and then, according to and , the statistical feature matrix is further obtained. For example, and can be concatenated to obtain the statistical feature matrix z'. The specific calculation method is shown in formula (1):

[0047] Formula (1)

[0048] where z nhc is an element in the output feature matrix z, and , the mean matrix , the variance matrix , .

[0049] Figure 4 shows a schematic diagram of the calculation method of the class discrimination index of the statistical feature matrix according to an embodiment of the present application. As Figure 4 shown, the normalization layer 302 may further include a first fully connected layer 3021, a ReLU activation function layer 3022, a second fully connected layer 3023, and a Softmax layer 3024 connected in sequence, and the number of neurons in the second fully connected layer 3023 is set to be equal to the number of classes in the training sample set.

[0050] Specifically, the statistical feature matrix z' is input into the first fully connected layer 3021, and the output of the Softmax layer 3024 is used as the probability matrix Pt of the probability that the training sample t belongs to each class in the training sample set. Then, based on the probability matrix Pt and the true annotation matrix y of the training sample, the class discrimination index r of the statistical feature matrix z' is calculated.

[0051] Only as an example, for example, the reciprocal of the cross entropy of the probability matrix Pt and the true annotation matrix y of the training sample can be used as the class discrimination index r of the statistical feature matrix z', as shown in formula (2):

[0052] Formula (2)

[0053] Furthermore, the first coefficient of the batch normalization weight can be determined based on the class discrimination index r according to the following formula (3): and the second coefficient of the sample normalization weight :

[0054] Formula (3)

[0055] Among them, r is the category discrimination index, is the first coefficient, is the second coefficient.

[0056] From formula (1) and formula (2), we can see that when the value of r is large, it means that the cross entropy is small, the statistical feature matrix has category discrimination, and the weight of batch normalization is Should be high; on the contrary, the cross entropy is large, the statistical feature matrix has no category distinction ability, then the weight of sample normalization In particular, as the denominator of r approaches 0 and the value of r approaches infinity, the weights of batch normalization Tends to 1, the weight of sample normalization On the contrary, when the denominator of r is large and the value of r is close to 0, the weight of batch normalization tends to 0, while the weight of sample normalization Then it tends to 1. In this way, it is not necessary to identify the features of the training sample set in advance. Instead, the weights of the two normalization processes are gradually trained to appropriate values ​​as the training process proceeds, so that the two normalization methods can be adaptively integrated to improve the classification accuracy based on whether the feature statistics have the ability to distinguish categories.

[0057] After the deep neural network is trained, it can be used to classify the data samples to be tested. Figure 5 A schematic diagram of a process of predicting the category to which a data sample to be tested belongs using a trained deep neural network according to an embodiment of the present application is shown. Figure 5 Still in use Figure 3 and Figure 4 The deep neural network in 300, but the difference is that Figure 5 The deep neural network 300 in is trained using each training sample t in the training data set and is used to predict the category to which the test data sample x belongs. Figure 5 As shown, the deep neural network 300 may include a pooling layer 303 and other applicable network layers (not shown) in addition to the convolution layer 301 and the normalization layer 302. The process of predicting the category of the test data sample x is as follows:

[0058] First, input the data sample x to be measured into the trained deep neural network 300. For example, take the data sample x to be measured as the input of the convolutional layer 301. After feature encoding, the feature matrix z is output from the convolutional layer 301. Next, send the output feature matrix z of the convolutional layer into the normalization layer 302 and process it in two paths: one path performs batch normalization on the output feature z of the convolutional layer to obtain the first normalized feature matrix of the data sample x to be measured ; the other path performs sample normalization on the output feature matrix z of the convolutional layer to obtain the second normalized feature matrix of the data sample x to be measured . Then, use the first coefficient of the batch normalization weights and the second coefficient of the sample normalization weights at the end of the training phase to perform weighted processing on the first normalized feature matrix and the second normalized feature matrix to obtain the output feature matrix of the normalization layer of the data sample x to be measured . Next, the output feature matrix of the normalization layer of the data sample x to be measured can be input into the pooling layer 303 to obtain the probability matrix of the category to which the data sample x to be measured belongs. Among them, each element in represents the probability value of each data sample belonging to each category. In some embodiments, the deep neural network 300 may include multiple convolutional modules, and each convolutional module includes a convolutional layer and a normalization layer. It should be noted that in each convolutional module, the calculation processes of the convolutional layer output feature matrix, the first normalized feature matrix, the second normalized feature matrix, and the output feature matrix of the normalization layer described above will be executed, but only the last convolutional module is connected to the pooling layer 303, and the output feature matrix of the normalization layer of the data sample x output by the last convolutional module is sent into the pooling layer 303 for subsequent calculations. It should be noted that each convolutional module needs to perform weighted processing on the first normalized feature matrix and the second normalized feature matrix actually calculated by this module according to the first coefficient of the batch normalization weights and the second coefficient of the sample normalization weights corresponding to this module after training to obtain the output feature matrix of the normalization layer of this module and input it into the next convolutional module. In some embodiments, the pooling layer 303 can use average pooling or any applicable structure, and the specific implementation method is not limited in this application. In other embodiments, based on , the category to which each data sample belongs can be determined. For example, the category corresponding to the maximum probability value can be determined as the classification result of the data sample, and this application does not make specific restrictions on this

[0059] In some embodiments, the deep neural network according to the embodiments of the present application can classify various types of data samples. Particularly for complex time series data with non-linear characteristics such as stock trading data, weather data, electrocardiogram (ECG) data, traffic flow data, satellite telemetry data, etc., the method according to the embodiments of the present application has a significant improvement in classification performance compared to the prior art. Only as an example, the data samples may include, for example, ECG time series data samples of a patient. An electrocardiogram is a typical time series data that reflects the health status of the heart. For a normal heart, its heart state corresponds to several fixed regular waveforms, while for an abnormal heart, in addition to the fixed regular waveforms, its heart state may also include various types of changing waveforms, and its changing pattern is unpredictable. For different patients or even different time periods of the same patient, its waveforms may exhibit completely different characteristics. During the experiment of the present application, experiments were carried out based on multiple ECG data sets respectively, and it was found that for some ECG data sets, the classification effect of using batch normalization alone is better than that of using sample normalization alone, while for some other ECG data sets, the opposite is true, that is, the classification effect of sample normalization is better than that of batch normalization. That is to say, even for the same ECG data, it is impossible to predict which normalization method will ensure a better classification effect. In the case of using the dynamic batch sample normalization method (DBIN) according to the embodiments of the present application, there is no need to pay attention to the specific characteristics and distribution of the samples in the sample set. As the training process progresses, the deep neural network will adaptively adjust the matching degree between its model and the training data set to the best. Therefore, it can be understood that when the data characteristics in the data set can be predicted and thus batch normalization or sample normalization can be clearly adopted, the DBIN method of the present application will have comparable classification performance with it. However, for the case where the data characteristics are complex and difficult to predict, and thus it is impossible to determine whether to use batch normalization or sample normalization, the DBIN method of the present application will be able to ensure acceptable and better classification performance, thereby avoiding worse classification performance caused by using an inappropriate normalization method. Therefore, the deep neural network according to the embodiments of the present application can perform more efficient and accurate automatic recognition and class prediction on ECG time series data, thereby providing more reliable auxiliary diagnosis for doctors.

[0060] Figure 8Performance analysis of a method for class prediction of data samples using a deep neural network according to an embodiment of the present application is shown. In the field of time series data mining, the publicly available dataset commonly used for validating classification methods is the UCR Archive (https: / / www.cs.ucr.edu / ~eamonn / time_series_data / ), which contains 85 time series datasets in different fields, including electrocardiograms, motion, sensors, spectral data, etc. The sizes, number of classes, data lengths, etc. of different datasets vary. Figure 8 Taking the deep neural network as FCN as an example, the comparison of the classification accuracies of the method according to an embodiment of the present application (i.e., DBIN), the normalization layer using batch normalization (BN), and instance normalization (IN) on 8 datasets such as ChlorineCon, UWaveX, InsectWingbeatSo, Lighting2, MoteStrain, ToeSeg2, ECG5000, and ECGFiveDays is exemplarily shown. Among them, ECG5000 is an electrocardiogram dataset of patients with severe heart failure, containing 5000 annotated electrocardiogram time series data samples, each sample having a length of 140 and a total of 5 classes, with 500 and 4500 samples in the training set and the test set respectively; ECGFiveDays is another electrocardiogram dataset, containing 884 samples in 2 classes, each sample having a length of 136, and the number of samples in the training set and the test set being 23 and 861 respectively.

[0061] From Figure 8 It can be seen that the experimental results show that compared with the normalization methods using a single BN or a single IN, the DBIN method according to an embodiment of the present application has improved classification accuracy in most cases, and is only comparable or equivalent to the IN normalization method on the Lighting2 dataset and the ECG5000 dataset. Thus, it can be seen that the method according to an embodiment of the present application has a wide range of applications, does not depend on the specific characteristics of the data samples, has good performance on various different datasets, and has higher accuracy when predicting the class of the data samples to be measured.

[0062] An embodiment of the present application further provides a device for class prediction of data samples. The device includes an interface and a processor. Among them, the interface is configured to obtain a training sample set and data samples to be measured; the processor is configured to execute the method for class prediction of data samples using a deep neural network according to various embodiments of the present application.

[0063] In some embodiments, the apparatus for predicting the category of data samples may be a special-purpose computer or a general-purpose computer. For example, the apparatus may be a customized computer that performs the function of predicting the electrocardiogram category of a patient or a server arranged in the cloud. The apparatus may also include a memory and a bus, etc., and the interface, the memory, and the processor may be connected to the bus and communicate with each other through the bus. This application does not make specific limitations on this.

[0064] The above interface may include, for example, a network cable connector, a cable connector, a serial connector, a USB connector, a parallel connector, high-speed data transmission adapters such as optical fibers, USB 3.0, Thunderbolt, etc., wireless network adapters such as WiFi adapters, and telecommunication (3G, 4G / LTE, etc.) adapters.

[0065] The above processor may be a processing device including one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be one or more dedicated processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), system-on-chips (SoCs), etc.

[0066] According to an embodiment of the present application, there is also provided a non-transitory computer-readable storage medium, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the steps of the method for predicting the category of data samples using a deep neural network as described in various embodiments of the present application are executed.

[0067] In some embodiments, the above non-transitory computer-readable medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a phase change random access memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), other types of random access memories (RAMs), a flash drive or other forms of flash memory, a cache, a register, a static memory, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD) or other optical memories, a cassette tape or other magnetic storage devices, or any other possible non-transitory medium used to store information or instructions that can be accessed by a computer device.

[0068] The foregoing description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. Other embodiments may be utilized by those of ordinary skill in the art upon reviewing the above description. Also, in the above detailed description, various features may be combined together to simplify the present application. This should not be construed as intending that the disclosed features not claimed are essential to any claim. Thus, the claims are hereby incorporated into the detailed description as examples or embodiments, where each claim stands on its own as a separate embodiment, and it is contemplated that these embodiments may be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the claims and the full scope of equivalents given to those claims.

Claims

1. A method for class prediction of data samples using a deep neural network, characterized in that: The data sample is an electrocardiogram time series of a patient, the deep neural network is used to predict the category to which the data sample belongs, and comprises a convolution layer and a normalization layer, and the method comprises: In the process of training the deep neural network based on the training sample set: Inputting the training samples in the training sample set into the deep neural network; Based on the output feature matrix of the convolutional layer, obtaining a statistical feature matrix of the training samples; Combined with the real label of the category to which the training sample belongs, determining the category discrimination index of the statistical feature matrix, specifically comprising: inputting the statistical feature matrix into the first fully connected layer in the normalization layer, and using the output of the Softmax layer in the normalization layer as the probability matrix of the probability that the training sample belongs to each category in the training sample set; using the inverse of the cross entropy between the probability matrix and the real label matrix of the training sample as the category discrimination index of the statistical feature matrix; wherein, the larger the category discrimination index, the stronger the category discrimination ability of the statistical feature matrix for the training samples in the training sample set; Determining a first coefficient of a batch normalization weight and a second coefficient of a sample normalization weight based on the class discrimination index, so that the larger the class discrimination index, the larger the first coefficient and the smaller the second coefficient; In the normalization layer, batch normalization is performed on the output feature matrix of the convolution layer to obtain a first normalized feature matrix of the training sample, and sample normalization is performed on the output feature matrix of the convolution layer to obtain a second normalized feature matrix of the training sample; Based on the first coefficient and the second coefficient, weighting the first normalized feature matrix of the training sample and the second normalized feature matrix of the training sample to obtain an output feature matrix of the normalized layer of the training sample; and Use the trained deep neural network to predict the category of the test data sample.

2. The method according to claim 1, characterized in that The deep neural network also includes a pooling layer, and the method of using the trained deep neural network to predict the category to which the test data sample belongs further includes: Inputting the data sample to be tested into a trained deep neural network; In the normalization layer, batch normalization is performed on the output feature matrix of the convolution layer to obtain a first normalized feature matrix of the data sample to be tested, and sample normalization is performed on the output feature matrix of the convolution layer to obtain a second normalized feature matrix of the data sample to be tested; Using the first coefficient and the second coefficient at the end of the training phase to perform weighted processing on the first normalized feature matrix and the second normalized feature matrix, so as to obtain an output feature matrix of the normalized layer of the data sample to be tested; The output feature matrix of the normalization layer of the data sample to be tested is input into the pooling layer to obtain a predicted probability matrix of the category to which the data sample to be tested belongs.

3. The method according to claim 1 or 2, characterized in that: The step of obtaining the statistical feature matrix of the training sample based on the output feature matrix of the convolutional layer specifically includes: Based on the output feature matrix of the convolutional layer, the mean and variance of each channel of the training sample are calculated respectively to obtain the mean matrix and variance matrix of the training sample, and the mean matrix and variance matrix are concatenated to obtain a statistical feature matrix.

4. The method according to claim 1 or 2, characterized in that: The normalization layer also includes a ReLU activation function layer and a second fully connected layer, the first fully connected layer, the ReLU activation function layer, the second fully connected layer and the Softmax layer are connected sequentially, and the number of neurons in the second fully connected layer is set to be equal to the number of categories in the training sample set.

5. The method according to claim 1 or 2, characterized in that: The determining of the first coefficient of the batch normalization weight and the second coefficient of the sample normalization weight based on the class discrimination index further includes calculating the first coefficient and the second coefficient according to formula (3): Formula (3) Among them, r is the category discrimination index, is the first coefficient, is the second coefficient.

6. The method according to claim 1 or 2, characterized in that: The deep neural network includes a fully convolutional neural network.

7. A device for predicting the category of a data sample, characterized in that: include: An interface configured to obtain a training sample set and a data sample to be tested; A processor configured to execute the method for predicting categories of data samples using a deep neural network as described in any one of claims 1-6.

8. A non-temporary computer-readable storage medium having computer-executable instructions stored thereon, wherein when the computer-executable instructions are executed by a processor, the method for predicting categories of data samples using a deep neural network as described in any one of claims 1-6 is performed.

Citation Information

Patent Citations

  • Unbalanced ship classification method based on deep convolutional neural network

    CN111461190A

  • Domain generalization normalization method based on statistical correlation clustering

    CN116824190A