Data feature discrimination method and system

By combining the attention mechanism and the data feature discrimination method of the bidirectional GRU network, the problem of large data volume and insignificant feature changes is solved, and the generalization ability and discrimination accuracy of the model are improved.

CN120067947APending Publication Date: 2025-05-30INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210345.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the data processing process, we face problems such as large amount of data and insufficient changes in characteristics, which creates difficulties in data analysis and artificial intelligence models.

Method used

The data feature discrimination method combining attention mechanism and bidirectional GRU network is adopted, and the feature information of time series data is extracted through the bidirectional gated loop unit, and the hidden state weighted calculation is used to complete effective feature screening. At the same time, the cross entropy loss function is improved, the model loss is calculated and the model weight is updated.

Benefits of technology

The generalization ability of the model and the accuracy of the discriminant model are improved, so that the trained model can better distinguish between normal and abnormal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067947A_ABST
    Figure CN120067947A_ABST
Patent Text Reader

Abstract

The invention discloses a data feature discrimination method and system, and belongs to the technical field of data processing and artificial intelligence. Data feature discrimination is realized in combination with an attention mechanism and a bidirectional GRU network, and feature information of time sequence data is extracted by using a bidirectional gating loop unit; carrying out weighted calculation on the hidden state by adopting an attention mechanism so as to complete effective feature screening; a cross entropy loss function is improved, model loss is calculated, and model weight updating is carried out, so that the trained model can better distinguish normal and abnormal data. According to the invention, the data features are discriminated based on the attention mechanism and the bidirectional GRU network, and the feature extraction capability is enhanced; an attention mechanism is introduced in a decoding stage, the influence of irrelevant features on a result is weakened, the generalization ability of the model is improved, a cross entropy loss function is improved, the cross entropy loss function is more suitable for input data with non-uniform positive and abnormal sample numbers, and the precision of the discrimination model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and artificial intelligence, and specifically relates to a method and system for discriminating data features. Background Art

[0002] For data, time is a very important dimension and attribute, and the accumulation of historical data is an important reason for the "bigness" of big data. Time series data exists in various fields, such as financial market trend prediction, risk assessment, log recording, and data annotation. Generally speaking, time series data is distributed in all walks of life. However, in the process of data processing, problems such as large data volume and unobvious feature changes are often faced, causing certain difficulties in data analysis, artificial intelligence model establishment, etc. Summary of the Invention

[0003] The technical task of the present invention is to provide a method and system for discriminating data features aiming at the above deficiencies, which can strengthen the feature extraction ability, improve the generalization ability of the model, and improve the accuracy of the discrimination model.

[0004] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0005] A method for discriminating data features combines an attention mechanism and a bidirectional GRU network to implement data feature discrimination. The feature information of time series data is extracted by using a bidirectional gated recurrent unit, and the attention mechanism is used to calculate the weighted hidden state to complete the effective feature screening; the cross-entropy loss function is improved, the model loss is calculated, and the model weights are updated; the trained model can better distinguish normal and abnormal data.

[0006] In recent years, with the rapid development of deep learning and the improvement of machine computing power, the combined application of recurrent neural networks dedicated to processing sequence data and data analysis provides a new idea for data feature extraction. As a special recurrent neural network, the bidirectional gated recurrent unit can consider forward information and backward information according to the chronological order of data. On this basis, an attention mechanism and an improved cross-entropy loss function are introduced to enhance the model's learning ability of features, analyze the importance of different features, and better use for data feature discrimination.

[0007] Further, the implementation manner of this method is as follows:

[0008] Initialize the initial model of the bidirectional gated recurrent unit according to the data length and the number of features of the time series data; slice the training data and input it into the respective forward GRU and backward GRU in sequence. Sum and average the output matrices of the forward GRU and backward GRU of each group of bidirectional GRUs to obtain the initial state vector, and obtain the anomaly label probability through the calculation of the Attention layer and the processing of the softmax layer;

[0009] Calculate the model loss using the improved cross-entropy loss function and update the model weights. For each training of the model, calculate the cross-entropy loss function once, and introduce the ratio of the number of anomaly samples to the number of normal samples in the positive sample loss term in the function to balance the quantity difference between positive and anomaly samples; execute the backpropagation algorithm, calculate the partial derivatives of each part in the discriminant model, and use the gradient descent algorithm to update the weights of each layer of the model;

[0010] Train the network iteratively until the model converges or reaches the preset maximum number of training times to obtain the final network model, and use this network model for data feature discrimination.

[0011] Furthermore, the specific steps for implementing this method include:

[0012] Step S1: Classify using a bidirectional gated recurrent unit with an attention mechanism introduced;

[0013] Step S2: Calculate the model loss using the improved cross-entropy loss function and update the model weights;

[0014] Step S3: Construct a data feature discrimination model.

[0015] Furthermore, the specific implementation process of step S1 is as follows:

[0016] (1) Build the initial model of the bidirectional gated recurrent unit. According to the data D = {d 1 , d 2 , d 3 … d l}, d i = {d i1 , d i2 , d i3 … d ia}, initialize the batch size b of the single training data of the model, the time length r according to the data length l and the number of features a, and set the value of the number of input variables to be equal to the number of features a, and set the value of the output dimension to 1;

[0017] (2) According to the data D = {d 1 , d 2 , d 3 … d l} The data length l and the model's single - training data batch size b are used to calculate the number of times c required for one - time model training, where c = l / b. Taking the time length r as the window size and setting the step size to 1, the input data D is sliced and processed into three - dimensional data D′={d′ 1 ,d′ 2 ,d′ 3 …d′ l-b+1}, where d′ i ={d i ,d i+1 ,d i+2 …d i+b-1};

[0018] (3) For each training, three - dimensional input data equal to the model's single - training data batch size is input into the network. For each group of two - dimensional data contained in this data, they are respectively input into their forward GRU and backward GRU in sequence. The output of each group of bidirectional GRUs is the output matrix S t + of the forward GRU and the output matrix S t - of the backward GRU. The sum is averaged to obtain the initial state vector S t . Through S t + and S t - , the formula for calculating is:

[0019]

[0020] (4) For the state vector S( i ) obtained after each group of BiGRU processes the data, through the calculation of the Attention layer, the target attention weight e( i ), the probability vector α (i) =(α 0 ,α 1 ,α 2 …α k ) are generated in sequence. The weighted sum operation is performed on all the state vectors S (i) and the probability vector α (i) to obtain the corresponding score vector Y (i) =(y 0 (i) ,y 1 (i) ,y 2 (i) …y k (i) ), where k represents k types of abnormal working conditions. Among them, the target attention weight e (i) , the probability vector α(i) and the fractional vector Y (i) is calculated as follows:

[0021] e (i) = ω i ·S (i) + b i

[0022]

[0023] where S (i) is the initial state vector of the i-th eigenvector, ω i represents the weight coefficient matrix of the i-th eigenvector, and b i represents the offset corresponding to the i-th eigenvector;

[0024] (5) Process the fractional vector Y (i) through a layer of softmax to process the score of the corresponding label into the probability that the current training data may be a certain label, and obtain the probability vector The calculation formula is as follows:

[0025]

[0026] where, represents the score of the i-th sample classified as the j-th working condition, represents the probability that the i-th sample is classified as the j-th working condition.

[0027] Furthermore, the above are respectively input into their forward GRU and backward GRU in sequence:

[0028] For the forward GRU, the output of the previous GRU serves as part of the input of the next GRU; for the backward GRU, the output of the next GRU serves as part of the input of the previous GRU.

[0029] Furthermore, the specific implementation process of step S2 is as follows:

[0030] (1) Traverse the input data D, count the number of normal samples and abnormal samples in the data, and calculate the ratio ω of the number of normal samples n + to the number of abnormal samples n - :

[0031]

[0032] where, y (i) represents the label of the i-th sample, and its value range is 0 - k, where 0 represents the normal working condition, represents the k-th abnormal working condition, and I(·) is the indicator function, which takes the value of 1 when · is true and 0 when · is false;

[0033] (2) For each training of the model, calculate the cross-entropy loss function once, and introduce a ratio ω at the positive sample loss term in the function to balance the quantity difference between positive and abnormal samples. The calculation formula of the improved cross-entropy loss function is as follows;

[0034]

[0035] where λ is a control parameter. When λ approaches 0, this formula is the standard cross-entropy loss function. When λ is large enough, the effect of distinguishing different abnormal working conditions becomes weak, and this formula becomes the loss function for solving binary classification problems, aiming to identify abnormal and normal data;

[0036] (3) Execute the backpropagation algorithm, calculate the partial derivatives of the loss function in the model with respect to each weight parameter, solve the optimization direction of each layer's weight parameter in the model along the direction of gradient descent, update the weights of each layer of the model using the gradient descent method, and the calculation formulas for the partial derivatives of each part of the loss function are as follows:

[0037]

[0038] where, is the l-th probability in the probability vector α calculated for the i-th sample in the attention layer t in.

[0039] Furthermore, the specific implementation manner of the step S3 is as follows:

[0040] Perform cyclic training on the network until the model converges or reaches the preset maximum number of training times to obtain the final network model; collect the data that has not participated in training as the test set to test the model, calculate the accuracy, precision, and recall parameters of the model, and use them to measure the specific performance of the model.

[0041] The present invention also claims to protect a data feature discrimination system, including:

[0042] A bidirectional gated recurrent unit for extracting the feature information of time series data and realizing effective feature screening in combination with the attention mechanism to obtain the abnormal label probability;

[0043] A cross-entropy loss function module for improving the cross-entropy loss function, calculating the model loss, and updating the model weights;

[0044] A data feature discrimination model training module for performing cyclic training on the network until the model converges or reaches the preset maximum number of training times to obtain the final network model, and using this network model for data feature discrimination;

[0045] This system realizes data feature discrimination through the above method.

[0046] The present invention also claims to protect a data feature discrimination device, including at least one memory and at least one processor;

[0047] The at least one memory is used for storing a machine-readable program;

[0048] The at least one processor is used for calling the machine-readable program to implement the above method.

[0049] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor implements the above method.

[0050] Compared with the prior art, a data feature discrimination method and system of the present invention have the following beneficial effects:

[0051] Based on the attention mechanism and the bidirectional GRU network, the present invention discriminates data features, constructs a bidirectional gated recurrent unit neural network diagnosis model, considers both the forward and reverse information of the data, and strengthens the feature extraction ability; the attention mechanism is introduced in the decoding stage, weakening the influence of irrelevant features on the result, improving the generalization ability of the model, and improving the cross-entropy loss function to make it more suitable for input data with uneven numbers of positive and abnormal samples, thereby improving the accuracy of the discrimination model. Description of the Drawings

[0052] Figure 1 It is a flowchart of the data feature discrimination method provided by an embodiment of the present invention. Detailed Embodiments

[0053] The present invention will be further described below in conjunction with specific embodiments.

[0054] An embodiment of the present invention provides a data feature discrimination method, in combination with the attached Figure 1As shown, initialize the initial model of the bidirectional gated recurrent unit according to the data length and the number of features of the time series data; slice the training data and input it into the respective forward GRU and backward GRU in sequence. Sum and average the output matrices of the forward GRU and backward GRU of each group of bidirectional GRUs to obtain the initial state vector. Calculate through the Attention layer and process through the softmax layer to obtain the abnormal label probability. Use the improved cross-entropy loss function to calculate the model loss and update the model weights. For each training of the model, calculate the cross-entropy loss function once, and introduce the ratio of the number of abnormal samples to the number of normal samples at the positive sample loss term in the function to balance the quantity difference between positive and abnormal samples; execute the backpropagation algorithm, calculate the partial derivatives of each part in the discriminant model, and use the gradient descent algorithm to update the weights of each layer of the model. Perform cyclic training on the network until the model converges or reaches the preset maximum number of training times to obtain the final network model, and use this network model for data feature discrimination.

[0055] This method combines the attention mechanism and the bidirectional GRU network to realize data feature discrimination. By using the bidirectional gated recurrent unit to extract the feature information of time series data, and using the attention mechanism to weight and calculate the hidden state to complete effective feature screening; improve the cross-entropy loss function so that the trained model can better distinguish normal and abnormal data. The specific steps of implementing this method include:

[0056] Step S1: Use a bidirectional gated recurrent unit with an attention mechanism for classification.

[0057] (1) Build the initial model of the bidirectional gated recurrent unit. According to the data D = {d 1 , d 2 , d 3 … d l}, d i = {d i1 , d i2 , d i3 … d ia}, initialize the batch size b of the single training data of the model, the time length r, and set the value of the number of input variables to be equal to the number of features a, and set the value of the output dimension to 1.

[0058] (2) According to the data length l of the data D = {d 1 , d 2 , d 3 … d l} and the batch size b of the single training data of the model, calculate the number of times c = l / b required for the model to train once. Use the time length r as the window size and the step size as 1 to slice the input data D and process it into three-dimensional data D′ = {d′ 1 , d′2 , d' 3 … d' l-b+1},where d' i = {d i , d i+1 , d i+2 … d i+b-1}}。

[0059] (3) For each training, input three - dimensional input data equal to the batch size of the model's single - training data into the network. For each group of two - dimensional data included in this data, input them into their respective forward GRU and backward GRU in sequence. For the forward GRU, the output of the previous GRU serves as part of the input of the next GRU. For the backward GRU, the output of the next GRU serves as part of the input of the previous GRU; The output of each group of bidirectional GRUs is the output matrix S t + of the forward GRU and the output matrix S t - of the backward GRU. Sum and average them to obtain the initial state vector S t , and calculate t + through S t - and S . The formula for

[0060]

[0061] (4) For the state vector S( i ) obtained after each group of BiGRU processes the data, through the calculation of the Attention layer, successively generate the target attention weight e( i ), the probability vector α (i) = (α 0 , α 1 , α 2 … α k ). Perform a weighted sum operation on all the state vectors S (i) and the probability vector α (i) to obtain the corresponding score vector Y (i) = (y 0 (i) , y 1 (i) , y 2 (i) … y k (i) ), where k represents there are k abnormal working conditions. Among them, the calculation methods of the target attention weight e (i) , the probability vector α (i) and the score vector Y (i) are as follows:

[0062] e (i) = ω i ·S (i) + b i

[0063]

[0064] Wherein, S (i) is the initial state vector of the i-th eigenvector, ω i represents the weight coefficient matrix of the i-th eigenvector, b i represents the offset corresponding to the i-th eigenvector.

[0065] (5) Process the score vector Y (i) through a layer of softmax, and process the score of the corresponding label into the probability that the current training data may be a certain label, to obtain the probability vector The calculation formula is as follows:

[0066]

[0067] Wherein, represents the score that the i-th sample is classified as the j-th working condition, represents the probability that the i-th sample is classified as the j-th working condition.

[0068] Step S2: Calculate the model loss using an improved cross-entropy loss function and update the model weights.

[0069] (1) Traverse the input data D, count the number of normal samples and abnormal samples in the data, and calculate the ratio ω of the number of normal samples n + to the number of abnormal samples n - :

[0070]

[0071]

[0072] Wherein, y (i) represents the label of the i-th sample, and the value range is 0-k, where 0 represents the normal working condition, represents the k-th abnormal working condition, I(·) is the indicator function, which takes the value of 1 when · is true and 0 when · is false.

[0073] (2) For each training of the model, calculate the cross-entropy loss function once, and introduce the ratio ω at the positive sample loss term in the function to balance the quantity difference between positive and abnormal samples. The calculation formula of the improved cross-entropy loss function is as follows;

[0074]

[0075] Among them, λ is a control parameter. When λ approaches 0, this formula is the standard cross-entropy loss function. When λ is large enough, the effect of distinguishing different abnormal working conditions becomes weak, and this formula becomes the loss function for solving the binary classification problem, aiming to identify abnormal and normal data.

[0076] (3) Execute the backpropagation algorithm, calculate the partial derivatives of the loss function in the model with respect to each weight parameter, solve along the direction of gradient descent to obtain the optimization direction of each layer's weight parameter in the model, and update the weights of each layer in the model by using the gradient descent method. The calculation formulas for the partial derivatives of each part of the loss function are as follows:

[0077]

[0078] Among them, is the probability vector α calculated for the i-th sample in the attention layer t and is the l-th probability in

[0079] Step S3: Construct a data feature discrimination model.

[0080] Train the network in a loop until the model converges or reaches the preset maximum number of training times to obtain the final network model; collect the data that has not participated in training as the test set to test the model, calculate the accuracy, precision, and recall parameters of the model, and use them to measure the specific performance of the model.

[0081] The embodiment of the present invention also provides a data feature discrimination system, including:

[0082] A bidirectional gated recurrent unit, which is used to extract the feature information of time series data and combine the attention mechanism to achieve effective feature screening to obtain the abnormal label probability.

[0083] Initialize the initial model of the bidirectional gated recurrent unit according to the data length and feature quantity of the time series data; slice the training data and input it into the respective forward GRU and backward GRU in sequence, sum and average the output matrices of the forward GRU and backward GRU of each group of bidirectional GRUs to obtain the initial state vector, and obtain the abnormal label probability through the calculation of the Attention layer and the processing of the softmax layer.

[0084] A cross-entropy loss function module, which improves the cross-entropy loss function, calculates the model loss, and updates the model weights.

[0085] The improved cross - entropy loss function is used to calculate the model loss and update the model weights. For each training of the model, the cross - entropy loss function is calculated once, and the ratio of the number of abnormal samples to the number of normal samples is introduced into the positive - sample loss term in the function to balance the quantity difference between positive and abnormal samples; the backpropagation algorithm is executed to calculate the partial derivatives of each part in the discriminant model, and the gradient descent algorithm is used to update the weights of each layer of the model.

[0086] The data feature discriminant model training module conducts cyclic training on the network until the model converges or reaches the preset maximum number of training times, obtaining the final network model, and uses this network model for data feature discrimination.

[0087] This system realizes data feature discrimination through the data feature discrimination method described in the above embodiments. The specific implementation steps are as follows:

[0088] Step S1: Classification is performed using a bidirectional gated recurrent unit with an attention mechanism introduced.

[0089] (1) Build the initial model of the bidirectional gated recurrent unit. According to the data length l, the number of features a of the data D = {d 1 , d 2 , d 3 … d l}, d i = {d i1 , d i2 , d i3 … d ia}, initialize the batch size b of the single - training data of the model, the time length r, and set the value of the number of input variables to be equal to the number of features a, and set the value of the output dimension to 1.

[0090] (2) According to the data length l of the data D = {d 1 , d 2 , d 3 … d l} and the batch size b of the single - training data of the model, calculate the number of times c = l / b required for the model to be trained once. Using the time length r as the window size and the step size set to 1, slice the input data D and process it into three - dimensional data D′ = {d′ 1 , d′ 2 , d′ 3 … d′ l-b+1}, where d′ i = {d i , d i+1 , d i+2 … d i+b-1}.

[0091] (3) For each training, input three-dimensional input data equal to the batch size of the model's single training data into the network. For each group of two-dimensional data contained in this data, input them into their respective forward GRU and backward GRU in sequence. For the forward GRU, the output of the previous GRU serves as part of the input of the next GRU. For the backward GRU, the output of the next GRU serves as part of the input of the previous GRU; the output of each group of bidirectional GRUs is the output matrix S of the forward GRU t + and the output matrix S of the backward GRU t - Sum and average them to obtain the initial state vector S t , through S t + and S t - Calculate to obtain The formula of is:

[0092]

[0093] (4) For the state vector S( i ) obtained after each group of BiGRU processes the data, calculate through the Attention layer, and sequentially generate the target attention weight e (i) , probability vector α (i) =(α 0 ,α 1 ,α 2 …α k ), perform a weighted sum operation on all the state vectors S (i) and the probability vector α (i) to obtain the corresponding score vector Y representing each label (i) =(y 0 (i) ,y 1 (i) ,y 2 (i) ...y k (i) ), k represents that there are k abnormal working conditions, where the target attention weight e (i) , probability vector α (i) and score vector Y (i) The calculation methods are as follows:

[0094] e (i) =ω i ·S (i) +b i

[0095]

[0096] Among them, S(i) is the initial state vector of the i-th eigenvector, ω i represents the weight coefficient matrix of the i-th eigenvector, b i represents the offset corresponding to the i-th eigenvector.

[0097] (5) Pass the score vector Y (i) through a layer of softmax for processing, and process the score of the corresponding label into the probability that the current training data may be a certain label, obtaining the probability vector The calculation formula is as follows:

[0098]

[0099] where, represents the score of the i-th sample classified as the j-th working condition, represents the probability that the i-th sample is classified as the j-th working condition.

[0100] Step S2: Calculate the model loss using an improved cross-entropy loss function and update the model weights.

[0101] (1) Traverse the input data D, count the number of normal samples and abnormal samples in the data, and calculate the ratio ω of the number of normal samples n + to the number of abnormal samples n - :

[0102]

[0103] where, y (i) represents the label of the i-th sample, and the value range is 0-k, where 0 represents the normal working condition, represents the k-th abnormal working condition, I(·) is the indicator function, which takes the value of 1 when · is true and 0 when · is false.

[0104] (2) For each training of the model, calculate the cross-entropy loss function once, and introduce the ratio ω at the positive sample loss term in the function to balance the quantity difference between positive and abnormal samples. The calculation formula of the improved cross-entropy loss function is as follows;

[0105]

[0106] where, λ is the control parameter. When λ approaches 0, this formula is the standard cross-entropy loss function. When λ is large enough, the effect of distinguishing different abnormal working conditions becomes weak, and this formula becomes the loss function for solving the binary classification problem, aiming to identify abnormal and normal data.

[0107] (3) Execute the backpropagation algorithm, calculate the partial derivatives of the loss function in the model with respect to each weight parameter, solve along the direction of gradient descent to obtain the optimization direction of the weight parameters of each layer in the model, and update the weights of each layer in the model using the gradient descent method. The calculation formulas for the partial derivatives of each part of the loss function are as follows:

[0108]

[0109] Among them, is the probability vector α calculated in the attention layer for the i-th sample t and is the l-th probability in

[0110] Step S3: Construct a data feature discrimination model.

[0111] Perform cyclic training on the network until the model converges or reaches the preset maximum number of training times to obtain the final network model; collect the data that has not participated in training as the test set to test the model, calculate the accuracy, precision, and recall parameters of the model, and use them to measure the specific performance of the model.

[0112] An embodiment of the present invention also provides a data feature discrimination device, including at least one memory and at least one processor;

[0113] The at least one memory is used to store machine-readable programs;

[0114] The at least one processor is used to call the machine-readable program to implement the data feature discrimination method described in the above embodiment.

[0115] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the data feature discrimination method described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.

[0116] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0117] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0118] Furthermore, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by a computer, but also by causing an operating system or the like operating on the computer to complete some or all of the actual operations based on the instructions of the program code.

[0119] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then the CPU or the like installed on the expansion board or the expansion unit is caused to execute some and all of the actual operations based on the instructions of the program code, thereby realizing the functions of any one of the above embodiments.

[0120] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.

Claims

1. A data feature identification method, characterized in that: The attention mechanism and the bidirectional GRU network are combined to realize data feature discrimination. The feature information of time series data is extracted by using a bidirectional gated recurrent unit, and the attention mechanism is used to weight the hidden state to complete effective feature screening. The cross entropy loss function is improved, the model loss is calculated and the model weights are updated, so that the trained model can better distinguish normal and abnormal data.

2. A data feature identification method according to claim 1, characterized in that: The implementation of this method is as follows: Initialize the initial model of the bidirectional gated recurrent unit according to the data length and number of features of the time series data; slice the training data and input them into the respective forward GRU and reverse GRU in order, sum and average the output matrices of the forward GRU and reverse GRU of each group of bidirectional GRU to obtain the initial state vector, and obtain the abnormal label probability through the Attention layer calculation and softmax layer processing; The improved cross entropy loss function is used to calculate the model loss and update the model weights. For each training of the model, the cross entropy loss function is calculated once, and the ratio of the number of abnormal samples to the number of normal samples is introduced in the positive sample loss term in the function to balance the difference in the number of positive and abnormal samples; the back propagation algorithm is executed to calculate the partial derivatives of each part in the discriminant model, and the gradient descent algorithm is used to update the weights of each layer of the model; The network is trained in a loop until the model converges or reaches the preset maximum number of training times, and the final network model is obtained, which is used for data feature discrimination.

3. A data feature identification method according to claim 1 or 2, characterized in that: The specific steps to implement this method include: Step S1: Classification is performed using a bidirectional gated recurrent unit with an attention mechanism. Step S2: Use the improved cross entropy loss function to calculate the model loss and update the model weight; Step S3: Construct a data feature discrimination model.

4. A data feature identification method according to claim 3, characterized in that: The specific implementation process of step S1 is as follows: (1) Build the initial model of bidirectional gated recurrent unit, based on the data D = {d1, d2, d3…d l }, d i ={d i1 ,d i2 ,d i3 …d ia }, the data length l, the number of features a, the single training data batch size b of the initialization model, the time length r, and set the value of the number of input variables equal to the number of features a, and the value of the output dimension is set to 1; (2) Based on data D = {d1, d2, d3…d l } data length l and the model single training data batch size b, calculate the number of times c = l / b required for model training once, take the time length r as the window size, set the step size to 1, slice the input data D, and process it into three-dimensional data D′ = {d′1, d′2, d′3…d′ l-b+1 }, where d′ i ={d i ,d i+1 ,d i+2 …d i+b-1 }; (3) For each training, three-dimensional input data equal to the batch size of the model's single training data is input into the network. For each set of two-dimensional data contained in the data, it is input into its respective forward GRU and reverse GRU in order. The output of each set of bidirectional GRU is the output matrix S of the forward GRU. t + and the output matrix S of the reverse GRU t - Sum and average to get the initial state vector S t , through S t + and S t - Calculated The formula is: (4) For each group of BiGRU, the state vector S is obtained after processing the data (i) , after the Attention layer calculation, the target attention weights e are generated in turn (i) , probability vector α (i) =(α0,α1,α2…α k ), the entire state vector S (i) and the probability vector α (i) Perform a weighted sum operation to obtain the corresponding score vector Y representing each label (i) =(y0 (i) ,y1 (i) ,y2 (i) …y k (i) ), k represents k abnormal working conditions, where the target attention weight e (i) , the probability vector α (i) and the score vector Y (i) The calculation method is as follows: e (i) =ω i ·S (i) +b i Among them, S (i) is the initial state vector of the i-th eigenvector, ω i represents the weight coefficient matrix of the i-th eigenvector, b i Represents the offset corresponding to the i-th eigenvector; (5) The score vector Y (i) After a layer of softmax processing, the score of the corresponding label is processed into the probability that the training data may be a certain label, and the probability vector is obtained. The calculation formula is as follows: in, represents the score of the i-th sample classified as the j-th working condition, It represents the probability that the i-th sample is classified as the j-th operating condition.

5. A data feature identification method according to claim 4, characterized in that: The above are input into their respective forward GRU and reverse GRU in order: For the forward GRU, the output of the previous GRU serves as part of the input of the next GRU; for the reverse GRU, the output of the next GRU serves as part of the input of the previous GRU.

6. A data feature identification method according to claim 3, characterized in that: The specific implementation process of step S2 is as follows: (1) Traverse the input data D, count the number of normal samples and the number of abnormal samples in the statistics, and calculate the number of normal samples n + and the number of abnormal samples n - The ratio ω: Among them, y (i) represents the label of the i-th sample, with a value range of 0-k, where 0 represents the normal working condition, represents the k-th abnormal working condition, and I(·) is the indicator function, which takes the value of 1 when · is true and takes the value of 0 when I is false; (2) For each training of the model, the cross entropy loss function is calculated once, and the ratio ω is introduced in the positive sample loss term in the function to balance the difference in the number of positive and abnormal samples. The improved cross entropy loss function calculation formula is as follows; Among them, λ is the control parameter. When λ tends to 0, the formula is the standard cross entropy loss function. When λ is large enough, the effect of distinguishing different abnormal conditions becomes weaker, and the formula becomes a loss function for solving the binary classification problem, identifying abnormal and normal data; (3) Execute the back propagation algorithm to calculate the partial derivatives of the loss function in the model with respect to each weight parameter, solve along the direction of gradient descent to obtain the optimization direction of the weight parameters of each layer in the model, and use the gradient descent method to update the weights of each layer of the model. The calculation formula for the partial derivatives of each part of the loss function is as follows: in, is the probability vector α calculated for the i-th sample in the attention layer t The lth probability in .

7. A data feature identification method according to claim 3, characterized in that: The specific implementation of step S3 is as follows: The network is trained in a loop until the model converges or reaches the preset maximum number of training times to obtain the final network model; data that has not participated in the training is collected and used as a test set to test the model, and the model's accuracy, precision, and recall parameters are calculated and used to measure the specific performance of the model.

8. A data feature identification system, characterized in that: include: Bidirectional gated recurrent unit, used to extract feature information of time series data, and combined with the attention mechanism to achieve effective feature screening and obtain abnormal label probability; The cross entropy loss function module improves the cross entropy loss function, calculates the model loss and updates the model weights; The data feature discrimination model training module performs cyclic training on the network until the model converges or reaches the preset maximum number of training times, and obtains the final network model, which is used for data feature discrimination; The system realizes data feature discrimination through the method described in any one of claims 1 to 7.

9. A data feature identification device, characterized in that: comprising at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, enable the processor to implement the method according to any one of claims 1 to 7.