A model training method and related device

By using the output of noise-free data and noise data to calculate the loss function when training the noise-reducing classification network, the problem of low classification recognition accuracy in sparse timing data in the prior art is solved, and higher classification recognition accuracy and more complete global feature learning are achieved.

CN114692667BActive Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011623266.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-06-10
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

In the prior art, the classification and identification accuracy of noisy sparse timing data is low, and it is difficult to ensure the normal identification of sparse timing data.

Method used

In the process of training the noise reduction classification network, the first and second loss functions are obtained based on the output of the noiseless data and the noise data, and the noise reduction classification network is trained based on these loss functions to obtain the target network.

Benefits of technology

Ensure that the noise reduction target and the classification accuracy target are consistent, so that the network can learn more complete global features during the noise reduction stage, enhance the network's ability to suppress local noise and disturbances, and improve the accuracy of classification recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692667B_ABST
    Figure CN114692667B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, which can be applied to the field of artificial intelligence. The method includes: obtaining a sample pair including noise data and noise-free data; inputting the noise-free data into a noise reduction classification network of a noise reduction network and a classification network to obtain first output data output by the noise reduction network and second output data output by the classification network; inputting the noise data into the noise reduction classification network to obtain third output data output by an intermediate layer of the noise reduction network and fourth output data output by the classification network. Determining a first loss function according to the first output data and the third output data; determining a second loss function according to the second output data and the fourth output data; training the noise reduction classification network at least according to the first loss function and the second loss function until a preset training condition is satisfied to obtain a target network. This solution can enhance the ability of the network to suppress local noise and perturbations, and improve the accuracy of the network for classification and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method and related devices. Background Art

[0002] Artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0003] Sparse time-series data is a set of non-dense data sequences arranged in chronological order. Common sparse time-series data includes, for example, human body bone key point data, electrocardiogram data, and inertial measurement unit (IMU) data. By classifying and identifying sparse time-series data, useful information can be obtained. For example, based on human body bone key point data, the posture and movement of the human body can be identified; based on electrocardiogram data, the physical condition of the human body can be diagnosed; based on IMU data on wearable devices, the movement state of the human body can be identified.

[0004] In a real environment, sparse time-series data is usually affected by various noises, resulting in the sparse time-series data collected by the device including noises. Based on this, in related technologies, usually, after denoising the sparse time-series data, the denoised sparse time-series data is classified and identified based on the original classification method.

[0005] However, the classification and identification accuracy of noisy sparse time-series data in related technologies is relatively low, and it is difficult to ensure the normal identification of sparse time-series data. Therefore, there is an urgent need for a method that can effectively classify and identify noisy sparse time-series data. Summary of the Invention

[0006] The present application provides a model training method and related devices. During the process of training a noise reduction classification network, a first loss function is obtained based on the output of the noise reduction network of the noise reduction classification network for noise-free data and the output of the intermediate layer of the noise reduction network for noisy data. A second loss function is obtained based on the outputs of the entire noise reduction classification network for noise-free data and noisy data. Then, the noise reduction classification network is trained based on the first loss function and the second loss function to obtain a target network. By training the noise reduction classification network and obtaining the loss function based on the outputs of the noise reduction network and the noise reduction classification network for noise-free data and noisy data, it can ensure that the noise reduction target and the classification accuracy target are consistent, enable the network to learn more complete global features during the noise reduction stage, enhance the network's ability to suppress local noise and perturbations, and improve the accuracy of the network for classification and recognition.

[0007] In the first aspect of the present application, a model training method is provided. The method includes: The terminal obtains a sample pair from a sample set including multiple sample pairs. The sample pair includes noisy data and the noise-free data corresponding to the noisy data, that is, the noise-free data corresponding to the noisy data in the sample pair is obtained by removing the noise from the noisy data. The terminal inputs the noise-free data into the noise reduction classification network to obtain first output data and second output data. Among them, the noise reduction classification network includes a noise reduction network and a classification network. The first output data is the output of the noise reduction network, and the second output data is the output of the classification network. The terminal inputs the noisy data into the noise reduction classification network to obtain third output data and fourth output data. The third output data is obtained based on the intermediate layer of the noise reduction network, and the fourth output data is the output of the classification network.

[0008] Then, the terminal determines a first loss function according to the first output data and the third output data. The first loss function is used to represent the difference between the first output data and the third output data. By obtaining the first loss function between the outputs of the noise reduction network for noise-free data and noisy data respectively, and training the noise reduction classification network based on the first loss function, the noise reduction classification network can learn more global features closer to the noise-free data during the noise reduction process, thereby enhancing the ability of the noise reduction network to suppress local noise and perturbations.

[0009] Secondly, the terminal determines a second loss function according to the second output data and the fourth output data. The second loss function is used to represent the difference between the second output data and the fourth output data and the true class label of the noise-free data. For example, the second loss function is obtained by calculating the sum of the first difference value between the second output data and the true class label of the noise-free data and the second difference value between the fourth output data and the true class label of the noise-free data.

[0010] Finally, the terminal trains the noise reduction classification network based on at least the first loss function and the second loss function until a preset training condition is met, and obtains a target network. Specifically, the terminal can obtain a total loss function based on the first loss function and the second loss function, and the terminal trains the noise reduction classification network based on the total loss function. The total loss function can be the sum of the first loss function and the second loss function, or can be obtained by adding the product of the first loss function and the first proportionality coefficient to the product of the second loss function and the second proportionality coefficient.

[0011] In this solution, by training the noise reduction classification network and obtaining the loss function based on the output of the noise reduction network and the output of the noise reduction classification network for the noise-free data and the noise data, it can ensure that the noise reduction target and the classification accuracy target are consistent, enable the network to learn more complete global features during the noise reduction stage, enhance the network's ability to suppress local noise and perturbations, and improve the accuracy of the network for classification and recognition.

[0012] Optionally, in a possible implementation, the noise reduction network is an autoencoder, and the autoencoder includes an encoder and a decoder. Among them, the autoencoder, also known as the automatic encoder, is an artificial neural network that can learn an efficient representation of the input data through unsupervised learning. In practical applications, by adding constraint conditions to the autoencoder, the autoencoder can perform noise reduction processing on the noise data. In the autoencoder, the encoder is used to perform compression encoding on the input data, and the decoder is used to reconstruct the data output by the encoder. The first output data is the output of the encoder, and the third output data is obtained based on the intermediate layer of the encoder. By using the autoencoder including the encoder and the decoder as the noise reduction network, it can reduce the modification to the existing technology and improve the practicality of the solution.

[0013] Optionally, in a possible implementation, the terminal inputs the noise data into the noise reduction classification network to obtain the third output data, including: the terminal inputs the noise data into the noise reduction classification network to obtain the feature data output by the intermediate layer of the noise reduction network. The terminal divides the feature data into multiple sub-feature data to obtain the third output data, and the third output data includes multiple sub-feature data. The terminal determines the difference value between each sub-feature data in the third output data and the first output data. The terminal determines the first loss function according to the difference value between each sub-feature data and the first output data.

[0014] In this solution, by dividing the feature data output by the intermediate layer of the noise reduction network for the noise data into multiple sub-feature data and establishing a loss function based on the sub-feature data and the first output data, it can guide the noise reduction network to learn rich global information and improve the noise reduction effect of the noise reduction network.

[0015] Optionally, in a possible implementation, the terminal evenly divides the feature data into multiple sub-feature data in chronological order to obtain third output data, where the length of the time period corresponding to each sub-feature data in the multiple sub-feature data is the same, and the noise data is time-series data.

[0016] Since the time-series data is coherent and adjacent time-series data has a certain correlation, constructing a loss function based on the local time-series features corresponding to the noise data and the global time-series features corresponding to the noise-free data can guide the denoising network to learn richer global time-series information, thereby enhancing the ability of the denoising network to suppress local noise and perturbations.

[0017] Optionally, in a possible implementation, since the first output data is the data output by the encoder in the denoising network and the third output data is the data output by the intermediate layer of the encoder in the denoising network, their dimensions are not the same. Therefore, before calculating the first loss function, a dimension alignment operation can be performed on the first output data and the third output data to make their dimensions the same, and then the difference value between them can be calculated.

[0018] Specifically, the terminal determines the difference value between each sub-feature data in the third output data and the first output data, including: the terminal performs a dimension alignment operation on each sub-feature data in the first output data and the third output data respectively to obtain the dimension-aligned first output data and third output data. The terminal determines the difference value between each sub-feature data in the dimension-aligned third output data and the dimension-aligned first output data. In practical applications, the terminal can pre-construct multiple dimension alignment sub-networks, and by inputting the first output data into one of the dimension alignment sub-networks and inputting each sub-feature data in the third output data into other corresponding dimension alignment sub-networks, the dimension-aligned first output data and third output data can be obtained.

[0019] Optionally, in a possible implementation, the terminal determines the second loss function according to the second output data and the fourth output data, including: the terminal determines the difference between the second output data and the true class label of the noise-free data to obtain a first difference value. The terminal determines the difference between the fourth output data and the true class label of the noise-free data to obtain a second difference value. The terminal obtains the second loss function according to the first difference value and the second difference value. Wherein, the second output data is the prediction result of multi-classification, which is used to represent the result predicted by the classification network.

[0020] Optionally, in a possible implementation, during the training process of the noise reduction classification network, a binary classifier can also be introduced. The binary classifier can perform binary classification prediction on the data input to the noise reduction classification network based on the features extracted by the noise reduction classification network, that is, predict whether the data input to the noise reduction classification network is noise data or noise-free data. Then, based on the binary classification result output by the binary classifier and the true binary classification label corresponding to the input data, the terminal determines a third loss function, which is used to obtain the total loss function together with the first loss function and the second loss function, that is, the third loss function is also used for the training of the noise reduction classification network.

[0021] Specifically, the method further includes: the terminal obtains a first feature and predicts the binary classification result corresponding to the noise-free data according to the first feature to obtain a first prediction result. The first feature is extracted by the classification network in the noise reduction classification network after the noise-free data is input to the noise reduction classification network. The terminal obtains a second feature and predicts the binary classification result corresponding to the noise data according to the second feature to obtain a second prediction result. The second feature is extracted by the classification network in the noise reduction classification network after the noise data is input to the noise reduction classification network. The terminal determines the third loss function according to the first prediction result and the true binary classification label of the noise-free data, and the second prediction result and the true binary classification label of the noise data. The terminal trains the noise reduction classification network at least according to the first loss function, the second loss function, and the third loss function. The binary classification result corresponding to the noise-free data is of the noise-free type or the noise type, and the binary classification result corresponding to the noise data is of the noise-free type or the noise type.

[0022] In this solution, by introducing a binary classifier in the training stage and predicting the binary classification result of the input data through the binary classifier based on the features extracted by the classification network in the noise reduction classification network, the loss function corresponding to the binary classification result is obtained. By introducing the loss function corresponding to the binary classification result on the basis of the original loss function, an additional evaluation dimension can be introduced, so that the trained noise reduction classification network can have an adaptive noise reduction classification scale for different types of input data, and the classification accuracy of the noise reduction classification network can be improved.

[0023] Optionally, in a possible implementation, the terminal trains a noise reduction classification network at least based on a first loss function and a second loss function, including: the terminal updates the parameters of the noise reduction classification network at least based on the first loss function and the second loss function through the error backpropagation algorithm. Briefly, the terminal can correct the magnitudes of the parameters in the initial noise reduction classification network during the training process of the noise reduction classification network through the error backpropagation algorithm, so that the reconstruction error loss of the noise reduction classification network becomes smaller and smaller. Specifically, the forward transmission of the input signal until the output will generate an error loss, and the initial parameters in the noise reduction classification network are updated by backpropagating the error loss information, so that the error loss converges.

[0024] Optionally, in a possible implementation, the noise data in the sample pair includes sparse time series data, and the sparse time series data includes skeleton point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data.

[0025] The second aspect of this application provides a noise reduction classification method, which includes: obtaining data to be classified; inputting the data to be classified into a target network to obtain a prediction result, where the prediction result is the classification result of the data to be classified; where the target network is used to perform noise reduction processing and classification on the data to be classified, and the target network is trained based on the method described in the first aspect.

[0026] The third aspect of this application provides a model training device, including: an acquisition unit and a processing unit. The acquisition unit is used to acquire a sample pair, where the sample pair includes noise data and noise-free data corresponding to the noise data; the processing unit is used to input the noise-free data into a noise reduction classification network to obtain first output data and second output data, where the noise reduction classification network includes a noise reduction network and a classification network, the first output data is the output of the noise reduction network, and the second output data is the output of the classification network; the processing unit is further used to input the noise data into the noise reduction classification network to obtain third output data and fourth output data, where the third output data is obtained based on an intermediate layer of the noise reduction network, and the fourth output data is the output of the classification network; the processing unit is further used to determine a first loss function according to the first output data and the third output data, where the first loss function is used to represent the difference between the first output data and the third output data; the processing unit is further used to determine a second loss function according to the second output data and the fourth output data, where the second loss function is used to represent the difference between the second output data and the fourth output data and the true class label of the noise-free data; the processing unit is further used to train the noise reduction classification network at least based on the first loss function and the second loss function until a preset training condition is met to obtain a target network.

[0027] Optionally, in a possible implementation, the noise reduction network includes an encoder and a decoder. The encoder is configured to perform compression encoding on input data, and the decoder is configured to perform data reconstruction on the data output by the encoder. The first output data is the output of the encoder, and the third output data is obtained based on an intermediate layer of the encoder.

[0028] Optionally, the processing unit is further configured to: input the noise data into the noise reduction classification network to obtain feature data output by an intermediate layer of the noise reduction network; divide the feature data into a plurality of sub-feature data to obtain the third output data, where the third output data includes the plurality of sub-feature data; determine a difference value between each sub-feature data in the third output data and the first output data; and determine the first loss function according to the difference value between each sub-feature data and the first output data.

[0029] Optionally, in a possible implementation, the processing unit is further configured to evenly divide the feature data into a plurality of sub-feature data in chronological order to obtain the third output data, where the length of the time period corresponding to each sub-feature data in the plurality of sub-feature data is the same; and the noise data is time series data.

[0030] Optionally, in a possible implementation, the processing unit is further configured to: respectively perform a dimension alignment operation on each sub-feature data in the first output data and the third output data to obtain dimension-aligned first output data and third output data; and determine a difference value between each sub-feature data in the dimension-aligned third output data and the dimension-aligned first output data.

[0031] Optionally, in a possible implementation, the processing unit is further configured to: determine a difference between the second output data and the true class label of the noise-free data to obtain a first difference value; determine a difference between the fourth output data and the true class label of the noise-free data to obtain a second difference value; and obtain the second loss function according to the first difference value and the second difference value; where the second output data is a multi-classification prediction result for representing the result predicted by the classification network.

[0032] Optionally, in a possible implementation, the obtaining unit is further configured to obtain a first feature, and predict a binary classification result corresponding to the noise-free data according to the first feature to obtain a first prediction result, where the first feature is extracted by a classification network in the noise reduction classification network after the noise-free data is input into the noise reduction classification network; the obtaining unit is further configured to obtain a second feature, and predict a binary classification result corresponding to the noise data according to the second feature to obtain a second prediction result, where the second feature is extracted by a classification network in the noise reduction classification network after the noise data is input into the noise reduction classification network; the processing unit is further configured to determine a third loss function according to the first prediction result and the true binary classification label of the noise-free data, and the second prediction result and the true binary classification label of the noise data; the processing unit is further configured to train the noise reduction classification network at least according to the first loss function, the second loss function, and the third loss function; where the binary classification result corresponding to the noise-free data is a noise-free type or a noise type, and the binary classification result corresponding to the noise data is a noise-free type or a noise type.

[0033] Optionally, in a possible implementation, at least according to the first loss function and the second loss function, the parameters of the noise reduction classification network are updated by using the error backpropagation algorithm.

[0034] Optionally, in a possible implementation, the noise data includes sparse time series data.

[0035] Optionally, in a possible implementation, the sparse time series data includes skeleton point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data.

[0036] A fourth aspect of this application provides a noise reduction classification device, which includes an obtaining unit and a processing unit. The obtaining unit is configured to obtain data to be classified. The processing unit is configured to input the data to be classified into a target network to obtain a prediction result, where the prediction result is a classification result of the data to be classified; where the target network is used to perform noise reduction processing and classification on the data to be classified, and the target network is trained based on the method described in the first aspect.

[0037] A fifth aspect of this application provides a model training device, which may include a processor. The processor is coupled to a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method described in the first aspect is implemented. For the steps executed by the processor in each possible implementation manner of the first aspect, reference may specifically be made to the first aspect, and details are not described herein again.

[0038] The sixth aspect of the present application provides a noise reduction classification device, which may include a processor. The processor is coupled to a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method described in the second aspect above is implemented. For the steps in each possible implementation manner of the second aspect executed by the processor, reference may specifically be made to the second aspect, and details are not elaborated herein.

[0039] The seventh aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, the computer is enabled to execute the method described in the first aspect or the second aspect above.

[0040] The eighth aspect of the present application provides a computer program product, in which a computer program is stored. When it runs on a computer, the computer is enabled to execute the method described in the first aspect or the second aspect above.

[0041] The ninth aspect of the present application provides a circuit system, which includes a processing circuit configured to execute the method described in the first aspect or the second aspect above.

[0042] The tenth aspect of the present application provides a chip, including one or more processors. Some or all of the processors are used to read and execute a computer program stored in a memory to execute the method in any possible implementation manner of any of the above aspects. Optionally, the chip may include a memory, and the memory is connected to the processor through a circuit or a wire. Optionally, the chip further includes a communication interface, and the processor is connected to the communication interface. The communication interface is used to receive data and / or information that needs to be processed. The processor obtains the data and / or information from the communication interface, processes the data and / or information, and outputs the processing result through the communication interface. The communication interface may be an input / output interface. The method provided by the present application may be implemented by one chip or by multiple chips working together. Description of the Drawings

[0043] Figure 1 It is a schematic structural diagram of an artificial intelligence main body framework provided by an embodiment of the present application;

[0044] Figure 2a It is a data processing system provided by an embodiment of the present application;

[0045] Figure 2b It is another data processing system provided by an embodiment of the present application;

[0046] Figure 2c It is a schematic diagram of related devices for data processing provided by an embodiment of the present application;

[0047] Figure 3aSchematic diagram of an architecture of system 100 provided by an embodiment of the present application;

[0048] Figure 3b Schematic diagram of an application scenario provided by an embodiment of the present application;

[0049] Figure 3c Schematic diagram of a specific application of time series data provided by an embodiment of the present application;

[0050] Figure 4 Schematic flowchart of a model training method provided by an embodiment of the present application;

[0051] Figure 5 Schematic diagram of the structure of a noise reduction classification network provided by an embodiment of the present application;

[0052] Figure 6 Schematic diagram of the structure of a local-global feature correlation module and a hybrid classifier provided by an embodiment of the present application;

[0053] Figure 7 Schematic flowchart of training a noise reduction classification network provided by an embodiment of the present application;

[0054] Figure 8 Schematic diagram of generating noise data provided by an embodiment of the present application;

[0055] Figure 9a Schematic flowchart of constructing a local-global feature correlation loss provided by an embodiment of the present application;

[0056] Figure 9b Schematic flowchart of constructing a local-global feature correlation loss and a hybrid classification loss provided by an embodiment of the present application;

[0057] Figure 10 Schematic diagram of comparison between an existing solution and the solution of the present application provided by an embodiment of the present application;

[0058] Figure 11 Schematic diagram of the structure of a model training device provided by an embodiment of the present application;

[0059] Figure 12 Schematic diagram of the structure of a noise reduction classification device provided by an embodiment of the present application;

[0060] Figure 13 Schematic diagram of a structure of an execution device provided by an embodiment of the present application;

[0061] Figure 14 Schematic diagram of a structure of a training device provided by an embodiment of the present application;

[0062] Figure 15A schematic structural diagram of a chip provided by an embodiment of the present application. Detailed implementation manners

[0063] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are only for explaining the specific embodiments of the present invention, and are not intended to limit the present invention.

[0064] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0065] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0066] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence theme framework will be described below from two dimensions: "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0067] (1) Infrastructure

[0068] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and this data is provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0069] (2) Data

[0070] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0071] (3) Data Processing

[0072] Data processing usually includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0073] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, training, etc. on data in a symbolic and formalized manner.

[0074] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formalized information for machine thinking and problem-solving according to the reasoning control strategy. The typical function is search and matching.

[0075] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0076] (4) General Capabilities

[0077] After the data undergoes the above-mentioned data processing, some general capabilities can be formed further based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0078] (5) Intelligent Products and Industry Applications

[0079] Intelligent products and industry applications refer to the products and applications of the artificial intelligence system in various fields. It is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, safe cities, etc.

[0080] Next, several application scenarios of the present application will be introduced.

[0081] Figure 2a A data processing system provided for an embodiment of the present application, the data processing system includes a user device and a data processing device. Among them, the user device includes intelligent terminals such as mobile phones, personal computers or information processing centers. The user device is the initiating end of data processing and is the initiator of the data noise reduction and classification request. Usually, the user initiates a request through the user device.

[0082] The above data processing device can be a device or server with data processing functions such as a cloud server, a network server, an application server, and a management server. The data processing device receives a data noise reduction and classification request from the intelligent terminal through an interaction interface, and then performs data processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through the memory for storing data and the processor for data processing. The memory in the data processing device can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.

[0083] In Figure 2a In the data processing system shown, the user device can receive the user's instruction. For example, the user device can obtain a set of data input / selected by the user, and then initiate a request to the data processing device, so that the data processing device executes the data noise reduction and classification application for the data obtained by the user device, thereby obtaining a corresponding processing result for the data. Shown in Figure 2a In, the data processing device can execute the model training method of the embodiment of the present application.

[0084] Figure 2b Another data processing system provided for an embodiment of the present application, in Figure 2b In, the user device directly serves as the data processing device. The user device can directly obtain the input from the user and directly process it by the hardware of the user device itself. The specific process is similar to Figure 2a the above description, and will not be elaborated here.

[0085] In Figure 2b In the data processing system shown, the user device can receive the user's instruction. For example, the user device can obtain a piece of data selected by the user in the user device, and then the user device itself executes a data processing application for the data, thereby obtaining a corresponding processing result for the data.

[0086] In Figure 2b In, the user device itself can execute the model training method of the embodiment of the present application.

[0087] Figure 2cIt is a schematic diagram of related devices for data processing provided by an embodiment of the present application.

[0088] The above Figure 2a and Figure 2b The user device in Figure 2c can specifically be the local device 301 or the local device 302 in Figure 2a The data processing device in Figure 2c can specifically be the execution device 210 in

[0089] Figure 2a and Figure 2b The processor in

[0090] Figure 3a is a schematic diagram of the architecture of a system 100 provided by an embodiment of the present application. In Figure 3a the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices. A user can input data to the I / O interface 112 through the client device 140. The input data in the embodiment of the present application can include: various tasks to be scheduled, callable resources, and other parameters.

[0091] During the preprocessing of the input data by the execution device 110, or during the process of the computing module 111 of the execution device 110 performing calculations and other related processes (such as implementing the functions of the neural network in the present application), the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.

[0092] Finally, the I / O interface 112 returns the processing result to the client device 140 for the user.

[0093] It should be noted that the training device 120 can generate corresponding target models / rules based on different training data for different targets or tasks. The corresponding target models / rules can be used to achieve the above targets or complete the above tasks, thereby providing the results required by the user. Among them, the training data can be stored in the database 130 and comes from the training samples collected by the data acquisition device 160.

[0094] In Figure 3a In the case shown, the user can manually give input data, and this manual input can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 in the client device 140, and the specific presentation form can be specific ways such as display, sound, and action. The client device 140 can also be used as a data acquisition end to collect the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as shown in the figure as new sample data and store them in the database 130. Of course, it is also possible not to collect through the client device 140, but directly store the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as shown in the figure as new sample data in the database 130.

[0095] It should be noted that Figure 3a is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 3a , the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110. As Figure 3a shown, a neural network can be trained according to the training device 120.

[0096] The embodiments of the present application also provide a chip, which includes a neural network processor NPU. The chip can be set in the execution device 110 as shown in Figure 3a to complete the computing work of the computing module 111. The chip can also be set in the training device 120 as shown in Figure 3a to complete the training work of the training device 120 and output the target model / rule.

[0097] The neural network processor NPU, the NPU is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are allocated by the main CPU. The core part of the NPU is the arithmetic circuit, and the controller controls the arithmetic circuit to extract data from the memory (weight memory or input memory) and perform arithmetic operations.

[0098] In some implementations, the arithmetic circuit includes multiple processing units (process engines, PEs) internally. In some implementations, the arithmetic circuit is a two-dimensional systolic array. The arithmetic circuit can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit is a general matrix processor.

[0099] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in an accumulator.

[0100] The vector calculation unit can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. For example, the vector calculation unit can be used for network calculations in non-convolutional / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.

[0101] In some implementations, the vector calculation unit can store the processed output vector into a unified buffer. For example, the vector calculation unit can apply a non-linear function to the output of the arithmetic circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as an activation input to the arithmetic circuit, such as for use in subsequent layers in a neural network.

[0102] The unified memory is used to store input data and output data.

[0103] The weight data directly transfers the input data in the external memory to the input memory and / or the unified memory, stores the weight data in the external memory into the weight memory, and stores the data in the unified memory into the external memory through a direct memory access controller (DMAC).

[0104] The bus interface unit (BIU) is used to interact between the main CPU, the DMAC, and the instruction fetch memory through a bus.

[0105] An instruction fetch buffer connected to the controller, which is used to store instructions used by the controller;

[0106] The controller is used to call the instructions cached in the instruction fetch memory to control the operation process of the arithmetic accelerator.

[0107] Generally, the unified memory, input memory, weight memory, and instruction fetch memory are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDRSDRAM), a high bandwidth memory (HBM), or other readable and writable memories.

[0108] Since the embodiments of this application involve a large number of neural network applications, for the sake of easy understanding, relevant terms and related concepts such as neural networks involved in the embodiments of this application will be introduced below.

[0109] (1) Neural network

[0110] A neural network can be composed of neural units. A neural unit can be an arithmetic unit with xs and intercept 1 as inputs. The output of this arithmetic unit can be:

[0111]

[0112] where s = 1, 2,... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0113] The work of each layer in the neural network can be expressed by the mathematical formula described as follows: The operation of each layer in a neural network at the physical level can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimension increase / dimension decrease; 2. Enlargement / reduction; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by , the operation of 4 is completed by +b, and the operation of 5 is implemented by a(). The reason for using the word "space" here is that the objects to be classified are not individual things, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training a neural network is to finally obtain the weight matrices of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially a process of learning the way to control space transformation, and more specifically, learning the weight matrix.

[0114] Because it is hoped that the output of the neural network is as close as possible to the value that is really wanted to be predicted, the weight vector of each layer of the neural network can be updated by comparing the predicted value of the current network and the really wanted target value, and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the predicted value of the network is too high, adjust the weight vector to make it predict lower, and keep adjusting until the neural network can predict the really wanted target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.

[0115] (2) Backpropagation algorithm

[0116] The neural network can use the back propagation (BP) algorithm to correct the magnitudes of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, forward propagating the input signal until the output will generate an error loss, and updating the parameters in the initial neural network model by backpropagating the error loss information, so as to converge the error loss. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0117] The method provided in this application will be described below from the training side of the neural network and the application side of the neural network.

[0118] The training method of the neural network provided in the embodiments of this application involves data processing, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as sample pairs in this application), and finally obtains a trained noise reduction classification model; moreover, the data noise reduction classification method provided in the embodiments of this application can use the above-trained noise reduction model, input the input data (such as the data to be processed in this application) into the trained noise reduction classification model, and obtain the output data. It should be noted that the model training method and the data noise reduction classification method provided in the embodiments of this application are inventions generated based on the same concept, and can also be understood as two parts of a system, or two stages of an overall process: such as the model training stage and the model application stage.

[0119] In daily life, sparse time series data is everywhere. Sparse time series data is a set of non-dense data sequences arranged in chronological order. Among them, the sparsity in sparse time series data is relative to density. Common dense data is, for example, an image, and common sparse time series data includes, for example, the human body bone point coordinate data of multiple frames input in a motion sensing game, human electrocardiogram data, and attitude data obtained by an IMU. In practical applications, a large amount of useful information can be obtained by classifying and identifying sparse time series data.

[0120] For example, the human body bone point coordinate data has very wide applications in the field of motion sensing interaction. In a motion sensing game scenario, it is usually achieved by using a motion sensing interaction device to obtain the human body bone point coordinate data and identify the hand bones to realize the virtual interaction of both hands.

[0121] For another example, human electrocardiogram (ECG) data is one of the important bases for doctors to diagnose cardiovascular diseases. With the popularization of artificial intelligence technology, the application of deep learning based on big data to achieve automated analysis and diagnosis can break through the limitations of the accuracy and application scope of traditional statistical models. The automated analysis and diagnosis technology based on classifying and identifying ECG data can more deeply interpret the massive data of patients and achieve accurate classification of patients.

[0122] For yet another example, IMUs have been very popular in intelligent terminals such as smartphones and tablets, as well as wearable devices such as virtual reality helmets. The IMU data collected by an IMU can be used as an important basis for identifying the motion state of the human body. For example, based on the IMU data collected by the gyroscope on a smart bracelet or smart watch, the motion state of the wearer can be identified.

[0123] The various application scenarios mentioned above are all based on the ability to perform classification and identification on time-series data. However, sparse time-series data in the real environment will be affected by various noises, which will cause a significant decrease in the classification accuracy of sparse time-series data. For example, for human skeleton point coordinate data, due to occlusion of the body, the coordinates of some key skeleton points of the human body will jitter or be missing, resulting in incorrect prediction of action categories. For ECG data, interference sources are everywhere, whether in hospitals, ambulances, airplanes, ships, clinics or at home. For IMU data, common systematic errors of IMUs include constant zero-bias errors, scale factor errors, non-coincidence and non-orthogonal errors, non-linear errors, and temperature errors after the IMU is powered on. These errors will affect the subsequent classification and identification tasks to varying degrees. Therefore, time-series data denoising is an issue that cannot be ignored in various application scenarios.

[0124] Taking human skeleton point coordinate data as an example, the action recognition method based on skeleton points takes the human skeleton point coordinates as the direct input. Due to its small data volume, obvious semantic features of the input data, and high robustness to complex environments, it has a wide range of applications in fields such as human-computer interaction, intelligent monitoring, and service robots. Similar to the noise in the field of image processing, the skeleton point coordinate data in the user scenario usually has noise problems such as missing or jittering of key skeleton points due to occlusion or lighting. However, existing human action recognition methods require the input information to be complete and without missing data. Therefore, when the above-mentioned skeleton point coordinate data with noise problems is used as the input, the phenomenon of incorrect human action recognition is likely to occur.

[0125] Based on this, in the related art, usually, after denoising the sparse time-series data, the denoised sparse time-series data is classified and recognized based on the original classification method. To achieve the denoising of the above-mentioned sparse time-series data, the existing data denoising methods usually achieve the denoising of the sparse time-series data by training a specific denoising network. For example, when the input is human body bone point coordinate data, a smoothing loss that maintains visual rationality is used to guide the update of network parameters during the training stage of the denoising network, and finally a network for realizing data denoising is obtained. Generally speaking, the denoising network in the related art usually aims to achieve the training of the denoising network based on methods such as data without pose ambiguity, time-series smoothing, and / or maintaining visual rationality. For an end-to-end denoising classification task, its ultimate goal is the classification accuracy of the data. Therefore, there is a gap between the optimization goal of the denoising network in the related art and the ultimate classification accuracy goal, resulting in a low classification and recognition accuracy of the denoised sparse time-series data in the related art, and it is difficult to ensure the normal recognition of the sparse time-series data.

[0126] In view of this, the embodiments of the present application provide a model training method and related device. During the process of training a denoising classification network, a first loss function is obtained based on the output of the denoising network of the denoising classification network for the noise-free data and the output of the noise data at the intermediate layer of the denoising network, a second loss function is obtained based on the outputs of the noise-free data and the noise data at the entire denoising classification network, and the denoising classification network is trained based on the first loss function and the second loss function to obtain a target network. By training the denoising classification network and obtaining the loss function based on the outputs of the noise-free data and the noise data at the denoising network and the denoising classification network, it is possible to ensure that the denoising goal and the classification accuracy goal are consistent, and enable the network to learn more complete global features during the denoising stage, enhance the network's ability to suppress local noise and perturbations, and improve the classification and recognition accuracy of the network for sparse time-series data.

[0127] The model training method provided by the embodiments of the present application can be applied to a terminal, which is a device capable of performing model training. After training a target network based on the model training method provided by the embodiments of the present application, the terminal can perform noise reduction classification on the acquired time series data based on the target network. Exemplarily, the terminal can be, for example, a smart TV, a personal computer (PC), a laptop, a server, a mobile phone, a tablet computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The terminal can be a device running an Android system, an IOS system, a Windows system, or other systems.

[0128] For ease of understanding, a specific application scenario provided by the embodiments of the present application will be introduced below with reference to the accompanying drawings. Refer to Figure 3b , Figure 3b which is a schematic diagram of an application scenario provided by the embodiments of the present application. As Figure 3b shown, a possible application scenario is that a smart TV classifies and recognizes the user's intention based on time series data such as human body posture, head posture, eye gaze direction, facial expression, gesture, or voice, and performs corresponding operations based on the classified and recognized user intention, thereby completing the interaction with the user.

[0129] Refer to Figure 3c , Figure 3cA schematic diagram of a specific application of time series data provided by an embodiment of this application. Specifically, data acquisition devices such as cameras and microphones can be installed on a smart TV. Through data acquisition devices such as cameras and microphones, the smart TV can obtain relevant data of users around the smart TV. For example, the smart TV obtains time series data such as hand key point data describing the user's gesture intention, human body pose data describing the user's body movement intention, head pose data describing the user's gaze direction intention, and / or face key point data describing the user's expression intention through the camera. For another example, the smart TV obtains audio data describing the user's voice intention through the microphone. In addition, a communication device can also be set on the smart TV, which can receive data sent by an external data acquisition device. For example, a Bluetooth module is set on the smart TV, which can receive inertial measurement unit data sent by an external smart watch or smart bracelet. Based on the obtained time series data for describing the user's intention, the smart TV can classify and identify the user's intention through a built-in classification model, obtain a classification prediction result of the user's intention, and thus complete corresponding interaction response operations.

[0130] It can be understood that in actual applications, the time series data obtained by the smart TV can be one or more types of the foregoing data, and the smart TV can classify and identify the user's intention through a classification model based on the one or more types of data obtained. Simply put, for the classification model pre-set in the smart TV, the input data of the classification model can be one or more types of data. For example, the input data of the classification model is these two types of data: the head pose data describing the user's gaze direction intention and the hand key point data describing the user's gesture intention mentioned above; for another example, the input data of the classification model is this one type of data: the audio data describing the user's voice intention. When the input data is multiple types of data, the classification model can classify and identify the user's intention based on the combination of multiple types of data, and obtain a classification prediction result of the user's intention.

[0131] Exemplarily, in the case where the input data is the head pose data describing the user's gaze direction intention and the hand key point data describing the user's gesture intention, when the head pose data specifically indicates that the user's gaze direction is the direction where the smart TV is located and the hand key point data specifically indicates waving to the right, the classification model can predict that the user's intention is to switch the TV channel.

[0132] Reference can be made to Figure 4 , Figure 4 A schematic flowchart of a model training method provided by an embodiment of this application. As Figure 4 shown, the model training method includes the following steps 401-406.

[0133] Step 401: Obtain a sample pair, where the sample pair includes noisy data and noiseless data corresponding to the noisy data.

[0134] In this embodiment, the terminal can obtain a sample set including multiple sample pairs, and each sample pair in the sample set includes a pair of noisy data and noiseless data. Among them, the noiseless data in the sample pair corresponds to the noisy data, that is, the noiseless data corresponding to the noisy data in the sample pair can be obtained by removing the noise from the noisy data in the sample pair, or in other words, the noisy data corresponding to the noiseless data can be obtained by adding noise to the noiseless data in the sample pair. Optionally, the noisy data in the sample pair can be sparse time-series data. For example, the noisy data is skeleton point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data. Among them, the fault diagnosis data can be, for example, device operation data or power grid operation data, such as data of voltage, current, frequency, or waveform generated in real time during the operation of the power grid.

[0135] Exemplarily, the noisy data in the sample pair can be one or more types of data. For example, the noisy data is only inertial measurement unit data, or the noisy data is head pose data and hand key point data. For the sake of easy description, the following will take the noisy data in the sample pair as one type of data as an example to introduce the model training method of the embodiment of the present application.

[0136] In practical applications, after obtaining the noiseless data, the terminal can add random noise to the noiseless data to obtain the corresponding noisy data, so as to construct a sample pair. During the training process, the terminal can successively obtain sample pairs from the sample set to implement the training of the noise reduction classification network.

[0137] In a possible example, the terminal for executing the model training method in this embodiment can be, for example, a server. By executing Figure 4 the corresponding model training method, a trained model can be obtained. Then, the trained model can be deployed in the smart TV before the smart TV leaves the factory; or after the smart TV leaves the factory, the smart TV can connect to the server through the network and obtain the trained model on the server by downloading or updating, so as to implement the deployment of the trained model on the smart TV.

[0138] Step 402: Input the noiseless data into the noise reduction classification network to obtain first output data and second output data.

[0139] In this embodiment, the noise reduction classification network includes a noise reduction network and a classification network. The noise reduction network is connected to the classification network, and the input of the noise reduction network is the input of the noise reduction classification network, and the output of the noise reduction network is the input of the classification network. The noise reduction network is used to perform noise reduction processing on the data input to the noise reduction classification network to obtain the noise-reduced data. After inputting the noise-free data into the noise reduction classification network, the first output data and the second output data can be obtained. Among them, the first output data is the output of the noise reduction network, and the second output data is the output of the classification network, that is, the second output data is the classification result corresponding to the noise-free data output by the classification network.

[0140] Optionally, the noise reduction network can be an autoencoder, and the autoencoder includes an encoder and a decoder. Among them, the autoencoder, also known as the automatic encoder, is an artificial neural network that can learn an efficient representation of the input data through unsupervised learning. In practical applications, by adding constraint conditions to the autoencoder, the autoencoder can perform noise reduction processing on the noise data. In the autoencoder, the encoder is used to perform compression encoding on the input data, and the decoder is used to perform data reconstruction on the data output by the encoder. By performing compression encoding on the input data by the encoder in the autoencoder, the key information of the input data is obtained, and then the decoder in the autoencoder performs data reconstruction on the compressed data to restore the input data based on the key information of the input data, and the restored data achieves noise elimination, that is, data noise reduction processing can be achieved based on the autoencoder. Exemplarily, the encoder and the decoder can be a recurrent neural network (RNN) including multiple convolutional layers.

[0141] Among them, the above-mentioned first output data can be the output data obtained after the encoder in the noise reduction network processes the noise-free data after inputting the noise-free data into the noise reduction classification network. Generally, the data processed by the encoder in the noise reduction network is also called the content vector. The content vector refers to the output of the encoder in the autoencoder, generally a set of high-dimensional feature vectors.

[0142] The classification network is used to obtain the data processed by the noise reduction network and output a set of probability value vectors corresponding to the data input into the classification network. Each element value in the vector is the probability of the input data corresponding to a category. Generally, the category with the highest probability is the category to which the input data belongs. That is to say, the classification network is used to classify the obtained data to obtain the corresponding category of the data. Exemplarily, when the data input into the noise reduction classification network is the IMU data collected by a smart bracelet, the classification network is used to classify the IMU data among the four categories of walking, running, cycling, and climbing stairs. For example, when the probability labels output by the classification network are {0.1, 0.7, 0.15, 0.05}, it can be determined that the category with a probability of 0.7 (i.e., running) is the category to which the IMU data belongs.

[0143] Among them, the above second output data may be the data output by the classification network after the noise-free data is input into the noise reduction classification network and processed by the noise reduction network and the classification network.

[0144] Step 403: Input the noise data into the noise reduction classification network to obtain a third output data and a fourth output data. The third output data is obtained based on the intermediate layer of the noise reduction network, and the fourth output data is the output of the classification network.

[0145] In this embodiment, after inputting the noise data corresponding to the above noise-free data into the noise reduction classification network, the terminal can obtain the third output data extracted from the intermediate layer of the noise reduction network in the noise reduction classification network. The third output data is the feature data extracted from the intermediate layer of the noise reduction network. In addition, the terminal can also obtain the fourth output data output by the classification network in the noise reduction classification network. The fourth output data is the classification result corresponding to the noise data output by the classification network.

[0146] Optionally, the third output data may be obtained based on the intermediate layer of the encoder in the noise reduction network. Exemplarily, the encoder may include one or more intermediate layers. The terminal can obtain the feature data output by each intermediate layer in the encoder and use these feature data as the third output data. For example, if the encoder is a recurrent neural network (RNN) including three convolutional layers, the intermediate layer of the encoder is the latter two convolutional layers in the encoder, and the data output by these two convolutional layers is the third output data.

[0147] It can be understood that there is no limitation on the execution order between step 402 and step 403. In practical applications, step 402 may be executed first, or step 403 may be executed first. This embodiment does not specifically limit the execution order of step 402 and step 403.

[0148] Step 404: Determine a first loss function based on the first output data and the third output data, where the first loss function is used to represent the difference between the first output data and the third output data.

[0149] Since both the first output data and the third output data are obtained based on the noise reduction network, after obtaining the first output data and the third output data, the first loss function can be determined based on the first output data and the third output data to characterize the difference between the first output data and the third output data. In this way, by obtaining the first loss function between the outputs of the noise-free data and the noisy data in the noise reduction network respectively, and training the noise reduction classification network based on this first loss function, the noise reduction classification network can learn global features closer to the noise-free data during the noise reduction process, thereby enhancing the ability of the noise reduction network to suppress local noise and disturbances.

[0150] Optionally, when the third output data is obtained based on the intermediate layer of the encoder in the noise reduction network, the third output data includes the data output by one or more intermediate layers of the encoder. In this case, the difference value between the data output by each intermediate layer and the first output data can be obtained, and the first loss function can be obtained by summing the difference values between the data output by multiple intermediate layers and the first output data.

[0151] Optionally, in a possible embodiment, in step 403, after the terminal inputs the noisy data into the noise reduction classification network and obtains the feature data output by the intermediate layer of the noise reduction network, the terminal can divide the obtained feature data into multiple sub-feature data to obtain the third output data, and the third output data includes multiple sub-feature data.

[0152] Exemplarily, when the noisy data is sparse time-series data, the terminal can evenly divide the feature data into multiple sub-feature data in chronological order to obtain the third output data, and the length of the time period corresponding to each sub-feature data in the multiple sub-feature data is the same. For example, when the feature data output by the intermediate layer of the noise reduction network is the data in the time period from T0 to T3, the terminal can divide the feature data in chronological order to obtain the first sub-feature data in the time period from T0 to T1, the second sub-feature data in the time period from T1 to T2, and the third sub-feature data in the time period from T2 to T3, where the lengths of the time periods corresponding to the first sub-feature data, the second sub-feature data, and the third sub-feature data are the same.

[0153] Then, after obtaining multiple sub-feature data through partitioning, the terminal calculates the difference value between each sub-feature data in the third output data and the first output data, and determines the first loss function based on the sum of the difference values between the multiple sub-feature data and the first output data. In this way, by extracting the feature data corresponding to the noise data output by different intermediate layers in the noise reduction network, evenly partitioning to obtain temporal local features, constructing temporal global features based on the feature data corresponding to the noise-free data output by the noise reduction network, and constructing a loss function based on the temporal local features and the temporal global features, it is possible to guide the noise reduction network to learn global temporal features more fully, thereby enhancing the ability of the noise reduction network to suppress local noise and disturbances.

[0154] For time series data, the time series data is coherent, and adjacent time series data has a certain correlation. Constructing a loss function based on the local temporal features corresponding to the noise data and the global temporal features corresponding to the noise-free data can guide the noise reduction network to learn richer global temporal information, thereby enhancing the ability of the noise reduction network to suppress local noise and disturbances. For example, when the noise data is time series data with a small period of missing data, constructing a loss function and training the noise reduction network based on the above method can enable the noise reduction network to restore the missing data based on the learned global temporal information, so as to achieve noise reduction of the noise data with a good noise reduction effect.

[0155] Optionally, in a possible embodiment, since the first output data is the data output by the encoder in the noise reduction network, and the third output data is the data output by the intermediate layer of the encoder of the noise reduction network, their dimensions are not the same. Therefore, before calculating the first loss function, a dimension alignment operation can be performed on the first output data and the third output data to make their dimensions the same, and then the difference value between them can be calculated.

[0156] Exemplarily, the terminal performs a dimension alignment operation on each sub-feature data in the first output data and the third output data respectively to obtain the dimension-aligned first output data and third output data. Specifically, the terminal can pre-construct multiple dimension alignment sub-networks. By inputting the first output data into one of the dimension alignment sub-networks and inputting each sub-feature data in the third output data into other corresponding dimension alignment sub-networks, the dimension-aligned first output data and third output data are obtained. Among them, the dimension alignment sub-network takes the gated recurrent unit (GRU) as the basic unit, and the dimension alignment sub-network is composed of multiple layers of GRUs. The dimension alignment sub-network can change the dimension of the input data. After the dimension alignment of the third output data and the first output data, the terminal determines the difference value between each sub-feature data in the dimension-aligned third output data and the dimension-aligned first output data.

[0157] Exemplarily, the process of determining the first loss function based on the first output data and the third output data can be shown as Formula 1 below.

[0158]

[0159] Wherein, represents the first loss function; l represents the l-th intermediate layer of the encoder; i represents the i-th temporal local feature; log() represents the logarithmic function; exp() represents the exponential function; represents the i-th temporal local feature obtained by evenly dividing the features output from the intermediate layer of the l-th layer of the encoder; ρ l represents the dimension alignment sub-network corresponding to the temporal local feature of the l-th intermediate layer of the encoder; represents the dimension alignment sub-network corresponding to the content vector output by the encoder, that is, the dimension alignment sub-network corresponding to the first output data.

[0160] Specifically, the dimension alignment sub-network using GRU as the basic unit can be formally expressed as GRU(x, num_layers, out_dim), where num_layers is the number of network layers and out_dim represents the expected output dimension.

[0161] Step 405, determine a second loss function according to the second output data and the fourth output data, where the second loss function is used to represent the difference between the second output data and the fourth output data and the true class label of the noise-free data.

[0162] Wherein, the second output data is the output of the noise reduction classification network after inputting the noise-free data, that is, the class result corresponding to the noise-free data predicted by the noise reduction classification network. The fourth output data is the output of the noise reduction classification network after inputting the noisy data, that is, the class result corresponding to the noisy data predicted by the noise reduction classification network. Actually, the true class labels corresponding to the noise-free data and the noisy data are the same. Therefore, the terminal can determine the second loss function by obtaining the first difference value between the second output data and the true class label of the noise-free data and the second difference value between the fourth output data and the true class label of the noise-free data. For example, by obtaining the sum of the first difference value between the second output data and the true class label of the noise-free data and the second difference value between the fourth output data and the true class label of the noise-free data, the second loss function is obtained.

[0163] Taking the noise-free data as the IMU data as an example, when the category corresponding to the IMU data is running, the true category label corresponding to the IMU data can be {0, 1, 0, 0}. Among them, the four element values in the true category label respectively represent the four categories of walking, running, cycling, and climbing stairs. Suppose that after inputting the noise-free data into the noise reduction classification network, the obtained second output data is {0.1, 0.7, 0.15, 0.05}. Then, the process of obtaining the first difference value between the second output data and the true category label of the noise-free data is actually the process of obtaining the difference value between the vector {0.1, 0.7, 0.15, 0.05} and the vector {0, 1, 0, 0}.

[0164] Exemplarily, the process of determining the second loss function based on the second output data and the fourth output data can be shown in the following formula 2.

[0165] L 2 = L mul (X normal ) + L mul (X noise ) Formula 2

[0166] Among them, L 2 represents the second loss function; L mul (X normal ) represents the cross-entropy loss between the second output data and the true category label of the noise-free data; L mul (X noise ) represents the cross-entropy loss between the fourth output data and the true category label of the noise-free data.

[0167] In formula 2,

[0168] where p(X noise ) represents the probability that the predicted category of the input sample X noise is y noise,i .

[0169] In formula 2,

[0170] where p(X normal ) represents the probability that the predicted category of the input sample X normal is y normal,i .

[0171] In practical applications, in addition to representing the difference value between the output data and the true category label by obtaining the cross-entropy loss, it can also be represented by other difference measurement methods. This embodiment does not make specific limitations on this.

[0172] Step 406: Train the noise reduction classification network based on at least the first loss function and the second loss function until a preset training condition is met, and obtain the target network.

[0173] After obtaining the first loss function and the second loss function, the terminal can obtain the total loss function based on the first loss function and the second loss function. The total loss function can be the sum of the first loss function and the second loss function, or it can be obtained by adding the product of the first loss function and the first proportionality coefficient to the product of the second loss function and the second proportionality coefficient. After obtaining the total loss function, the terminal trains the noise reduction classification network based on the total loss function. The process of the terminal training the noise reduction classification network based on the total loss function includes: the terminal adjusts the parameters in the noise reduction classification network (including the noise reduction network and the classification network) based on the value of the total loss function, and repeats steps 401 - 406, so as to continuously adjust the parameters in the noise reduction classification network until the obtained total loss function is less than a preset threshold, then it can be determined that the preset training condition is met, and the target network is obtained. The target network is the trained noise reduction classification network and can be used for subsequent data noise reduction and classification.

[0174] Optionally, during the training process of the noise reduction classification network, the terminal can update the parameters of the noise reduction classification network through the error backpropagation algorithm. Briefly speaking, the terminal can, through the error backpropagation algorithm, correct the magnitudes of the parameters in the initial noise reduction classification network during the training process of the noise reduction classification network, so that the reconstruction error loss of the noise reduction classification network becomes smaller and smaller. Specifically, forward-propagate the input signal until an output error loss is generated, and update the parameters in the initial noise reduction classification network by backpropagating the error loss information, so as to make the error loss converge. The backpropagation algorithm is a reverse propagation movement dominated by the error loss, aiming to obtain the parameters of the optimal neural network model, such as the weight matrix.

[0175] Optionally, in a possible embodiment, during the training process of the noise reduction classification network, a binary classifier can also be introduced. The binary classifier can perform binary classification prediction on the data input into the noise reduction classification network based on the features extracted by the noise reduction classification network, that is, predict whether the data input into the noise reduction classification network is noise data or noise-free data. Briefly speaking, the binary classification result output by the binary classifier is a probability value vector, and the vector includes two element values, respectively representing the probabilities of the input data belonging to the noise-free data category and the noise data category. Then, based on the binary classification result output by the binary classifier and the true binary classification label corresponding to the input data, determine the third loss function, and the third loss function is used to obtain the total loss function together with the first loss function and the second loss function, that is, the third loss function is also used for the training of the noise reduction classification network.

[0176] Specifically, the terminal obtains a first feature and predicts a binary classification result corresponding to the noise-free data based on the first feature, obtaining a first prediction result. Among them, the first feature is extracted by the classification network in the noise reduction classification network after the noise-free data is input into the noise reduction classification network. For example, after the noise-free data is input into the noise reduction classification network, the classification network in the noise reduction classification network extracts features from the input data and performs multi-class prediction based on the extracted features to classify the noise-free data. Then, the terminal can obtain the first feature by acquiring the features extracted by the classification network. Then, the terminal inputs the obtained first feature into a binary classifier to predict the binary classification result corresponding to the noise-free data.

[0177] The terminal obtains a second feature and predicts a binary classification result corresponding to the noisy data based on the second feature, obtaining a second prediction result. Among them, the second feature is extracted by the classification network in the noise reduction classification network after the noisy data is input into the noise reduction classification network. Similarly, after the noisy data is input into the noise reduction classification network, the classification network in the noise reduction classification network extracts features from the input data and performs multi-class prediction based on the extracted features. The terminal can obtain the second feature by acquiring the features extracted by the classification network. Then, the terminal inputs the obtained second feature into a binary classifier to predict the binary classification result corresponding to the noisy data.

[0178] After obtaining the first prediction result and the second prediction result, the terminal determines a third difference value between the first prediction result and the true binary classification label of the noise-free data, and a fourth difference value between the second prediction result and the true binary classification label of the noisy data, and determines a third loss function based on the third difference value and the fourth difference value.

[0179] Finally, the terminal trains the noise reduction classification network at least according to the first loss function, the second loss function, and the third loss function. Among them, the total loss function can be the sum of the first loss function, the second loss function, and the third loss function, or the product of the first loss function and the first proportionality coefficient, the product of the second loss function and the second proportionality coefficient, and the sum of the third loss function and the third proportionality coefficient. In practical applications, the first proportionality coefficient, the second proportionality coefficient, and the third proportionality coefficient can be adjusted according to the accuracy requirements of noise reduction classification, which are not specifically limited here.

[0180] Exemplarily, the process of determining the third loss function based on the first prediction result and the second prediction result can be shown in Formula 3 below.

[0181] L 3 =L bin (X normal )+L bin (X noise ) Formula 3

[0182] Among them, L 3 represents the third loss function; L bin (X normal ) represents the cross-entropy loss between the first prediction result and the true binary classification label of the noise-free data; L bin (X noise ) represents the cross-entropy loss between the second prediction result and the true binary classification label of the noisy data.

[0183] In Equation 3,

[0184] Among them, p(X normal ) represents the probability that the predicted label for the corresponding input X normal is b i , and N is the number of batch samples.

[0185] In Equation 3,

[0186] Among them, p(X noise ) represents the probability that the predicted label for the corresponding input X noise is b i , and N is the number of batch samples.

[0187] In this embodiment, by introducing a binary classifier in the training stage and based on the features extracted by the classification network in the noise reduction classification network, the binary classifier is used to predict the binary classification result of the input data, and the loss function corresponding to the binary classification result is obtained. By introducing the loss function corresponding to the binary classification result on the basis of the original loss function, an additional evaluation dimension can be introduced, so that the trained noise reduction classification network can have an adaptive noise reduction classification scale for different types of input data, and the classification accuracy of the noise reduction classification network can be improved. For example, for noise-free data and noisy data, the noise reduction classification network can learn different noise reduction classification scales, so that whether the input data is noise-free data or noisy data, the trained noise reduction classification network can have a high noise reduction classification accuracy.

[0188] For ease of understanding, the model training method provided by the embodiments of the present application will be described below with specific examples.

[0189] Reference can be made to Figure 5 , Figure 5 which is a schematic structural diagram of a noise reduction classification network provided by an embodiment of the present application. As shown in Figure 5As shown in the figure, a training set 501 and a test set 502 are stored on the server 500, and a noise reduction classification network 503 is also deployed on the server 500. Among them, the training set 501 includes multiple sample pairs for training the noise reduction classification network 503; the test set 502 also includes multiple sample pairs for testing the performance of the trained noise reduction classification network.

[0190] The noise reduction classification network 503 includes a noise reduction network 5031 and a classification network 5032. The noise reduction network 5031 includes an encoder 50311 and a decoder 50312. The encoder 50311 is used to perform compression encoding on the input data, and the decoder 50312 is used to reconstruct the data output by the encoder 50311. Based on the encoder 50311 and the decoder 50312, noise reduction processing of the input data can be achieved. The classification network 5032 includes a feature extraction module 50321 and a classifier 50322. The feature extraction module 50321 is used to extract features from the data output by the decoder 50312, and the classifier 50322 is used to perform multi-class prediction based on the features extracted by the feature extraction module 50321 to obtain a prediction result, that is, the predicted class corresponding to the input noise reduction classification network 503.

[0191] In addition, to implement the training of the noise reduction classification network 503, a local-global feature association module 504 and a hybrid classifier 505 are also deployed on the server 500. Among them, the local-global feature association module 504 is used to obtain the local-global feature association loss based on the output of the noise-free data in the noise reduction network 5031 and the output of the noisy data in the noise reduction network 5031, that is, the above-mentioned first loss function.

[0192] Specifically, reference can be made to Figure 6 , Figure 6 which is a schematic structural diagram of the local-global feature association module and the hybrid classifier provided by the embodiments of the present application. As Figure 6 shown, the local-global feature association module 504 includes a feature division module 5041 and a dimension alignment sub-network 5042. The feature division module 5041 is used to obtain the output data (i.e., feature data) corresponding to the intermediate layer of the encoder 50311 of the noisy data, and uniformly divide the obtained feature data in chronological order to obtain sub-feature data. The dimension alignment sub-network 5042 is used to perform dimension alignment on the uniformly divided sub-feature data and the output data of the noise-free data in the encoder 50311. In this way, the local-global feature association module 504 can obtain the local-global feature association loss based on the dimension-aligned data, that is, the above-mentioned first loss function.

[0193] The hybrid classifier 505 is used to calculate the multi-class classification loss, i.e., the second loss function above, based on the output of the classifier 50322 and the true class labels corresponding to the original input data of the noise reduction classification network 503. The hybrid classifier 505 is also used to calculate the binary classification loss, i.e., the third loss function above, based on the output of the feature partitioning module 5041 and the true binary classification labels corresponding to the original input data of the noise reduction classification network 503. As Figure 6 shown, the hybrid classifier 505 includes a binary classifier 5051 and a multi-class classifier 5052. The binary classifier 5051 is used to calculate the binary classification loss, and the multi-class classifier 5052 is used to calculate the multi-class classification loss. Exemplarily, the binary classifier 5051 can be composed of two fully connected layers plus a softmax layer, and the multi-class classifier 5052 can be a human action recognition (HAR) classifier composed of a spatio-temporal graph neural network.

[0194] During the training process, the server 500 updates the parameters in the noise reduction classification network 503 through the error backpropagation algorithm based on the first loss function obtained by the local-global feature association module 504 and the second and third loss functions obtained by the hybrid classifier 505 until the preset training conditions are met, and a target network is obtained.

[0195] The following will be based on Figure 5 the noise reduction classification network shown, and will introduce in detail the process of training the noise reduction classification network. Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a method for training a noise reduction classification network provided by an embodiment of the present application. As Figure 7 shown, in the experiment, the human skeleton key point coordinate data is input into the noise reduction classification network 503. The corresponding local-global feature loss and hybrid classification loss are obtained based on the local-global feature association module 504 and the hybrid classifier 505, and the total loss function is calculated according to the local-global feature loss and the hybrid classification loss. Then, the network parameter update module 802 updates the parameters of the noise reduction network 5031 and the classification network 5032 in the noise reduction classification network 503 based on the calculated total loss function by using the error backpropagation algorithm, thereby realizing the training of the noise reduction classification network 503.

[0196] Specifically, the process of training the noise reduction classification network will be described in detail below.

[0197] 1. Before the training starts, prepare the training set and the test set.

[0198] Before starting to train the noise reduction classification network, it is necessary to prepare the dataset on the server, namely the training set and the test set. Among them, both the training set and the test set include multiple sample pairs, and each sample pair includes noise-free data and noisy data. Moreover, the noise-free data and the noisy data in the sample pair are both time series data. In practical applications, the dataset can be prepared according to the scenario where the noise reduction classification network is applied. For example, when the noise reduction classification network is used for human action classification and recognition, the training set and the test set composed of bone point coordinate data are prepared.

[0199] Specifically, the process of constructing the training set and the test set can be as Figure 8 shown. Figure 8 This is a schematic diagram of generating noisy data provided by an embodiment of the present application. The terminal can add random noise, such as noise of different degrees, to the noise-free data after obtaining the noise-free data to obtain the corresponding noisy data, thereby constructing a sample pair. After constructing multiple sample pairs, a part of the sample pairs can be divided into the training set, and the other part of the sample pairs can be divided into the test set. Among them, the ratio of the sample pairs included in the training set and the test set can be, for example, 4:1.

[0200] II. Construct the local-global feature association loss.

[0201] In the training stage, the server obtains the feature data output by the middle layer of the encoder when the noisy data in the sample pair is input, and evenly divides the obtained feature data to obtain time series local features. The server also obtains the content vector output by the encoder when the noise-free data in the sample pair is input, that is, the time series global feature. Then, the server aligns the dimension of the divided time series local features and the content vector corresponding to the noise-free data through the feature dimension alignment sub-network, and constructs the local-global feature association loss, that is, the above-mentioned first loss function. Among them, the server can construct the local-global feature association loss based on the above formula 1, and specifically can refer to the above formula 1.

[0202] Exemplarily, reference can be made to Figure 9a , Figure 9a This is a schematic flowchart of constructing the local-global feature association loss provided by an embodiment of the present application. As Figure 9a shown, X normal represents the noise-free data in the sample pair, X noise represents the noisy data in the sample pair, E1 represents the encoder, g represents the content vector output by the encoder, and D1 represents the decoder. After inputting the noise-free data, the server obtains the content vector g output by the encoder. After inputting the noisy data, the server then obtains the feature data output by the middle layer of the encoder (that is, Figure 9a(the intermediate layer time series features before partitioning), and evenly partition the obtained feature data to obtain sub-feature data (i.e., Figure 9a (the intermediate layer time series features after partitioning). Then, the server performs dimension alignment on the intermediate layer time series features after partitioning and the content vector g based on the dimension alignment sub-network, and constructs a local-global feature correlation loss.

[0203] III. Construct a hybrid classification loss.

[0204] The hybrid classifier constructed by the server includes a multi-class classifier and a binary classifier. For the noise-free data and noise data in the sample pair, the hybrid classifier can calculate the hybrid losses corresponding to the noise-free data X normal and the noise data X noise respectively, that is, the multi-class loss and the binary classification loss. Specifically, the process by which the server calculates the hybrid classification losses corresponding to the noise-free data X normal and the noise data X noise can be shown in the following formulas 4 and 5.

[0205] L hc1 = L bin (X normal ) + aL mul (X normal ) Formula 4

[0206] L hc2 = L bin (X noise ) + αL mul (X noise ) Formula 5

[0207] Among them, L hc1 is the hybrid classification loss corresponding to the noise-free data X normal ; L hc2 is the hybrid classification loss corresponding to the noise data X noise ; L bin (X normal ) represents the cross-entropy loss between the first prediction result and the true binary classification label of the noise-free data; α is a proportionality coefficient; L bin (X noise ) represents the cross-entropy loss between the second prediction result and the true binary classification label of the noise data; L mul (X normal ) represents the cross-entropy loss between the second output data and the true class label of the noise-free data; L mul (X noise ) represents the cross-entropy loss between the fourth output data and the true class label of the noise-free data.

[0208] It can be referred to Figure 9b ,Figure 9b Schematic diagram of a process for constructing local-global feature correlation loss and hybrid classification loss provided by an embodiment of the present application. X normal represents the noise-free data in the sample pair, X noise represents the noisy data in the sample pair, E1 represents the encoder, g represents the content vector output by the encoder, and D1 represents the decoder. After inputting the noise-free data, the server obtains the content vector g output by the encoder. After inputting the noisy data, the server obtains the feature data output by the intermediate layer of the encoder, and evenly divides the obtained feature data to obtain sub-feature data Then, the server aligns the divided intermediate layer temporal features and the content vector g based on the dimension alignment sub-network, and constructs the local-global feature correlation loss L ln . In addition, after the denoising network denoises the noise-free data and the noisy data, the feature extraction module in the classification network continues to perform feature extraction processing on the data output by the denoising network, and the extracted feature data are respectively input into the multi-class classifier (binary classifier C1) and the binary classifier (HAR classifier C2) to obtain the hybrid classification loss L of the noise-free data hc1 and the corresponding hybrid classification loss L of the noisy data hc2 .

[0209] IV. Construct the total loss function of the entire denoising classification network and complete the parameter update of the denoising classification network.

[0210] After obtaining the local-global feature correlation loss and the hybrid classification loss based on Steps III and IV, the server uses the sum of the ratios between the local-global feature correlation loss and the hybrid classification loss as the total loss function of the entire denoising classification network, and uses the error backpropagation algorithm to complete the parameter update of the denoising classification network. Among them, the total loss function of the entire denoising classification network can be shown in Formula 6 below.

[0211] L total = L ln + λ(L hc1 + L hc2 ) Formula 6

[0212] Among them, L total is the total loss function of the denoising classification network; L ln is the local-global feature correlation loss; λ is the proportionality coefficient; L hc1 is the hybrid classification loss corresponding to the noise-free data X normal ; L hc2 is the hybrid classification loss corresponding to the noisy data X noise .

[0213] After the noise reduction classification network is trained, the obtained target network can be deployed on terminals such as servers or smartphones to implement the noise reduction classification application of time series data.

[0214] Reference can be made to Figure 10 , Figure 10 which is a comparison schematic diagram of the existing solution provided in the embodiment of the present application and the solution of the present application. As Figure 10 shown, Figure 10 is a comparison chart of the training situations of the classification network in the existing solution and the noise reduction classification network trained by the method of the present application. Among them, Figure 10 the abscissa in represents the number of iterations in the training stage, and its unit is 103; the ordinate is the comparison of the loss function curves of the classification network in the existing solution and the noise reduction classification network trained by the method of the present application on the validation set during the training stage. According to Figure 10 it can be seen that when the number of iterations of the noise reduction classification network trained by the method of the present application is 20000, its validation set loss is already lower than that of the classification network in the existing solution. When the number of iterations is greater than 40000, the noise reduction classification network trained by the method of the present application begins to converge.

[0215] Reference can be made to Table 1, which is a comparison of the classification accuracies of the human action classification model ST-GCN in the existing solution and the noise reduction classification network trained by the model training method provided in the embodiment of the present application.

[0216] Table 1

[0217] Noise level = 0 Noise level = 1 Noise level = 3 Noise level = 5 Existing solution 81.57% 73.78% 57.76% 42.73% Solution of this application 84.49% 84.11% 83.28% 82.20%

[0218] In Table 1, the noise level n (n = 0, 1, 3, 5) means randomly selecting n bone joints on each frame of bone points in the normal human bone point coordinate data and adding random spatial translation noise. Among them, the spatial coordinates of the bone points after adding noise are restricted within the minimum bounding box determined by the normal bone point spatial coordinates. When the noise level = 0, the noise reduction classification network trained by the model training method provided in the embodiment of the present application can play a role in data augmentation and improve the classification accuracy of the model. As the noise level increases, the gap between the noise reduction classification network of the solution of the present application and the classification network in the existing solution also gradually increases.

[0219] The embodiments of the present application further provide a noise reduction and classification method, which includes: obtaining data to be classified; inputting the data to be classified into a target network to obtain a prediction result, where the prediction result is the classification result of the data to be classified. The data to be classified includes sparse time series data, and the sparse time series data includes skeletal point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data. Among them, the target network is used to perform noise reduction processing and classification on the data to be classified, and the target network is trained based on the model training method described in the above embodiments. For specific details, please refer to the introduction of the above embodiments and will not be elaborated here. Optionally, the target network may be deployed on a smart TV to perform noise reduction classification on the data obtained by the smart TV to obtain a classification prediction result, and the classification prediction result is specifically used to represent the user's intention. In this way, the smart TV can perform an interactive response operation related to the user's intention based on the obtained classification prediction result, such as performing a channel switching operation.

[0220] The above describes the model training method and the noise reduction and classification method provided by the embodiments of the present application. The following will introduce the devices for executing the methods mentioned in the above embodiments.

[0221] It can be referred to Figure 11 , Figure 11 is a schematic structural diagram of a model training device provided by the embodiments of the present application. As Figure 11As shown in the figure, the model training device includes: an acquisition unit 1101 and a processing unit 1102. The acquisition unit 1101 is configured to acquire a sample pair, where the sample pair includes noise data and noise-free data corresponding to the noise data; the processing unit 1102 is configured to input the noise-free data into a noise reduction classification network to obtain first output data and second output data, the noise reduction classification network includes a noise reduction network and a classification network, the first output data is the output of the noise reduction network, and the second output data is the output of the classification network; the processing unit 1102 is further configured to input the noise data into the noise reduction classification network to obtain third output data and fourth output data, the third output data is obtained based on an intermediate layer of the noise reduction network, and the fourth output data is the output of the classification network; the processing unit 1102 is further configured to determine a first loss function according to the first output data and the third output data, and the first loss function is used to represent the difference between the first output data and the third output data; the processing unit 1102 is further configured to determine a second loss function according to the second output data and the fourth output data, and the second loss function is used to represent the difference between the second output data and the fourth output data and the true class label of the noise-free data; the processing unit 1102 is further configured to train the noise reduction classification network at least according to the first loss function and the second loss function until a preset training condition is met to obtain a target network.

[0222] Optionally, in a possible implementation manner, the noise reduction network includes an encoder and a decoder, the encoder is configured to perform compression encoding on input data, and the decoder is configured to perform data reconstruction on the data output by the encoder; the first output data is the output of the encoder, and the third output data is obtained based on an intermediate layer of the encoder.

[0223] Optionally, the processing unit 1102 is further configured to: input the noise data into the noise reduction classification network to obtain feature data output by an intermediate layer of the noise reduction network; divide the feature data into multiple sub-feature data to obtain the third output data, and the third output data includes the multiple sub-feature data; determine a difference value between each sub-feature data in the third output data and the first output data; and determine the first loss function according to the difference value between each sub-feature data and the first output data.

[0224] Optionally, in a possible implementation manner, the processing unit 1102 is further configured to evenly divide the feature data into multiple sub-feature data in chronological order to obtain the third output data, and the length of the time period corresponding to each sub-feature data in the multiple sub-feature data is the same; where the noise data is time series data.

[0225] Optionally, in a possible implementation, the processing unit 1102 is further configured to: perform a dimension alignment operation on each sub-feature data in the first output data and the third output data respectively, to obtain dimension-aligned first output data and third output data. Determine the difference value between each sub-feature data in the dimension-aligned third output data and the dimension-aligned first output data.

[0226] Optionally, in a possible implementation, the processing unit 1102 is further configured to: determine the difference between the second output data and the true class label of the noise-free data, to obtain a first difference value; determine the difference between the fourth output data and the true class label of the noise-free data, to obtain a second difference value; obtain the second loss function according to the first difference value and the second difference value; wherein, the second output data is the prediction result of multi-classification, and is used to represent the result predicted by the classification network.

[0227] Optionally, in a possible implementation, the obtaining unit 1101 is further configured to obtain a first feature, and predict the binary classification result corresponding to the noise-free data according to the first feature, to obtain a first prediction result, where the first feature is extracted by the classification network in the noise reduction classification network after the noise-free data is input into the noise reduction classification network; the obtaining unit 1101 is further configured to obtain a second feature, and predict the binary classification result corresponding to the noise data according to the second feature, to obtain a second prediction result, where the second feature is extracted by the classification network in the noise reduction classification network after the noise data is input into the noise reduction classification network; the processing unit 1102 is further configured to determine a third loss function according to the first prediction result and the true binary classification label of the noise-free data, and the second prediction result and the true binary classification label of the noise data; the processing unit 1102 is further configured to train the noise reduction classification network at least according to the first loss function, the second loss function, and the third loss function; wherein, the binary classification result corresponding to the noise-free data is noise-free type or noise type, and the binary classification result corresponding to the noise data is noise-free type or noise type.

[0228] Optionally, in a possible implementation, update the parameters of the noise reduction classification network by the error backpropagation algorithm at least according to the first loss function and the second loss function.

[0229] Optionally, in a possible implementation, the noise data includes sparse time series data.

[0230] Optionally, in a possible implementation, the sparse time series data includes skeleton point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data.

[0231] Reference can be made to Figure 12 , Figure 12 which is a schematic structural diagram of a noise reduction and classification device provided by an embodiment of the present application. An embodiment of the present application provides a noise reduction and classification device, which includes: an acquisition unit 1201 and a processing unit 1202. The acquisition unit 1201 is configured to acquire data to be classified. The processing unit 1202 is configured to input the data to be classified into a target network to obtain a prediction result, where the prediction result is the classification result of the data to be classified; wherein, the target network is used to perform noise reduction processing and classification on the data to be classified, and the target network is trained based on the model training method described in the above embodiments.

[0232] Next, an execution device provided by an embodiment of the present application will be introduced. Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of an execution device provided by an embodiment of the present application. The execution device 1300 may specifically be a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a server, etc., which is not limited herein. Among them, a Figure 13 data processing device described in the corresponding embodiment may be deployed on the execution device 1300 to implement Figure 13 the functions of data processing in the corresponding embodiment. Specifically, the execution device 1300 includes: a receiver 1301, a transmitter 1302, a processor 1303, and a memory 1304 (where the number of processors 1303 in the execution device 1300 may be one or more, Figure 13 and one processor is taken as an example herein). Among them, the processor 1303 may include an application processor 13031 and a communication processor 13032. In some embodiments of the present application, the receiver 1301, the transmitter 1302, the processor 1303, and the memory 1304 may be connected through a bus or other means.

[0233] The memory 1304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1303. A part of the memory 1304 may further include a non-volatile random access memory (NVRAM). The memory 1304 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0234] The processor 1303 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, which may include, in addition to a data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.

[0235] The method disclosed in the embodiments of the present application described above can be applied to the processor 1303 or implemented by the processor 1303. The processor 1303 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1303 or the instructions in the form of software. The above-mentioned processor 1303 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1303 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1304, and the processor 1303 reads the information in the memory 1304 and combines its hardware to complete the steps of the above method.

[0236] The receiver 1301 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1302 can be used to output digital or character information through the first interface; the transmitter 1302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1302 may further include a display device such as a display screen.

[0237] In the embodiments of the present application, in one case, the processor 1303 is used to execute Figure 4 the training method of the noise reduction model executed by the execution device in the corresponding embodiment.

[0238] The embodiments of the present application also provide a training device. Please refer toFigure 14 , Figure 14 FIG. Figure 14 is a schematic structural diagram of a training device provided by an embodiment of the present application. Specifically, the training device 1400 is implemented by one or more servers. The training device 1400 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1414 (for example, one or more processors) and a memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) for storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage media 1430 may be transient storage or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1414 may be configured to communicate with the storage media 1430 and execute a series of instruction operations in the storage media 1430 on the training device 1400.

[0239] The training device 1400 may further include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, and one or more input / output interfaces 1458; or, one or more operating systems 1441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0240] An embodiment of the present application further provides a computer program product, which when running on a computer, causes the computer to execute the steps performed by the foregoing execution device, or causes the computer to execute the steps performed by the foregoing training device.

[0241] An embodiment of the present application further provides a computer-readable storage medium, which stores a program for signal processing. When running on a computer, the program causes the computer to execute the steps performed by the foregoing execution device, or causes the computer to execute the steps performed by the foregoing training device.

[0242] The execution device, training device, or terminal device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute the computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiments, or to cause the chip in the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0243] Specifically, please refer to Figure 15 , Figure 15 which is a schematic structural diagram of the chip provided in the embodiments of the present application. The chip may be embodied as a neural network processor NPU 1500. The NPU 1500 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1503, and the arithmetic circuit 1503 is controlled by the controller 1504 to extract matrix data from the memory and perform multiplication operations.

[0244] In some implementations, the arithmetic circuit 1503 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1503 is a two-dimensional systolic array. The arithmetic circuit 1503 may also be a one-dimensional systolic array or other electronic circuits that can perform mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1503 is a general matrix processor.

[0245] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1501 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 1508.

[0246] The unified memory 1506 is used to store input data and output data. The weight data directly passes through the Direct Memory Access Controller (DMAC) 1505 and is transferred to the weight memory 1502 by the DMAC. The input data is also transferred to the unified memory 1506 by the DMAC.

[0247] The BIU is the Bus Interface Unit, that is, the Bus Interface Unit (BIU) 1510, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1509.

[0248] The bus interface unit 1510 is used for the instruction fetch buffer 1509 to obtain instructions from the external memory, and is also used for the storage unit access controller 1505 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0249] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1506, or transfer the weight data to the weight memory 1502, or transfer the input data to the input memory 1501.

[0250] The vector calculation unit 1507 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit 1503, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.

[0251] In some implementations, the vector calculation unit 1507 can store the processed output vector in the unified memory 1506. For example, the vector calculation unit 1507 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1503, such as performing linear interpolation on the feature plane extracted by the convolutional layer, or, for another example, a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 1507 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1503, for example, for use in subsequent layers in the neural network.

[0252] The instruction fetch buffer 1509 connected to the controller 1504 is used to store the instructions used by the controller 1504;

[0253] The unified memory 1506, the input memory 1501, the weight memory 1502, and the fetch memory 1509 are all On-Chip memories. The external memory is private to the NPU hardware architecture.

[0254] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.

[0255] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0256] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0257] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0258] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partly generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A model training method, characterized in that, it includes: Obtain sample pairs, where the sample pairs include noise data and noise-free data corresponding to the noise data. The noise data includes sparse time-series data, and the sparse time-series data includes skeleton point coordinate data, electrocardiogram data, inertial measurement unit data, or fault diagnosis data; Input the noise-free data into a denoising classification network to obtain first output data and second output data. The denoising classification network includes a denoising network and a classification network. The first output data is the output of the denoising network, and the second output data is the output of the classification network; Input the noise data into the denoising classification network to obtain third output data and fourth output data. The third output data is obtained based on the intermediate layer of the denoising network, and the fourth output data is the output of the classification network; Determine a first loss function according to the first output data and the third output data. The first loss function is used to represent the difference between the first output data and the third output data; Determine a second loss function according to the second output data and the fourth output data. The second loss function is used to represent the difference between the second output data and the fourth output data and the true class label of the noise-free data; Train the denoising classification network at least according to the first loss function and the second loss function until a preset training condition is satisfied to obtain a target network.

2. The method according to claim 1, characterized in that, the denoising network includes an encoder and a decoder. The encoder is used to perform compression encoding on the input data, and the decoder is used to perform data reconstruction on the data output by the encoder; the first output data is the output of the encoder, and the third output data is obtained based on the intermediate layer of the encoder.

3. The method according to claim 1 or 2, characterized in that, inputting the noise data into the denoising classification network to obtain third output data includes: inputting the noise data into the denoising classification network to obtain the feature data output by the intermediate layer of the denoising network; dividing the feature data into multiple sub-feature data to obtain the third output data. The third output data includes the multiple sub-feature data; determining a first loss function according to the first output data and the third output data includes: determining the difference value between each sub-feature data in the third output data and the first output data; determining the first loss function according to the difference value between each sub-feature data and the first output data.

4. The method according to claim 3, characterized in that, the dividing the feature data into multiple sub-feature data to obtain the third output data includes: uniformly dividing the feature data into multiple sub-feature data in chronological order to obtain the third output data. The length of the time period corresponding to each sub-feature data in the multiple sub-feature data is the same; wherein, the noise data is time-series data.

5. The method according to claim 3, characterized in that, Determining the difference value between each sub-feature data in the third output data and the first output data includes: Performing a dimension alignment operation on each sub-feature data in the first output data and the third output data respectively to obtain dimension-aligned first output data and third output data; Determining the difference value between each sub-feature data in the dimension-aligned third output data and the dimension-aligned first output data.

6. The method according to any one of claims 1-2, wherein, Determining the second loss function according to the second output data and the fourth output data includes: Determining the difference between the second output data and the true class label of the noise-free data to obtain a first difference value; Determining the difference between the fourth output data and the true class label of the noise-free data to obtain a second difference value; Obtaining the second loss function according to the first difference value and the second difference value; wherein, the second output data is a multi-classification prediction result for representing the result predicted by the classification network.

7. The method according to any one of claims 1-2, wherein, The method further includes: Obtaining a first feature and predicting a binary classification result corresponding to the noise-free data according to the first feature to obtain a first prediction result, where the first feature is extracted by the classification network in the noise reduction classification network after the noise-free data is input into the noise reduction classification network; Obtaining a second feature and predicting a binary classification result corresponding to the noisy data according to the second feature to obtain a second prediction result, where the second feature is extracted by the classification network in the noise reduction classification network after the noisy data is input into the noise reduction classification network; Determining a third loss function according to the first prediction result and the true binary classification label of the noise-free data, the second prediction result and the true binary classification label of the noisy data; Training the noise reduction classification network at least according to the first loss function and the second loss function includes: Training the noise reduction classification network at least according to the first loss function, the second loss function and the third loss function; wherein, the binary classification result corresponding to the noise-free data is noise-free type or noise type, and the binary classification result corresponding to the noisy data is noise-free type or noise type.

8. The method according to any one of claims 1-2, wherein, Training the noise reduction classification network at least according to the first loss function and the second loss function includes: Updating the parameters of the noise reduction classification network at least according to the first loss function and the second loss function through the error backpropagation algorithm.

9. A terminal, wherein, including a memory and a processor; the memory stores code, and the processor is configured to execute the code, and when the code is executed, the terminal executes the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, wherein, Comprising computer-readable instructions which, when run on a computer, cause the computer to perform the method according to any one of claims 1 to 8.

11. A computer program product, characterized in that it comprises computer-readable instructions which, when run on a computer, cause the computer to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent health monitoring

    US20200388287A1

  • Continuous monitoring of a user's health with a mobile device

    WO2019071201A1