Variable batch size multi-layer domain adaptive pattern recognition method based on small sample
Patent Information
- Application Number
- CN202410636434.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-05-22
AI Technical Summary
然而,由于源域数据与目标域数据之间存在分布差异,如果直接进行深度迁移学习,会影响故障诊断模型的训练效果,从而无法保证故障诊断结果的准确性
[0019]借由上述技术方案,本申请提供的一种基于小样本的变批量多层域自适应模式识别方法及装置,在模型训练的过程中,通过同时度量源域数据和目标域数据在特征提取层和分类层上的特征分布差异,能够纠正训练数据的跨域偏差,更好地将从源域数据学到的知识应用于目标域,从而能够在小样本情况下保证故障诊断模型的训练效果,提高故障诊断精度。
Smart Images

Figure CN118535972B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of rotating machinery fault diagnosis technology, and in particular to a method and apparatus for adaptive pattern recognition based on small sample size and variable batch size and multi-domain. Background Technology
[0002] Industrial equipment is increasingly trending towards larger scale, automation, and intelligence. Rotating machinery such as bearings and gears are widely used in aerospace, transportation, and other fields. Their operational status is directly related to the health of various mechanical equipment. Failures can range from affecting normal equipment operation and causing economic losses to serious safety accidents. Therefore, real-time and efficient fault diagnosis of rotating machinery such as bearings has significant practical implications.
[0003] In recent years, deep transfer learning has received widespread attention and research in the field of intelligent fault diagnosis. To address the problem of limited sample data in rotating machinery fault diagnosis scenarios, deep transfer learning techniques are often directly used to train fault diagnosis models. However, due to the distributional differences between source and target domain data, directly applying deep transfer learning can negatively impact the training performance of the fault diagnosis model, thus compromising the accuracy of the fault diagnosis results. Summary of the Invention
[0004] In view of this, this application provides a method and apparatus for adaptive pattern recognition based on small sample size and variable batch size of multi-layer domain, the main purpose of which is to ensure the training effect of the fault diagnosis model and improve the accuracy of fault diagnosis.
[0005] According to a first aspect of this application, a method for adaptive pattern recognition based on small sample size and variable batch size in a multi-domain domain is provided, the method comprising:
[0006] Acquire source domain data and target domain data of rotating machinery, as well as an initial fault diagnosis model, wherein the initial fault diagnosis model includes a feature extraction layer and a classification layer;
[0007] Based on the initial fault diagnosis model, the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer are measured.
[0008] Based on the first feature distribution difference and the second feature distribution difference, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates. In the process of iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as part of the influencing factors of the loss function.
[0009] Based on the fault diagnosis model updated in this round, the process of repeatedly measuring the difference in feature distribution and iteratively updating the model is repeated until the preset convergence condition is met, and then the trained preset fault diagnosis model is output.
[0010] Pattern recognition is performed using the preset fault diagnosis model.
[0011] According to a second aspect of this application, a small-sample, variable-batch, multi-domain adaptive pattern recognition device is provided, the device comprising:
[0012] An acquisition unit is used to acquire source domain data and target domain data of a rotating machinery device, as well as an initial fault diagnosis model, wherein the initial fault diagnosis model includes a feature extraction layer and a classification layer;
[0013] A measurement unit is used to measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, based on the initial fault diagnosis model.
[0014] The update unit is used to iteratively update the initial fault diagnosis model based on the first feature distribution difference and the second feature distribution difference to obtain the fault diagnosis model after this round of updates. In the process of iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as part of the influencing factors of the loss function.
[0015] The iterative unit is used to repeatedly measure the difference in feature distribution and iteratively update the model based on the fault diagnosis model updated in this round, until the preset convergence condition is met, and then outputs the trained preset fault diagnosis model.
[0016] The identification unit is used to perform pattern recognition using the preset fault diagnosis model.
[0017] According to a third aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described small-sample, variable-batch, multi-domain adaptive pattern recognition method.
[0018] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described small-sample variable-batch multi-domain adaptive pattern recognition method.
[0019] By employing the above technical solutions, this application provides a method and apparatus for adaptive pattern recognition based on small sample size and variable batch size of multi-domain data. During model training, by simultaneously measuring the feature distribution differences between source domain data and target domain data in the feature extraction layer and classification layer, it can correct cross-domain bias in training data and better apply the knowledge learned from source domain data to the target domain. This ensures the training effect of the fault diagnosis model under small sample conditions and improves the accuracy of fault diagnosis.
[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0022] Figure 1 The diagram illustrates a flowchart of a variable-batch, multi-domain adaptive pattern recognition method based on small samples, according to an embodiment of this application.
[0023] Figure 2 This paper illustrates a flowchart of another small-sample, variable-batch, multi-domain adaptive pattern recognition method provided in an embodiment of this application.
[0024] Figure 3 This illustration shows a schematic diagram of the overall model training process provided in an embodiment of this application;
[0025] Figure 4 The diagram shows a structural schematic of a variable batch multi-domain adaptive pattern recognition device based on small samples provided in an embodiment of this application. Detailed Implementation
[0026] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0027] Because of the distribution differences between the source domain data and the target domain data, directly performing deep transfer learning will affect the training effect of the fault diagnosis model, thus failing to guarantee the accuracy of the fault diagnosis results.
[0028] To address the aforementioned problems, embodiments of the present invention provide a variable-batch, multi-domain adaptive pattern recognition method based on small samples, such as... Figure 1 As shown, the method includes:
[0029] Step 101: Obtain source domain data and target domain data of the rotating machinery device, as well as the initial fault diagnosis model.
[0030] The rotating mechanical device includes bearings, gears, etc., but the rotating mechanical device in this embodiment is not limited to bearings and gears, and can also be other rotating machinery. The fault type of the source domain data is known, while the fault type of the target domain data is unknown. The initial fault diagnosis model includes a feature extraction layer and a classification layer.
[0031] This invention is primarily applicable to scenarios where fault diagnosis models are trained based on few samples. The executing entity of this invention is a device or equipment capable of training fault diagnosis models based on few samples, specifically, it can be located on the server side.
[0032] In the fault diagnosis of rotating machinery, it is necessary to measure signals, specifically including one or more of the following: vibration signals, temperature signals, sound wave signals, and current signals. Among these signals, vibration signals are easy to measure and can provide the most intrinsic information for mechanical fault diagnosis. Therefore, this embodiment of the invention preferentially selects to use vibration signals to identify the health status of rotating machinery, that is, starting from vibration signals, exploring the relationship between fault signals and fault types. Specific fault types include pitting, wear, etc.
[0033] It should be noted that in the fault diagnosis of rotating machinery, the selected measurement signal is not limited to vibration signal, but can also be other types of signal. The embodiments of the present invention do not make specific limitations on this.
[0034] Furthermore, to successfully acquire vibration signals, an experimental platform needs to be built, along with vibration sensors and a data acquisition card. The vibration simulation experimental platform or industrial field equipment is fundamental for data acquisition; therefore, a comprehensive understanding of the platform's system composition and the functions of each component is required, along with proficiency in its use to conduct fault simulation experiments on rotating machinery. Vibration sensors need to be installed in appropriate locations to maximize the optimal representation of the current signal and provide a solid foundation for subsequent analysis and processing. The data acquisition card ensures stable data acquisition; specifically, the vibration sensor needs to be connected to the acquisition card, which then uploads the measured data to a computer client.
[0035] After the hardware equipment is set up, raw vibration signals under different operating conditions are collected and divided into source domain data and target domain data. The method for this process includes: acquiring the raw vibration signals of the rotating machinery; and based on the raw vibration signals, determining source domain data with known fault types and target domain data with unknown fault types.
[0036] For example, multiple sets of raw vibration signals of a bearing under different speeds and forces are collected and divided into source domain data and target domain data. The fault types in the source domain data are known and have fault type labels, while the fault types in the target domain data are unknown and do not have fault type labels. During sampling, a sampling step size of 1024 is preferred, with no overlap.
[0037] In some embodiments, all source domain data can be used as training data, while the target domain data can be randomly divided into two parts: one part as training data and the other part as test data. For example, 20% of the target domain data can be used as test data, and 80% of the data can be used as training data.
[0038] In some embodiments, the method for constructing an initial fault diagnosis model includes: using a preset neural network algorithm to construct an initial fault diagnosis model corresponding to the rotating machinery device.
[0039] Specifically, the initial fault diagnosis model can be a neural network model, such as an ML-MM network model, or other neural network models; this embodiment of the invention does not impose specific limitations on this. Before training the model, it is necessary to set the model parameters, learning rate, sample batch size for different iteration stages, preset convergence conditions (such as the number of iterations), initialization bias, and weight parameters, etc.
[0040] Step 102: Based on the initial fault diagnosis model, measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer.
[0041] In this embodiment of the invention, to ensure the training effect of the fault diagnosis model, it is necessary to simultaneously measure the difference in the first feature distribution of the source domain data and the target domain data at the feature extraction layer and the difference in the second feature distribution at the classification layer. Specifically, the source domain data and the target domain data can be input into the initial fault diagnosis model respectively to obtain the multidimensional features output by the feature extraction layer of the source domain data and the target domain data. Then, the difference in the feature distribution of the source domain data and the target domain data at the feature extraction layer can be determined. Similarly, the multidimensional features output by the classification layer of the source domain data and the target domain data can be obtained, and then the difference in the feature distribution of the source domain data and the target domain data at the classification layer can be determined.
[0042] In some embodiments, a portion of the target domain data is used for training and another portion is used for testing. Based on this, step 102 specifically includes: dividing the target domain data into first target domain data for training and second target domain data for testing; and based on the initial fault diagnosis model, measuring the difference in first feature distribution between the source domain data and the first target domain data on the feature extraction layer and the difference in second feature distribution on the classification layer.
[0043] The ratio between the first target domain data and the second target domain data can be set according to actual business needs. For example, 80% of the first target domain data can be used for training, and 20% of the second target domain data can be used for testing.
[0044] Step 103: Based on the first feature distribution difference and the second feature distribution difference, iteratively update the initial fault diagnosis model to obtain the fault diagnosis model after this round of updates.
[0045] Specifically, when iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as partial influencing factors of the loss function.
[0046] In this embodiment of the invention, after determining the first feature distribution difference between the source domain data and the target domain data on the feature extraction layer and the second feature distribution difference on the classification layer, a loss function can be constructed based on the first feature distribution difference and the second feature distribution difference, that is, the overall feature distribution difference is minimized. Then, based on the loss function, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates.
[0047] Step 104: Based on the fault diagnosis model updated in this round, repeat the process of measuring feature distribution differences and iteratively updating the model until the preset convergence condition is met, and output the trained preset fault diagnosis model.
[0048] The preset convergence condition can be the number of iterations, that is, when a certain number of iterations is reached, training stops.
[0049] In this embodiment of the invention, the fault diagnosis model is iteratively trained based on the constructed loss function. When a certain number of iterations are reached, the trained preset fault diagnosis model is output. Then, the trained preset fault diagnosis model can be tested using test data from the target domain data to ensure the accuracy of the fault diagnosis model's diagnostic results.
[0050] Step 105: Perform pattern recognition using the preset fault diagnosis model.
[0051] In this embodiment of the invention, pattern recognition is fault diagnosis. Through pattern recognition, the fault type of rotating mechanical device, such as wear and pitting, can be determined.
[0052] This invention provides a small-sample, variable-batch, multi-domain adaptive pattern recognition method. During model training, by simultaneously measuring the feature distribution differences between source domain data and target domain data in the feature extraction layer and classification layer, it can correct cross-domain bias in the training data and better apply the knowledge learned from the source domain data to the target domain. This ensures the training effect of the fault diagnosis model under small-sample conditions and improves the accuracy of fault diagnosis.
[0053] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, and to fully illustrate the implementation of this embodiment, this embodiment also provides another method for adaptive pattern recognition based on small sample size and variable batch size, such as... Figure 2 As shown, the method includes:
[0054] Step 201: Obtain source domain data and target domain data of the rotating machinery device, as well as the initial fault diagnosis model.
[0055] The initial fault diagnosis model includes a feature extraction layer and a classification layer.
[0056] In this embodiment of the invention, the process of obtaining source domain data, target domain data and initial fault diagnosis model is exactly the same as step 101, and will not be repeated here.
[0057] Step 202: Based on the initial fault diagnosis model, measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer.
[0058] In this embodiment of the invention, to measure the difference in the first feature distribution of source domain data and target domain data at the feature extraction layer and the difference in the second feature distribution at the classification layer, step 202 specifically includes: inputting the source domain data into the initial fault diagnosis model for fault diagnosis to obtain the first multidimensional feature output by the feature extraction layer and the first multidimensional feature output by the classification layer; inputting the target domain data into the initial fault diagnosis model for fault diagnosis to obtain the second multidimensional feature output by the feature extraction layer and the second multidimensional feature output by the classification layer; and based on the first and second multidimensional features output by the feature extraction layer and the first and second multidimensional features output by the classification layer, respectively measuring the difference in the first feature distribution of source domain data and target domain data at the feature extraction layer and the difference in the second feature distribution at the classification layer.
[0059] Further, the step of measuring the first feature distribution difference on the feature extraction layer and the second feature distribution difference on the classification layer based on the first and second multidimensional features output by the feature extraction layer and the first and second multidimensional features output by the classification layer, respectively, includes: calculating the multidimensional feature distance corresponding to the feature extraction layer based on the first and second multidimensional features output by the feature extraction layer; calculating the multidimensional feature distance corresponding to the classification layer based on the first and second multidimensional features output by the classification layer; and measuring the first feature distribution difference and the second feature distribution difference based on the multidimensional feature distance corresponding to the feature extraction layer and the multidimensional feature distance corresponding to the classification layer, respectively. The specific calculation process for the multidimensional feature distances corresponding to the feature extraction layer and the classification layer is as follows:
[0060]
[0061] Among them, L MMD This represents the multidimensional feature distance corresponding to the feature extraction layer and the multidimensional feature distance corresponding to the classification layer. This multidimensional feature distance is used to reflect the difference in feature distribution between the source domain data and the target domain data in the feature extraction layer and the classification layer. The larger the multidimensional feature distance, the greater the difference in feature distribution. s X represents the first multidimensional feature of the source domain data in the feature extraction layer or classification layer. t This represents the second multidimensional feature of the target domain data at the feature extraction layer or classification layer. Based on the first and second multidimensional features corresponding to the feature extraction layer, the multidimensional feature distance corresponding to the feature extraction layer is calculated. Similarly, based on the first and second multidimensional features corresponding to the classification layer, the multidimensional feature distance corresponding to the classification layer is calculated. Finally, the multidimensional feature distances corresponding to the feature extraction layer and the classification layer are added together to obtain L. MMD .
[0062] Step 203: Obtain the actual fault type corresponding to the rotating mechanical device in the source domain data, and input the source domain data into the initial fault diagnosis model for fault diagnosis to obtain the predicted fault type corresponding to the source domain data.
[0063] In this embodiment of the invention, the source domain data contains fault type labels, which can be used to determine the actual fault type of the rotating mechanical device.
[0064] Step 204: Based on the actual fault type and predicted fault type corresponding to the source domain data, as well as the first feature distribution difference and the second feature distribution difference, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates.
[0065] To ensure the training accuracy of the fault diagnosis model under small sample conditions, this embodiment of the invention divides the training process of the fault diagnosis model into two stages. In the first training stage, when constructing the loss function, only the fault diagnosis error corresponding to the source domain data and the feature distribution differences on the feature extraction layer and the classification layer are considered. Based on the constructed loss function, the target fault diagnosis model is trained. In the second training stage, using the target fault diagnosis model obtained in the first training stage, the pseudo-fault type (pseudo-label) corresponding to the target domain data is determined. When constructing the loss function, not only the fault diagnosis error corresponding to the source domain data and the feature distribution differences on the feature extraction layer and the classification layer are considered, but also the fault diagnosis error corresponding to the target domain data is considered. The overall algorithm flow is as follows: Figure 3 As shown.
[0066] For the update and iteration process in the first training phase, step 204 specifically includes: calculating the fault type difference corresponding to the source domain data based on the actual fault type and the predicted fault type corresponding to the source domain data; calculating the total feature distribution difference between the fully connected layer and the classification layer based on the first feature distribution difference and the second feature distribution difference; constructing a first loss function based on the fault type difference corresponding to the source domain data and the total feature distribution difference; and iteratively updating the initial fault diagnosis model based on the first loss function to obtain the fault diagnosis model after this round of updates. The specific formula for the first loss function is as follows:
[0067] L1 = L s +λL MMD
[0068]
[0069] Where L1 represents the first loss function. MMD L represents the multidimensional feature distance between the feature extraction layer and the classification layer, i.e., the difference in the total feature distribution between the source domain data and the target domain data at the feature extraction layer and the classification layer. s This represents the difference in fault types corresponding to the source domain data. This represents the actual fault type corresponding to the source domain data. The predicted fault type represents the source domain data, j represents the dimension, and the total dimension of the predicted fault type or the actual fault type is k, n s The amount of data representing the source domain data.
[0070] After constructing the first loss function L1, the initial fault diagnosis model is updated and iterated for the first time based on the first loss function L1 to obtain the fault diagnosis model after this round of updates.
[0071] Step 205: Based on the fault diagnosis model updated in this round, repeat the process of measuring feature distribution differences and iteratively updating the model until the preset convergence condition is met, and output the trained preset fault diagnosis model.
[0072] The preset convergence conditions include the first preset number of iterations corresponding to the first training phase and the second preset number of iterations corresponding to the second training phase.
[0073] In this embodiment of the invention, after the first round of update iteration is completed in the first training phase, the fault diagnosis model is updated and iterated based on the constructed first preset loss function. The process of measuring the first feature distribution difference between the source domain data and the target domain data on the feature extraction layer and the second feature distribution difference on the hierarchical class is repeated, as well as the process of iteratively updating the model, until the first preset number of iterations is reached, at which point the iterative update is stopped, the target fault diagnosis model is output, and the first phase of training is completed.
[0074] For the second stage of the training process, the method includes: determining the pseudo-fault type (pseudo-label) corresponding to the target domain data based on the target fault diagnosis model; inputting the target domain data into the initial fault diagnosis model for fault diagnosis to obtain the predicted fault type corresponding to the target domain data; iteratively updating the initial fault diagnosis model based on the fault type difference corresponding to the source domain data, the pseudo-fault type and the predicted fault type corresponding to the target domain data, and the total feature distribution difference, until a second preset number of iterations is reached, and then outputting the trained preset fault diagnosis model.
[0075] Further, based on the fault type difference corresponding to the source domain data, the pseudo-fault type and predicted fault type corresponding to the target domain data, and the total feature distribution difference, the initial fault diagnosis model is iteratively updated until a second preset number of iterations is reached, at which point a trained preset fault diagnosis model is output. This includes: calculating the fault type difference corresponding to the target domain data based on the pseudo-fault type and predicted fault type; constructing a second loss function based on the fault type difference corresponding to the source domain data, the fault type difference corresponding to the target domain data, and the total feature distribution difference; and iteratively updating the initial fault diagnosis model based on the second loss function until a second preset number of iterations is reached, at which point a trained preset fault diagnosis model is output. The specific calculation formula for the second loss function is as follows:
[0076] L2 = L s +λL MMD +L t
[0077]
[0078] Here, L2 represents the second loss function. MMD L represents the multidimensional feature distance between the feature extraction layer and the classification layer, i.e., the difference in the total feature distribution between the source domain data and the target domain data at the feature extraction layer and the classification layer. s L represents the difference in fault types corresponding to the source domain data. t This represents the difference in fault types corresponding to the target domain data. This represents the pseudo-fault type corresponding to the target domain data. The predicted fault type represents the target domain data, j represents the dimension, and the total dimension of the predicted fault type or the actual fault type is k, n t The amount of data representing the target domain.
[0079] After constructing the second loss function L2, the initial fault diagnosis model is iteratively trained in the second stage based on this second loss function L2 until the second preset number of iterations is reached. At this point, the iterative update stops, and the trained preset fault diagnosis model is output. To test the diagnostic accuracy of the preset fault diagnosis model, test data from the target domain can be used to test the model.
[0080] In practical applications, while small batches of samples can accelerate training, the fault diagnosis types within these small batches are imbalanced. For example, there might be seven sets of data for pitting corrosion, but only one set for wear. This could lead to each iteration updating only within the results of the small batch, resulting in unstable diagnostic results. To address this issue, this embodiment of the invention employs a variable batch training strategy to update and iterate the initial fault diagnosis model, where the sample batch size differs at different iteration stages.
[0081] For example, the sum of the first preset number of iterations and the second preset number of iterations is 500. The first 100 iterations use 20 sets of training data, the second 100 iterations use 100 sets of training data, the third 100 iterations use 150 sets of training data, the fourth 100 iterations use 200 sets of training data, and the fifth 100 iterations use 250 sets of training data.
[0082] Therefore, it can be seen that by adopting a variable batch training strategy, the present invention uses a small batch of training samples at the beginning of training to accelerate the convergence speed of the model, and gradually increases the sample batch in the later stage, which can ensure that the sample data of each fault type is more balanced, thereby making the diagnostic results of the fault diagnosis model more stable and the accuracy will not fluctuate.
[0083] Step 206: Perform pattern recognition using the preset fault diagnosis model.
[0084] In this embodiment of the invention, pattern recognition is fault diagnosis. Through pattern recognition, the fault type of rotating mechanical device, such as wear and pitting, can be determined.
[0085] Another method for adaptive pattern recognition based on small sample size and variable batch size of multi-domain data provided in this invention can correct cross-domain bias in training data by simultaneously measuring the feature distribution differences between source domain data and target domain data in the feature extraction layer and classification layer during model training. This allows for better application of knowledge learned from source domain data to the target domain, thereby ensuring the training effect of the fault diagnosis model under small sample conditions and improving the accuracy of fault diagnosis.
[0086] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a variable-batch, multi-domain adaptive pattern recognition device based on small samples, such as... Figure 4 As shown, the device includes: an acquisition unit 31, a measurement unit 32, an update unit 33, an iteration unit 34, and an identification unit 35.
[0087] The acquisition unit 31 can be used to acquire source domain data and target domain data of the rotating machinery device, as well as an initial fault diagnosis model, wherein the initial fault diagnosis model includes a feature extraction layer and a classification layer.
[0088] The measurement unit 32 can be used to measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, based on the initial fault diagnosis model.
[0089] The update unit 33 can be used to iteratively update the initial fault diagnosis model based on the first feature distribution difference and the second feature distribution difference to obtain the fault diagnosis model after this round of updates. When iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as part of the influencing factors of the loss function.
[0090] The iterative unit 34 can be used to repeatedly measure the difference in feature distribution and iteratively update the model based on the fault diagnosis model updated in this round, until the preset convergence condition is met, and then output the trained preset fault diagnosis model.
[0091] The identification unit 35 can be used to perform pattern recognition using the preset fault diagnosis model.
[0092] In some embodiments, the acquisition unit 31 may be specifically used to acquire the original vibration signal of the rotating mechanical device; and based on the original vibration signal, to determine source domain data of known fault types and target domain data of unknown fault types.
[0093] In some embodiments, the acquisition unit 31 may also be specifically used to construct an initial fault diagnosis model corresponding to the rotating machinery device using a preset neural network algorithm.
[0094] In some embodiments, the measurement unit 32 may be specifically used to divide the target domain data into first target domain data for training and second target domain data for testing; based on the initial fault diagnosis model, to measure the difference in first feature distribution between the source domain data and the first target domain data on the feature extraction layer and the difference in second feature distribution on the classification layer.
[0095] In some embodiments, the measurement unit 32 includes a first determination module and a measurement module.
[0096] The first determining module can be used to input the source domain data into the initial fault diagnosis model for fault diagnosis, and obtain the first multidimensional feature output by the feature extraction layer and the first multidimensional feature output by the classification layer.
[0097] The first determining module can also be used to input the target domain data into the initial fault diagnosis model for fault diagnosis, and obtain the second multidimensional feature output by the feature extraction layer and the second multidimensional feature output by the classification layer.
[0098] The measurement module can be used to measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, respectively, based on the first multidimensional feature and the second multidimensional feature output by the feature extraction layer and the first multidimensional feature and the second multidimensional feature output by the classification layer.
[0099] In some embodiments, the measurement module may be specifically used to calculate the multidimensional feature distance corresponding to the feature extraction layer based on the first multidimensional feature and the second multidimensional feature output by the feature extraction layer; calculate the multidimensional feature distance corresponding to the classification layer based on the first multidimensional feature and the second multidimensional feature output by the classification layer; and measure the first feature distribution difference and the second feature distribution difference respectively based on the multidimensional feature distance corresponding to the feature extraction layer and the multidimensional feature distance corresponding to the classification layer.
[0100] In some embodiments, the iteration unit 34 can also be used to update and iterate the initial fault diagnosis model using a variable batch training strategy, wherein the sample batches corresponding to different iteration stages are different.
[0101] In some embodiments, the update unit 33 includes: an acquisition module, a second determination module, and an update module.
[0102] The acquisition module can be used to acquire the actual fault type corresponding to the rotating mechanical device in the source domain data.
[0103] The second determining module can be used to input the source domain data into the initial fault diagnosis model for fault diagnosis, and obtain the predicted fault type corresponding to the source domain data.
[0104] The update module can be used to iteratively update the initial fault diagnosis model based on the actual fault type and predicted fault type corresponding to the source domain data, as well as the first feature distribution difference and the second feature distribution difference, to obtain the fault diagnosis model after this round of updates.
[0105] In some embodiments, the update module may be specifically configured to: calculate the fault type difference corresponding to the source domain data based on the actual fault type and the predicted fault type corresponding to the source domain data; calculate the total feature distribution difference between the fully connected layer and the classification layer based on the first feature distribution difference and the second feature distribution difference; construct a first loss function based on the fault type difference corresponding to the source domain data and the total feature distribution difference; and iteratively update the initial fault diagnosis model based on the first loss function to obtain the fault diagnosis model after this round of updates.
[0106] In some embodiments, the preset convergence condition includes a first preset number of iterations and a second preset number of iterations. The iteration unit 34 includes an iteration module and a third determination module.
[0107] The iteration module can be used to repeatedly measure the difference in feature distribution and iteratively update the model based on the fault diagnosis model updated in this round, until the first preset number of iterations is reached, and then output the target fault diagnosis model.
[0108] The third determining module can be used to determine the pseudo-fault type corresponding to the target domain data based on the target fault diagnosis model.
[0109] The third determining module can also be used to input the target domain data into the initial fault diagnosis model for fault diagnosis, and obtain the predicted fault type corresponding to the target domain data.
[0110] The iteration module can also be used to iteratively update the initial fault diagnosis model based on the fault type difference corresponding to the source domain data, the pseudo fault type and predicted fault type corresponding to the target domain data, and the total feature distribution difference, until the second preset number of iterations is reached, and then output the trained preset fault diagnosis model.
[0111] In some embodiments, the iteration module may be specifically used to calculate the fault type difference corresponding to the target domain data based on the pseudo fault type and predicted fault type corresponding to the target domain data; construct a second loss function based on the fault type difference corresponding to the source domain data, the fault type difference corresponding to the target domain data, and the total feature distribution difference; and iteratively update the initial fault diagnosis model based on the second loss function until a second preset number of iterations is reached, and output the trained preset fault diagnosis model.
[0112] It should be noted that other corresponding descriptions of the functional units involved in the small-sample, variable-batch, multi-domain adaptive pattern recognition device provided in this embodiment of the invention can be found in the following references. Figure 1 and Figure 3 The corresponding description in [the document] will not be repeated here.
[0113] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown is a variable batch multi-domain adaptive pattern recognition method based on small samples.
[0114] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0115] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 4 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method shown is a variable batch multi-domain adaptive pattern recognition method based on small samples.
[0116] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0117] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0118] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0120] This invention, by simultaneously measuring the feature distribution differences between source domain data and target domain data in the feature extraction layer and classification layer, can correct cross-domain bias in training data, better apply the knowledge learned from source domain data to the target domain, thereby ensuring the training effect of the fault diagnosis model and improving fault diagnosis accuracy even with small sample sizes.
[0121] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0122] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A variable-batch, multi-domain adaptive pattern recognition method based on small samples, characterized in that, include: Acquire source domain data and target domain data of rotating machinery, as well as an initial fault diagnosis model, wherein the initial fault diagnosis model includes a feature extraction layer and a classification layer; Based on the initial fault diagnosis model, the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer are measured. Based on the first feature distribution difference and the second feature distribution difference, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates. In the process of iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as part of the influencing factors of the loss function. Based on the fault diagnosis model updated in this round, the process of repeatedly measuring the difference in feature distribution and iteratively updating the model is repeated until the preset convergence condition is met, and then the trained preset fault diagnosis model is output. Pattern recognition is performed using the preset fault diagnosis model; This includes acquiring source domain data and target domain data of the rotating machinery device, including: Obtain the original vibration signal of the rotating mechanical device; Based on the original vibration signal, determine the source domain data for known fault types and the target domain data for unknown fault types; Obtain the initial fault diagnosis model, including: An initial fault diagnosis model for the rotating mechanical device is constructed using a preset neural network algorithm. The step of measuring the difference in the first feature distribution of the source domain data and the target domain data at the feature extraction layer and the difference in the second feature distribution at the classification layer, based on the initial fault diagnosis model, includes: The target domain data is divided into a first target domain data for training and a second target domain data for testing; Based on the initial fault diagnosis model, the difference in the first feature distribution of the source domain data and the first target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer are measured. The step of measuring the difference in the first feature distribution of the source domain data and the first target domain data at the feature extraction layer and the difference in the second feature distribution at the classification layer based on the initial fault diagnosis model includes: The source domain data is input into the initial fault diagnosis model for fault diagnosis, and the first multidimensional feature output by the feature extraction layer and the first multidimensional feature output by the classification layer are obtained. The first target domain data is input into the initial fault diagnosis model for fault diagnosis, and the second multi-dimensional feature output by the feature extraction layer and the second multi-dimensional feature output by the classification layer are obtained. Based on the first and second multidimensional features output by the feature extraction layer, and the first and second multidimensional features output by the classification layer, the differences in the first feature distribution of the source domain data and the first target domain data on the feature extraction layer and the differences in the second feature distribution on the classification layer are respectively measured. The method further includes: The initial fault diagnosis model is updated and iterated using a variable batch training strategy, wherein the sample batches are different for different iteration stages.
2. The method according to claim 1, characterized in that, The first and second multidimensional features output by the feature extraction layer, and the first and second multidimensional features output by the classification layer, respectively measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, including: Based on the first and second multidimensional features output by the feature extraction layer, the multidimensional feature distance corresponding to the feature extraction layer is calculated. Based on the first and second multidimensional features output by the classification layer, the multidimensional feature distance corresponding to the classification layer is calculated. The first feature distribution difference and the second feature distribution difference are measured based on the multidimensional feature distance corresponding to the feature extraction layer and the multidimensional feature distance corresponding to the classification layer, respectively.
3. The method according to claim 1, characterized in that, The step of iteratively updating the initial fault diagnosis model based on the first feature distribution difference and the second feature distribution difference to obtain the updated fault diagnosis model includes: Obtain the actual fault type corresponding to the rotating mechanical device in the source domain data; The source domain data is input into the initial fault diagnosis model for fault diagnosis to obtain the predicted fault type corresponding to the source domain data. Based on the actual fault type and predicted fault type corresponding to the source domain data, as well as the first feature distribution difference and the second feature distribution difference, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates.
4. The method according to claim 3, characterized in that, Based on the actual fault type and predicted fault type corresponding to the source domain data, as well as the first feature distribution difference and the second feature distribution difference, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates, including: Calculate the fault type difference corresponding to the source domain data based on the actual fault type and the predicted fault type corresponding to the source domain data; Based on the first feature distribution difference and the second feature distribution difference, calculate the total feature distribution difference between the fully connected layer and the classification layer; A first loss function is constructed based on the difference in fault types corresponding to the source domain data and the difference in the total feature distribution. Based on the first loss function, the initial fault diagnosis model is iteratively updated to obtain the fault diagnosis model after this round of updates.
5. The method according to claim 4, characterized in that, The preset convergence conditions include a first preset number of iterations and a second preset number of iterations. The process of repeatedly measuring feature distribution differences and iteratively updating the model based on the updated fault diagnosis model continues until the preset convergence conditions are met, at which point the trained preset fault diagnosis model is output. This includes: Based on the fault diagnosis model updated in this round, the process of repeatedly measuring feature distribution differences and iteratively updating the model is repeated until the first preset number of iterations is reached, at which point the target fault diagnosis model is output. Based on the target fault diagnosis model, determine the pseudo-fault type corresponding to the target domain data; The target domain data is input into the initial fault diagnosis model to perform fault diagnosis, and the predicted fault type corresponding to the target domain data is obtained. Based on the fault type difference corresponding to the source domain data, the pseudo fault type and predicted fault type corresponding to the target domain data, and the total feature distribution difference, the initial fault diagnosis model is iteratively updated until the second preset number of iterations is reached, at which point the trained preset fault diagnosis model is output.
6. The method according to claim 5, characterized in that, Based on the fault type difference corresponding to the source domain data, the pseudo-fault type and predicted fault type corresponding to the target domain data, and the difference in the total feature distribution, the initial fault diagnosis model is iteratively updated until a second preset number of iterations is reached, at which point a trained preset fault diagnosis model is output, including: Calculate the fault type difference corresponding to the target domain data based on the pseudo fault type and predicted fault type corresponding to the target domain data; A second loss function is constructed based on the fault type difference corresponding to the source domain data, the fault type difference corresponding to the target domain data, and the total feature distribution difference. Based on the second loss function, the initial fault diagnosis model is iteratively updated until the second preset number of iterations is reached, at which point the trained preset fault diagnosis model is output.
7. A variable-batch, multi-domain adaptive pattern recognition device based on small samples, characterized in that, include: An acquisition unit is used to acquire source domain data and target domain data of a rotating machinery device, as well as an initial fault diagnosis model, wherein the initial fault diagnosis model includes a feature extraction layer and a classification layer; A measurement unit is used to measure the difference in the first feature distribution of the source domain data and the target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, based on the initial fault diagnosis model. The update unit is used to iteratively update the initial fault diagnosis model based on the first feature distribution difference and the second feature distribution difference to obtain the fault diagnosis model after this round of updates. In the process of iteratively updating the initial fault diagnosis model, the first feature distribution difference and the second feature distribution difference are used as part of the influencing factors of the loss function. The iterative unit is used to repeatedly measure the difference in feature distribution and iteratively update the model based on the fault diagnosis model updated in this round, until the preset convergence condition is met, and then outputs the trained preset fault diagnosis model. The identification unit is used to perform pattern recognition using the preset fault diagnosis model; The acquisition unit is specifically used to acquire the original vibration signal of the rotating machinery; and based on the original vibration signal, to determine source domain data of known fault types and target domain data of unknown fault types. The acquisition unit is also specifically used to construct an initial fault diagnosis model corresponding to the rotating mechanical device using a preset neural network algorithm; The metric unit is specifically used to divide the target domain data into first target domain data for training and second target domain data for testing; based on the initial fault diagnosis model, it measures the difference in first feature distribution of the source domain data and the first target domain data on the feature extraction layer and the difference in second feature distribution on the classification layer. The measurement unit includes: a first determination module and a measurement module; The first determining module is used to input the source domain data into the initial fault diagnosis model for fault diagnosis, and obtain the first multi-dimensional feature output by the feature extraction layer and the first multi-dimensional feature output by the classification layer; and to input the first target domain data into the initial fault diagnosis model for fault diagnosis, and obtain the second multi-dimensional feature output by the feature extraction layer and the second multi-dimensional feature output by the classification layer. The measurement module is used to measure the difference in the first feature distribution of the source domain data and the first target domain data on the feature extraction layer and the difference in the second feature distribution on the classification layer, respectively, based on the first multidimensional feature and the second multidimensional feature output by the feature extraction layer and the first multidimensional feature and the second multidimensional feature output by the classification layer. The iterative unit is also used to update and iterate the initial fault diagnosis model using a variable batch training strategy, wherein the sample batches corresponding to different iteration stages are different.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Long-term service elevator guide rail fault diagnosis method based on transfer learning and data driving
CN117312962A