Model training method, equipment trend detection method, device and equipment
By generating a missing dataset and using comparative learning between the teacher model and the initial learning module, fusion features are extracted and trained to train the device trend detection module. This solves the problem of modality missing in device trend recognition, achieves accurate detection in the case of missing data, and improves the efficiency and reliability of device operation and maintenance.
Patent Information
- Application Number
- CN202511052508.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, equipment trend recognition methods suffer from insufficient fault tolerance for modality loss and low recognition accuracy in multimodal data processing, making it difficult to adapt to data incompleteness and dynamic changes in complex industrial environments.
By collecting multimodal historical data to generate a missing dataset, and by using comparative learning between the teacher model and the initial learning module, the fusion features of the missing data are extracted, and the device trend detection module is trained to form a device trend detection model, thereby enhancing the robustness of the model in the case of missing data.
In the absence of data, it can accurately detect the operating trend of power equipment, improve the efficiency and reliability of equipment operation and maintenance, and avoid the decline in detection performance.
Smart Images

Figure CN120974182A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of safety management, and particularly relates to a model training method, a device trend detection method, device and apparatus. BACKGROUND
[0002] With the rapid development of industrial automation and intelligent manufacturing, device state monitoring and trend analysis have become key technologies to ensure production safety and efficiency. During the operation of the device, multi-modal data are generated, including electrical parameters, mechanical vibration, temperature and humidity environmental information, images and videos, etc. These data contain rich health state and fault warning information. Through multi-modal data, device state monitoring and warning are performed to ensure the smooth operation of the device and improve the stability and safety of production.
[0003] In the related art, life prediction and resource optimization are achieved through multi-source data acquisition for anomaly detection, but there is a problem of insufficient fault tolerance for missing modalities and low identification accuracy.
[0004] Therefore, there is an urgent need for a device trend identification method with high identification accuracy. SUMMARY
[0005] The present application provides a model training method, a device trend detection method, device and apparatus to improve the accuracy of device trend identification.
[0006] In a first aspect, the present application provides a model training method, comprising:
[0007] During the operation of the power device, a plurality of historical data of the power device are collected to obtain a first power data set; each historical data has a different modality;
[0008] The data in the first power data set are modified and / or deleted to obtain a second power data set, which is a missing data set corresponding to the first power data set;
[0009] The second power data set is used as the input of an initial learning module, and the output of a teacher model is used as the training target of the initial learning module. Based on an evaluation function and an optimization algorithm, the corresponding parameters in the initial learning module are adjusted to obtain a feature learning module; wherein the teacher model is trained by the first power data set;
[0010] The first fusion feature corresponding to each historical data in the second power data set is extracted by the feature learning module;
[0011] Based on each historical data and the first fusion feature corresponding to each historical data, a device trend detection module is trained to obtain a trained device trend detection module;
[0012] The integrated feature learning module and the device trend detection module obtain a device trend detection model.
[0013] In a possible implementation, the second power data set is taken as an input of an initial learning module, an output of a teacher model is taken as a training target of the initial learning module, corresponding parameters in the initial learning module are adjusted based on an evaluation function and an optimization algorithm, and a feature learning module is obtained, including:
[0014] The first fusion features corresponding to each piece of historical data in the second power data set are extracted by a feature learning unit in the initial learning module;
[0015] The second fusion features corresponding to each piece of historical data in the first power data set are extracted by the teacher model;
[0016] The error between the first fusion features and the second fusion features is determined based on the evaluation function, and a loss value is obtained;
[0017] If the loss value is greater than or equal to a loss value threshold, the corresponding parameters in the initial learning module are adjusted based on the loss value by using the optimization algorithm until the loss value is less than the loss value threshold;
[0018] If the loss value is less than the loss value threshold, the feature learning module is obtained.
[0019] In a possible implementation, the feature learning unit includes a plurality of encoders, a feature switch and a gated fusion unit, the first fusion features corresponding to each piece of historical data in the second power data set are extracted by the feature learning unit in the initial learning module, including:
[0020] Each piece of historical data in the second power data set is encoded by the plurality of encoders sharing convolution kernel weights, and first power features corresponding to different modalities in each piece of historical data are obtained;
[0021] The first power features corresponding to the target modality in each piece of historical data are exchanged by the feature switch, and exchanged features corresponding to the target modality in each piece of historical data are obtained;
[0022] The exchanged features corresponding to different modalities are selectively transmitted by the gated fusion unit, and the first fusion features corresponding to each piece of historical data are obtained.
[0023] In a possible implementation, the first power features corresponding to the target modality in each piece of historical data are exchanged by the feature switch, and the exchanged features corresponding to the target modality in each piece of historical data are obtained, including:
[0024] The first power features corresponding to the target modality in each piece of historical data are exchanged by the feature switch, and initial exchanged features corresponding to the target modality in each piece of historical data are obtained;
[0025] The distance between the initial exchanged features corresponding to the target modality in each piece of historical data is determined by pixel-level contrast learning, and a contrast loss is obtained.
[0026] If the contrast loss is greater than or equal to a contrast loss threshold, the exchange mask tensor in the feature switcher is adjusted, and the first power feature corresponding to the target modality in each piece of historical data is re-exchanged until the contrast loss is less than the contrast loss threshold.
[0027] If the contrast loss is less than the contrast loss threshold, the initial exchanged feature is determined as the exchanged feature corresponding to the target modality.
[0028] In a possible implementation, the process of training the teacher model by the first power data set includes:
[0029] Each piece of historical data in the first power data set is encoded by the plurality of feature extractors in the initial model to obtain second power features corresponding to different modalities in each piece of historical data.
[0030] The contribution of the first power features corresponding to different modalities is modeled by a modality attention mechanism and a missing simulation strategy to obtain learning weights corresponding to each modality.
[0031] The second power features corresponding to each modality are dynamically weighted according to the learning weights corresponding to each modality to obtain first fusion features.
[0032] Based on the first fusion features, an overall loss of the initial model is determined, and the overall loss includes a modality consistency loss, an accuracy loss, and a trend loss.
[0033] If the overall loss does not meet a preset requirement, the parameters of the modality attention mechanism are adjusted and the learning weights corresponding to each modality are re-determined until the overall loss meets the preset requirement.
[0034] If the overall loss meets the preset requirement, a teacher model is obtained.
[0035] In a possible implementation, during the operation of the power equipment, a plurality of pieces of historical data of the power equipment are collected to obtain a first power data set, including:
[0036] Raw data information of the power equipment in operation is collected by sensors deployed in the power equipment, and the raw data information includes electrical information, environmental information, and image information.
[0037] The collected raw data information is preprocessed to obtain data information, and the data preprocessing at least includes data synchronization processing, data alignment processing, data cleaning processing, abnormality rejection processing, and normalization processing.
[0038] The data information is divided by a preset length time window to obtain a first power data set representing the running state of the power equipment; the first power data set includes a plurality of historical data.
[0039] In a second aspect, the application provides a device trend detection method, comprising
[0040] Real-time acquisition of power data of different modalities of the target device;
[0041] Inputting the power data into the device trend detection model to obtain the running trend of the target device in a future preset time; wherein the device trend detection model is trained according to the model training method of the first aspect and / or various possible embodiments of the first aspect;
[0042] Performing management measures corresponding to the trend result on the target device.
[0043] In a third aspect, the application provides a model training device, comprising:
[0044] The acquisition module is configured to collect a plurality of historical data of the power equipment during the running process of the power equipment to obtain a first power data set; each historical data has a different modality; the data in the first power data set is modified and / or deleted to obtain a second power data set, which is a missing data set corresponding to the first power data set;
[0045] The training module is configured to input the second power data set into an initial learning module as an input, input the output of a teacher model into the initial learning module as a training target, adjust the corresponding parameters in the initial learning module based on an evaluation function and an optimization algorithm, and obtain a feature learning module; wherein the teacher model is trained by the first power data set; the feature learning module is used to extract the first fusion feature corresponding to each historical data in the second power data set; the device trend detection module is trained based on each historical data and the first fusion feature corresponding to each historical data to obtain a trained device trend detection module; and the feature learning module and the device trend detection module are integrated to obtain a device trend detection model.
[0046] In a possible implementation, the training module is specifically configured to:
[0047] The feature learning unit in the initial learning module is used to extract the first fusion feature corresponding to each historical data in the second power data set;
[0048] The teacher model is used to extract the second fusion feature corresponding to each historical data in the first power data set;
[0049] The evaluation function is used to determine the error between the first fusion feature and the second fusion feature to obtain a loss value;
[0050] If the loss value is greater than or equal to the loss value threshold, based on the loss value, the corresponding parameters in the initial learning module are adjusted by an optimization algorithm until the loss value is less than the loss value threshold.
[0051] If the loss value is less than the loss value threshold, the feature learning module is obtained.
[0052] In a possible implementation, the feature learning unit includes a plurality of encoders, a feature switch, and a gated fusion unit, and the training module is further configured to:
[0053] The plurality of encoders share convolution kernel weights, and each piece of historical data in the second power dataset is encoded to obtain first power features corresponding to different modalities in each piece of historical data.
[0054] The feature switch exchanges the first power features corresponding to the target modality in each piece of historical data to obtain exchanged features corresponding to the target modality in each piece of historical data.
[0055] The gated fusion unit selectively transmits the exchanged features corresponding to different modalities to obtain first fusion features corresponding to each piece of historical data.
[0056] In a possible implementation, the training module is further configured to:
[0057] The feature switch exchanges the first power features corresponding to the target modality in each piece of historical data to obtain initial exchanged features corresponding to the target modality in each piece of historical data.
[0058] Pixel-level contrastive learning is performed to determine distances between the initial exchanged features corresponding to the target modality in each piece of historical data, respectively, to obtain a contrastive loss.
[0059] If the contrastive loss is greater than or equal to a contrastive loss threshold, the exchange mask tensor in the feature switch is adjusted to re-exchange the first power features corresponding to the target modality in each piece of historical data until the contrastive loss is less than the contrastive loss threshold.
[0060] If the contrastive loss is less than the contrastive loss threshold, the initial exchanged features are determined as the exchanged features corresponding to the target modality.
[0061] In a possible implementation, during training of the teacher model based on the first power dataset, the training module is further configured to:
[0062] The plurality of feature extractors in the initial model encode each piece of historical data in the first power dataset to obtain second power features corresponding to different modalities in each piece of historical data.
[0063] The contribution degrees of the first power features corresponding to different modalities are modeled through a modal attention mechanism and a missing simulation strategy, to obtain learning weights corresponding to each modality;
[0064] The second power features corresponding to each modality are dynamically weighted according to the learning weights corresponding to each modality, to obtain first fusion features;
[0065] Based on the first fusion features, an overall loss of an initial model is determined, the overall loss including: a modality consistency loss, an accuracy loss and a trend loss;
[0066] If the overall loss does not meet a preset requirement, parameters of the modal attention mechanism are adjusted and the learning weights corresponding to each modality are re-determined, until the overall loss meets the preset requirement;
[0067] If the overall loss meets the preset requirement, a teacher model is obtained.
[0068] In a possible implementation, the obtaining module is specifically configured to:
[0069] Raw data information of the power equipment in operation is collected through sensors deployed in the power equipment, the raw data information including: electrical information, environmental information and image information;
[0070] The collected raw data information is preprocessed to obtain data information; the data preprocessing at least includes data synchronization processing, data alignment processing, data cleaning processing, abnormality elimination processing and normalization processing;
[0071] The data information is divided into a first power dataset representing the running state of the power equipment in a preset length time window; the first power dataset includes multiple historical data.
[0072] In a fourth aspect, the present application provides a device trend detection apparatus, comprising:
[0073] An obtaining module is configured to obtain power data of different modalities of a target device in real time;
[0074] A detection module is configured to input the power data into a device trend detection model to obtain a running trend of the target device in a preset future time; wherein the device trend detection model is trained according to the model training method of the first aspect and / or various possible implementation manners of the first aspect;
[0075] A management module is configured to perform a management measure corresponding to the trend result on the target device.
[0076] In a fifth aspect, the present application provides an electronic device, comprising: a memory, a processor;
[0077] The memory stores computer execution instructions;
[0078] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above, and / or, the second aspect and / or various possible implementations of the second aspect as described above.
[0079] In a sixth aspect, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect, and / or the second aspect and / or various possible implementations of the second aspect.
[0080] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect, and / or the second aspect and / or various possible implementations of the second aspect.
[0081] The model training method, equipment trend detection method, device, and equipment provided in this application simulate incomplete data conditions by collecting multimodal historical data and generating a second power dataset with missing data, thus providing diverse training data for subsequent model training. By utilizing comparative learning between the teacher model and the initial learning module, knowledge transfer from complete data to missing data is achieved, enhancing the model's robustness under data loss conditions. The feature learning module extracts fusion features from the missing data and trains the equipment trend detection module, enabling it to accurately predict equipment operating trends. Integrating the feature learning module and the trend detection module forms a complete equipment trend detection model. This model can accurately detect power equipment operating trends even with missing data, avoiding performance degradation due to incomplete data and improving the efficiency and reliability of equipment operation and maintenance. Attached Figure Description
[0082] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0083] Figure 1 A schematic diagram of a scenario for the device trend detection method provided in the embodiments of this application;
[0084] Figure 2 A flowchart illustrating the model training method provided in the embodiments of this application. Figure 1 ;
[0085] Figure 3 A flowchart illustrating the model training method provided in the embodiments of this application. Figure 2 ;
[0086] Figure 4 A flowchart of a device trend detection method provided for an embodiment of the present application is shown in the figure;
[0087] Figure 5 A structural diagram of a model training device provided for an embodiment of the present application is shown in the figure;
[0088] Figure 6 A structural diagram of a device trend detection device provided for an embodiment of the present application is shown in the figure;
[0089] Figure 7 A structural diagram of an electronic device provided for an embodiment of the present application is shown in the figure.
[0090] The specific embodiments of the present application have been shown in the above figures, and will be described in more detail hereinafter. These figures and the written description are not intended to limit the scope of the present application concept in any way, but to illustrate the present application concept to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0091] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to indicate the same or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.
[0092] With the rapid development and intelligentization of power systems, real-time monitoring and trend analysis of the operating state of power equipment become particularly important. Traditional monitoring methods often rely on a single data source or simple data processing techniques, making it difficult to fully capture the complex correlations and dynamic changes between multi-modal data of power equipment. In addition, common data incompleteness problems in practical applications, such as sensor failures and data loss, further increase the difficulty of state monitoring.
[0093] In related technologies, the multi-modal fusion method still has deficiencies in handling data incompleteness and dynamic weight adjustment, limiting the adaptability and generalization performance of the model in complex industrial environments. There is a lack of effective processing mechanism for data incompleteness, which leads to a decline in model performance. The influence of data from different modalities on the state of the equipment varies under different conditions. Using a static weight allocation method cannot adapt to the dynamic changes in the device operating environment and data quality. In the process of multi-modal data fusion, the complementary information between different modalities cannot be effectively utilized, resulting in low accuracy of device trend identification.
[0094] The model training method provided by the embodiment of the application realizes knowledge transfer from complete data to missing data through contrast learning between the teacher model and the initial learning module, and enhances the robustness of the model under the condition of data missing. The fusion features of the missing data are extracted through the feature learning module, and the equipment trend detection module is trained, so that the equipment trend detection module can accurately predict the operation trend of the equipment. The feature learning module and the trend detection module are integrated to form a complete equipment trend detection model. The equipment trend detection model can accurately detect the operation trend of the power equipment under the condition of data missing, avoid the performance degradation caused by incomplete data, and improve the efficiency and reliability of equipment operation and maintenance.
[0095] Figure 1 The scene schematic diagram of the equipment trend detection method provided by the embodiment of the application is shown in FIG. 1. Figure 1 As shown in the figure, the specific application scenario of the application includes a target equipment 11 and a control center 12, wherein:
[0096] The target equipment 11 is a power equipment that needs to be subjected to trend detection. The monitoring of the operation state of the target equipment 11 is crucial for preventing faults and optimizing maintenance plans. A sensor network 13 is deployed on the target equipment 11, and the sensor network 13 is responsible for collecting real-time multi-modal power data. The sensor network 13 can transmit the collected power data to the control center 12 through network connection.
[0097] The control center 12 is the hub of data processing and analysis, responsible for receiving data from the target equipment 11, performing trend detection, and making decisions based on the results of trend detection. The control center 12 is usually equipped with high-performance computing devices for processing and analyzing power data collected from the target equipment 11.
[0098] The technical solutions of the application and how the technical solutions of the application solve the above technical problems will be described in detail in the following specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the application will be described below with reference to the accompanying drawings.
[0099] Figure 2 The flowchart of the model training method provided by the embodiment of the application is shown in FIG. 2. Figure 1 As shown in the figure, the method comprises the following steps. Figure 2
[0100] S201, in the operation process of the power equipment, a plurality of historical data of the power equipment are collected to obtain a first power data set; each historical data has different modalities.
[0101] During the operation of the power equipment, multi-modal data of the power equipment is collected through a sensor network. The power equipment is equipment in the power system for power generation, power transmission and power distribution. Optionally, the power equipment includes one or more of a generator, a transformer, a switching device and other devices in the power system. Optionally, the sensor network includes any one or more of a current sensor, a voltage sensor, a temperature sensor and a vibration sensor.
[0102] The multi-modal data includes data obtained from different sensor types or measurement methods. The multi-modal data includes electrical parameters, environmental parameters and mechanical parameters. The electrical parameters can further include voltage and current; the environmental parameters can further include temperature and humidity; and the mechanical parameters can further include vibration frequency and vibration amplitude.
[0103] Each piece of multi-modal data obtained is recorded and stored in a specific storage space to obtain a plurality of historical data. One piece of historical data includes data records generated by one power equipment during operation at different time points. Each piece of historical data can be continuous data or discrete data. Then, according to the plurality of historical data, a first power dataset is formed. The first power dataset includes complete modal data collected for the power equipment. The modal data of different modalities in the historical data carries information indicating the type of the modal data. For example, if the historical power data includes multi-modal data such as current, voltage, temperature and vibration, and the current is 100 A, the voltage is 220 V, the temperature is 50°C and the vibration is 0.1 g; the historical data can be represented as: {current: 100 A; voltage: 220 V; temperature: 50°C; vibration: 0.1 g}.
[0104] For example, assuming that historical data needs to be collected for power equipment A, first, a plurality of sensors installed on the power equipment A collect the historical data at a predetermined collection time and at a preset collection frequency to form a first power dataset. The predetermined collection time can be any one of 1 minute, 5 minutes, 1 hour, 1.5 hours, 1 day, 5 days, 1 month or any other feasible time. The preset collection frequency can be any one of 1 second / time, 30 seconds / time, 1 minute / time, 5 minutes / time, 1 hour / time, 2.5 hours / time, 1 day / time, 5 days / time, 1 month / time or any other feasible collection frequency.
[0105] S202, modifying and / or deleting the data in the first power dataset to obtain a second power dataset, the second power dataset being a missing dataset corresponding to the first power dataset.
[0106] Selecting a specific data point from the first power dataset and modifying the selected data point. Optionally, replacing the specific data point with an outlier or noise value. Selecting a specific data point from the first power dataset and deleting the selected data point, simulating a sensor failure or data loss situation. The first power dataset is modified and deleted to form a second power dataset.
[0107] For example, in the first power dataset, a specific temperature data point is selected and modified to simulate a temperature sensor failure or data noise pollution. If the original data of the specific temperature data point is 50℃, it is modified to an outlier value of 100℃.
[0108] S203, taking the second power dataset as the input of the initial learning module, taking the output of the teacher model as the training target of the initial learning module, adjusting the corresponding parameters in the initial learning module based on the evaluation function and the optimization algorithm, obtaining the feature learning module; wherein the teacher model is obtained by training the first power dataset.
[0109] First, use the teacher model to process the first power dataset to obtain the standard features output by the teacher model, and take the standard features as the training target of the initial learning module. The teacher model is a feature extraction model trained by the first dataset, which can extract standard features with good expression ability from the first dataset.
[0110] Input the second power dataset into the initial learning module; calculate the difference between the features output by the initial learning module and the standard features output by the teacher model through the evaluation function, and adjust the parameters of the initial learning module using the optimization algorithm to minimize the difference. The initial learning module is a machine learning model or neural network used to learn feature representation from the input second power dataset. The goal of the initial network training is to extract useful features from incomplete data. The evaluation function is a function used to evaluate the performance of the initial learning module. Optionally, the evaluation function can be a mean square error function or a cross-entropy loss function.
[0111] Repeat the above steps until the performance of the initial learning module reaches the expected value, and finally obtain the feature learning module. The feature learning module obtained by training and adjusting the initial learning module can extract effective feature representation from missing data.
[0112] S204, extracting the first fusion feature corresponding to each historical data in the second power dataset through the feature learning module.
[0113] The trained feature learning module is used for processing each piece of historical data in the second power data set, and first fusion features corresponding to each piece of data are extracted. The first fusion features fuse information of different modalities of each piece of historical data in the second power data set, and can effectively represent feature information of each piece of historical data in the second power data set.
[0114] In S205, the device trend detection module is trained based on each piece of historical data and the first fusion features corresponding to each piece of historical data, and a trained device trend detection module is obtained.
[0115] The device trend detection module is initialized, each piece of historical data in the second power data set and the first fusion features corresponding to each piece of historical data are taken as inputs of the device trend detection module, and actual running trend labels corresponding to each piece of historical data are taken as expected outputs, the device trend detection module is trained, parameters in the device trend detection module are adjusted through an optimization algorithm to minimize differences between predicted values and the actual running trend labels, and the trained device trend detection module is obtained. The performance of the trained device trend detection module is evaluated through a verification set or cross-validation corresponding to the second power data set, and it is ensured that the device trend detection module has good generalization ability.
[0116] In S206, the feature learning module and the device trend detection module are integrated, and a device trend detection model is obtained.
[0117] The output of the feature learning module is taken as the input of the device trend detection module, an end-to-end model architecture is constructed, and the device trend detection model is obtained. The integrated device trend detection model is adjusted and optimized as a whole, and it is ensured that the feature learning module and the device trend detection module can work cooperatively. The trained device trend detection model is deployed to actual applications, and is used for monitoring the running state of the power equipment in real time.
[0118] The model training method provided in the embodiments of the present application collects multi-modal historical data and generates a second power data set with missing data, simulates the case of incomplete data, and provides diversified training data for subsequent model training. Through contrast learning between the teacher model and the initial learning module, knowledge transfer from complete data to missing data is realized, and the robustness of the model in the case of missing data is enhanced. The feature learning module extracts fusion features of the missing data, and the device trend detection module is trained, so that the device trend detection module can accurately predict the running trend of the equipment. The feature learning module and the trend detection module are integrated to form a complete device trend detection model. The device trend detection model can accurately detect the running trend of the power equipment in the case of missing data, avoid performance degradation caused by incomplete data, and improve the efficiency and reliability of equipment operation and maintenance.
[0119] Figure 3A flowchart of a model training method provided by an embodiment of the present application Figure 2 As shown in Figure 3 the embodiment, the model training method is described in detail based on the embodiment, and the method comprises the following steps. Figure 2
[0120] In a possible implementation, the step S201 further comprises the following steps.
[0121] S2011, collecting original data information of the power equipment in operation by a sensor deployed in the power equipment, wherein the original data information comprises electrical information, environmental information and image information.
[0122] The sensor network is deployed at key positions of the power equipment, and the sensor network comprises a plurality of different sensors. The original data information of the power equipment in operation is collected in real time by the sensor network.
[0123] The original data information is raw data collected directly from the sensor without processing, and comprises electrical information, environmental information and image information. The electrical information comprises electrical data such as current and voltage; the environmental information comprises data such as temperature and humidity; and the image data comprises data such as the appearance and internal structure of the power equipment.
[0124] S2012, performing data preprocessing on the collected original data information to obtain data information; the data preprocessing at least comprises data synchronization processing, data alignment processing, data cleaning processing, abnormality elimination processing and normalization processing.
[0125] The data synchronization processing of the original data information comprises aligning different sampling frequency data in the collected original data information to the same frequency. The different modal data in the original data information are aligned in time stamp to realize data alignment processing. The missing data points of the original data information on the time axis are completed by linear interpolation or spline interpolation based on a set reference sampling period, and the modal consistency is maintained by using zero-order hold or sliding window filling for the discontinuous data stream. Meanwhile, the image data in the original data information is mapped with the time stamp and frame number, and is synchronized to the time line of the structured data to ensure that the data of different modal corresponds to a unique time index in each time slice.
[0126] The original data information is subjected to data cleaning processing. The missing data in the original data information is interpolated or filled to ensure the integrity of the data in the original data information. The abnormal data segments in the original data information that cannot be repaired are subjected to abnormal elimination processing to ensure the accuracy of the data. Through the data cleaning processing, the data defects caused by sensor abnormalities, signal noise or communication failures that may occur in the collection process are eliminated. For example, the legal value range is set to detect illegal values and null values in the original data information, and the outlier data and / or mutation anomaly in the original data information are obtained. Then, the outlier data or mutation anomaly is processed by using the Z-score method. For small range missing data segments, the adjacent value interpolation can be used, and for serious error data segments that cannot be repaired, the serious error data segments are directly eliminated to improve the quality of the data information obtained by processing.
[0127] The data of different dimensions in the abnormal elimination processing is subjected to normalization processing, so that the dimensions of the original data information are unified. In order to eliminate the scale difference between different modalities and sensor types, the original data information needs to be subjected to unified normalization or standardization processing. The time series data in the original data information is mapped to a unified interval by using the Z-score standardization, so as to enhance the fusion ability of different modalities in the feature space. For the image data in the original data information, the resolution size of the image is uniformly adjusted, and basic image enhancement processing such as pixel normalization, brightness enhancement and contrast adjustment is performed, so as to provide more rich data information.
[0128] S2013, the data information is divided into a first power data set representing the running state of the power equipment by using a time window with a preset length; the first power data set includes a plurality of historical data.
[0129] The continuous data stream in the preprocessed data information is divided according to a preset window length, and the data in each window is taken as an independent sample. Each sample includes a continuous data stream in the data information in a time period, and can represent the running state of the power equipment in the time period. All independent samples are combined to form a first power data set representing the running state of the power equipment. The preset window length can be any one of 1 second, 5 seconds, 10 seconds, 1 minute, 5 minutes, 1 hour or any other feasible window length.
[0130] In a possible implementation, the step S203 can further include:
[0131] S2031, the first fusion feature corresponding to each historical data in the second power data set is extracted by using the feature learning unit in the initial learning module.
[0132] The initial learning module includes a feature learning unit, which can extract a feature representation corresponding to the data output by the initial learning module. By processing each piece of historical data in the second power data set through the feature learning unit in the initial learning module, a first fusion feature corresponding to each piece of historical data is extracted.
[0133] S2032, through the teacher model, a second fusion feature corresponding to each piece of historical data in the first power data set is extracted.
[0134] Each piece of historical data in the first power data set is identified by the teacher model, and a second fusion feature corresponding to each piece of historical data is extracted. The second fusion feature can effectively represent the feature information of each piece of historical data in the first data set.
[0135] For example, assume that the historical data A in the first power data set includes current, voltage, temperature and vibration data. The teacher model identifies the feature of the historical data A and extracts a low-dimensional feature as the second fusion feature.
[0136] S2033, based on the evaluation function, the error between the first fusion feature and the second fusion feature is determined, and a loss value is obtained.
[0137] The evaluation function can measure the error between the first fusion feature and the second fusion feature. The evaluation function can be mean square error; or cross-entropy loss function.
[0138] For example, the evaluation function can be:
[0139]
[0140] wherein L D2 (F,F PO ) represents the loss value between the first fusion feature F and the second fusion feature F PO ; F i and F d represent the features of different modalities in the historical data in the first fusion feature F, and represent the features of different modalities in the historical data in the second fusion feature; i∈{4,5} is the difference between F i with index position 4 and F with index position 5. Those skilled in the art should understand that i∈{4,5} is only an example for understanding, and cannot be understood as a limitation on the range of i, which can be limited according to actual working conditions. For example, according to actual working conditions, the range of i is limited to i∈{2,5}.
[0141] S2034, if the loss value is greater than or equal to the loss value threshold, adjusting the corresponding parameters in the initial learning module based on the loss value by using an optimization algorithm until the loss value is less than the loss value threshold.
[0142] First, it is determined whether the loss value is greater than or equal to the loss value threshold. The loss value threshold is used to determine whether the initial learning module has reached the expected performance. If the loss value is greater than or equal to the loss value threshold, it means that the performance of the initial learning module needs to be further optimized.
[0143] If the loss value is greater than or equal to the loss value threshold, the parameters of the initial learning module are adjusted based on the loss value using an optimization algorithm to reduce the loss value until the loss value is less than the loss value threshold.
[0144] For example, when adjusting the parameters of the initial learning module, the optimization algorithm used can be a gradient descent algorithm and / or an Adam optimizer.
[0145] S2035, if the loss value is less than the loss value threshold, the feature learning module is obtained.
[0146] If the loss value is less than the loss value threshold, it means that the performance of the initial learning module has met the expectations, and the current initial learning module is determined as the feature learning module.
[0147] In one possible implementation, the feature learning unit includes a plurality of encoders, a feature switch and a gated fusion unit, and the above step S2031 can further include:
[0148] Step 1, encoding each piece of historical data in the second power dataset through a plurality of encoders sharing convolution kernel weights to obtain first power features corresponding to different modalities in each piece of historical data.
[0149] Each piece of historical data in the second power dataset is input into a plurality of encoders sharing convolution kernel weights. Then, each encoder uses shared convolution kernel weights to perform convolution operation on the data of the corresponding modality in the historical data to extract the power features of the corresponding modality in the historical data. The encoder is used to convert each piece of input historical data into a feature representation. Each encoder is responsible for processing data of one modality.
[0150] The power features of different modalities in the historical data are summarized to obtain the first power features corresponding to different modalities in each piece of historical data.
[0151] Encoding through a plurality of encoders sharing convolution kernel weights can enable the data of different modalities to use the same parameters in the convolution operation, realize parameter sharing, reduce the complexity of the feature learning unit and improve the feature learning efficiency.
[0152] Step 2, exchange the first power features corresponding to the target modality in each piece of historical data through the feature exchanger to obtain exchanged features corresponding to the target modality in each piece of historical data.
[0153] First, the first power features of different modalities in each piece of historical data are input into the feature exchanger. The feature exchanger determines the target modality that needs to be exchanged according to the feature exchange strategy. Then, the first power features corresponding to the target modality in each piece of historical data are exchanged according to the feature exchange strategy to obtain exchanged features corresponding to the target modality in each piece of historical data.
[0154] Optionally, the feature exchanger can determine the feature exchange strategy according to a preset rule, or learn the differences between features according to machine learning to obtain the feature exchange strategy.
[0155] Step 3, selectively transmit the exchanged features corresponding to different modalities through the gated fusion unit to obtain first fusion features corresponding to each piece of historical data.
[0156] The exchanged features of different modalities in each piece of historical data are input into the gated fusion unit. The gated fusion unit dynamically and selectively transmits feature information of different modalities according to the quality and importance of the current data to obtain first fusion features corresponding to each piece of historical data. The gated fusion unit can selectively transmit the exchanged features corresponding to different modalities and dynamically adjust the weights of the exchanged features.
[0157] The gated fusion unit is a unit for fusing multi-modal features. The structure of the gated fusion unit includes multiple convolution layers and gated neurons. The gated fusion unit can automatically determine the activation strength of different modalities, thereby dynamically adjusting the weights of the modalities, eliminating the need for manual weight adjustment, and enabling the gated fusion unit to flexibly balance the contributions of different modalities when processing multi-modal data, effectively avoiding information loss caused by forced alignment.
[0158] For example, if the exchanged features corresponding to two modalities are F I and F O , the sum of the exchanged features corresponding to the two modalities is F I+O =F I +F O . The gated fusion unit first performs convolution transformation and nonlinear mapping on F I , F O and F I+O , respectively, to calculate the hidden representation h:
[0159] h1=tanh(W1*F I )
[0160] h2=tanh(W2*F O )
[0161] h3 = tanh(W3 * F I+O )
[0162] wherein h1, h2 and h3 represent hidden representations; W1, W2 and W3 are corresponding convolution kernel weights; * represents a convolution operation; tanh is a hyperbolic tangent activation function.
[0163] Then, the gating neuron weights are calculated for each feature corresponding branch:
[0164]
[0165] wherein Z1 represents the gating neuron weights corresponding to the hidden representation h1; W1 ′ is a convolution weight for gating, [·] represents a channel concatenation operation of the feature, and σ is a Sigmoid activation function, represents a convolution operation. Similarly, the gating neuron weights Z2 corresponding to the hidden representation h2 and the gating neuron weights Z3 corresponding to the hidden representation h3 are also calculated in the same way:
[0166]
[0167] wherein W2 ′ and W3 ′ are both convolution weights for gating.
[0168] Finally, the first fusion feature S GSU corresponding to each piece of historical data is obtained through the gating fusion unit:
[0169] S GSU = Z1 ⊙ h1 + Z2 ⊙ h2 + Z3 ⊙ h3
[0170] wherein ⊙ represents an element-level multiplication operation. The first fusion feature S GSU can not only retain high-frequency detailed information of different modalities, but also realize balance and selection between modalities through dynamic gating weights, avoiding information loss caused by forced alignment.
[0171] In a possible implementation, the first power feature corresponding to the target modality in each piece of historical data is exchanged through a feature switcher to obtain an exchanged feature corresponding to the target modality in each piece of historical data, including:
[0172] Step a, exchanging the first power feature corresponding to the target modality in each piece of historical data through a feature switcher to obtain an initial exchanged feature corresponding to the target modality in each piece of historical data.
[0173] The first power feature of the target modality in each piece of historical data is input into the feature switch. The feature switch determines the target modality that needs to be switched according to the feature switching strategy. Further, the feature switch determines the initial switching feature corresponding to the target modality in each piece of historical data according to the feature switching strategy.
[0174] Step b, respectively determine the distance between the initial switching features corresponding to the target modality in each piece of historical data by pixel-level contrast learning, and obtain a contrast loss.
[0175] The distance between the initial switching features corresponding to the target modality in different historical data is quantified by pixel-level contrast learning, and a contrast loss is obtained. The contrast loss can compare the similarity and difference of the initial switching features corresponding to the target modality in each piece of historical data at the pixel level.
[0176] The distance between the initial switching features corresponding to the target modality in different historical data is usually quantified by cosine similarity or Euclidean distance.
[0177] Optionally, in the pixel-level contrast learning, the contrast loss L is determined by the contrast loss determination formula CL The corresponding contrast loss determination formula is:
[0178]
[0179] Wherein, C(i) represents the feature vector e i The category or modality to which it belongs; C(j) represents the feature vector e j The category or modality to which it belongs; e i Represents the initial switching feature corresponding to the target modality in the i-th piece of historical data; e j Represents the initial switching feature corresponding to the target modality in the j-th piece of historical data; r(e i ,e j ) represents the similarity measure between the feature vectors e i And e j , usually using cosine similarity or Euclidean distance; N p Represents the number of feature vectors of the same category or modality as e i ; k represents an index variable, used to traverse all e i The feature vectors of the same category or modality as e k .
[0180] Step c, if the contrast loss is greater than or equal to the contrast loss threshold, adjust the switching mask tensor in the feature switch, and re-switch the first power feature corresponding to the target modality in each piece of historical data until the contrast loss is less than the contrast loss threshold.
[0181] determining whether the contrast loss is greater than or equal to a contrast loss threshold. If the contrast loss is greater than or equal to the threshold, adjusting a swap mask tensor in the feature swapper to change a manner of feature swapping. Re-swapping the first power features corresponding to the target modality in each piece of historical data using the adjusted swap mask tensor. Repeating the above steps until the contrast loss is less than the contrast loss threshold. The swap mask tensor is a parameter tensor used to control feature swapping in the feature swapper, and determines the target modality and the initial swapped features corresponding to the target modality, and a manner of swapping the initial swapped features corresponding to different target modalities.
[0182] Optionally, the swap mask tensor can be:
[0183]
[0184] wherein n represents the number of pieces of historical data processed at one time; c represents the number of channels; h represents the height of the target modality; and w represents the width of the target modality. When w is even, w mod 2 = 0, and M(n, c, h, w) = 0, no swapping is performed. When w is odd, w mod 2≠0, and M(n, c, h, w) = 1, swapping is performed.
[0185] For example, if the target modality includes modality A and modality B, the first power features corresponding to modality A in each piece of historical data are [I1, I2, I3, I4]; the first power features corresponding to modality B in each piece of historical data are [T1, T2, T3, T4]; according to the swap mask tensor, feature swapping is performed at positions where w is odd, to obtain the first power features corresponding to the swapped modality A, which are [I1, T2, I3, T4]; and the first power features corresponding to the swapped modality B, which are [T1, I2, T3, I4].
[0186] Step d, if the contrast loss is less than the contrast loss threshold, determining the initial swapped features as the swapped features corresponding to the target modality.
[0187] determining whether the currently calculated contrast loss is less than the contrast loss threshold. If the contrast loss is less than the threshold, it indicates that the swapping effect of the feature swapper has met the optimization requirement, and the current initial swapped features are determined as the swapped features corresponding to the target modality.
[0188] In one possible implementation, the process of training the teacher model through the first power dataset includes:
[0189] Step 1, encoding each piece of historical data in the first power dataset through a plurality of feature extractors in the initial model, to obtain second power features corresponding to different modalities in each piece of historical data.
[0190] Each piece of historical data in the first power dataset is input into a plurality of feature extractors in the initial model. Each feature extractor encodes the data of different modalities in each piece of historical data, extracts features, and obtains second power features corresponding to different modalities in each piece of historical data.
[0191] The initial model is a machine learning model or a neural network model used to train the teacher model. The feature extractor is capable of extracting features from data of a specific modality.
[0192] Step 2, the contribution of the first power features corresponding to different modalities is modeled through a modality attention mechanism and a missing simulation strategy, and learning weights corresponding to each modality are obtained.
[0193] The second power features of different modalities in each piece of historical data are input into the modality attention mechanism. Through the modality attention mechanism, learning weights are assigned to each modality according to the quality and importance of the second power features; and the missing simulation strategy is used to enhance the adaptability of the initial model to incomplete data. The missing simulation strategy helps the initial model learn how to more effectively utilize information of other available modalities when some modality data is missing, thereby improving the robustness of the model. The modality attention mechanism is used to dynamically adjust the weights of features of different modalities and assign weights according to the quality and importance of the data. The learning weights reflect the importance and contribution of different modalities in overall trend prediction.
[0194] Step 3, the second power features corresponding to each modality are dynamically weighted according to the learning weights corresponding to each modality, and first fusion features are obtained.
[0195] The second power features of different modalities in each piece of historical data and the learning weights corresponding thereto are input into a weighting module. The features of each modality are weighted and summed according to the learning weights, and first fusion features are generated. The first fusion features are output and used for subsequent trend detection. The features of different modalities are weighted and summed according to the learning weights assigned by the modality attention mechanism, and first fusion features are generated. The first fusion features integrate multi-modal information, emphasize more important modal information, and have high expression ability.
[0196] For example, if each piece of power data includes two modalities of current and voltage; the second power feature of the current modality is A, and the learning weight is 0.6; the second power feature of the voltage modality is B, and the learning weight is 0.3. Then, through dynamic weighting processing, the first fusion features are 0.6×A+0.3×B.
[0197] Step 4, based on the first fusion features, the overall loss of the initial model is determined, and the overall loss includes modality consistency loss, accuracy loss, and trend loss.
[0198] According to the first fusion feature, the initial model is used for prediction to obtain the corresponding device trend prediction result. The consistency loss is obtained by measuring the difference between the device trend prediction result and the actual device trend. The accuracy loss of the initial model is determined between the category corresponding to the device trend prediction result and the category corresponding to the actual device trend. The trend loss is obtained by measuring the prediction ability of the initial model for the trend according to the device trend prediction result. Then, the consistency loss, the accuracy loss and the trend loss are weighted and summed to obtain the overall loss. The overall loss can balance the influence of the modal consistency loss, the accuracy loss and the trend loss on the training of the initial model, and ensure that the initial model can achieve good performance in multiple aspects.
[0199] Step 5, if the overall loss does not meet the preset requirement, the parameters of the modal attention mechanism are adjusted and the learning weight corresponding to each modal is re-determined until the overall loss meets the preset requirement.
[0200] Firstly, it is judged whether the overall loss meets the preset requirement. If the overall loss does not meet the preset requirement, the parameters in the initial model are adjusted and the initial model is re-trained. The above steps are repeated until the overall loss meets the preset requirement. For example, the parameters of the modal attention mechanism are adjusted and the learning weight corresponding to each modal is re-determined. The adjustment of the parameters of the modal attention mechanism causes the learning weight of different modal to change. By optimizing the learning weight, the capture ability of the initial model for key information can be improved.
[0201] Step 6, if the overall loss meets the preset requirement, the teacher model is obtained.
[0202] If the overall loss meets the preset requirement, the current initial model is determined as the teacher model.
[0203] Figure 4 A flowchart of a device trend detection method provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the method comprises the following steps. Figure 4
[0204] S401, real-time acquisition of power data of different modalities of a target device.
[0205] A sensor network is deployed on the target device to collect power data of different modalities of the target device in real time. The target device is a power device that needs to be detected for trend. Optionally, the target device can be one or more of a transformer, a generator or other power devices.
[0206] S402, inputting the power data into a device trend detection model to obtain the running trend of the target device in a future preset time; wherein the device trend detection model is trained according to the model training method of the above embodiments or any possible implementation manner of the above embodiments.
[0207] The real-time acquired power data is input into the device trend detection model to predict the running trend of the target device in a future preset time. The running trend can be a state that the target device can have in the future preset time. The running trend can be normal, abnormal, or failure. The future preset time is a time range for which the device trend detection model makes a prediction. Optionally, the future preset time can be 1 minute, 5 minutes, 1 hour, or other feasible times.
[0208] S403, performing a management measure corresponding to the running trend on the target device.
[0209] The corresponding maintenance, optimization, or early warning management measures are taken according to the running trend. The management measures include on-site regular maintenance, on-site preventive maintenance, and on-site troubleshooting.
[0210] For example, if the running trend shows that the target device is normal in the next 24 hours and the risk level is low, no special management measures need to be taken, and the running state of the target device continues to be monitored remotely.
[0211] For example, if the running trend shows that the target device can have an abnormality in the next 24 hours and the risk level is medium, preventive maintenance is performed. The maintenance plan of the operation and maintenance personnel is arranged according to the running trend, and the device is checked in advance to avoid failure.
[0212] The device trend detection method provided by the embodiments of the present application provides real-time input for the device trend detection model by acquiring the multi-modal power data of the target device in real time. The trained device trend detection model is used to predict the running trend of the target device in a future preset time according to the power data. According to the running trend of the device in the future preset time, corresponding management measures are formulated and executed to ensure the running of the target device and improve the reliability and operation and maintenance efficiency of the target device.
[0213] Figure 5 The structure schematic diagram of the model training device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the model training device 50 provided by the embodiments of the present application comprises: Figure 5
[0214] The acquisition module 501 is configured to collect a plurality of historical data of the power device during the running of the power device to obtain a first power data set; each historical data has a different mode; the data in the first power data set is modified and / or deleted to obtain a second power data set, and the second power data set is a missing data set corresponding to the first power data set;
[0215] The training module 502 is configured to take the second power data set as an input of an initial learning module, take an output of a teacher model as a training target of the initial learning module, adjust corresponding parameters in the initial learning module based on an evaluation function and an optimization algorithm, and obtain a feature learning module; the teacher model is obtained by training the first power data set; the feature learning module is configured to extract first fusion features corresponding to each piece of historical data in the second power data set; and the device trend detection module is trained based on each piece of historical data and the first fusion features corresponding to each piece of historical data, so as to obtain a trained device trend detection module; and the feature learning module and the device trend detection module are integrated to obtain the device trend detection model.
[0216] In a possible implementation, the training module 502 is specifically configured to: extract, by a feature learning unit in the initial learning module, first fusion features corresponding to each piece of historical data in the second power data set; extract, by the teacher model, second fusion features corresponding to each piece of historical data in the first power data set; determine, based on an evaluation function, an error between the first fusion features and the second fusion features, to obtain a loss value; if the loss value is greater than or equal to a loss value threshold, adjust, based on the loss value, corresponding parameters in the initial learning module by an optimization algorithm, until the loss value is less than the loss value threshold; and if the loss value is less than the loss value threshold, obtain the feature learning module.
[0217] In a possible implementation, the feature learning unit includes a plurality of encoders, a feature switch, and a gated fusion unit, and the training module 502 is further configured to: encode, by the plurality of encoders sharing convolution kernel weights, each piece of historical data in the second power data set, to obtain first power features corresponding to different modalities in each piece of historical data; exchange, by the feature switch, the first power features corresponding to a target modality in each piece of historical data, to obtain exchanged features corresponding to the target modality in each piece of historical data; and selectively transmit, by the gated fusion unit, the exchanged features corresponding to different modalities, to obtain the first fusion features corresponding to each piece of historical data.
[0218] In a possible implementation, the training module 502 is further configured to: exchange, by the feature switch, the first power features corresponding to the target modality in each piece of historical data, to obtain initial exchanged features corresponding to the target modality in each piece of historical data; respectively determine, by pixel-level contrast learning, distances between the initial exchanged features corresponding to the target modality in each piece of historical data, to obtain a contrast loss; if the contrast loss is greater than or equal to a contrast loss threshold, adjust an exchange mask tensor in the feature switch, and re-exchange the first power features corresponding to the target modality in each piece of historical data, until the contrast loss is less than the contrast loss threshold; and if the contrast loss is less than the contrast loss threshold, determine the initial exchanged features as the exchanged features corresponding to the target modality.
[0219] In a possible implementation, during the process of training the teacher model by the first power dataset, the training module 502 is further configured to: encode each piece of historical data in the first power dataset by the plurality of feature extractors in the initial model to obtain second power features corresponding to different modalities in each piece of historical data; model the contribution degrees of the first power features corresponding to different modalities by the modality attention mechanism and the missing simulation strategy to obtain learning weights corresponding to each modality; perform dynamic weighting processing on the second power features corresponding to each modality according to the learning weights corresponding to each modality to obtain first fusion features; and determine an overall loss of the initial model based on the first fusion features, where the overall loss includes a modality consistency loss, an accuracy loss, and a trend loss; if the overall loss does not meet a preset requirement, adjust parameters of the modality attention mechanism and re-determine the learning weights corresponding to each modality until the overall loss meets the preset requirement; and if the overall loss meets the preset requirement, obtain the teacher model.
[0220] In a possible implementation, the acquisition module 501 is specifically configured to: collect original data information of the power equipment in operation by sensors deployed in the power equipment, where the original data information includes electrical information, environmental information, and image information; perform data preprocessing on the collected original data information to obtain data information, where the data preprocessing at least includes data synchronization processing, data alignment processing, data cleaning processing, abnormality elimination processing, and normalization processing; divide the data information into time windows with a preset length to obtain a first power dataset representing the running state of the power equipment, where the first power dataset includes a plurality of pieces of historical data.
[0221] The model training apparatus provided in this embodiment can execute the method provided in the method embodiments, and has similar implementation principles and technical effects, which will not be described here in detail.
[0222] Figure 6 A structural schematic diagram of the device trend detection apparatus provided in this embodiment is shown in FIG. 6. Figure 6 As shown in FIG. 6, the device trend detection apparatus 60 provided in this embodiment includes:
[0223] The acquisition module 601 is configured to acquire power data of different modalities of a target device in real time.
[0224] The detection module 602 is configured to input the power data to a device trend detection model to obtain a running trend of the target device within a preset time in the future, where the device trend detection model is trained according to the model training method in the above embodiments or any possible implementation of the above embodiments.
[0225] The management module is configured to perform a management measure corresponding to the trend result on the target device.
[0226] The device trend detection apparatus provided in this embodiment can execute the method provided in the method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.
[0227] Figure 7 A structural schematic diagram of an electronic device is provided in this embodiment. As shown in the figure, Figure 7 The electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. The processor 701, the memory 702 and the communication component 703 are connected through a bus 704.
[0228] In the specific implementation process, the at least one processor 701 executes the computer execution instructions stored in the memory 702, so that the at least one processor 701 executes the method described above.
[0229] The specific implementation process of the processor 701 can refer to the method embodiments described above, and has similar implementation principles and technical effects. Details are not described herein again.
[0230] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as execution completed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0231] The memory can include a random access memory (RAM), and can also include a non-volatile memory (NVM), for example, at least one disk memory.
[0232] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the embodiments of the application is not limited to only one bus or one type of bus.
[0233] The embodiment of the present application further provides a computer program product comprising a computer program, which, when executed by a processor, implements the method described above.
[0234] The embodiment of the present application further provides a computer readable storage medium, which stores computer execution instructions, and when the computer execution instructions are executed, any method described above is implemented.
[0235] The readable storage medium described above can be implemented by any type of volatile or nonvolatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-only Memory (EPROM), Programmable Read-only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0236] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0237] The division of units is only a logical function division, and in actual implementation, there can be another division mode. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0238] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0239] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0240] If the function is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0241] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The aforementioned program can be stored in a computer readable storage medium. The program executes the steps including the above-mentioned method embodiments when executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, and various media that can store program codes.
[0242] Finally, it should be noted that those skilled in the art, after considering the specification and practicing the disclosed application, will easily think of other embodiments of the present application. The present application is intended to cover any variations, uses or adaptations of the present application that follow the general principles of the present application and include common knowledge or conventional technical means in the art that are not disclosed in the present application, and is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is only limited by the appended claims.
Claims
1. A model training method, characterized in that, include: During the operation of the power equipment, multiple historical data points of the power equipment are collected to obtain the first power dataset; Each piece of historical data possesses a different modality; Modify and / or delete the data in the first power dataset to obtain a second power dataset, which is a missing dataset corresponding to the first power dataset; The second power dataset is used as the input to the initial learning module, and the output of the teacher model is used as the training objective of the initial learning module. Based on the evaluation function and optimization algorithm, the corresponding parameters in the initial learning module are adjusted to obtain the feature learning module; wherein, the teacher model is trained using the first power dataset. The feature learning module extracts the first fusion feature corresponding to each historical data in the second power dataset. Based on each piece of historical data and the first fusion feature corresponding to each piece of historical data, the device trend detection module is trained to obtain the trained device trend detection module; By integrating the feature learning module and the device trend detection module, a device trend detection model is obtained.
2. The method according to claim 1, characterized in that, The step involves using the second power dataset as input to the initial learning module, using the output of the teacher model as the training objective of the initial learning module, and adjusting the corresponding parameters in the initial learning module based on the evaluation function and optimization algorithm to obtain the feature learning module, including: The first fusion feature corresponding to each historical data in the second power dataset is extracted through the feature learning unit in the initial learning module. Using the teacher model, extract the second fusion feature corresponding to each historical data point in the first power dataset; Based on the evaluation function, the error between the first fusion feature and the second fusion feature is determined, and the loss value is obtained; If the loss value is greater than or equal to the loss value threshold, the corresponding parameters in the initial learning module are adjusted based on the loss value using the optimization algorithm until the loss value is less than the loss value threshold. If the loss value is less than the loss value threshold, the feature learning module is obtained.
3. The method according to claim 2, characterized in that, The feature learning unit includes multiple encoders, feature switches, and gated fusion units. The step of extracting the first fused feature corresponding to each historical data point in the second power dataset through the feature learning unit in the initial learning module includes: By using the multiple encoders that share convolutional kernel weights, each piece of historical data in the second power dataset is encoded to obtain the first power feature corresponding to different modalities in each piece of historical data; By exchanging the first power feature corresponding to the target mode in each piece of historical data through the feature exchange, the exchange feature corresponding to the target mode in each piece of historical data is obtained; The gating fusion unit selectively transmits the exchange features corresponding to different modalities to obtain the first fusion feature corresponding to each piece of historical data.
4. The method according to claim 3, characterized in that, The step of exchanging the first power feature corresponding to the target mode in each piece of historical data through the feature exchange to obtain the exchange feature corresponding to the target mode in each piece of historical data includes: By exchanging the first power feature corresponding to the target mode in each piece of historical data through the feature exchange, the initial exchange feature corresponding to the target mode in each piece of historical data is obtained; By learning pixel-level contrast, the distance between the initial exchange features corresponding to the target modality in each piece of historical data is determined, and the contrast loss is obtained. If the contrast loss is greater than or equal to the contrast loss threshold, adjust the switching mask tensor in the feature switch, and re-swap the first power feature corresponding to the target mode in each piece of historical data until the contrast loss is less than the contrast loss threshold. If the contrast loss is less than the contrast loss threshold, the initial exchange feature is determined as the exchange feature corresponding to the target mode.
5. The method according to any one of claims 1-4, characterized in that, The process of training the teacher model using the first power dataset includes: By using multiple feature extractors in the initial model, each piece of historical data in the first power dataset is encoded to obtain the second power feature corresponding to different modes in each piece of historical data; By using a modal attention mechanism and a missing feature simulation strategy, the contribution of the first power feature corresponding to the different modalities is modeled to obtain the learning weight corresponding to each modal. The second power feature corresponding to each mode is dynamically weighted according to the learning weight corresponding to each mode to obtain the first fused feature; Based on the first fusion feature, the overall loss of the initial model is determined, and the overall loss includes: modality consistency loss, accuracy loss, and trend loss; If the overall loss does not meet the preset requirement, adjust the parameters of the modal attention mechanism and redetermine the learning weights corresponding to each modality until the overall loss meets the preset requirement; If the overall loss reaches the preset requirement, the teacher model is obtained.
6. The method according to any one of claims 1-4, characterized in that, During the operation of the power equipment, multiple historical data points of the power equipment are collected to obtain a first power dataset, including: By using sensors deployed in the power equipment, raw data information of the power equipment during operation is collected, including electrical information, environmental information, and image information; The collected raw data information is preprocessed to obtain data information; the data preprocessing includes at least data synchronization processing, data alignment processing, data cleaning processing, anomaly removal processing, and normalization processing. The data information is divided into time windows of a preset length to obtain the first power dataset representing the operating status of the power equipment; the first power dataset includes the multiple historical data.
7. A method for detecting equipment trends, characterized in that, include: Real-time acquisition of power data for different modes of the target device; The power data is input into the equipment trend detection model to obtain the operating trend of the target equipment within a preset future time period; wherein, the equipment trend detection model is trained by the model training method according to any one of claims 1-6; Implement management measures corresponding to the operating trend on the target device.
8. A model training device, characterized in that, include: The acquisition module is used to collect multiple historical data of the power equipment during its operation to obtain a first power dataset. Each piece of historical data possesses a different modality; Modify and / or delete the data in the first power dataset to obtain a second power dataset, which is a missing dataset corresponding to the first power dataset; The training module is used to take the second power dataset as input to the initial learning module and the output of the teacher model as the training objective of the initial learning module. Based on the evaluation function and optimization algorithm, the corresponding parameters in the initial learning module are adjusted to obtain the feature learning module. The teacher model is trained using the first power dataset. The feature learning module extracts the first fusion feature corresponding to each historical data point in the second power dataset. Based on each piece of historical data and the first fusion feature corresponding to each piece of historical data, the device trend detection module is trained to obtain the trained device trend detection module; the feature learning module and the device trend detection module are integrated to obtain the device trend detection model.
9. A device for detecting equipment trends, characterized in that, include: The acquisition module is used to acquire power data of the target device in different modes in real time; A detection module is used to input the power data into an equipment trend detection model to obtain the operating trend of the target equipment within a preset time period; wherein, the equipment trend detection model is trained by the model training method according to any one of claims 1-6; The management module is used to implement management measures corresponding to the trend results on the target device.
10. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6, and / or the method as described in claim 7.