Online detection method for multi-modal defect and anomaly detection
By acquiring multimodal real-time detection data, extracting cross-modal correlation feature parameters, building a detection evaluation model and deducing detection status data, determining key detection areas, adjusting detection frequency and generating early warning processes, the shortcomings of multimodal detection methods are addressed and efficient, accurate and dynamically adaptive detection effects are achieved.
Patent Information
- Application Number
- CN202511016549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing online detection methods for multimodal defect and anomaly detection have shortcomings in multimodal data fusion, feature extraction, real-time performance, model accuracy, and detection strategy adjustment. They are difficult to fully and accurately reflect the status of the detection object, and lack dynamic adaptability and real-time performance, resulting in irrational allocation of detection resources and affecting detection efficiency and effectiveness.
By acquiring multimodal real-time detection data, extracting cross-modal associated characteristic parameters, building a detection evaluation model and deducing the detection status data backward, calculating the state change, determining the key detection areas, generating detection adjustment instructions, adjusting the detection frequency according to historical status records, and performing comparison and early warning process judgment based on the latest status data.
The accuracy, real-time performance and adaptability of multimodal defect and anomaly detection have been improved, which can better reflect the state changes of the detection objects, rationally allocate detection resources, improve the flexibility and reliability of detection, and reduce the occurrence of accidents.
Smart Images

Figure CN120804994A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal defect and anomaly detection, in particular to an online detection method for multi-modal defect and anomaly detection. BACKGROUND
[0002] In many fields such as industrial production and equipment operation, defect and anomaly detection is crucial. With the development of technology, single-modal detection methods gradually reveal limitations. Traditional single-modal detection often only obtains information from a single dimension, making it difficult to comprehensively and accurately reflect the actual status of the detected object. For example, relying solely on visual images for detection may be disturbed by factors such as lighting and angle, leading to missed or false detections; while relying solely on sensor data for detection may not be able to capture some complex structural defects.
[0003] With the emergence of multi-modal detection technology, although it has made up for the shortcomings of single-modal detection to some extent, existing multi-modal defect and anomaly detection online detection methods still have many problems. The fusion processing of multi-modal data is not efficient enough, and the extraction of associated features between different modal data is not sufficient, which affects the accuracy of the detection model. For example, after obtaining multi-modal real-time detection data, the potential relationship between each modal data is not effectively mined, resulting in insufficient extraction of feature parameters and inability to accurately reflect the status of the detected object.
[0004] Existing detection methods have defects in real-time performance and dynamic adaptability. When facing complex and variable detection scenarios, they cannot adjust the detection strategy in a timely manner according to the changes in real-time data. For example, when the status of the detected object changes suddenly, it is not possible to quickly determine the key detection area, nor to adjust the detection frequency in a timely manner, which may result in the inability to discover and handle abnormal situations in a timely manner.
[0005] In addition, the existing detection evaluation model lacks sufficient use and scientific verification of historical data in the construction and deduction process, resulting in low prediction accuracy of the model. The method is not scientific and reasonable in calculating the state change quantity and determining the key detection area, which is prone to misjudgment. At the same time, when triggering the early warning process according to the comparison result, it lacks comprehensive consideration and reasonable instruction generation mechanism, affecting the reliability and effectiveness of the detection.
[0006] Moreover, in terms of management of detection nodes and adjustment of detection frequency, the existing method does not fully consider the actual situation of different regions and nodes, and the adjustment strategy is not fine enough. For example, when dividing the detection area and determining the detection frequency, it does not conduct scientific analysis according to the historical state record, resulting in unreasonable allocation of detection resources and affecting the detection efficiency and effect.
[0007] The existing multi-modal defect and anomaly detection online detection method has many deficiencies in multi-modal data fusion, feature extraction, real-time performance, model accuracy, detection strategy adjustment and the like, and an online detection method that is more efficient, accurate and dynamically adaptive is urgently needed to solve these problems. SUMMARY
[0008] The present application aims to provide an online detection method for multi-modal defect and anomaly detection to solve the problems raised in the background art.
[0009] To achieve the above object, the present application provides the following technical solution: an online detection method for multi-modal defect and anomaly detection, comprising: acquiring multi-modal real-time detection data, extracting cross-modal associated feature parameters, the feature parameters including instantaneous features and continuous features; constructing a detection evaluation model according to the extracted parameters and deducing the detection state data of several time periods backward to obtain initial detection state data; calculating the state change amount of adjacent time periods in the initial detection state data, and when the change amount reaches a set threshold, determining a key detection area according to the change amount and the corresponding relationship of the multi-modal data types; collecting the historical state records of a plurality of detection nodes within a set coverage range according to the key detection area; generating a detection adjustment instruction according to the historical state records, and delivering the adjustment instruction to the plurality of detection nodes, the adjustment instruction being used to indicate that the detection frequency of a specific time period is to be improved according to the historical state records; based on the detection evaluation model, supplementing the actual state data of the latest time period and acquiring updated detection state data of the next time period, comparing the initial detection state data and the updated detection state data of the next time period to generate a comparison result, and determining whether to trigger a warning process according to the comparison result.
[0010] Preferably, the acquisition of multi-modal real-time detection data and the extraction of cross-modal associated feature parameters comprise the following steps: acquiring real-time detection data of multi-modal equipment within a set time range; selecting all time periods that are in the same cycle as the current detection time period based on the set time range; extracting cross-modal associated feature parameters of all time periods in the real-time detection data.
[0011] Preferably, the construction of a detection evaluation model according to the extracted parameters and the deduction of the detection state data of several time periods backward to obtain initial detection state data comprise the following steps: constructing a detection evaluation model according to the cross-modal feature parameters of all time periods, and generating a first matching curve that meets a first matching condition; According to the first matching curve, the detection state data of several time periods is deduced backward to obtain initial detection state data.
[0012] Preferably, the state change amount of adjacent time periods in the initial detection state data is calculated, and when the change amount reaches a set threshold, a key detection area is determined according to the change amount and the corresponding relationship of the multi-modal data type, which comprises the following steps: Obtaining the data type characteristics of a single detection point in the state change amount; According to the data type characteristics and the total number of detection points in the state change amount, a key detection area is determined, which includes specific positions that need to be strengthened.
[0013] Preferably, the detection adjustment instruction is generated according to the historical state record, and the adjustment instruction is transmitted to the plurality of detection nodes, which comprises the following steps: Taking the key detection area as the center, the set coverage range is segmented in order from near to far to obtain a plurality of sub-regions within the set sub-range; Selecting detection nodes in each sub-region, determining the basic detection frequency of each sub-region within the error range allowed according to the historical state record of the detection nodes in each sub-region; Determining the detection nodes in each sub-region that meet the basic detection frequency to obtain a plurality of detection nodes, wherein the sum of the basic detection frequencies of the plurality of detection nodes is greater than or equal to the target detection frequency; Based on the plurality of detection nodes and the basic detection frequency of each detection node, a detection adjustment instruction is generated and issued; The historical state record of the plurality of detection nodes within the set coverage range is collected according to the key detection area, which comprises the following steps: Taking the key detection area as the center, the set coverage range is divided into three layers of regions with increasing radius; For the detection nodes in each layer region, the multi-modal historical data of the same period within the past six months is collected to form the historical state record of each node.
[0014] Preferably, the actual state data of the latest period is supplemented based on the detection evaluation model, and the updated detection state data of the next period is obtained, and the comparison between the initial detection state data and the updated detection state data is performed to generate a comparison result, which comprises the following steps: According to the state data of the latest period, the detection evaluation model is updated to generate a second matching curve that meets the second matching condition; The updated detection state data of the next period in the second matching curve is identified, and the fluctuation comparison between the initial detection state data and the updated detection state data is performed to generate a comparison result.
[0015] Preferably, the fluctuation comparison according to the initial detection state data and the updated detection state data generates a comparison result, and the comparison result comprises the following steps: extracting the predicted feature value of the next period in the initial detection state data and the actual feature value of the next period in the updated detection state data; respectively calculating the deviation degree of the feature predicted value and the actual value and the deviation degree of the other type of feature predicted value and the actual value, and taking the average value of the two deviation degrees as the fluctuation comparison result.
[0016] Preferably, the judging whether to trigger the early warning process according to the comparison result comprises the following steps: when the difference value of the detection data of the next period in the updated detection state data and the initial detection state data is within the allowable difference value, generating a first instruction to maintain the current detection frequency to maintain the detection frequency of the specific period; when the updated detection state data is less than the detection data of the next period in the initial detection state data, and the difference value between the two is greater than the allowable difference value, generating a partial node enhanced detection and early warning coexistence instruction according to the distribution of the basic detection frequency in the plurality of detection nodes, so that at least one node in the plurality of detection nodes detects other nodes except itself, and the detection frequency of the at least one node is maintained; when the updated detection state data is greater than the detection data of the next period in the initial detection state data, and the difference value between the two is greater than the allowable difference value, generating a second instruction to comprehensively enhance the detection, so as to completely improve the detection frequency of the specific period.
[0017] Preferably, the detection evaluation model is constructed according to the cross-modal feature parameters of all periods, and the first matching curve meeting the first matching condition is generated, and the first matching curve comprises the following steps: arranging the cross-modal feature parameters of all periods in time sequence to construct a training data set; using the historical data verification method to train the training data set, adjusting the model parameters until the fitting error of the model output is less than the first matching threshold, and generating the first matching curve.
[0018] Preferably, the state change amount of the adjacent period in the initial detection state data is calculated, and the state change amount comprises the following steps: selecting the cross-modal feature parameters of the continuous period in the initial detection state data to calculate the feature difference value between the current period and the previous period; taking the sum of the absolute values of the feature difference values as the state change amount, and the state change amount is used to reflect the defect detection fluctuation degree of the adjacent period.
[0019] Compared with the prior art, the present application has the following advantages: By acquiring multi-modal real-time detection data and extracting cross-modal correlated feature parameters, including instantaneous features and continuous features, the internal relationship between different modal data can be comprehensively and deeply mined, the defects of single modal detection information in traditional methods are overcome, and the actual state of the detected object can be more accurately reflected, thereby improving the comprehensiveness and accuracy of detection.
[0020] In the process of constructing the detection evaluation model and deducing the initial detection state data, a training data set is constructed using all the cross-modal feature parameters of all time periods, and a historical data verification method is used for model training. The model parameters are adjusted until the fitting error is less than the first matching threshold, and the first matching curve that meets the first matching condition is generated. This scientific model construction and verification method greatly improves the prediction accuracy of the model, making the deduced initial detection state data more reliable and laying a solid foundation for subsequent detection and analysis.
[0021] The state change amount of adjacent time periods in the initial detection state data is calculated, and when the change amount reaches a set threshold, the key detection area is determined according to the corresponding relationship between the change amount and the multi-modal data type. This method can quickly and accurately locate the area that may have defects or abnormalities. Compared with traditional methods, it avoids blind detection, improves the pertinence and efficiency of detection, and makes the detection resources more reasonably allocated to key areas.
[0022] The key detection area is taken as the center to divide the area and collect the historical state records of multiple detection nodes. According to these records, detection adjustment instructions are generated to indicate the increase of detection frequency in a specific period. This fine-tuned adjustment strategy based on historical data fully considers the actual situation of different areas and nodes, and can adjust the detection frequency according to the change rule of historical state, ensuring the effectiveness of detection and avoiding the waste of detection resources, improving the flexibility and adaptability of detection Based on the detection evaluation model, the actual state data of the latest period is supplemented, the updated detection state data of the next period is obtained, and the comparison result is generated by comparing the updated detection state data with the initial detection state data. According to the comparison result, it is judged whether to trigger the early warning process. This dynamic comparison and early warning mechanism can monitor the state change of the detection object in real time and discover abnormal situations in time. When the difference between the updated detection state data and the initial detection state data is within different ranges, different instructions are generated, such as maintaining the current detection frequency, coexisting of partial node strengthening detection and early warning, and comprehensive strengthening of detection, realizing the dynamic adjustment of detection strategy, improving the real-time and reliability of detection, and effectively preventing and reducing the occurrence of accidents.
[0023] By extracting the predicted characteristic values and the actual characteristic values in the initial detection state data and the updated detection state data, calculating the deviation degree and taking the average value as the fluctuation comparison result, the comparison result is more objective and accurate, and the state fluctuation of the detection object can be more sensitively reflected, and the detection accuracy and reliability are further improved.
[0024] The method of the application improves the accuracy, real-time performance, pertinence and adaptability of multi-modal defect and anomaly detection through multi-modal data fusion, scientific model construction, accurate key area positioning, refined detection frequency adjustment and dynamic early warning mechanism, and can better meet the high requirements of industrial production and equipment operation for defect and anomaly detection, and has significant economic and social benefits. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 The working principle diagram of the online detection method of multi-modal defect and anomaly detection described in the application; Figure 2 The flowchart of detection adjustment instruction generation and historical data acquisition; Figure 3 The flowchart of detection state data updating and comparison; Figure 4 The flowchart of fluctuation comparison result generation. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0027] Please refer to Figures 1-4 The online detection method of multi-modal defect and anomaly detection related to the application comprises the following specific implementation steps: Obtain multi-modal real-time detection data, extract cross-modal associated feature parameters, and the feature parameters include instantaneous features and continuous features. Specifically, obtain the real-time detection data of multi-modal equipment within a set time range, select all time periods in the same cycle as the current detection period based on the set time range, and then extract the cross-modal associated feature parameters of all time periods in the real-time detection data.
[0028] A detection evaluation model is constructed according to the extracted parameters, and the detection state data of several time periods is deduced backward to obtain initial detection state data. That is, a detection evaluation model is constructed according to the cross-modal feature parameters of all time periods, a training data set is constructed by arranging the cross-modal feature parameters of all time periods in time sequence, the model training is performed on the training data set by using the historical data verification method, the model parameters are adjusted until the fitting error of the model output is less than a first matching threshold, a first matching curve meeting the first matching condition is generated, and then the detection state data of several time periods is deduced backward according to the first matching curve to obtain the initial detection state data.
[0029] The state change amount of adjacent time periods in the initial detection state data is calculated, and when the change amount reaches a set threshold, a key detection area is determined according to the change amount and the corresponding relationship of the multi-modal data types. The specific steps are as follows: the cross-modal feature parameters of continuous time periods in the initial detection state data are selected, the feature difference value between the current time period and the previous time period is calculated, and the sum of the absolute values of each type of feature difference value is taken as the state change amount, which is used to reflect the defect detection fluctuation degree of adjacent time periods. When the state change amount reaches the set threshold, the data type characteristics of a single detection point in the state change amount are obtained, and according to the data type characteristics and the total number of detection points in the state change amount, the key detection area is determined, which contains the specific position that needs to be strengthened.
[0030] The historical state records of a plurality of detection nodes in a set coverage range are collected according to the key detection area. The set coverage range is divided into three layers of areas with increasing radiuses with the key detection area as the center, and the multi-modal historical data of the same time period in the past six months is collected for the detection nodes in each layer of area to form the historical state records of each node.
[0031] Detection adjustment instructions are generated according to the historical state records, and the adjustment instructions are transmitted to a plurality of detection nodes, which are used to indicate that the detection frequency of a specific time period is improved according to the historical state records. Specifically, the set coverage range is divided into segments in order from near to far with the key detection area as the center to obtain a plurality of sub-areas in the set sub-range, the detection nodes in each sub-area are selected, the basic detection frequency of each sub-area in which the error is within the allowable range is determined according to the historical state records of the detection nodes in each sub-area, the detection nodes in each sub-area that meet the basic detection frequency are determined, and a plurality of detection nodes are obtained, wherein the sum of the basic detection frequencies of the plurality of detection nodes is greater than or equal to the target detection frequency, and the detection adjustment instructions are generated and issued based on the plurality of detection nodes and the basic detection frequency of each detection node.
[0032] The actual state data of the latest period is supplemented based on the detection evaluation model, and the updated detection state data of the next period is obtained. The comparison of the next period is performed according to the initial detection state data and the updated detection state data, and the comparison result is generated. According to the comparison result, it is judged whether to trigger the early warning process. Specifically, the detection evaluation model is updated according to the state data of the latest period, a second matching curve meeting a second matching condition is generated, the updated detection state data of the next period in the second matching curve is identified, the predicted feature value of the next period in the initial detection state data and the actual feature value of the next period in the updated detection state data are extracted, the deviation degrees of the feature predicted value and the actual value and the deviation degrees of the other type of feature predicted value and the actual value are calculated respectively, and the average value of the two deviation degrees is taken as the fluctuation comparison result. According to the comparison result, it is judged whether to trigger the early warning process. When the difference between the updated detection state data and the initial detection state data of the next period is within the allowed difference, a first instruction to maintain the current detection frequency is generated to maintain the detection frequency of the specific period; when the updated detection state data is less than the detection data of the next period in the initial detection state data, and the difference between them is greater than the allowed difference, according to the distribution of the basic detection frequency in the plurality of detection nodes, a part of the node strengthening detection and early warning coexistence instruction is generated, so that at least one node in the plurality of detection nodes detects the other nodes except itself, and the detection frequency of the at least one node is maintained; when the updated detection state data is greater than the detection data of the next period in the initial detection state data, and the difference between them is greater than the allowed difference, a second instruction to comprehensively strengthen the detection is generated to completely improve the detection frequency of the specific period.
[0033] In embodiment 1, multi-modal real-time detection data is obtained, and cross-modal associated feature parameters are extracted, which include instantaneous features and continuous features. In the process of obtaining multi-modal real-time detection data and extracting cross-modal associated feature parameters, the time range is explicitly set. The determination of this time range needs to be combined with the specific detection scene and the operation law of the equipment. For example, for production equipment with periodic operation characteristics, a complete production cycle may be used as the set time range, such as 24 hours, 12 hours, etc. After determining the time range, the multi-modal equipment collects the real-time detection data in the time range. The multi-modal equipment includes multiple types, such as visual detection equipment, which can obtain image information of objects through cameras, etc., for detecting defects, shape abnormalities, etc. on the surface of the object; infrared detection equipment, which can sense the temperature distribution of the object, for discovering temperature abnormal areas; and ultrasonic detection equipment, which can detect defects inside the object using ultrasonic waves, etc. These devices continuously collect data within the set time range to form a multi-modal real-time detection data set.
[0034] Based on the set time range, all time periods in the same cycle as the current detection period need to be selected. The key of this step is to analyze the periodic characteristics of the device operation. For example, if the current detection period is 9 am on a certain day, and the device has similar running state at around 9 am every day, it belongs to the same cycle period, then the 9 am period of the previous several days (such as the previous 7 days, the previous 15 days, etc.) needs to be selected. When selecting, it needs to be ensured that the selected period is indeed in the same cycle position, which may need to be verified by analyzing the historical data, such as observing whether the device running parameters, production data, etc. of the same period on different dates have similar change trends or distribution characteristics.
[0035] After completing the time period selection, cross-modal correlation feature parameters of all time periods are extracted from real-time detection data. These feature parameters mainly include two categories: instantaneous features and continuous features. Instantaneous features are features that can reflect the state of the device or the detection object at a specific moment, such as the temperature peak value of a certain point detected by an infrared detection device at a certain moment, which can directly indicate the temperature abnormality of the point at that moment; the edge features of a certain area in the image obtained by the visual detection device can be used to judge whether the area has shape defects. Continuous features are features extracted based on the change of data within a period of time, such as the temperature change trend over time, by analyzing the rising or falling rate of temperature within a period of time, it can be judged whether the device has overheating or cooling abnormality; the area change rate of the defect area in the image, by calculating the area change of the defect area at different moments, the development trend of the defect can be understood.
[0036] When extracting cross-modal correlation feature parameters, the correlation between different modal data needs to be considered. For example, when visual detection finds that there is a surface defect in a certain area, the temperature feature of the area under the infrared mode is also checked, and whether there is a correlation between the two, such as whether the defect area is accompanied by temperature anomaly, this cross-modal correlation analysis can more comprehensively reflect the state of the detection object, and improve the accuracy of defect and anomaly detection.
[0037] In order to ensure the reliability and effectiveness of the extracted feature parameters, the collected real-time detection data also needs to be preprocessed. The preprocessing process includes noise removal processing, as the detection device may be affected by environmental noise, electromagnetic interference, etc., resulting in noise in the collected data, the noise is removed by filtering algorithm, etc., so that the data is more smooth and accurate; normalization processing, different modal data may have different dimensions and value ranges, normalization processing can unify the data to the same scale, which is convenient for subsequent analysis and model construction.
[0038] In extracting instantaneous features, it is necessary to accurately capture key data points at each time. For example, for temperature peaks, the maximum value of temperature and its corresponding position and time need to be found in each time data within a certain time range, and the difference between the peak value and the surrounding temperature is also recorded. For edge features of visual images, edge detection algorithms such as Canny algorithm can be used to identify the edges of objects in the image and extract features such as length, shape, and direction of the edges.
[0039] In extracting continuous features, statistical analysis and trend fitting of data within a period of time are needed. For example, for temperature change trend, a time window (such as 1 hour, 2 hours, etc.) is selected, and the temperature data points within the window are collected. Through linear regression or other fitting methods, the change slope of temperature is calculated, and the change trend feature of temperature is obtained. For defect area change rate, the defect area is segmented and the area is calculated at different times, and then the difference between the areas of adjacent times and the ratio of the area of the previous time are calculated to obtain the area change rate.
[0040] In the whole feature extraction process, the time sequence consistency of data needs to be ensured, and the feature parameters of each period of time and the corresponding time points are accurately corresponding. At the same time, for different modal data, a unified time stamp needs to be established to facilitate cross-modal feature correlation and analysis.
[0041] In addition, the extraction of feature parameters also needs to be adjusted and optimized according to the specific detection target and application scene. For example, when detecting defects of metal parts, more attention may be paid to the edge features of visual modal and the temperature features of infrared modal; while detecting plastic parts, more consideration may be given to the color features and surface texture features of visual modal, etc.
[0042] Through the above detailed steps, cross-modal correlation feature parameters containing instantaneous features and continuous features are extracted from multi-modal real-time detection data. These parameters can comprehensively and accurately reflect the state of the detection object at different times and different modalities, laying a solid data foundation for subsequent construction of detection evaluation model, state deduction and anomaly detection. The whole process strictly follows the logical order of data collection, time selection, feature extraction and data preprocessing, ensuring the accuracy and reliability of each link, so as to ensure the effectiveness of the subsequent detection method.
[0043] Example 2, construction of detection evaluation model In constructing the detection evaluation model and deducing the initial detection state data backward, the cross-modal feature parameters of all time periods are arranged in time sequence to construct the training data set. All time periods here refer to all historical time periods in the same cycle as the current detection time period selected when obtaining real-time multi-modal detection data. For example, if the current detection time period is 10 o'clock in the morning, and the cycle is a week, the training data set may include the cross-modal feature parameters at 10 o'clock in the morning of each day in the previous 4 weeks. The feature parameters of each time period need to integrate multi-modal data such as vision, infrared, and ultrasound. For example, the feature parameter set of a certain time period can include image edge complexity, temperature gradient distribution, ultrasonic reflection intensity, etc., and each index needs to correspond to an accurate timestamp to ensure the continuity and accuracy of the time sequence.
[0044] When constructing the training data set, the multi-modal feature parameters need to be structured. For example, the image features of the vision modality are converted into numerical feature vectors, such as defect area texture feature values extracted by a convolutional neural network; the temperature data of the infrared modality are converted into temperature values and spatial distribution features of each detection point; the echo data of the ultrasound modality are converted into defect depth, size, etc. After these parameters are arranged in chronological order, a multi-dimensional time series data set is formed, where each row represents a combination of multi-modal features of a time period, and each column corresponds to a specific feature index of different modalities.
[0045] The training data set is trained using the historical data validation method. The selection of the model needs to be combined with the characteristics of multi-modal data. Supervised learning models in machine learning, such as random forest, gradient boosting tree, etc., or recurrent neural networks (RNN), long short-term memory networks (LSTM) in deep learning, etc., can be used to process the time sequence dependence in time series data. In the training process, the training data set needs to be divided into a training set and a validation set, usually in a ratio of 7:3 or 8:2. The model parameters are optimized through the training set, and the generalization ability of the model is evaluated using the validation set.
[0046] The core of model training is to adjust the parameters so that the fitting error of the model output is less than the first matching threshold. Taking the LSTM model as an example, the parameters to be adjusted include the number of hidden layer neurons, the learning rate, the number of iterations, the dropout rate, etc. During the training process, the model predicts the feature value of the next time based on the input historical cross-modal feature parameters, and calculates the error with the actual value in the validation set. The commonly used error indicators are mean square error (MSE) or mean absolute error (MAE). The model weights are continuously updated through the backpropagation algorithm until the error on the validation set does not decrease significantly for several iterations, and the final error value is less than the pre-set first matching threshold (such as 0.1 or 0.05 relative error standard according to the actual detection scene).
[0047] When the model training meets the first matching condition, a first matching curve is generated. The curve is the fitting result of the model to the historical multi-modal feature parameter time series, which can reflect the trend of each feature parameter changing with time and the mutual relationship. For example, in the temperature-time curve, the first matching curve can reflect the periodic fluctuation range of temperature under normal working conditions; in the defect area-time curve, it can reflect the normal rate range of defect development. The generation of the curve needs to be based on the overall fitting of the historical data by the model, rather than the prediction of a single time point, so it can more comprehensively capture the dynamic change law of multi-modal data.
[0048] When the first matching curve is used to deduce the detection state data of several time periods backward, the number of time periods to be deduced and the time interval need to be determined. The setting of the number of time periods needs to balance the real-time of detection and the accuracy of prediction, and usually 3-5 time periods can be deduced backward, and the time interval of each time period is consistent with the data acquisition frequency, such as the data acquisition frequency is once an hour, then each time period represents 1 hour. In the deduction process, the model is based on the first matching curve, and the feature parameter values of future time periods are extrapolated according to the change trend of historical data, for example, through the recursive calculation of the LSTM model, the predicted feature values of the next time period, the next time period, etc. are generated in turn, and the prediction of each time period depends on the prediction result of the previous time period, forming a continuous prediction sequence.
[0049] In the process of constructing the detection evaluation model and deducing the initial detection state data backward, when there is an association between multi-modal feature parameters, a cross-modal association mechanism preset in the model needs to be used for systematic adjustment to ensure the consistency of multi-modal prediction data. This process closely combines the internal relationship of multi-modal data, relies on the model's learning of historical association rules, and realizes the coordinated change of each modal data when deducing the future state.
[0050] The systematic adjustment is that, when constructing the detection evaluation model and deducing the initial detection state data backward, the learning of the model on the historical multi-modal feature parameter correlation is relied on to realize the collaborative change of each modal data. Taking the metal plate rolling detection as an example, the model masters the correlation among the infrared modal temperature, the visual modal scratch and the ultrasonic modal stress through the historical data. When deducing that the temperature of a certain period of time rises from 60℃ to 65℃, the corresponding region scratch length will be adjusted from 4mm to 6mm and the internal stress will be adjusted from 80MPa to 95MPa according to the correlation, so that the changes of each modal parameter are echoed. If an abnormal fluctuation of a single modal parameter occurs in the deduction, such as the visual scratch suddenly increases to 8mm but the temperature and stress do not change, and the historical data shows that the scratch exceeds 7mm, which requires the temperature to rise by more than 5℃ and the stress to increase by more than 10MPa, the model will start cross-modal verification, adjust the temperature by 6℃ and the stress by 12MPa, and correct the scratch to 7.5mm to eliminate the contradiction. In the multi-region scene such as large generator set detection, the model considers the spatial correlation. When deducing that the vibration amplitude of a certain region of the shell increases, the model will simultaneously increase the predicted value of the temperature of the corresponding stator winding. If there is a heat dissipation correlation in the adjacent region, it will also drive the surrounding temperature to rise slightly to form a distribution conforming to the heat conduction law. Moreover, the model will dynamically adjust the correlation strength according to the working condition, such as weakening the influence of temperature on scratch during low-speed rolling and strengthening it during high-speed rolling. Through incremental learning, new correlation data is included. When the material quality of the plate changes, the correlation law changes, and the model is updated to ensure that the parameter adjustment in the subsequent deduction reflects the latest characteristics, thereby ensuring the consistency of the multi-modal prediction data.
[0051] Taking the detection of the rolling process of metal plates in industrial production as an example, three types of data are involved in this scene, namely, the visual modal (detecting plate surface scratches), the infrared modal (detecting plate temperature distribution) and the ultrasonic modal (detecting plate internal stress), and there is a significant correlation among them: when the temperature of the plate abnormally rises, scratches are more likely to appear on the surface, and the internal stress will also increase. When constructing the training data set, the feature parameters of the three modalities are arranged in the time sequence of rolling, including scratch length, depth, etc. for the visual modal, temperature values for the infrared modal, stress peak values for the ultrasonic modal, and each parameter corresponds to a specific rolling period, ensuring the continuity of the time sequence.
[0052] In the model training stage, a deep learning model capable of capturing cross-modal correlation is used to learn the quantitative relationship between temperature rise, scratch deepening and stress increase through learning of historical data. For example, the model learns that when the temperature of a certain region exceeds 60℃, the probability of the occurrence of a scratch with a length exceeding 5mm in that region increases by 30%, and the internal stress value increases by an average of 15MPa. During the training process, the model parameters are adjusted so that these correlation rules are fully learned until the fitting error of the model output is less than the first matching threshold, generating a first matching curve that can reflect the multi-modal correlation.
[0053] When inferring the initial detection state data of the future period, the model will collaboratively adjust the parameters of each modality based on the first matching curve and the cross-modality correlation mechanism. Assuming that according to the historical data inference, the plate temperature will rise to 65℃ at a certain period in the future (such as the period of rolling speed increase), the model will not only adjust the temperature parameter of the infrared modality, but also will adjust the prediction values of the visual modality and the ultrasonic modality simultaneously according to the learned correlation law: in the area where the temperature rises, the prediction value of the scratch length of the visual modality increases accordingly, and the prediction value of the internal stress of the ultrasonic modality also rises synchronously. For example, when the model predicts that the temperature rises from 60℃ to 65℃, it will automatically adjust the prediction value of the scratch length in this area from 4mm to 6mm, and the prediction value of the internal stress from 80MPa to 95MPa, ensuring that the changes of the three conform to the historical correlation law.
[0054] If abnormal fluctuations occur in a modality parameter during the inference process, the model will correct it through the cross-modality verification mechanism. For example, if the model predicts that the scratch length of the visual modality suddenly increases to 8mm at a certain period, but the temperature of the infrared modality and the stress of the ultrasonic modality do not change significantly, which does not conform to the historical correlation law. At this time, the model will call the stored historical correlation data for verification and find that when the scratch length exceeds 7mm, the temperature rises by at least 5℃ and the stress increases by 10MPa. Therefore, it is determined that the prediction is contradictory, and then the parameters of each modality are corrected: the temperature prediction value is increased by 6℃, the stress prediction value is increased by 12MPa, and the scratch length is corrected to 7.5mm, so that the changes of the three remain consistent.
[0055] In the scene involving multi-region collaborative detection, the model will also consider the cross-modality correlation in the spatial dimension. For example, in the detection of large generator sets, the visual modality detects the vibration amplitude of the generator shell, and the infrared modality detects the temperature of the stator winding. There is a corresponding relationship between the two in space - the area with larger shell vibration amplitude often corresponds to the position with higher winding temperature. When the model infers that the vibration amplitude of a certain area increases, it will simultaneously increase the prediction value of the winding temperature of the corresponding area, and if there is a heat dissipation correlation in the adjacent area, it will also drive the temperature prediction value of the surrounding area to rise slightly, forming a temperature field distribution that conforms to the heat conduction law, avoiding the contradictory prediction of violent vibration but stable temperature.
[0056] In addition, the model will dynamically adjust according to the differences in correlation strength under different working conditions. In the above metal plate rolling scenario, when the rolling speed is in the low speed interval, the correlation strength between temperature and scratch is weak, while at high speed rolling, the correlation strength is significantly enhanced. The model learns the historical data in different speed intervals, and when deducing the low speed period, it allows small changes in temperature while keeping the scratch parameter stable; but when deducing the high speed period, it will strengthen the influence weight of temperature change on the scratch parameter, to ensure that the predicted data conforms to the correlation characteristics under actual working conditions.
[0057] As new detection data is continuously collected, the model will regularly incorporate the latest cross-modal correlation data for updating to adapt to slow changes in equipment state. For example, when the material of the metal plate is slightly adjusted, the correlation between temperature and stress may change, and the model will integrate the new correlation data into the first matching curve through incremental learning, so that the parameter adjustment in the subsequent deduction process can reflect the latest correlation characteristics, avoiding inconsistent prediction data due to outdated correlation rules.
[0058] In the correlation data analysis and adjustment process, when new detection data is continuously collected, for the scenario of slow changes in equipment state (such as slight adjustment of the material of the metal plate), continuously collect multi-modal detection data after adjustment, and focus on the real-time changes of temperature and stress and other correlation characteristic parameters. By comparing and analyzing the historical data before and after adjustment, identify the change trend of the temperature and stress correlation rule, for example, originally temperature increases by 10℃ corresponds to stress increase by 15MPa, after the material is adjusted, the actual detection shows that when the temperature increases by 10℃, the stress only increases by 12MPa, thus determining the new correlation characteristics. The model uses incremental learning to integrate these newly collected correlation data into the latest part of the training data set in time sequence, without retraining all historical data, only adjusting the model parameters for new data segments. In the training process, the influence weight parameter of temperature features on stress features is optimized, and the fitting error of the model for new correlation data is iteratively adjusted to gradually reduce to below the set threshold through historical data verification method. At the same time, update the first matching curve according to the new correlation rule, and correct the corresponding slope of temperature and stress change in the curve, for example, adjust the slope of the temperature-stress correlation segment in the curve from the original 1.5 (MPa / ℃) to 1.2 (MPa / ℃), to ensure that the curve accurately reflects the correlation characteristics under the current material. After completing the parameter and curve adjustment, when deducing the initial detection state data in the subsequent process, the model will adjust the parameters according to the updated first matching curve, and when the predicted temperature changes, the corresponding output stress prediction value conforms to the new correlation rule, avoiding inconsistent prediction data due to outdated correlation rules caused by material adjustment, and ensuring that the multi-modal prediction data always conforms to the actual correlation characteristics of the current equipment state.
[0059] Through this collaborative adjustment mechanism based on historical correlation rules, the model can ensure that the changes of each modality parameter are mutually echoed and consistent with the actual physical rules when deducing the initial detection state data backward, thereby realizing the consistency of multi-modal prediction data and providing reliable basic data for subsequent state change calculation and key detection area determination.
[0060] The obtained initial detection state data contains multi-modal feature prediction values of each deduced period, such as temperature distribution prediction values, image defect feature prediction values, and ultrasonic echo feature prediction values of the next 3 periods. These data are stored in the form of structured tables or time series arrays, and the prediction data of each period contains feature parameters of all modalities, facilitating subsequent state change calculation and anomaly analysis.
[0061] In addition, during the construction of the model and the deduction process, the training data set needs to be updated regularly. With the continuous acquisition of new detection data, the latest historical period data can be included in the training data set, the model can be retrained and new first matching curves can be generated to adapt to the slow changes of the equipment operating state or the subtle adjustments of the detection environment, ensuring that the prediction accuracy of the model does not decrease over time. For example, the training data set is updated once a month, the detection data of the past month is included, and the oldest month of data is eliminated to maintain the timeliness of the data set.
[0062] The entire process needs to strictly follow the logical flow of data preprocessing, model training, curve generation, and state deduction, and each link needs to ensure the correctness of the time sequence of the data and the correlation of the multi-modal features. The historical data verification method is used to ensure the accurate fitting of the model to the historical rules, and the initial detection state data with time continuity is generated by backward deduction to provide a reliable prediction benchmark for subsequent calculation of the state change of adjacent periods and determination of the key detection area, thereby realizing online detection and early warning of multi-modal defects and anomalies.
[0063] Example 3, calculation of the state change of adjacent periods in the initial detection state data In calculating the state change of adjacent periods in the initial detection state data and determining the key detection area, the cross-modality feature parameters of consecutive periods in the initial detection state data need to be selected. The initial detection state data here is the prediction data of future periods obtained by backward deduction of the detection evaluation model, for example, T1, T2, and T3 are deduced backward, and then the consecutive periods may be T1 and T2, T2 and T3. For each selected consecutive period, the feature parameters of all modalities in that period need to be obtained, such as defect area of visual modality, temperature value of infrared modality, and echo intensity of ultrasonic modality. These parameters are all numerical values predicted by the model, and each parameter corresponds to a specific detection point position and time identifier.
[0064] The feature difference between the current period and the previous period is calculated. Taking a certain type of feature at a certain detection point as an example, assuming that the feature value of the previous period (such as T1) is , and the feature value of the current period (such as T2) is , then the difference between the two periods is . The feature difference needs to be calculated for each type of feature at each detection point, such as the temperature difference between adjacent periods for each temperature detection point, the edge complexity difference for each edge detection point, etc.
[0065] The sum of the absolute values of the feature differences is taken as the state change. Assuming that there are N detection points in a certain adjacent period, and each detection point contains M types of feature parameters, the calculation formula of the state change is:
[0066] where represents the state change, which reflects the degree of fluctuation of the adjacent period defect detection; represents the th detection point ( ); represents the th feature parameter ( ); represents the difference between the current period and the previous period of the th feature of the th detection point. This formula accumulates the absolute values of the feature differences of all detection points, and comprehensively reflects the overall change amplitude of the multi-modal features in the adjacent period. The larger the value, the more significant the fluctuation of the detection state.
[0067] After calculating the state change, it needs to be compared with the set threshold value. The set threshold value is a reference value determined in advance according to the specific detection scene and equipment running characteristics, for example, by analyzing the state change distribution of historical normal operation data, the threshold value is set to the mean value plus several times the standard deviation, to ensure that within the normal fluctuation range, the determination of the key detection area will not be triggered. When the state change reaches or exceeds the set threshold value, it indicates that the detection state has a significant fluctuation, and the key detection area needs to be further determined.
[0068] When the threshold condition is triggered, the data type characteristics of the individual detection points in the state change are first obtained. The data type characteristics refer to the modal type and specific feature category corresponding to each detection point, such as a certain detection point belonging to the temperature detection type of the infrared modal, its feature being the temperature value; a certain detection point belonging to the edge detection type of the visual modal, its feature being the edge length, etc. These data type characteristics need to be associated with the location information of the detection points in order to determine the specific physical location later.
[0069] According to the total number of detection points in the data type characteristics and the state change amount, the key detection area is determined. Here, the total number of detection points N refers to the number of all detection points involved in the calculation of the state change amount, and the analysis of the data type characteristics needs to pay attention to which type of characteristics contributes more in the state change amount. For example, if the sum of the absolute values of the temperature characteristic difference of the infrared modal in the state change amount accounts for a high proportion, and these temperature detection points are concentrated in a certain area of the device (such as near the heating module), it can be preliminarily judged that there may be defects or abnormalities related to temperature abnormalities in that area.
[0070] When determining the key detection area, the data type characteristics of the detection points, the size of the characteristic difference, and the spatial distribution of the detection points need to be considered comprehensively. For detection points with the same data type characteristics and large characteristic differences, the aggregation of these detection points in the physical space is analyzed. If these detection points are concentrated in a certain part of the device, that part is determined as the key detection area. For example, the characteristic difference of multiple temperature detection points exceeds the normal range, and these points are all located in the motor shell area of the device. Therefore, the motor shell is the specific position that needs to be strengthened for detection.
[0071] In the process of determining the key detection area, the correlation between different modal data also needs to be considered. For example, when a certain area of the visual modal detects shape abnormalities, if the same area also appears temperature abnormalities under the infrared modal, it can be more convincing that there are defects or abnormalities in that area, thereby determining it as the key detection area. This cross-modal correlation analysis can improve the accuracy of the determination of the key detection area and avoid misjudgment that may be caused by single modal data.
[0072] In addition, the range of the key detection area needs to be adjusted according to the distribution density of the detection points and the influence degree of the characteristic difference. If the detection points are highly concentrated in a small area and the characteristic difference is significant, the key detection area can be set as the small area. If the detection points are scattered but all belong to the same type of abnormal characteristics, the key detection area may need to be expanded to a larger range that includes these scattered points to ensure that all areas that may have abnormalities are covered.
[0073] Throughout the process, from the calculation of the characteristic difference to the generation of the state change amount, and then to the determination of the key detection area, each step needs to be based on the time series data and spatial distribution information of the multi-modal characteristic parameters to ensure logical coherence and sufficient basis.
[0074] Embodiment 4, collecting historical state records of multiple detection nodes within the coverage range according to the key detection area In the process of generating detection adjustment instructions and delivering them to multiple detection nodes according to historical state records, the set coverage range is divided into several sub-regions in order from near to far, with the key detection area as the center. For example, assuming that the key detection area is the motor part of a production equipment, and the set coverage range is a circular area with the motor as the center and a radius of 5 meters, the range can be divided into three layers of sub-regions in an increasing radius manner: the inner layer is the area 0-1 meters away from the motor center, the middle layer is the area 1-3 meters away from the motor center, and the outer layer is the area 3-5 meters away from the motor center. This division method makes the sub-regions closer to the key detection area more concerned, facilitating the subsequent allocation of different detection resources according to the distance.
[0075] After completing the division, detection nodes in each sub-region need to be selected. Detection nodes can be various sensors installed at different parts of the equipment, such as temperature sensors, vibration sensors distributed near the motor, and visual detection cameras installed in the surrounding area. Taking the three-layer sub-regions as an example, the inner layer region may contain 5 temperature sensors and 2 vibration sensors, the middle layer region has 3 temperature sensors, 1 vibration sensor, and 1 visual camera, and the outer layer region has 2 temperature sensors and 1 visual camera. Each detection node has a unique identifier and specific installation location for subsequent data collection and instruction delivery.
[0076] According to the historical state records of the detection nodes in each sub-region, the basic detection frequency of each sub-region within the allowable error range is determined. The collection of historical state records is based on the key detection area as the center, dividing the set coverage range into three layers of regions with increasing radius. For the detection nodes in each layer, multi-modal historical data is collected at the same time period every day for the past six months. For example, temperature, vibration, image, and other data of each detection node are collected at 9 am every day to form the historical state records of each node. Taking a temperature sensor in the inner layer region as an example, by analyzing its historical data for the past six months, it is found that when the detection frequency is once every 10 minutes, the detection data error of the sensor is within the allowable range of ±2%, so the basic detection frequency of the sensor is determined to be once every 10 minutes. The vibration sensor in another middle layer region has an error within the allowable range when the detection frequency is once every 15 minutes, so its basic detection frequency is once every 15 minutes.
[0077] When determining the detection nodes in each sub-region that meet the basic detection frequency, it is necessary to ensure that the sum of the basic detection frequencies of these nodes is greater than or equal to the target detection frequency. The target detection frequency is pre-set according to the detection requirements and the importance of the equipment, for example, for the inner sub-region where the key detection area is located, the target detection frequency is set to 10 times per hour. Assuming that the inner region has 5 temperature sensors and 2 vibration sensors, 3 of which have a basic detection frequency of once every 10 minutes (i.e. 6 times per hour), 2 of which have a basic detection frequency of once every 15 minutes (4 times per hour), and 2 of which have a basic detection frequency of once every 20 minutes (3 times per hour), then the sum of the basic detection frequencies of these detection nodes is 3x6+2x4+2x3=18+8+6=32 times / hour, which is much greater than the target detection frequency of 10 times / hour, meeting the requirements. At this time, the 5 temperature sensors and 2 vibration sensors are determined as detection nodes that meet the basic detection frequency.
[0078] Based on the multiple detection nodes and the basic detection frequency of each detection node, detection adjustment instructions are generated and issued. The detection adjustment instructions need to clearly specify the detection frequency that each detection node needs to improve in a specific period. For example, in the inner region, the temperature sensor with a basic detection frequency of once every 10 minutes may be instructed to increase the detection frequency to once every 5 minutes in a specific period; the temperature sensor with a basic detection frequency of once every 15 minutes is increased to once every 10 minutes, and the vibration sensor is increased to once every 10 minutes. In this way, by increasing the detection frequency of each detection node, the total detection frequency of the entire inner region is further increased, thereby strengthening the detection intensity of the key detection area.
[0079] When collecting historical state records, dividing three layers of regions centered on the key detection area can ensure that the collected data is targeted and hierarchical. For example, for the motor as the key detection area, the detection nodes in the inner region are directly installed on or near the motor surface, which can most directly reflect the running state of the motor; the detection nodes in the middle region are distributed on the equipment supports around the motor, which can detect the influence of the motor operation on the surrounding components; the detection nodes in the outer region are installed in a more distant position, which are used to monitor the overall state of the entire equipment system. Collecting data at the same time period every day within the past six months is to exclude the influence of different time periods on the data due to different equipment operating loads, to ensure the consistency and comparability of the historical data.
[0080] Taking a specific application scenario as an example, assume that the rolling mill of a certain factory determines the roll part as the key detection area through the previous steps during operation, and sets the coverage range as a region with a radius of 8 meters centered on the roll. Divide this range into three sub-regions: inner layer (0-2 meters), middle layer (2-5 meters), and outer layer (5-8 meters). The inner layer has 8 temperature sensors and 4 vibration sensors, the middle layer has 5 temperature sensors, 3 vibration sensors, and 2 visual cameras, and the outer layer has 3 temperature sensors and 1 visual camera. Collect the temperature, vibration, and image data of each detection node at 2 pm every day for nearly six months to form historical state records. By analyzing these records, it is determined that the basic detection frequency of the temperature sensors in the inner layer is once every 15 minutes, and the vibration sensors are once every 20 minutes; the temperature sensors in the middle layer are once every 20 minutes, the vibration sensors are once every 30 minutes, and the visual cameras are once every hour; the temperature sensors in the outer layer are once every 30 minutes, and the visual cameras are once every 2 hours. Set the target detection frequency of the inner layer to 12 times per hour, the middle layer to 8 times per hour, and the outer layer to 4 times per hour. Calculate the sum of the basic detection frequencies of the detection nodes in each layer, and find that the inner layer is 8x4 (once every 15 minutes, i.e., 4 times per hour) + 4x3 (once every 20 minutes, i.e., 3 times per hour) = 32 + 12 = 44 times per hour, which meets the target; the middle layer is 5x3 + 3x2 + 2x1 = 15 + 6 + 2 = 23 times per hour, which meets the target; and the outer layer is 3x2 + 1x0.5 = 6 + 0.5 = 6.5 times per hour, which meets the target. Therefore, generate a detection adjustment instruction, which requires the temperature sensors in the inner layer to increase their detection frequency to once every 10 minutes and the vibration sensors to increase their detection frequency to once every 15 minutes in a specific time period; the temperature sensors in the middle layer to increase their detection frequency to once every 15 minutes, the vibration sensors to increase their detection frequency to once every 20 minutes, and the visual cameras to increase their detection frequency to once every 30 minutes; and the temperature sensors in the outer layer to increase their detection frequency to once every 20 minutes and the visual cameras to increase their detection frequency to once every 1 hour. In this way, the detection frequency of each layer is increased, enabling more intensive monitoring of the state of the roll and its surrounding area, and timely detection of possible defects and abnormalities.
[0081] Throughout the process, the division of zones, the selection of detection nodes, the determination of basic detection frequencies, and the generation of detection adjustment instructions are all closely related to the key detection area, making full use of data information in historical state records, and reasonably allocating detection resources according to the distance of different sub-regions from the key detection area and the distribution of detection nodes. This approach can maximize the detection accuracy and abnormality detection capability of the key area without increasing excessive detection costs, achieving efficient online detection of multi-modal defects and abnormalities.
[0082] Example 5, based on the detection evaluation model to supplement the latest period of actual state data and comparison to determine early warning In the detection evaluation model based on the supplement of the latest period of actual state data and comparison to determine early warning, the actual state data of the latest period is first obtained. For example, in a certain industrial equipment detection scene, the latest period is set to the current hour (such as 14:00-15:00), and the detection data of each part of the equipment in this period is collected in real time through multi-modal equipment (such as visual camera, infrared sensor, vibration detector), including component surface image, temperature distribution, vibration frequency, etc. These data need to match the input format of the detection evaluation model, such as converting image data into feature vectors, structuring temperature data according to detection point position, extracting peak value, mean value and other feature parameters of vibration data.
[0083] After obtaining the data, the actual state data of the latest period is supplemented to the detection evaluation model, and the model is updated. For example, the original model is trained based on the historical data of the previous 72 hours, and after supplementing the actual data of the current 1 hour, the model training data set is updated to the data of the previous 72 hours plus the current hour. During the model updating process, incremental learning or retraining is used to adjust the parameters, so that the model can more accurately reflect the current running state of the equipment. Taking a neural network model as an example, the weights are updated through the back propagation algorithm, so that the prediction error of the model for the latest data is reduced, and a second matching curve that meets the second matching condition is generated. Compared with the first matching curve, the curve can better fit the running trend of the current equipment, for example, if the temperature baseline of the equipment rises slightly due to wear in the recent period, the second matching curve will adjust the baseline value of the temperature prediction accordingly.
[0084] After updating the model, the updated detection state data of the next period in the second matching curve is identified. For example, the current period is 15:00-16:00, and the next period is 16:00-17:00. The model predicts the feature values of each modality in this period based on the updated curve, such as predicting that the average temperature of a certain bearing in 16:00-17:00 is 65℃, the peak value of vibration acceleration is 2.3g, and the image edge wear of the gear is 0.12mm. At the same time, the predicted feature values of the next period in the initial detection state data are extracted, i.e. the same period data predicted by the model before updating based on the first matching curve, such as the original predicted average temperature of the bearing is 62℃, the peak value of vibration acceleration is 2.0g, and the wear amount is 0.10mm.
[0085] Next, the deviation of the two types of characteristic values is extracted through fluctuation comparison. For example, the deviation of the temperature prediction value 62°C and the actual value 65°C is 3°C, the deviation of the vibration prediction value 2.0g and the actual value 2.3g is 0.3g, and the deviation of the wear prediction value 0.10mm and the actual value 0.12mm is 0.02mm. These deviations are classified according to the characteristic type, and the mean value of the deviation of the same type of characteristics is calculated, such as the mean value of the temperature deviation is 3°C, and the mean value of the vibration deviation is 0.3g. Then, the mean values of the deviations of different types of characteristics are considered comprehensively to form the overall fluctuation comparison result. This process needs to pay attention to the associated deviation of cross-modal characteristics, for example, whether the temperature rise is accompanied by vibration anomaly. If the deviation of both exceeds the normal range, it may indicate that the equipment has a potential fault.
[0086] According to the comparison result, it is judged whether to trigger the early warning process. It is assumed that the allowed difference value is set according to the characteristics of the equipment, such as the allowed difference value of temperature is 2°C, the allowed difference value of vibration is 0.2g, and the allowed difference value of wear is 0.01mm. When the difference value between the updated detection state data and the initial detection state data is within the allowed range, such as the temperature difference value is 1°C and the vibration difference value is 0.1g, a first instruction is generated to maintain the current detection frequency. For example, the current detection frequency of a certain detection node is once every 30 minutes, and the instruction maintains the frequency unchanged, and continues to collect data at the normal frequency.
[0087] When the updated detection state data is less than the initial detection state data and the difference value exceeds the allowed range, such as the temperature prediction value is 62°C and the actual value is 58°C, the difference value is -4°C (absolute value 4°C>2°C), at this time, part of the node needs to be strengthened according to the distribution of the basic detection frequency of the detection node. For example, the basic detection frequency of the temperature sensor near the cooling system of the equipment is once every 15 minutes, and the detection frequency of the sensor in other areas is once every 30 minutes. At this time, the instruction requires the sensor near the cooling system to be upgraded to once every 10 minutes, and at the same time, the early warning is triggered to prompt the risk of cooling efficiency decline, and a certain sensor is specified to detect the surrounding nodes in coordination, such as the temperature data of the adjacent components is collected synchronously by the main cooling sensor, and the original detection frequency is maintained to ensure the data correlation.
[0088] When the updated detection state data is greater than the initial detection state data and the difference value exceeds the allowed range, such as the actual value of the temperature is 67°C (the difference value is 5°C>2°C), a second instruction of comprehensive strengthening detection is generated. For example, the detection frequency of all detection nodes in the high temperature area of the equipment is upgraded from once every 30 minutes to once every 15 minutes, including temperature sensors, vibration sensors and visual cameras, and the scanning frequency of infrared thermal imaging is increased to ensure comprehensive coverage of the abnormal area and timely capture of the defect development trend.
[0089] Taking a specific application scenario as an example, in the multi-modal detection system of a certain automobile engine cylinder body production line, the initial detection state data predicts that the temperature mean value of the cylinder body surface in the next time period (15:00-16:00) is 110℃ after the 14:00-15:00 time period ends, while the second matching curve generated by the model update predicts that the temperature mean value in the next time period is 118℃ after supplementing the actual detection data of this time period. Extracting the initial prediction value 110℃ and the updated prediction value 118℃, the difference is 8℃. If the system sets the temperature allowed difference to be 5℃, then 8℃ exceeds the allowed range, and the updated value is greater than the initial value, triggering the comprehensive strengthening detection instruction. At this time, the detection frequency of all temperature sensors in the cylinder detection area is increased from every 20 minutes to every 10 minutes, the shooting frequency of the visual detection camera on the high temperature area is increased from 5 times per hour to 10 times per hour, and additional ultrasonic detection nodes are started to scan the cylinder body to ensure that cracks or material deformation defects caused by abnormal temperature rise can be found in time.
[0090] During the whole process, the model update needs to ensure the continuity of the data time sequence, and the latest time period data needs to be integrated into the historical data set in chronological order to avoid disrupting the logic of the original time sequence. When comparing fluctuations, both single modal deviation and cross-modal correlation deviation need to be considered, for example, if the temperature abnormally rises while the vibration frequency is abnormal, the deviation degree of both needs to be evaluated comprehensively, rather than independently. The triggering conditions of the early warning process need to be reasonably set according to factors such as equipment operation characteristics and historical fault data, for example, the allowed difference of precision instruments is usually less than that of ordinary industrial equipment, in order to balance the false positive rate and the false negative rate. In addition, the adjustment of detection frequency needs to consider the equipment load and data processing capacity to avoid system overload due to too high frequency, for example, when the instruction requires to increase the detection frequency, the data transmission link and storage resources need to be optimized synchronously to ensure the real-time and integrity of the detection data. Through this dynamic adjustment mechanism, the system can adaptively adjust the detection strategy according to the comparison results of the latest detection data and the prediction data, realize accurate early warning of multi-modal defects and abnormalities, and improve the reliability and effectiveness of online detection.
[0091] It should be noted that in this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0092] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. An online detection method for multimodal defect and anomaly detection, characterized in that: The method comprises the following steps: Acquire multimodal real-time detection data and extract cross-modal associated feature parameters, wherein the feature parameters include instantaneous features and continuous features; A detection evaluation model is constructed based on the extracted parameters, and the detection status data of several time periods are deduced backward to obtain the initial detection status data; Calculating the state change amount of adjacent time periods in the initial detection state data, and when the change amount reaches a set threshold, determining the key detection area based on the corresponding relationship between the change amount and the multimodal data type; Collect historical status records of multiple detection nodes within the coverage area according to the key detection area; Generating a detection adjustment instruction based on the historical status record, and transmitting the adjustment instruction to multiple detection nodes, wherein the adjustment instruction is used to instruct to increase the detection frequency of a specific time period based on the historical status record; Based on the detection and evaluation model, the actual status data of the latest period is supplemented, and the updated detection status data of the next period is obtained. The next period is compared with the initial detection status data and the updated detection status data to generate a comparison result. Based on the comparison result, it is determined whether to trigger the early warning process.
2. The online detection method for multimodal defect and anomaly detection according to claim 1, characterized in that: The method of acquiring multimodal real-time detection data and extracting cross-modal associated feature parameters includes the following steps: Obtain real-time detection data of multimodal devices within a set time range; Based on the set time range, select all time periods in the same cycle as the current detection period; Extract cross-modal correlation feature parameters of all time periods in real-time detection data.
3. The online detection method for multimodal defect and anomaly detection according to claim 2, characterized in that: The construction of the detection evaluation model based on the extracted parameters and the backward deduction of the detection status data of several time periods to obtain the initial detection status data includes the following steps: Building a detection and evaluation model based on the cross-modal feature parameters of all time periods, and generating a first matching curve that meets the first matching condition; The detection state data of several time periods are deduced backward according to the first matching curve to obtain the initial detection state data.
4. The online detection method for multimodal defect and anomaly detection according to claim 3, characterized in that: The step of calculating the state change amount in adjacent time periods of the initial detection state data and, when the change amount reaches a set threshold, determining the key detection area based on the correspondence between the change amount and the multimodal data type comprises the following steps: Obtain the data type characteristics of a single detection point in the state change; According to the data type characteristics and the total number of detection points in the state change amount, a key detection area is determined, and the key detection area includes specific locations that need to be strengthened in detection.
5. The online detection method for multimodal defect and anomaly detection according to claim 1, characterized in that: Generating a detection adjustment instruction based on the historical status record and transmitting the adjustment instruction to multiple detection nodes comprises the following steps: With the key detection area as the center, the set coverage range is divided into sections in the order from near to far, and several sub-areas within the set sub-range are obtained; Select the detection nodes in each sub-area, and determine the basic detection frequency in each sub-area so that the error is within the allowable range based on the historical status records of the detection nodes in each sub-area; Determine the detection nodes that meet the basic detection frequency in each sub-area, and obtain multiple detection nodes, wherein the sum of the basic detection frequencies of the multiple detection nodes is greater than or equal to the target detection frequency; Generate and issue detection adjustment instructions based on multiple detection nodes and the basic detection frequency of each detection node; The method of collecting historical status records of multiple detection nodes within the coverage area according to the key detection area includes the following steps: With the key detection area as the center, the set coverage area is divided into three layers of areas with increasing radius; For the detection nodes in each layer area, multimodal historical data are collected at the same time every day in the past six months to form a historical status record of each node.
6. The online detection method for multimodal defect and anomaly detection according to claim 3, characterized in that: The method of supplementing the actual status data of the latest period based on the detection evaluation model and obtaining the updated detection status data of the next period, performing a comparison of the next period based on the initial detection status data and the updated detection status data, and generating a comparison result includes the following steps: The detection evaluation model is updated according to the state data of the latest period to generate a second matching curve that meets the second matching condition; Identify updated detection state data of the next time period in the second matching curve, perform fluctuation comparison based on the initial detection state data and the updated detection state data, and generate a comparison result.
7. The online detection method for multimodal defect and anomaly detection according to claim 6, characterized in that: The performing of fluctuation comparison based on the initial detection state data and the updated detection state data to generate a comparison result comprises the following steps: Extracting the predicted characteristic value of the next period in the initial detection state data and updating the actual characteristic value of the next period in the detection state data; The deviation between the feature prediction value and the actual value, and the deviation between the other type of feature prediction value and the actual value are calculated respectively, and the average of the two deviations is used as the fluctuation comparison result.
8. The online detection method for multimodal defect and anomaly detection according to claim 1 or 7, characterized in that: The process of determining whether to trigger an early warning based on the comparison results includes the following steps: When a difference between the detection data of the next period in the updated detection state data and the initial detection state data is within an allowable difference, a first instruction for maintaining the current detection frequency is generated to fully maintain the detection frequency of the specific period; When the updated detection status data is less than the detection data of the next period in the initial detection status data, and the difference between the two is greater than the allowed difference, based on the distribution of basic detection frequencies among multiple detection nodes, generate instructions for some nodes to strengthen detection and early warning coexistence, so that at least one node among the multiple detection nodes can coordinate detection with other nodes except itself, and maintain the detection frequency of at least one node; When the updated detection status data is greater than the detection data of the next period in the initial detection status data, and the difference between the two is greater than the allowed difference, a second instruction for comprehensively strengthening detection is generated to increase the detection frequency of the specific period.
9. The online detection method for multimodal defect and anomaly detection according to claim 3, characterized in that: The step of constructing a detection and evaluation model based on the cross-modal feature parameters of all time periods and generating a first matching curve that meets the first matching condition comprises the following steps: Arrange the cross-modal feature parameters of all time periods in time series to construct a training dataset; The historical data verification method is used to train the model on the training data set, and the model parameters are adjusted until the fitting error of the model output is less than the first matching threshold, thereby generating a first matching curve.
10. The online detection method for multimodal defect and anomaly detection according to claim 4, characterized in that: Calculating the state change amount in adjacent time periods in the initial detection state data comprises the following steps: Select cross-modal feature parameters of consecutive time periods in the initial detection state data and calculate the feature difference between the current time period and the previous time period; The sum of the absolute values of the characteristic differences of each type is taken as the state change amount, which is used to reflect the degree of fluctuation of defect detection in adjacent time periods.