Optimization processing method and system for oil detection data based on multi-dimensional feature fusion
Patent Information
- Application Number
- CN202511488627.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-10-17
AI Technical Summary
[0004]本申请的目的是提供多维特征融合的油品检测数据优化处理方法及系统,用于解决现有技术油品检测多源异构数据存在的噪声干扰和信息冗余,导致油品质量检测结果的准确度和可靠性低的技术问题
本申请实施例提供的方法通过利用自动化分样装置将目标油品样品按照预设分装要求进行分装,并送入多个检测设备中进行检测,获得多个分装样品检测数据集合,其中,所述预设分装要求为每个检测设备至少划分两个分装样品;遍历所述多个分装样品检测数据集合进行数据清洗预处理,获得多个预处理分装样品检测数据集合;对所述多个预处理分装样品检测数据集合分别进行集合内数据交叉增强处理,确定多个增强分装样品检测数据;基于所述多个增强分装样品检测数据进行多维特征跨模态融合解析,获得目标油品样品的初始油品样品检测结果;根据所述初始油品样品检测结果进行设备可靠性验证,当设备可靠性验证通过时,将所述初始油品样品检测结果作为目标油品样品检测结果。达到有效提升多源异构油品检测数据的融合精度,进而提高油品质量检测结果准确性和可靠性的技术效果。
Smart Images

Figure CN121328834B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for optimizing oil product testing data through multi-dimensional feature fusion. Background Technology
[0002] As a crucial medium for lubrication of industrial equipment, energy supply, and chemical production, the physical and chemical properties of oil directly impact equipment operating efficiency and service life. High-quality oil can improve combustion efficiency, reduce equipment wear, and lower exhaust pollution. Existing multi-dimensional feature fusion methods for oil detection have improved the comprehensiveness of testing quality. However, in comprehensively assessing oil deterioration, contamination, and performance changes by collecting multi-dimensional feature information such as spectral, chemical, and physical parameters, there are challenges. These challenges include complex data sources and high redundancy and noise interference among multi-dimensional features, which affect the accuracy and reliability of oil quality testing results.
[0003] Existing technologies for oil product testing suffer from noise interference and information redundancy in multi-source heterogeneous data, leading to low accuracy and reliability of oil product quality testing results. Summary of the Invention
[0004] The purpose of this application is to provide a method and system for optimizing oil product testing data through multi-dimensional feature fusion, in order to solve the technical problem that noise interference and information redundancy exist in the existing multi-source heterogeneous data of oil product testing, resulting in low accuracy and reliability of oil product quality testing results.
[0005] In view of the above problems, this application provides a method and system for optimizing oil product testing data by fusing multi-dimensional features.
[0006] The first aspect of this application provides a method for optimizing and processing oil product testing data through multi-dimensional feature fusion. The method includes: using an automated sampling device to package a target oil product sample according to preset packaging requirements, and sending the packaged sample to multiple testing devices for testing to obtain multiple sets of packaged sample testing data, wherein the preset packaging requirements divide each testing device into at least two packaged samples; traversing the multiple sets of packaged sample testing data for data cleaning and preprocessing to obtain multiple sets of preprocessed packaged sample testing data; performing intra-set data cross-enhancement processing on each of the multiple preprocessed packaged sample testing data sets to determine multiple enhanced packaged sample testing data; performing multi-dimensional feature cross-modal fusion analysis based on the multiple enhanced packaged sample testing data to obtain initial oil product sample testing results for the target oil product sample; and verifying equipment reliability based on the initial oil product sample testing results. When the equipment reliability verification passes, the initial oil product sample testing results are used as the target oil product sample testing results.
[0007] Optionally, a first pre-processed and packaged sample test data set is extracted from multiple pre-processed and packaged sample test data sets; the first pre-processed and packaged sample test data set is cross-correlated without omission or duplication to obtain a first data cross-correlation group set; each data cross-correlation group in the first data cross-correlation group set is cross-enhanced to obtain a first initial enhanced packaged sample test data set set; the mean of the first initial enhanced packaged sample test data set set is calculated to obtain the first enhanced packaged sample test data, and the first enhanced packaged sample test data is added to the multiple enhanced packaged sample test data sets.
[0008] Optionally, similarity analysis is performed on the same type of data in each data cross-association group in the first data cross-association group set to obtain a first similarity group set; intra-group similarity normalization is performed on the first similarity group set, and the processing results are added to an initially empty matrix to obtain a first cross-enhancement processing matrix set; cross-enhancement is performed on the first data cross-association group set using the first cross-enhancement processing matrix set to obtain the first initial enhanced packaged sample detection data group set.
[0009] Optionally, a convolutional enhancement network layer is pre-constructed; the first data cross-correlation group corresponding to the first cross-enhancement processing matrix set and the first data cross-correlation group set is transmitted to the convolutional enhancement network layer for analysis to obtain the first initial enhanced packaged sample detection data set.
[0010] Optionally, based on the feature extraction mechanism, features are extracted from different modal data of each enhanced repackaged sample detection data according to different detection types to construct multiple multimodal features; attention mechanism is used to perform cross-modal weighted coupling on the multiple multimodal features to obtain multimodal coupled features; detection and analysis are performed on the multimodal coupled features to obtain the initial oil sample detection results.
[0011] Optionally, the feature extraction mechanism involves performing principal component analysis on spectral data, extracting statistical features from physical data, and performing standardized index analysis on chemical features.
[0012] Optionally, M anomalies are extracted from the initial oil sample test results, where M is a positive integer; historical test logs of the same type of oil are extracted from the M testing devices corresponding to the M anomalies to obtain a set of M historical test logs; clustering trend identification is performed on the set of M historical test logs to determine M clustering trend features; it is determined whether the similarity between the M anomaly features of the M anomalies and the M clustering trend features meets a preset similarity threshold. If not, the device reliability verification is passed.
[0013] Optionally, the mean of the M historical detection log sets is calculated to obtain M mean drift centers; the mean drift algorithm is used in combination with the M mean drift centers to identify clustering trends in the M historical detection log sets and determine M clustering trend features.
[0014] Optionally, the multiple sets of subpackaged sample test data are traversed and structured parsed to transform them into a standardized data structure, thereby obtaining multiple standardized sets of subpackaged sample test data; outlier removal or baseline correction is performed on the multiple standardized sets of subpackaged sample test data respectively to obtain multiple preprocessed sets of subpackaged sample test data.
[0015] A second aspect of this application provides a multi-dimensional feature fusion-based oil product testing data optimization and processing system. The system includes: a testing data acquisition module, used to use an automated sampling device to package a target oil product sample according to preset packaging requirements and send it to multiple testing devices for testing, obtaining multiple sets of packaged sample testing data, wherein the preset packaging requirements divide each testing device into at least two packaged samples; a data cleaning module, used to traverse the multiple sets of packaged sample testing data for data cleaning and preprocessing, obtaining multiple preprocessed sets of packaged sample testing data; a data processing module, used to perform intra-set data cross-enhancement processing on each of the multiple preprocessed sets of packaged sample testing data to determine multiple enhanced packaged sample testing data; an initial testing result acquisition module, used to perform multi-dimensional feature cross-modal fusion analysis based on the multiple enhanced packaged sample testing data to obtain initial oil product sample testing results for the target oil product sample; and an oil product testing result acquisition module, used to perform equipment reliability verification based on the initial oil product sample testing results, and when the equipment reliability verification passes, use the initial oil product sample testing results as the target oil product sample testing results.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: The method provided in this application embodiment utilizes an automated sampling device to package a target oil sample according to preset packaging requirements, and sends it to multiple testing devices for testing, obtaining multiple sets of packaged sample testing data. The preset packaging requirements stipulate that each testing device must divide the sample into at least two packages. The method then traverses these multiple sets of packaged sample testing data for data cleaning and preprocessing, obtaining multiple preprocessed sets of packaged sample testing data. Each of these preprocessed sets of packaged sample testing data undergoes intra-set data cross-enhancement processing to determine multiple enhanced packaged sample testing data. Based on these enhanced packaged sample testing data, multi-dimensional feature cross-modal fusion analysis is performed to obtain the initial oil sample testing result for the target oil sample. Finally, the device reliability is verified based on the initial oil sample testing result. If the device reliability verification passes, the initial oil sample testing result is used as the target oil sample testing result. This method effectively improves the fusion accuracy of multi-source heterogeneous oil testing data, thereby enhancing the accuracy and reliability of oil quality testing results.
[0017] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the oil product testing data optimization processing method based on multi-dimensional feature fusion provided in this application.
[0020] Figure 2 A schematic diagram of the structure of the oil product testing data optimization processing system with multi-dimensional feature fusion provided in this application.
[0021] Figure labeling: Detection data acquisition module 11, data cleaning module 12, data processing module 13, initial detection result acquisition module 14, oil product detection result acquisition module 15. Detailed Implementation
[0022] This application provides a method and system for optimizing oil product testing data through multi-dimensional feature fusion. This method addresses the technical problem of low accuracy and reliability in oil product quality testing results due to noise interference and information redundancy in existing multi-source heterogeneous oil product testing data. The goal is to effectively improve the fusion accuracy of multi-source heterogeneous oil product testing data, thereby enhancing the accuracy and reliability of oil product quality testing results.
[0023] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. It should also be noted that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, not all of them.
[0024] Example 1, as Figure 1 As shown, this application provides a method for optimizing oil product testing data through multi-dimensional feature fusion. The method includes: An automated sampling device is used to package the target oil sample according to a preset packaging requirement, and then send it to multiple testing devices for testing to obtain multiple sets of test data for the packaged samples. The preset packaging requirement is that each testing device must divide at least two packaged samples.
[0025] Specifically, the target oil sample is the oil-based medium to be tested, such as industrial equipment lubricating oil, fuel oil, and synthetic oil. Industrial equipment lubricating oils include gear oil, hydraulic oil, and lubricating grease diluent. Fuel oils include diesel, gasoline, and aviation kerosene. The target oil samples are obtained from actual equipment operating sites and oil storage tanks. Pre-defined repackaging requirements are determined based on the oil testing objectives. These requirements refer to the rules for allocating repackaging volumes, quantities, and testing channels during the sample repackaging process, including but not limited to: each testing device must be divided into at least two repackaging samples to avoid random errors, ensure the stability of test results, and ensure the consistency of the volume of each repackaging sample. The oil testing objectives are divided into spectral data detection, physical characteristic data detection, and chemical component data detection, including oil moisture content determination, kinematic viscosity determination, acid value determination, dissolved gas component content determination, dielectric loss factor and volume resistivity determination, and antioxidant content determination. An automated sample dispensing device refers to an intelligent liquid dispensing device with quantitative dispensing and multi-channel distribution functions. It includes a sample input section, a dispensing control section, a dispensing execution section, and a sample delivery section. The sample input section is used to receive the target oil sample. The dispensing control section accurately controls the sample distribution according to the preset dispensing requirements. The dispensing execution section performs the specific dispensing operation and distributes the sample into different containers. The sample delivery section delivers the dispensed sample to various testing devices.
[0026] The target oil sample is thoroughly stirred and then fed into an automated dispensing device for dispensing according to preset dispensing requirements. The sample is dispensed into different containers, such as chromatograms, colorimetric tubes, beakers, and viscosity tubes. Multiple dispensed samples are then sent to multiple testing devices for analysis. For example, after dispensing, the moisture content is determined using a trace moisture analyzer; the kinematic viscosity is determined using a fully automated kinematic viscometer; after manual chemical pretreatment, the acid value is determined using a fully automated potentiometric titrator; the dissolved gas component content is determined using a gas chromatograph; the media loss factor (volume resistivity) is determined manually using a test cell method; and after manual chemical pretreatment, the antioxidant content is determined spectrophotometrically. Multiple dispensing samples are analyzed using multiple testing devices to obtain multiple sets of test data, with each set corresponding to one dispensing sample.
[0027] Automated sampling can efficiently and accurately complete the dispensing and testing of oil samples. Each testing device is divided into at least two dispensing samples, which can avoid random errors and ensure that the test data of at least two dispensing samples are mapped to each other, thereby improving testing efficiency and the reliability of results.
[0028] The multiple sets of test data for the dispensed samples are traversed for data cleaning and preprocessing to obtain multiple sets of preprocessed test data for the dispensed samples.
[0029] Furthermore, the multiple sets of subpackaged sample test data are traversed and structured parsed to transform them into standardized data structures, thereby obtaining multiple standardized sets of subpackaged sample test data; outlier removal or baseline correction is performed on the multiple standardized sets of subpackaged sample test data respectively to obtain multiple preprocessed sets of subpackaged sample test data.
[0030] Specifically, different testing devices may output data in different formats, and multiple datasets of sample testing may contain various data formats, such as CSV, TDMS, text logs, binary spectral files, and database records. The process involves traversing multiple datasets of sample testing and performing structured parsing on each dataset. Based on the data source and data type, appropriate parsing methods are used to parse the raw testing data into processable intermediate data structures, such as data tables or arrays. The parsed data undergoes field mapping and standardization, including naming conventions for the same fields, data type conversion, and timestamp format standardization. This transforms the sample testing data from different testing devices into a unified standard data structure, such as CSV format, ensuring consistent naming and formatting for fields in each dataset.
[0031] Through structured analysis, multiple standardized datasets of aliquoted sample detection are obtained. These datasets are then cleaned and preprocessed. Specifically, statistical methods such as mean, standard deviation, and quartiles are used to identify outliers. Any data point exceeding three times the standard deviation of the mean is considered an outlier and removed. Alternatively, baseline correction is performed on the standardized datasets. For spectral detection data, wavelet denoising or Savitzky-Golay filters are used for smoothing / denoising. For physical or time-series data, baseline drift correction is employed. For long-term series data, high-pass filtering or polynomial fitting is used to remove slow drift, and moving averages or Savitzky-Golay smoothing is used. If high-frequency electrical noise is present, low-pass filters or wavelet denoising can be used. For example, infrared spectral analyzer data may exhibit baseline drift. By fitting a baseline (e.g., polynomial fitting) and subtracting this baseline from the original spectral data, corrected spectral data can be obtained, ensuring the stability and accuracy of the feature data.
[0032] By removing outliers or correcting baselines for standardized datasets of multiple repackaged samples, the quality of the original test data is optimized, improving the quality and consistency of the test data, thereby enhancing the accuracy and stability of oil product testing results.
[0033] The multiple pre-processed and packaged sample test data sets are subjected to in-set data cross-enhancement processing to determine multiple enhanced packaged sample test data sets.
[0034] Furthermore, the multiple pre-processed and packaged sample test data sets are subjected to intra-set data cross-enhancement processing to determine multiple enhanced packaged sample test data sets, including: extracting a first pre-processed and packaged sample test data set from the multiple pre-processed and packaged sample test data sets; performing data-complete and non-repetitive cross-association on the first pre-processed and packaged sample test data set to obtain a first data cross-association group set; performing cross-enhancement processing on each data cross-association group in the first data cross-association group set to obtain a first initial enhanced packaged sample test data group set; calculating the mean of the first initial enhanced packaged sample test data group set to obtain the first enhanced packaged sample test data, and adding the first enhanced packaged sample test data to the multiple enhanced packaged sample test data sets.
[0035] Specifically, a first preprocessed and packaged sample detection dataset is randomly selected from multiple preprocessed and packaged sample detection datasets. The "first" in this dataset does not represent any particular order but refers to any one of the multiple datasets. By traversing the first preprocessed and packaged sample detection dataset, each data point is extracted. Then, for each data point, a cross-association is performed with other data points that have not been combined with it, ensuring no omissions or duplications. This results in a first set of cross-association groups, which contains the detection data of multiple packaged samples. "No omissions" means that each sample participates in at least one cross-association, ensuring that all data in the first preprocessed and packaged sample detection dataset are cross-associated. "No duplications" means that each sample combination is not calculated repeatedly, avoiding redundancy enhancement. For example, the first preprocessed and packaged sample detection dataset is D1={d1, d2, d3…dn}, where di represents the i-th data point, corresponding to the preprocessed sample detection measurement value under a certain detection device. For data point di, it is paired with di+1, di+2…dn sequentially to generate new data sets (d1, d1), (d1, d3)…(d1, dn). Next, data point d2 is taken; since d2 has already been paired with d1, it is paired with d3, d4…dn sequentially, similarly generating new data sets. This process is repeated for each data point in the first pre-processed sample detection data set, ensuring that each data point is combined with every other data point exactly once. This achieves comprehensive and efficient cross-association, ensuring no omissions or duplications. To further ensure the accuracy and efficiency of cross-association, a unique index tag can be assigned to each data point to control the pairing process.
[0036] For each data cross-association group in the first data cross-association group set, cross-enhancement processing is performed based on the similarity between the data cross-association groups. Based on the cross-enhancement processing results, a first initial enhanced sub-packaged sample detection data set is obtained. The average value of the first initial enhanced sub-packaged sample detection data set is calculated to generate the first enhanced sub-packaged sample detection data. The above steps are repeated to perform intra-set data cross-enhancement processing on multiple preprocessed sub-packaged sample detection data sets to obtain multiple enhanced sub-packaged sample detection data sets. The enhanced data is then added to the multiple enhanced sub-packaged sample detection data sets.
[0037] Through cross-correlation and enhancement processing, the resulting data not only retains the original information but also undergoes feature enhancement and redundancy optimization. Mean calculation reduces the impact of noise or occasional errors in individual samples on the overall features, thereby effectively improving the stability and quality of oil product repackaging sample testing data.
[0038] Furthermore, cross-enhancement processing is performed on each data cross-association group in the first data cross-association group set to obtain a first initial enhanced packaged sample detection data set, including: performing similarity analysis on the same type of data in each data cross-association group in the first data cross-association group set to obtain a first similarity group set; performing intra-group similarity normalization processing on the first similarity group set respectively, and adding the processing results into an initially empty matrix to obtain a first cross-enhancement processing matrix set; and using the first cross-enhancement processing matrix set to perform cross-enhancement on the first data cross-association group set respectively to obtain the first initial enhanced packaged sample detection data set.
[0039] Specifically, for each data cross-association group in the first data cross-association set, similarity analysis is performed on the data of the same type within it to quantify the degree of similarity between two data points. Similarity analysis can be calculated using Euclidean distance, cosine similarity, or Pearson correlation coefficient. Data of the same type refers to data in each data cross-association group where two data points belong to the same detection index. For example, if one data point represents viscosity and another data point also represents viscosity, then these two data points are of the same type. Through similarity calculation, the similarity values of the data of the same type in each data cross-association group are obtained, forming the first similarity set. Using a normalization formula, each similarity value in the first similarity set is normalized, mapping the similarity value to the 0-1 interval. The normalized similarity values are then added to an initially empty matrix, forming the first cross-enhancement processing matrix set. This provides a weight matrix for the cross-enhancement of the first data cross-association set, allowing the enhancement processing to adaptively adjust according to sample similarity. The first set of cross-enhancement processing matrices is used to perform cross-enhancement on the first set of cross-association data groups. Specifically, the normalized similarity in the first set of cross-enhancement processing matrices is used as weights to perform weighted enhancement on each cross-association group in the first set of cross-association data groups. The weights are controlled by the similarity matrix, with similar samples having lower weights and dissimilar samples having higher weights. Through weighted enhancement, an enhanced feature vector set for each cross-association group is obtained, which is the first initial enhanced set of sample detection data groups. Cross-enhancement processing amplifies the differences between data points, making abnormal data more prominent and improving the overall quality of oil sample detection data.
[0040] Furthermore, the first data cross-association group set is cross-enhanced using the first cross-enhancing processing matrix set to obtain the first initial enhanced packaged sample detection data set, including: pre-constructing a convolutional enhancement network layer; and transmitting the corresponding first data cross-association group in the first cross-enhancing processing matrix set and the first data cross-association group set to the convolutional enhancement network layer for analysis to obtain the first initial enhanced packaged sample detection data set.
[0041] Specifically, a pre-constructed convolutional augmentation network layer is constructed. First, the network architecture is designed, including an input layer, convolutional layers, activation layers, fully connected layers, and an output layer. The output layer receives the input data, which consists of a first set of cross-enhancement processing matrices and a first set of data cross-correlation groups. The convolutional layers use convolution operations to extract local features from the data. Multiple convolutional layers can be designed, each containing multiple convolutional kernels to capture features at different scales and orientations. The activation layers use the ReLU activation function to increase the non-linearity of the convolutional augmentation network layer. Pooling layers reduce the dimensionality of features, decreasing computation while preserving important features. The fully connected layers integrate the features extracted by the convolutional layers, and the output layer outputs the final augmented data. Then, the parameters of the convolutional augmentation network layer are initialized, including weight initialization and bias initialization. Weight initialization uses methods such as Xavier initialization or He initialization to initialize the weights of the convolutional kernels and fully connected layers, while bias initialization is set to 0. Next, a loss function and optimizer are defined. For example, the mean squared error loss function is chosen to measure the difference between the output of the convolutional augmentation network layer and the target data. The Adam optimizer is selected to update the network parameters and minimize the loss function. Historical data, including both input and target data, is used as training data and fed into the convolutional augmentation network layer for forward propagation. The network output is calculated, and the difference between the network output and the target data is calculated using the loss function. The gradient is calculated using the backpropagation algorithm to update the network parameters. The above steps are repeated until the network converges or reaches the predetermined number of training epochs, resulting in the constructed convolutional augmentation network layer. This convolutional augmentation network layer is a lightweight forward network used for feature enhancement. Although it is a lightweight forward network, it still needs to be trained with historical data to determine the weights of the convolutional kernels and fully connected layers, enabling the convolutional augmentation network layer to have feature enhancement capabilities. After training, the network can be directly used for online or batch feature enhancement.
[0042] The first cross-enhancement processing matrix set and the corresponding first data cross-association group in the first data cross-association group set are transmitted to the convolutional enhancement network layer for analysis. The convolutional enhancement network layer outputs enhanced data. Each cross-association group is processed by the convolutional enhancement network to obtain enhanced data, forming the first initial enhanced sample detection data set. By pre-constructing the convolutional enhancement network layer, automated data processing is achieved, reducing manual intervention and improving processing efficiency.
[0043] Based on the detection data of the multiple enhanced repackaged samples, multi-dimensional feature cross-modal fusion analysis is performed to obtain the initial oil sample detection results of the target oil sample.
[0044] Furthermore, based on the multiple enhanced repackaging sample detection data, multi-dimensional feature cross-modal fusion analysis is performed to obtain the initial oil sample detection result of the target oil sample, including: based on the feature extraction mechanism, extracting features from different modal data of each enhanced repackaging sample detection data according to different detection types to construct multiple multi-modal features; using an attention mechanism to perform cross-modal weighted coupling of the multiple multi-modal features to obtain multi-modal coupled features; and performing detection analysis on the multi-modal coupled features to obtain the initial oil sample detection result.
[0045] Furthermore, the feature extraction mechanism involves performing principal component analysis on spectral data, extracting statistical features from physical data, and performing standardized index analysis on chemical features.
[0046] Specifically, based on a feature extraction mechanism, features are extracted from different modal data of each enhanced packaged sample detection data according to different detection types. The feature extraction mechanism involves principal component analysis (PCA) of the spectral data. PCA is a dimensionality reduction technique that projects high-dimensional data into a low-dimensional space while retaining the main variation information. By performing PCA on the spectral data, the main component features are extracted, reducing data dimensionality while preserving key feature information. Statistical features are extracted from the physical data, calculating statistics such as mean, variance, standard deviation, maximum, and minimum values to reflect changes in physical properties. Standardization index analysis is performed on the chemical features, normalizing the chemical data to a uniform numerical range. By extracting and integrating features of different detection types from each enhanced packaged sample detection data, multiple multimodal features are formed.
[0047] An attention mechanism is used to perform cross-modal weighted coupling of the multiple multimodal features. Specifically, an attention mechanism is employed to analyze multiple multimodal features. Based on the correlation between each modal feature and the global feature distribution, the attention mechanism automatically calculates the importance weight of each modality. These importance weights can be generated through a self-attention mechanism. Each modal feature is multiplied by its corresponding importance weight and then summed in a weighted manner to obtain the fused multimodal coupled features. These multimodal coupled features integrate spectral, physical, and chemical data from oil product testing.
[0048] The multimodal coupling features are then detected to obtain the initial oil sample detection results. Detection of multimodal coupling features can employ neural networks, discriminative models, or regression models. The detection model is constructed according to actual needs, and the multimodal coupling features are detected. For example, a regression / classification detection model based on a multilayer perceptron: the input layer receives the multimodal coupling feature vectors; the feature interaction layer uses a fully connected structure to perform nonlinear interactions on features of different dimensions; the pattern recognition layer extracts oil state change patterns; the output layer selects different structures according to the task type; the regression output is used to output oil index values; and the Softmax classification outputs oil grade or state recognition. The activation function uses ReLU or LeakyReLU to improve feature representation ability. The optimizer is Adam. For index prediction, the mean squared error is used as the loss function; for state prediction, cross-entropy loss is used. Historical oil detection results are used as label data for supervised training. After training, the multimodal coupling features are input into the detection model for detection and analysis to obtain the initial oil sample detection results. The initial oil sample detection results include several detection indicators and a normal / abnormal label for each indicator.
[0049] By employing feature extraction mechanisms and cross-modal weighted coupling, the system strengthens the feature representation that plays a crucial role in determining oil quality, while filtering out redundant information and cross-interference, effectively improving the fusion accuracy of multi-source heterogeneous oil testing data. Through a detection model, the fused high-dimensional feature data is optimized for detection, enabling precise analysis and processing of oil testing data. This, in turn, improves the accuracy and reliability of initial oil sample testing results, reduces the false positive rate, and meets the high-precision requirements of industrial-grade oil quality testing.
[0050] The equipment reliability is verified based on the initial oil sample test results. When the equipment reliability verification is passed, the initial oil sample test results are used as the target oil sample test results.
[0051] Furthermore, based on the initial oil sample test results, equipment reliability verification is performed. When the equipment reliability verification passes, the initial oil sample test results are used as the target oil sample test results. This includes: extracting M anomalies from the initial oil sample test results, where M is a positive integer; extracting historical test logs of the same type of oil from the M test devices corresponding to the M anomalies to obtain M sets of historical test logs; identifying clustering trends in the M sets of historical test logs to determine M clustering trend features; and determining whether the similarity between the M anomaly features of the M anomalies and the M clustering trend features meets a preset similarity threshold. If not, the equipment reliability verification passes.
[0052] Specifically, after obtaining the initial oil sample test results, M anomalies are extracted from the initial oil sample test results, where M is a positive integer. These anomalies are determined based on the detection indicators and anomaly markers for each detection quality in the initial oil sample test results. When an indicator exceeds a preset threshold, or its corresponding marker is an anomaly marker, it is identified as an anomaly. Each anomaly is associated with a corresponding detection equipment marker. Based on the M anomalies, the corresponding M detection equipment are located, and historical oil detection logs for the same type of oil tested by that equipment are extracted from the oil testing database, resulting in a set of M historical detection logs. In other words, based on the anomalies in the initial oil sample test results, the corresponding detection equipment is determined, and the historical detection logs of these equipment are extracted. This ensures that the log data is the same type of oil being tested. The historical detection logs include detection time, detection type, and detection result value. During extraction, the latest record within the time window is prioritized, and abnormal data during equipment maintenance or calibration is filtered out to ensure the representativeness and stability of the log data.
[0053] Clustering trend identification is performed on M sets of historical detection logs to analyze the result change patterns of each detection device during historical detection processes. Specifically, each historical detection log is standardized. A mean-shift clustering algorithm or a time series clustering algorithm based on dynamic time warping is used to perform cluster analysis on each detection data in the historical detection logs, identifying the main detection trends and drift patterns of the detection devices within the historical period. Through clustering, M clustering trend features are obtained, including centrality features, slope of change, and fluctuation range, which are used to characterize the stability characteristics of the detection devices during historical detection processes.
[0054] After obtaining M clustering trend features, the similarity between each anomaly feature of the M outliers and the clustering trend features of their corresponding detection devices is calculated. This similarity calculation can be based on methods such as Euclidean distance or cosine similarity. The anomaly features corresponding to the outliers are extracted and their similarity is calculated with each clustering trend feature. If the similarity between an anomaly feature and any clustering trend feature is higher than a preset similarity threshold, it indicates that the anomaly is similar to the historical drift characteristics of the detection device, possibly stemming from a problem with the device itself. Conversely, if the similarity is lower than the preset similarity threshold, it indicates that the anomaly does not conform to the historical behavior pattern of the device and is more likely due to changes in the actual characteristics of the oil. The preset similarity threshold is used to determine whether the similarity between the anomaly features and the clustering trend features is within a reasonable range. The preset similarity threshold can be determined through statistical analysis of the similarity distribution in historical stable period data, or adaptively adjusted according to the dynamic changes in the similarity distribution during the detection task. In static scenarios, the standard deviation of the average similarity during the stable period minus a preset multiple is used as the threshold. In dynamic scenarios, the adaptive threshold is calculated based on the kernel density estimation or quantile interval of the historical distribution, thereby realizing the adaptive judgment of reliability under changes in equipment status.
[0055] Finally, the reliability of the testing equipment is verified based on the similarity judgment results. When the similarity between the abnormal features of M anomalies and the clustering trend features of their corresponding equipment does not reach a preset similarity threshold, the reliability of the testing equipment is deemed to have passed the verification, indicating that the testing equipment is operating normally, and the initial oil sample test result is used as the target oil sample test result. If one or more anomalies are highly similar to the equipment's clustering trend features, the equipment reliability verification is deemed to have failed, the corresponding equipment is automatically marked as abnormal equipment, and equipment maintenance and oil sample retesting procedures are performed to ensure the accuracy and reliability of the final oil test results. By identifying and verifying anomalies, potential problems with the equipment can be discovered in a timely manner, the testing process can be optimized, the reliability and accuracy of the target oil sample test results can be improved, the false judgment rate can be reduced, and the high-precision requirements of industrial-grade oil quality testing can be met.
[0056] Furthermore, determining M clustering trend features includes: calculating the mean of the M historical detection log sets to obtain M mean drift centers; using the mean drift algorithm in combination with the M mean drift centers to identify clustering trends in the M historical detection log sets and determine the M clustering trend features.
[0057] Specifically, the mean of the detection index is calculated for each of the M historical detection log sets, resulting in M mean drift centers. These M mean drift centers reflect the overall detection level and central trend of each detection device within the historical detection period. Using the mean drift algorithm in conjunction with the M mean drift centers, clustering trends are identified for each historical detection log set. The kernel function type and bandwidth parameters, such as a Gaussian kernel function, are determined based on the distribution characteristics of the detection index. Using the mean drift centers as initial estimation points, the density distribution of the data within the M historical detection log sets is estimated. In each iteration, the sample points are gradually moved to the point of maximum density according to the kernel density weighted average principle until all sample points converge to stable cluster centers. After iterative convergence, samples whose points converge to the same position in the clustering results are grouped into the same cluster. The center position, sample distribution variance, slope of change, and fluctuation characteristics of each cluster are extracted to obtain the clustering trend characteristics of each detection device. These clustering trend characteristics are used to describe the typical behavior patterns and performance change laws of the detection devices in historical testing cycles, reflecting the long-term stability and drift trend of the device's testing results. The clustering trend characteristics determined for all M detection devices are then organized to obtain M sets of clustering trend characteristics.
[0058] By extracting stable trend features from the equipment's historical testing logs, accurate and quantifiable references are provided for comparing the similarity between anomalies and the equipment's historical behavior, thereby achieving effective and reliable equipment reliability verification and improving the reliability and validity of oil testing data.
[0059] Example 2, based on the same inventive concept as the oil product testing data optimization processing method using multi-dimensional feature fusion in the foregoing examples, such as... Figure 2 As shown, this application provides a multi-dimensional feature fusion-based oil product testing data optimization processing system, wherein the multi-dimensional feature fusion-based oil product testing data optimization processing system includes: The detection data acquisition module 11 is used to use an automated sampling device to package the target oil sample according to preset packaging requirements and send it to multiple testing devices for testing to obtain multiple sets of packaged sample detection data. The preset packaging requirements are that each testing device should divide at least two packages. The data cleaning module 12 is used to traverse the multiple sets of packaged sample detection data to perform data cleaning and preprocessing to obtain multiple sets of preprocessed packaged sample detection data. The data processing module 13 is used to perform intra-set data cross-enhancement processing on the multiple sets of preprocessed packaged sample detection data to determine multiple enhanced packaged sample detection data. The initial detection result acquisition module 14 is used to perform multi-dimensional feature cross-modal fusion analysis based on the multiple enhanced packaged sample detection data to obtain the initial oil sample detection result of the target oil sample. The oil detection result acquisition module 15 is used to perform equipment reliability verification based on the initial oil sample detection result. When the equipment reliability verification is passed, the initial oil sample detection result is used as the target oil sample detection result.
[0060] Furthermore, the data processing module 13 is also configured to: extract a first pre-processed and packaged sample detection data set from multiple pre-processed and packaged sample detection data sets; perform cross-association on the first pre-processed and packaged sample detection data set without data omission or duplication to obtain a first data cross-association group set; perform cross-enhancement processing on each data cross-association group in the first data cross-association group set to obtain a first initial enhanced packaged sample detection data set; calculate the mean of the first initial enhanced packaged sample detection data set to obtain the first enhanced packaged sample detection data, and add the first enhanced packaged sample detection data into the multiple enhanced packaged sample detection data sets.
[0061] Furthermore, the data processing module 13 is also used to: perform similarity analysis on the same type of data in each data cross-association group in the first data cross-association group set to obtain a first similarity group set; perform intra-group similarity normalization processing on the first similarity group set respectively, and add the processing results into an initially empty matrix to obtain a first cross-enhancement processing matrix set; and use the first cross-enhancement processing matrix set to perform cross-enhancement on the first data cross-association group set respectively to obtain the first initial enhanced packaged sample detection data group set.
[0062] Furthermore, the data processing module 13 is also used to: pre-build a convolutional enhancement network layer; transmit the first data cross-association group corresponding to the first cross-enhancement processing matrix set and the first data cross-association group set to the convolutional enhancement network layer for analysis, and obtain the first initial enhanced packaged sample detection data set.
[0063] Furthermore, the initial detection result acquisition module 14 is also used to: extract features from different modal data of each enhanced repackaged sample detection data according to different detection types based on the feature extraction mechanism, and construct multiple multimodal features; use an attention mechanism to perform cross-modal weighted coupling on the multiple multimodal features to obtain multimodal coupling features; and perform detection analysis on the multimodal coupling features to obtain the initial oil sample detection result.
[0064] Furthermore, the initial detection result acquisition module 14 is also used for: the feature extraction mechanism is to perform principal component analysis on spectral data, perform statistical feature extraction on physical data, and perform standardized index analysis on chemical features.
[0065] Furthermore, the oil product testing result acquisition module 15 is also used to: extract M anomalies from the initial oil product sample testing results, where M is a positive integer; extract historical testing logs of the same type of oil from the M testing devices corresponding to the M anomalies to obtain a set of M historical testing logs; perform clustering trend identification on the set of M historical testing logs to determine M clustering trend features; determine whether the similarity between the M anomaly features of the M anomalies and the M clustering trend features meets a preset similarity threshold; if not, the device reliability verification is passed.
[0066] Furthermore, the oil product testing result acquisition module 15 is also used to: calculate the mean of the M historical testing log sets to obtain M mean drift centers; and use the mean drift algorithm in combination with the M mean drift centers to perform clustering trend identification on the M historical testing log sets to determine M clustering trend features.
[0067] Furthermore, the data cleaning module 12 is also used to: traverse the multiple sub-packaged sample test data sets for structured parsing, transform them into standardized data structures, and obtain multiple sub-packaged sample test standardized data sets; and perform outlier removal or baseline correction on the multiple sub-packaged sample test standardized data sets respectively to obtain the multiple pre-processed sub-packaged sample test data sets.
[0068] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0069] Obviously, those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for optimizing oil product testing data through multi-dimensional feature fusion, characterized in that, The method includes: The target oil sample is packaged according to the preset packaging requirements using an automated sampling device and sent to multiple testing devices for testing to obtain multiple sets of test data for the packaged samples. The preset packaging requirements are that each testing device must divide at least two packaged samples. The data cleaning and preprocessing of the multiple sets of test data for the sub-packaged samples are traversed to obtain multiple sets of preprocessed test data for the sub-packaged samples. The multiple pre-processed and packaged sample test data sets are subjected to in-set data cross-enhancement processing to determine multiple enhanced packaged sample test data sets, including: Extract the first pre-processed and packaged sample test data set from multiple pre-processed and packaged sample test data sets; Perform a cross-correlation of the first pre-processed and packaged sample detection data set without omissions or duplications to obtain a first data cross-correlation group set; Perform cross-enhancement processing on each data cross-association group in the first data cross-association group set to obtain the first initial enhanced sample detection data group set; Calculate the mean of the first initial enhanced subpackaged sample test data set to obtain the first enhanced subpackaged sample test data, and add the first enhanced subpackaged sample test data into the plurality of enhanced subpackaged sample test data; Based on the detection data of the multiple enhanced repackaged samples, multi-dimensional feature cross-modal fusion analysis is performed to obtain the initial oil sample detection results of the target oil sample; Based on the initial oil sample test results, equipment reliability verification is performed. When the equipment reliability verification passes, the initial oil sample test results are used as the target oil sample test results, including: Extract M outliers from the initial oil sample test results, where M is a positive integer; Extract historical detection logs of the same type of oil from the M detection devices corresponding to the M anomalies to obtain a set of M historical detection logs. Clustering trend identification is performed on the M historical detection log sets to determine M clustering trend features; Determine whether the similarity between the M anomaly features of the M anomaly items and the M clustering trend features meets a preset similarity threshold. If not, the device reliability verification passes. Specifically, the steps for identifying clustering trends in the M historical detection log sets to determine the M clustering trend features are as follows: Calculate the mean of the M historical detection log sets to obtain M mean drift centers; The mean shift algorithm is used in conjunction with the M mean shift centers to identify clustering trends in the M historical detection log sets, thereby determining M clustering trend features.
2. The method for optimizing oil product testing data through multi-dimensional feature fusion as described in claim 1, characterized in that, Cross-enhancement processing is performed on each data cross-association group in the first data cross-association group set to obtain a first initial enhanced sample detection data set, including: Similarity analysis is performed on the same type of data in each data cross-association group in the first data cross-association group set to obtain the first similarity group set; The first similarity group set is subjected to intra-group similarity normalization, and the processing results are added to an initially empty matrix to obtain the first cross-enhancement processing matrix set. The first set of cross-enhancement processing matrices is used to perform cross-enhancement on the first set of cross-correlation data groups to obtain the first set of enhanced subpackaged sample detection data groups.
3. The method for optimizing oil product testing data through multi-dimensional feature fusion as described in claim 2, characterized in that, The first set of cross-enhancement processing matrices is used to perform cross-enhancement on the first set of cross-correlation data groups to obtain the first set of enhanced packaged sample detection data groups, including: Pre-constructed convolutional enhancement network layers; The first data cross-association group corresponding to the first cross-enhancement processing matrix set and the first data cross-association group set is transmitted to the convolutional enhancement network layer for analysis to obtain the first initial enhanced packaged sample detection data set.
4. The method for optimizing oil product testing data through multi-dimensional feature fusion as described in claim 1, characterized in that, Based on the multidimensional feature cross-modal fusion analysis of the multiple enhanced repackaged sample detection data, the initial oil sample detection results of the target oil sample are obtained, including: Based on the feature extraction mechanism, features are extracted from different modal data of each enhanced sub-packaged sample detection data according to different detection types to construct multiple multimodal features; The attention mechanism is used to perform cross-modal weighted coupling on the multiple multimodal features to obtain multimodal coupled features; The multimodal coupling characteristics are detected and analyzed to obtain the detection results of the initial oil sample.
5. The method for optimizing oil product testing data through multi-dimensional feature fusion as described in claim 4, characterized in that, The feature extraction mechanism involves principal component analysis for spectral data, statistical feature extraction for physical data, and standardized index analysis for chemical features.
6. The method for optimizing oil product testing data through multi-dimensional feature fusion as described in claim 1, characterized in that, The multiple sets of test data for packaged samples are traversed for data cleaning and preprocessing to obtain multiple preprocessed sets of test data for packaged samples, including: The multiple sets of test data for the subpackaged samples are traversed and structured parsed to transform them into a standardized data structure, thereby obtaining a standardized set of test data for the multiple subpackaged samples. Outlier removal or baseline correction is performed on the multiple standardized test datasets of the subpackaged samples to obtain the multiple preprocessed test datasets of the subpackaged samples.
7. A multi-dimensional feature fusion-based oil product testing data optimization processing system, characterized in that, The steps for implementing the oil product testing data optimization processing method based on multi-dimensional feature fusion according to any one of claims 1 to 6 include: The detection data acquisition module is used to use an automated sampling device to package the target oil sample according to the preset packaging requirements and send it to multiple testing devices for testing to obtain multiple sets of test data for packaged samples. The preset packaging requirements are that each testing device should divide at least two packaged samples. The data cleaning module is used to traverse the multiple sets of test data for packaged samples to perform data cleaning and preprocessing, and obtain multiple sets of preprocessed test data for packaged samples. The data processing module is used to perform intra-set data cross-enhancement processing on the multiple pre-processed and packaged sample test data sets respectively to determine multiple enhanced packaged sample test data. The initial detection result acquisition module is used to perform multi-dimensional feature cross-modal fusion analysis based on the detection data of the multiple enhanced sub-packaged samples to obtain the initial oil sample detection result of the target oil sample; The oil product testing result acquisition module is used to verify the reliability of the equipment based on the initial oil product sample testing results. When the equipment reliability verification is passed, the initial oil product sample testing results are used as the target oil product sample testing results.
Citation Information
Patent Citations
Sample analysis method and system based on multi-dimensional gas chromatograph data fusion algorithm
CN119783038A
Oil product inspection and analysis method and device, computer equipment and storage medium
CN119829981A