A data screening method and system for recording the distance of a convolutionalized trajectory
By linearizing the data of multiple batches of products and convolution-based trajectory distance algorithm, abnormal data products and abnormal performance indicators are screened out, which solves the problem of difficult to identify data abnormalities of multiple batches of products in the existing technology, and achieves fast and accurate data screening and analysis.
Patent Information
- Application Number
- CN202510127488.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-05
AI Technical Summary
It is difficult to effectively identify data abnormalities in multiple batches, small batches, and high complexity products, especially the average deviation abnormalities caused by unstable batch product quality, and the multi-batch data analysis method is not fixed, which is due to the technical ability and subjective factors of the analyst.
By linearizing the data of preset batch multi-products, a data trajectory matrix is formed, and a convolution-based trajectory distance algorithm is used to calculate the difference between a single product and other products and the abnormal impact of performance data, and filter out abnormal data products and abnormal performance indicators based on the preset abnormal threshold.
It realizes the rapid screening of abnormal data from multiple batches of products, solves the problem of difficulty in forming unified data analysis standards for different categories of products, and uniformly describes complex data in multiple time, multiple volumes, and multi-dimensionality, reducing the subjectivity of the analysis process.
Smart Images

Figure CN119557606B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of product screening, and in particular relates to a data screening method and system for recording convolutional trajectory distances. Background Art
[0002] With the development of the existing technology, the performance of products and the complexity of testing have gradually increased. There is a lack of a specific and operable method for screening various large-scale data for multi-batch, small-batch, and high-complexity products, and it is difficult to identify data anomalies caused by factors such as unstable product quality in batches. For the screening of data anomalies for multi-batch, small-batch, and high-complexity products, currently it mainly relies on product designers. According to factors such as design requirements, component conditions, and development levels, independent anomaly judgment intervals are set for each indicator. And in actual data analysis, the anomaly status of a single indicator is mainly concerned. When all indicators are not in the anomaly interval, the product as a whole is judged to be normal, ignoring the anomaly that the indicator data of a small number of products deviates from the average value of the corresponding indicator data of the majority of products. When the deviation anomalies of the above-mentioned small number of products are superimposed on the averaging process of multi-batch product indicators, the anomaly deviations of the small number of products are easily diluted by the averaging process and are difficult to detect. On the other hand, the current methods for multi-batch data analysis are not fixed. The process and quality of batch anomaly screening are greatly affected by the technical capabilities and subjective factors of analysts. The data analysis of intermittent batch products lacks a unified dimension and cannot effectively support the multi-data analysis work of the entire batch. Summary of the Invention
[0003] The technical problem solved by the present invention is: overcoming the existing technology, providing a data screening method and system for recording convolutional trajectory distances, which can effectively identify batch anomalies and indicator anomalies in multi-batch data.
[0004] The object of the present invention is achieved through the following technical solutions: A data screening method for recording convolutional trajectory distances, comprising: performing linearization processing on the data of preset batch multi-products to obtain a data trajectory matrix; distinguishing the trajectory matrix of the same test data of multi-batch products in the data trajectory matrix from the trajectory matrix of single-product single-test data; obtaining the difference amount between a single product and other products according to the trajectory matrix of the same test data of multi-batch products; screening out abnormal data products in multi-batch products according to the difference amount between a single product and other products and a preset product anomaly threshold; obtaining the abnormal influence amount of performance data according to the trajectory matrix of single-product single-test data; screening out abnormal performance indicators according to the abnormal influence amount of performance data and a preset performance indicator anomaly threshold; obtaining abnormal products according to the abnormal data products and abnormal performance indicators in multi-batch products and locating specific abnormal performance indicators.
[0005] In the above data screening method for recording convolution trajectory distance, the linearization process of the preset batch multi-product data to obtain the data trajectory matrix includes: collecting the performance data of each index of each product, and calibrating the collected performance data of each index according to time to form a data table; according to the time series, each performance test data in the data table forms a trajectory of time-performance data; the product performance data under the same test process is placed in the same coordinate system to form a data trajectory matrix.
[0006] In the above data screening method for recording convolution trajectory distance, based on the trajectory matrix of the same test data of multiple batches of products, the trajectory distance algorithm based on convolution is used to obtain the difference amount between a single product and other products.
[0007] In the above data screening method for recording convolution trajectory distance, the difference amount between a single product and other products is obtained through the following formula:
[0008] ;
[0009] Where, is the difference amount between the th product and other products, is the minimum similarity of the trajectory of the th product in the data table in the th column index and the same index of other products except the th product in this batch of products in the th column, is the index value at the th column index and the rd moment, is the index value at the th column index and the th moment, is the first moment, is the total number of products in this batch, is the second moment, is the number of columns of the total performance index of the product, is the product serial number.
[0010] In the above data screening method for recording convolution trajectory distance, is obtained through the following formula:
[0011]
[0012] Where, is the similarity between the performance considering time shift error and the performance, is the th product at Set of column index values For the products in this batch except the th product, the other products are in the Set of column index values Is to introduce a displacement to the sequence Is an adjustable displacement threshold Is the index value at the th moment of the column index Is the index value at the th moment of the column index Is the index value at the th moment of the column index Is the index value at the th moment of the column index Is the third moment, m is the th product in the total time for column index testing
[0013] In the above data screening method for recording convolution trajectory distance, if the difference between a single product and other products is greater than the preset product anomaly threshold, then this product is an abnormal data product
[0014] In the above data screening method for recording convolution trajectory distance, the abnormal influence amount of performance data is obtained through the following formula
[0015] ;
[0016] Where Is the abnormal influence amount of performance data Is the start time of the time period partition Is the end time of the time period partition Is the column position serial number where the performance index is located Is the total number of columns of the performance index Is the preset time within the time partition Is column index value at the Is column to index average value at the Is the total number of products in this batch Is the fourth moment Is the product performance index matrix at the It is the average value matrix of product performance indicators.
[0017] A data screening system for recording convolutional trajectory distances, comprising: a first module for linearizing the data of preset batch multi-products to obtain a data trajectory matrix; a second module for distinguishing the trajectory matrix of the same test data of multi-batch products and the trajectory matrix of single-product single-test data in the data trajectory matrix; a third module for obtaining the difference amount between a single product and other products according to the trajectory matrix of the same test data of multi-batch products; a fourth module for screening out abnormal data products within multi-batch products according to the difference amount between a single product and other products and a preset product abnormality threshold; a fifth module for obtaining the abnormal influence amount of performance data according to the trajectory matrix of single-product single-test data; a sixth module for screening out abnormal performance indicators according to the abnormal influence amount of performance data and a preset performance indicator abnormality threshold; a seventh module for obtaining abnormal products according to the abnormal data products and abnormal performance indicators within multi-batch products and locating specific abnormal performance indicators.
[0018] In the above data screening system for recording convolutional trajectory distances, the linearization process of the data of preset batch multi-products to obtain a data trajectory matrix includes: collecting the performance data of each index of each product, calibrating the collected performance data of each index according to time to form a data table; according to the time series, each performance test data in the data table forms a trajectory of time-performance data; the product performance data under the same test process is placed in the same coordinate system to form a data trajectory matrix.
[0019] In the above data screening system for recording convolutional trajectory distances, according to the trajectory matrix of the same test data of multi-batch products, the difference amount between a single product and other products is obtained by using a convolutional-based trajectory distance algorithm.
[0020] The present invention has the following beneficial effects compared with the prior art:
[0021] (1) The present invention only performs correlation analysis through the provided data, quickly screens abnormal data in the testing process of multi-batch, small-batch, and high-complexity products, and solves the problem that it is difficult to form a unified data analysis standard for different types of products;
[0022] (2) Through the idea of data trajectoryization, the present invention can abstract each column index and time in the data into a two-dimensional trajectory, and abstract the data of multiple indicators, multiple products, and multiple batches into a high-dimensional space trajectory for description, solving the problem of unified description of complex data with multiple times, multiple individuals, and multiple dimensions;
[0023] (3) By applying the trajectory distance algorithm, the present invention performs multi-batch distance calculations on the data abstracted as multi-dimensional trajectories, and quantifies the relationships and trends of the test data using distances, thereby solving the problem of unifying the dimensions of different metrics.
[0024] (4) By supplementing the dynamic class convolution algorithm, the present invention performs round-robin comparisons during the process of the trajectory distance algorithm, and takes the minimum value of the trajectory distances within a certain period of time as the final result, which can offset the time shift error caused by the misalignment of the time dimension and solve the problem of false errors caused by the time misalignment of the test data of different individuals and batches of products. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0026] Figure 1 is a flowchart of the data screening method for recording class convolution trajectory distances provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0028] Figure 1 is a flowchart of the data screening method for recording class convolution trajectory distances provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0029] Perform linearization processing on the preset batch multi-product data to obtain a data trajectory matrix;
[0030] Distinguish the trajectory matrix of the same test data of multiple batches of products in the data trajectory matrix from the trajectory matrix of the single product's single test data;
[0031] Obtain the difference amount between a single product and other products based on the trajectory matrix of the same test data of multiple batches of products;
[0032] Screen out the abnormal data products in multiple batches of products according to the difference amount between a single product and other products and the preset product abnormality threshold;
[0033] Obtain the abnormal influence amount of performance data according to the single product single test data trajectory matrix;
[0034] According to the abnormal influence amount of performance data and the preset abnormal threshold of performance indicators Screen out the abnormal performance indicators;
[0035] Obtain the abnormal products according to the abnormal data products and abnormal performance indicators in multiple batches of products and locate the specific abnormal performance indicators.
[0036] Perform linearization processing on the data of the preset batch of multiple products to obtain a data trajectory matrix, including: collecting the performance data of each indicator of each product, and calibrating the collected performance data of each indicator according to time to form a data table; according to the time series, each performance test data in the data table forms a trajectory of time-performance data; the product performance data corresponding to the same test process is placed in the same coordinate system to form a data trajectory matrix.
[0037] According to the trajectory matrix of the same test data of multiple batches of products, use the convolutional-based trajectory distance algorithm to obtain the difference amount between a single product and other products.
[0038] The difference amount between a single product and other products is obtained through the following formula:
[0039] ;
[0040] Among them, is the difference amount between the th product and other products, is the th product's minimum similarity of the trajectory of the th column indicator in the data table with the same indicator of other products except the th product in this batch, is the trajectory minimum similarity of the same indicator of other products except the th product in this batch, is the indicator value at the th moment of the th column indicator, is the indicator value at the th moment of the th column indicator, is the first moment, is the second moment, is the number of columns of the total product performance indicators, is the product serial number.
[0041] Obtained by the following formula:
[0042]
[0043] Wherein, is the similarity between the performance considering the time shift error and the performance, is the set of column index values of the th product in the th position, is the set of column index values of other products in this batch of products except the th product, is the displacement introduced to the sequence, is the adjustable displacement threshold, is the index value at the th moment of the th column index, is the index value at the th moment of the th column index, is the index value at the th moment of the th column index, is the index value at the th moment of the th column index, is the index value at the th moment of the th column index, m is the total time for testing the th product in the th column index.
[0044] If the difference between a single product and other products is greater than the preset product anomaly threshold, then this product is an abnormal data product.
[0045] The abnormal influence amount of performance data is obtained by the following formula:
[0046] ;
[0047] Wherein, is the abnormal influence amount of performance data, is the start time of the time period partition, is the end time of the time period partition, is the column position serial number where the performance index is located, is the total number of columns of the performance index, is the preset time within the time partition, is the th column index value at the th moment, Column to the average value of the indicators at the moment, is the total number of products in this batch, is the fourth moment, is the product performance index matrix at the moment, is the average value matrix of product performance indicators.
[0048] Specifically, the method includes the following steps:
[0049] S1. Input product data.
[0050] S2. Linearly process the product data to obtain a complex trajectory matrix of performance descriptions in units of single products.
[0051] S21. During different stages such as in-house simulation, debugging, and environmental testing during development, data of each indicator is usually collected according to time, and the collected data is calibrated according to time to form a data table in.csv format, as shown in the following table:
[0052] Table 1 Collection format of data to be screened
[0053]
[0054] S22. According to the time series, each performance test data can form a trajectory of time-performance data, and the distance between trajectories can describe the difference between indicator data;
[0055] S23. Product performance data under the same test process can be placed in the same coordinate system to form a data trajectory matrix;
[0056] S3. Distinguish the trajectory matrix of the same test data of multiple batches of products from the trajectory matrix of single product single test data.
[0057] S4. Apply the trajectory distance algorithm to the trajectory matrix of the same test data of multiple batches of products to calculate the difference between a single product and other products.
[0058] S41. For the same type of test data of different products with a batch total of p, the data information is saved in the form of a matrix list in.csv format, and the test data of each product ( ) is stored in the performance index (column) matrix at the m moment (row), the total number of performance indicators (columns) of a single product in this batch, and the difference between the matrix test data of product ( ) and other products can be described by anomalies ; described;
[0059] S42, quantitatively describe the difference between each two columns of data. First, the selected two columns of performance indicator data are defined as:
[0060] ;
[0061] in, is the value at time 1 (row 1) Column indicator value, For time 2 (row 2) Column indicator value, is the value at time 3 (row 3) Column indicator value, is the value at time m (row m) Column indicator value, is the value at time 1 (row 1) Column indicator value, For time 2 (row 2) Column indicator value, is the value at time 3 (row 3) Column indicator value, is the value at time m (row m) Column indicator value.
[0062] S43, data may also face time shift errors in the performance trajectory curve, which can be eliminated by integrating the convolution calculation process. The sequence is processed to eliminate the potential time shift error caused by time misalignment. Sequence-Introduced Displacement , and set the displacement threshold make , then the displacement is included New sequence It can be defined as:
[0063]
[0064] S44, therefore, in the case of the above time-shift error, if the calculation is simply based on the trajectory distance, a gray inherent trajectory error will appear, but this error does not mean that the two curves have different trajectory forms. In order to eliminate the time-shift error, the present invention introduces displacement to the data and uses a convolution-like round-robin method to find the minimum trajectory distance. Specifically, the interval is the shift threshold. In the interval, the new sequence In a convolutional form, Calculate the similarity:
[0065]
[0066] S45, The minimum similarity with can be defined as:
[0067]
[0068] S46. In the actual screening of abnormal products, it is more desirable to locate the specific performance indicators with problems. Therefore, after grouping based on different attributes, slicing data by time sequence, calculating the mean square deviation within the same attribute and the same slice, and then normalizing, the abnormal device attributes with relatively large changes in a short time are found and marked as abnormal. Specifically, the overall difference of the data matrix between a single product ( ) and other products is defined, then:
[0069]
[0070] S5. Abnormal data products are screened out through the product anomaly threshold.
[0071] S51. Set the product anomaly judgment threshold as E according to the specific product characteristics;
[0072] S52. Through the preset product anomaly threshold E, when , the product ( ) can be judged as an abnormal product;
[0073] S6. Calculate the abnormal impact amount of performance data for the performance data trajectory matrix of a single product test.
[0074] S61. Since the product test time is uncontrollable, in order to reduce the computational complexity, the test matrix is partitioned by rows through time period , then the impact amount caused by the abnormal performance of a certain q (column) is defined as ; ;
[0075] S62. The specific performance indicators that cause the overall anomaly of the test data matrix between a single product ( ) and other products are The impact amount can be calculated by the following formula:
[0076]
[0077] S7. Screen out the specific abnormal performance indicators within a certain time period through the performance anomaly threshold.
[0078] S71. Set the performance indicator anomaly threshold according to the specific product characteristics;
[0079] S72. Therefore, through the preset performance indicator anomaly threshold , when When it can be determined that it is in the area The performance index of q (column) is the main reason for the abnormality of the product ( ).
[0080] S8. Summarize the screening information of abnormal products, determine the abnormal products and locate the specific abnormal performance indicators, and screen out the abnormal data.
[0081] This embodiment also provides a data screening system for recording the distance of convolutionalized trajectories of classes. The system includes: a first module for linearizing the data of preset batch multi-products to obtain a data trajectory matrix; a second module for distinguishing the trajectory matrix of the same test data of multi-batch products in the data trajectory matrix from the trajectory matrix of single-product single-test data; a third module for obtaining the difference amount between a single product and other products according to the trajectory matrix of the same test data of multi-batch products; a fourth module for screening out abnormal data products in multi-batch products according to the difference amount between a single product and other products and a preset product abnormality threshold; a fifth module for obtaining the abnormal influence amount of performance data according to the trajectory matrix of single-product single-test data; a sixth module for screening out abnormal performance indicators according to the abnormal influence amount of performance data and a preset performance index abnormality threshold; a seventh module for obtaining abnormal products according to the abnormal data products and abnormal performance indicators in multi-batch products and locating the specific abnormal performance indicators.
[0082] This embodiment only performs correlation analysis through the provided data, quickly screens out abnormal data in the testing process of multi-batch, small-batch, and high-complexity products, and solves the problem that it is difficult to form a unified data analysis standard for different categories of products; through the idea of data trajectory, this embodiment abstracts each column index and time in the data into a two-dimensional trajectory, and abstracts the data of multiple indicators, multiple products, and multiple batches into a high-dimensional space trajectory for description, solving the problem of unified description of complex data with multiple times, multiple individuals, and multiple dimensions; by applying the trajectory distance algorithm, this embodiment calculates the distances of the data abstracted into multi-dimensional trajectories in multiple batches, and quantifies the relationship and trend of the test data using the distance, solving the problem of unified dimension for different indicators; by supplementing the dynamic class convolution algorithm, this embodiment performs round-robin comparison during the process of the trajectory distance algorithm, and takes the minimum value of the trajectory distance within a certain time as the final result, which can offset the time shift error caused by the misalignment of the time dimension, and solves the problem of false error caused by the misalignment of the test data time of different individuals and batch products.
[0083] Although the present invention has been disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical content disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.
Claims
1. A data screening method for recording convolutional trajectory distances, characterized in that include: Linearize the data of preset batches of multiple products to obtain a data trajectory matrix; Distinguish between the trajectory matrix of the same test data of multiple batches of products and the trajectory matrix of the single test data of a single product in the data trajectory matrix; The difference between a single product and other products is obtained based on the trajectory matrix of the same test data of multiple batches of products; Filter out abnormal data products in multiple batches of products based on the difference between a single product and other products and the preset product abnormality threshold; Obtain the abnormal impact of performance data based on the single test data trajectory matrix of a single product; Filter out abnormal performance indicators based on the abnormal impact of performance data and the preset abnormal threshold of performance indicators; Obtain abnormal products and locate specific abnormal performance indicators based on abnormal data products and abnormal performance indicators in multiple batches of products; The difference between a single product and other products is obtained by the following formula: ; in, For the The difference between a product and other products, For the The products are in the data table The indicators listed are the same as those of this batch of products except Other products besides products are in List the minimum similarity of trajectories of the same indicator, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For the first moment, is the total number of products in this batch, For the second moment, is the number of columns of the total performance index of the product, is the product serial number; It is obtained by the following formula: in, To consider the time shift error Performance and The similarity between performance, For the Products in A collection of column indicator values, For this batch of products except Other products besides products are in A collection of column indicator values, For The sequence introduces displacement, is the adjustable displacement threshold, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, is the third moment, m is the Products in The total time of the column metric test.
2. The data screening method for recording convolutional trajectory distance according to claim 1 is characterized in that: The data trajectory matrix obtained by linearizing the preset batch multi-product data includes: Collect the performance data of each indicator of each product, and calibrate the collected performance data of each indicator according to time to form a data table; According to the time series, each performance test data in the data table forms a track of time-performance data; The product performance data corresponding to the same test process are placed in the same coordinate system to form a data trajectory matrix.
3. The data screening method for recording convolutional trajectory distance according to claim 1, characterized in that: If the difference between a single product and other products is greater than the preset product anomaly threshold, the product is an abnormal data product.
4. The data screening method for recording convolutional trajectory distance according to claim 1, characterized in that: The abnormal impact of performance data is obtained by the following formula: ; in, is the abnormal impact of performance data, is the starting time of the time period partition, is the end time of the partition according to the time period, is the column position number of the performance indicator, is the total number of columns of performance indicators, Preset time in time zone. for List The indicator value at the moment, for List arrive The average value of the indicator at the time, is the total number of products in this batch, For the fourth moment, For the The product performance indicator matrix at all times, It is the average value matrix of product performance indicators.
5. A data screening system for recording convolutional trajectory distances, characterized in that include: The first module is used to perform linear processing on the data of preset batches of multiple products to obtain a data trajectory matrix; The second module is used to distinguish between the trajectory matrix of the same test data of multiple batches of products and the trajectory matrix of the single test data of a single product in the data trajectory matrix; The third module is used to obtain the difference between a single product and other products based on the trajectory matrix of the same test data of multiple batches of products; The fourth module is used to screen out abnormal data products in multiple batches of products based on the difference between a single product and other products and a preset product abnormality threshold; The fifth module is used to obtain the abnormal impact of performance data based on the single test data trajectory matrix of a single product; The sixth module is used to filter out abnormal performance indicators according to the abnormal impact of performance data and the preset abnormal threshold of performance indicators; The seventh module is used to obtain abnormal products and locate specific abnormal performance indicators based on abnormal data products and abnormal performance indicators in multiple batches of products; The difference between a single product and other products is obtained by the following formula: ; in, For the The difference between a product and other products, For the The products are in the data table The indicators listed are the same as those of this batch of products except Other products besides products are in List the minimum similarity of trajectories of the same indicator, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For the first moment, is the total number of products in this batch, For the second moment, is the number of columns of the total performance index of the product, is the product serial number; It is obtained by the following formula: in, To consider the time shift error Performance and The similarity between performance, For the Products in A collection of column indicator values, For this batch of products except Other products besides products are in A collection of column indicator values, For The sequence introduces displacement, is the adjustable displacement threshold, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, For Column indicator The indicator value at the moment, is the third moment, m is the Products in The total time of the column metric test.
6. The data screening system for recording convolutional trajectory distance according to claim 5, characterized in that: The data trajectory matrix obtained by linearizing the preset batch multi-product data includes: Collect the performance data of each indicator of each product, and calibrate the collected performance data of each indicator according to time to form a data table; According to the time series, each performance test data in the data table forms a track of time-performance data; The product performance data corresponding to the same test process are placed in the same coordinate system to form a data trajectory matrix.
Citation Information
Patent Citations
Method and device for screening defective battery cells and electronic equipment
CN115061043A
Equipment anomaly detection method, equipment, storage medium and program product
CN116955103A
Method and system for predicting operation state of power distribution network based on digital twinning
CN117639228A