Detection data and sampling inspection strategy model analysis method based on big data technology
By building a knowledge graph for random inspections of power materials and a differentiated random inspection rule base, combined with big data and artificial intelligence, the problems of batch imbalance and supplier quality adjustment in the random inspection strategy of power materials have been solved, achieving more scientific and efficient random inspection management, and improving the safety and stability of the power grid.
Patent Information
- Application Number
- CN202510817391.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing power material sampling inspection strategy, the judgment probability of different product batches is uneven, and the sampling quantity is not adjusted according to the supplier's historical supply quality, resulting in a high probability of misjudgment of unqualified batches and a lack of scientificity and rationality.
Big data analysis technology is used to construct a knowledge graph for random inspections of power materials. Combined with suppliers' historical quality data and material characteristics, a differentiated random inspection rule base is established. Random inspection strategies are generated through artificial intelligence algorithms and risk assessment models, and model parameters are dynamically optimized.
It realizes differentiated sampling inspection strategy management, reduces reliance on manual experience, increases the detection rate of defective products, reduces inspection costs, and improves the safety and stability of power grid operation.
Smart Images

Figure CN120706972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power quality control, and in particular to a method for analyzing detection data and sampling strategy models based on big data technology. Background Art
[0002] With socioeconomic development and accelerated urbanization, the demand for power resources for social production activities is also increasing. Power grid companies must purchase a large amount of distribution network materials annually, specifically equipment and materials with voltage levels of 35kV and below. Therefore, pre-purchase sampling of distribution network materials is a key process in procurement quality control and a crucial task for achieving quality control of power grid materials and ensuring the safe and stable operation of the distribution network.
[0003] While the introduction and investment of numerous new technologies, products, and equipment have enhanced the safety and stability of power grid operations, suppliers' manufacturing and R&D costs have also continued to rise, resulting in corresponding declines in profit margins. This imbalance in cost and profit can lead to a host of problems. For example, some suppliers prioritize quantity over quality in exchange for profit, resulting in frequent product failures during operation, impacting power grid security and reducing the quality of power supply.
[0004] Against this backdrop, power equipment quality inspection, as a key control measure for power material management, is becoming increasingly important for power construction, grid security, and operation and maintenance. Power companies are urgently required to optimize distribution network equipment inspection techniques and models, guided by inspection technical standards, supply contracts, and relevant national standards. This allows them to more effectively conduct inspections of distribution network equipment, thereby gaining real-time insights into high-risk areas for quality risks across various materials, enabling effective supervision and control of relevant suppliers and materials. This ensures the continued need for quality management and management across all stages of the process, from product production to sales, installation, and subsequent operation and testing. This ensures the quality of materials entering the grid and further improves the safety and stability of grid operations.
[0005] Currently, research on differentiated sampling inspection strategies for power supplies, both domestically and internationally, has generally focused on conventional analysis methods. For example, a factory sampling inspection strategy based on classification evaluation selects and inspects a certain proportion of current supplies, while also allowing for secondary adjustments based on material and supplier dimensions to ultimately achieve differentiated material sampling inspections.
[0006] For example, the number of samples collected during random inspections at State Grid varies based on its coverage requirements. For a long time, key rural distribution network materials, such as distribution transformers, overhead insulated wires, power cables, distribution boxes, cement poles, and iron accessories, have been subject to strict inspections based on four criteria: all winning bid batches, all winning suppliers, all material categories, and all incoming shipments. The inspection percentage must not be less than 5%.
[0007] The current single-percentage sampling method uses the same threshold for all product batches, regardless of batch size, and samples are drawn from the batch in the same proportion. This approach is unscientific and presents two problems: 1) the severity of the sampling scheme varies significantly for different batches, increasing the probability of misclassifying unqualified batches as qualified when the batch size is large; and 2) the sampling size is not adjusted based on the supplier's historical quality feedback from random inspections.
[0008] Based on the relevant literature, there are currently few applications of material sampling inspection strategy decision-making systems based on big data analysis models. This project uses the design and decision-making theories of big data combined with artificial intelligence to create a supplier evaluation knowledge graph by integrating data such as suppliers' historical material procurement information, supplier bid winning status, and historical quality control. It also establishes a material sampling inspection strategy rule base. Through an intelligent decision support system, the supplier evaluation index scores are automatically matched with the material sampling inspection strategy rule base to output the sampling inspection strategy.
[0009] The analysis and research on the power material random inspection strategy model based on big data analysis can effectively solve the problem of the random inspection strategy being too dependent on experts, realize differentiated random inspection strategy management, make material quality management more controllable, cost-saving, and decision-making more reasonable. Summary of the Invention
[0010] The present invention aims to solve the problems of poor adaptability, low communication reliability and rigid control architecture in water quality monitoring and control of traditional phase-shifting external cooling water systems, and provides a detection data and sampling strategy model analysis method based on big data technology.
[0011] To achieve the above object, the present invention provides the following technical solutions:
[0012] A method for analyzing test data and sampling strategy models based on big data technology includes the following steps:
[0013] Constructing a knowledge graph for power material spot checks: Extracting entities and relationships from multiple dimensions, including material suppliers, equipment information, and operating status, to construct a knowledge graph with an "entity-relationship-entity" triple structure. This knowledge graph is constructed using a top-down approach, including ontology construction, knowledge extraction, fusion, and storage steps.
[0014] Data collection and data pool management: Multi-source data from testing equipment is collected through wired Ethernet ports, serial ports, or USB interfaces, and stored in a data pool after cleaning and conversion. The data pool integrates a storage engine, a computing engine, and a data governance module.
[0015] Establish a material sampling inspection strategy rule base: define inspection items, quantity, and organization sampling inspection strategy elements, and formulate differentiated sampling inspection rules based on the supplier's historical quality data and material characteristics;
[0016] Sampling inspection strategy decision support system modeling: Based on the knowledge graph and rule base, the sampling inspection strategy is generated using artificial intelligence algorithms, data mining methods and risk assessment models;
[0017] Scenario application and strategy optimization: Apply sampling strategies in scenarios such as low-qualification supplier analysis and new supplier testing, and dynamically optimize model parameters based on real-time data.
[0018] As a further technical solution of the present invention: in the knowledge graph construction step, the structured, semi-structured and unstructured data are preprocessed by ETL tools, uniformly converted into structured data and then loaded into the knowledge ontology library.
[0019] As a further technical solution of the present invention: the data acquisition module supports real-time data acquisition of transformers, lightning arrester equipment and materials, as well as text / database file parsing of cable protection pipes and line hardware materials.
[0020] As a further technical solution of the present invention: the differentiated sampling inspection rules include: increasing the sampling inspection volume for suppliers with many historical quality problems, adding inspection items for new suppliers or materials using new technologies, and designing a counting-adjusted sampling inspection plan based on the MIL-STD-1916 standard.
[0021] As a further technical solution of the present invention: the artificial intelligence algorithm includes a binary tree search algorithm, which realizes intelligent matching with the rule base by converting the knowledge graph tree structure into a binary tree.
[0022] As a further technical solution of the present invention: the data mining method includes:
[0023] Decision tree algorithm, used to classify and filter features of material detection data;
[0024] Canonical correlation coefficient analysis is used to calculate the correlation between test items;
[0025] The quadratic exponential smoothing method is used to predict the material failure rate.
[0026] As a further technical solution of the present invention: the risk assessment model is constructed based on the box plot, and the data is divided into reasonable areas, warning areas and abnormal areas by analyzing the minimum, quartile and maximum values. The probability density and weight density of the test parameters are introduced to establish a total probability function to achieve risk level division.
[0027] As a further technical solution of the present invention: the scenario application includes:
[0028] Analysis scenario for suppliers with low pass rates: Comprehensive quality problem list and sampling pass rate data, dynamically adjusting inspection levels;
[0029] Testing agency selection scenario: Allocate random inspection tasks based on the capabilities and carrying capacity of the testing center.
[0030] As a further technical solution of the present invention: the data pool management module adopts MySQL and ClickHouse databases to support data service output.
[0031] As a further technical solution of the present invention: the outlier detection algorithm adopts efficient distributed computing to identify abnormal points in the detection data by defining the neighborhood cardinality and the distance threshold.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. Intelligent decision-making: Through big data analysis and artificial intelligence algorithms, reduce dependence on manual experience and improve the scientific nature of sampling strategies.
[0034] 2. Precision management and control: Based on the knowledge graph and rule base, differentiated management of suppliers and materials is achieved, and the detection rate of defective products is improved.
[0035] 3. Efficient operation: Data pool and distributed computing technology improve data processing efficiency, shorten sampling cycle, and reduce testing costs.
[0036] 4. Risk visualization: Box plot models and outlier detection enable intuitive display of quality risks, facilitating timely control measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a schematic diagram of the process of building a knowledge graph for power material suppliers;
[0038] Figure 2 This is the overall architecture diagram of data collection;
[0039] Figure 3 This is a diagram of the process of converting a tree into a binary tree;
[0040] Figure 4 It is a schematic diagram of the risk assessment model based on the box plot;
[0041] Figure 5 This is the data collection flow chart for the intelligent gateway business system;
[0042] Figure 6 This is a diagram of the data pool management architecture. DETAILED DESCRIPTION
[0043] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0044] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention. The technology, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0045] like Figure 1-6 As shown, the present invention proposes a method for analyzing detection data and sampling strategy models based on big data technology, and the specific method is as follows:
[0046] (1) Constructing a knowledge graph for power materials:
[0047] Leveraging State Grid Corporation's existing information resources, we will build a knowledge graph for power supplies, collecting attribute information for specific entities from various information sources. This will lay the foundation for future data pool management and material sampling inspection rules. We can construct relevant knowledge graphs from perspectives such as material suppliers, equipment information, and equipment operating status.
[0048] (2) Construct data acquisition module and data pool management module:
[0049] Integrate and assetize the massive, multi-source, and heterogeneous data of electric power materials across the entire region to achieve enterprise-level data sharing and capability reuse, and provide data resources and capability support for other business front-ends to achieve data-driven refined operations.
[0050] (3) Establish a material sampling inspection strategy rule base:
[0051] Establish a supplier sampling inspection strategy rule base. Use the inference engine algorithm to achieve intelligent matching between the knowledge graph library and the sampling inspection strategy rule base. Implement intelligent output of sampling inspection strategies.
[0052] (4) Research on modeling of sampling strategy decision support system:
[0053] Based on big data analysis technology, a multi-category inspection strategy support model is constructed, which is supported by a knowledge graph library and a policy rule library. Based on the inspection items, multi-dimensional and differentiated inspection strategy management is achieved.
[0054] (5) Research on application scenarios of sampling strategy decision support:
[0055] The application development of the material sampling inspection system realizes the development of a series of application functions from the issuance and execution of sampling inspection plans to strategy configuration.
[0056] Here’s how it works:
[0057] 1. How to build a knowledge graph for random inspection of power materials Figure 1 As shown in the figure, in response to the current situation of diversity of data sources, semi-structured and unstructured data are first preprocessed and uniformly converted into structured data. Then, ETL tools are used to extract, transform, and clean the structured data, and finally the data is loaded into the knowledge ontology library.
[0058] The basic unit of the knowledge graph is the triple consisting of "Entity - Relationship - Entity", which is also the core of the knowledge graph. For example, the evaluation of power material suppliers is a core knowledge graph. The index system includes five aspects for evaluation: supplier qualifications and capabilities, performance evaluation, contract fulfillment evaluation, supply and demand dependence, and supplier negative evaluation. Each specific indicator includes multiple sub-evaluation indicators. According to the knowledge graph construction method, the supplier's annual material procurement situation, supplier's winning bid situation, historical quality control, and supplier material sampling inspection information can be used to establish a supplier evaluation knowledge graph. As an important part of the knowledge base of the sampling inspection decision support system, the knowledge graph is one of the cores of the decision-making system.
[0059] 2. Build a data collection module and data pool. To meet the needs of project construction, you can first build a data collection module and simultaneously build a data pool management module. Later, according to the growth of data volume, timely expansion will be carried out, including the construction of data collection module and data pool.
[0060] Data Collection Module Construction: When State Grid's material testing equipment is conducting experiments, the data gateway device collects the test results of the testing equipment through its supported physical peripherals (serial port, USB port, wired Ethernet port) and performs related business or algorithm processing. Finally, the collected data is connected to the private network through the 4G IoT card physical peripheral and transmitted to the IoT management platform in the form of a JSON string. The data gateway device accesses the State Grid's material testing results data through its supported physical peripherals: serial port, USB port, wired Ethernet port.
[0061] Data collection and processing process: It includes two categories: gateway module, serial port and Ethernet data collection module. The data collection gateway module processes the received data according to the data message protocol of each device and forwards the data according to the application platform protocol access method (MQTT). The serial port and Ethernet data collection module monitors the serial port data of the on-site detection equipment through the gateway collection device and transmits the data to the business service module. The data collection process of the intelligent gateway business system is as follows: Figure 5 .
[0062] Data pool management: Data storage management is achieved by establishing a data pool, including copies of raw data generated by the original system and converted data generated for various tasks. Based on the data pool, a simple computing engine and service engine are implemented, one or two knowledge graphs are implemented, and some basic data governance functions are realized.
[0063] Data pool management architecture Figure 6 ,include:
[0064] Data cleaning module: Since the data cleaning module is relatively independent in function and not strongly related to the business, it can be built in the early stages of the project. It mainly includes deviation detection, data correction, data conversion, and data extraction modules. Deviation detection and data correction are mainly used to detect and correct abnormal data, while data conversion is used to convert data in different formats. For example, date data has different collection formats in different systems, including yy-mm-dd, mm-dd, mm-dd-yy, and other formats, which require unified conversion. The data extraction module selectively loads a portion of data into the data mart according to business needs to control the data volume of the data mart and ensure access performance.
[0065] Storage Engine and Compute Engine: As core additional modules for data pool management, to achieve data pool construction goals, the storage engine and compute engine can initially handle only millions of data accesses. To this end, the storage engine can use the MySQL database, and the compute engine uses the ClickHouse column-based database. Both meet the performance requirements of lightweight data collection, storage, and conversion. Furthermore, their simple backend architecture allows for rapid deployment and facilitates daily operations and maintenance.
[0066] Data Governance Module: As data scale expands, manual management of data definitions, statistical calibers, and relationships between data becomes impossible. Automated management is required through the Data Governance Module to ensure data consistency and integrity. This module primarily includes metadata management.
[0067] Service Engine: This engine's key advantage is that it transforms direct database access into a data service, refining data granularity and increasing data access flexibility, enabling faster response to business changes. Therefore, during the data pool construction phase, a service engine is necessary, but it can only implement basic management, such as service registration and management.
[0068] Application level: Considering the construction cycle and data volume, it is recommended to select 1-2 material sampling and analysis areas for development during the construction phase. This will not only ensure controllable costs and risks, but also realize a complete business process, initially meet business needs, and accumulate experience for subsequent special analysis.
[0069] 3. Establish a material sampling inspection strategy rule base: Power material sampling inspection strategies primarily include adjusting inspection items, inspection content, inspection agencies, inspection quantity, and inspection standards. Before establishing a sampling inspection strategy base, you must first establish sampling inspection measures. Each sampling inspection item can have multiple measures, and a single sampling inspection strategy can include a combination of multiple sampling inspection measures.
[0070] Dynamic sampling inspection plan design (based on MIL-STD-1916):
[0071] Parameter definition:
[0072] Quality level: Select "Level I" (the acceptance probability is close to the middle value of the traditional 5% sampling plan); Sample code mapping:
[0073] Batch size 20-170: Sample code A, number of samples = batch size × 5% (rounded down, minimum 1);
[0074] Batch size 171-200: Sample code B, number of samples = batch size × 8% (round down).
[0075] Transfer rules:
[0076] 5 consecutive batches of qualified products → Relaxation of random inspection (reduction of sample size by 30%);
[0077] 1 batch failed → Stricter inspection (increase sample size by 50%).
[0078] 4. Decision support system algorithm implementation:
[0079] Binary tree search algorithm (rule matching):
[0080] Steps to convert a tree to a binary tree:
[0081] Add line: add horizontal lines between sibling nodes of the same level (such as "Supplier Qualification" and "Performance Evaluation");
[0082] Remove line: Only the line connecting the parent node and the first child node is retained (e.g. "Supplier Evaluation" → "Qualification and Capability"), and other child nodes are converted to the right subtree;
[0083] Level adjustment: Rotate clockwise with the root node as the center to form a binary tree structure (see example Figure 3 ).
[0084] Matching logic: input parameters (such as "supplier qualification rate 65%) → traverse the binary tree → match the "low qualification rate supplier sampling ratio strategy" → output sampling quantity increased by 40%.
[0085] The basic formula of the quadratic exponential smoothing method is shown in formulas (1) and (2).
[0086] yt+m=(2+am / (1-a))yt'-(1+am / (1-a))yt=(2yt'-yt)+m(yt'-yt)a / (1-a) Formula (1)
[0087] Where, 2yt'–yt is the intercept; (yt'–yt)a / (1–a) is the slope; a is the smoothing parameter; and t is the number of days to be predicted.
[0088] St=αSt+(1-α)St-1Yt+T=at+btT at=2St-St bt=(α / 1-α)(St-St)
[0089] St=aSt=(1-a)St-1Yt+T=at+btTat=2St-Stbt=(a / 1-a)(St-St)
[0090] Formula (2)
[0091] In the formula, St is the first exponential smoothing value of the t-th period; St–1 is the second exponential smoothing value of the t-th period; α is the smoothing coefficient; Yt+T is the predicted value of the t+T-th period; T is the number of periods moving backward from the t-th period.
[0092] After screening and filtering, the preprocessed data undergoes data descriptive analysis, classification, and prediction. Assuming there are six unqualified items, treat the six unqualified item labels as a vector of 6 dimensions. Perform PCA dimensionality reduction on all remaining vectors and map them onto a two-dimensional plane to form a two-dimensional image. If no obvious clustering effect is observed, this indicates that the correlations between the data items are weak and the connections are not strong. If the two-dimensional image and clustering effect are clearly visible after normalization of the original data, and the clustering effect is stronger than before normalization, but the classification effect is still poor and patterns are not clear, the CCA algorithm can be used to further process and improve the data after decision tree processing.
[0093] Efficient distributed outlier detection algorithm based on big data:
[0094] Outlier detection is mainly to mine data to make relevant work more effective. Usually, using this detection method will discover relevant specific behavior data, which will improve the efficiency of relevant work and reduce the time for unnecessary data exploration. According to the specific definition of outliers, an outlier is a corresponding observation point. If the deviation of an outlier from other observation points is large, there is reason to suspect that it is caused by a difference in mechanism. These deviated data and non-conforming data can be unified with a name, that is, outliers. Outliers can also be called isolated points or abnormal points. Outlier mining is also outlier detection. The efficient distributed outlier detection algorithm based on big data has low time complexity and high clustering accuracy in specific operations, and can gather different types of data together. The ultimate goal is to mine data clusters.
[0095] (1) Design of efficient distributed outlier detection algorithm:
[0096] Generally, given a data set P with d-dimensional attributes, the number of data points in the data set is |P|. For any data point p in P, p includes d measurable attribute values, denoted as p = <p[0], p[1], …, p[d - 1]> (for convenience of description, it is considered in the following text that the attribute values of each dimension of the data point are not less than 0)[3]. Then the distance between points p1 and p2 is
[0097]
[0098] Definition 1 Let it be the Q neighborhood. For any real number Q ≥ 0, the neighborhood of the data object P1 can be expressed as Q(P2 - P1), then it is defined as:
[0099] Q(P1, P2) = {P < I} (2)
[0100] Definition 2 Q(P1, P2) outliers. Set a positive integer i. If the r-neighborhood cardinality of the data point q is less than k, then q is a Q(P1, P2) outlier. Based on the calculation of distance-based outliers, relatively accurate data results can be calculated according to the specific discussion of the above formula, which can improve the work efficiency to a certain extent and reduce the process of repeatedly verifying the results. In this paper, real data is used for specific operations to detect whether the new algorithm is more real and effective compared with the traditional algorithm, which can ensure the rationality of the test effect to a certain extent.
[0101] (2) Implement distributed outlier detection:
[0102] If at least p objects in the dataset are farther away from object o than DT, then object o is a distance-based outlier with respect to parameters p and DT, i.e., DB(p, DT)-Outlier (outlier set). The definition here essentially applies to the global outliers of all datasets. If k is the same number of outliers expected by the user, then the deviation will be maximized. If k objects are outliers, the detection strategy is as follows: First, determine k clusters and n data. Then, describe s outliers so that outlierSet = K. The relative outlier set is assigned to the empty set, and the cluster set output by Definition 2 is KCo. When OKCo = KCo, a set of candidate microclusters containing outliers can be stored. Based on the calculated result, which is the information entropy of the cluster, the object with the largest deviation, i.e., Doli, is calculated, or the objects within the microcluster are displayed in descending order of deviation.
[0103] Then, each element is taken out in turn, starting from the first element. The next step is to calculate the information in the remaining data set, that is, the entropy value. The next step is to determine whether the value of the information entropy is within the threshold σ. If the calculated value is less than σ, it means that the result does not contain outliers, so this type of cluster can be excluded. Otherwise, the corresponding outliers can be correspondingly saved in the outlierSet. Finally, the s outliers in the outlierSet are output. Then, the efficient distributed outlier detection algorithm based on big data is used in the clusters that may appear in the outliers, and the outliers are placed in the outlierSet. After analyzing the global and local outliers, based on the real-time feedback of the distributed outlier detection algorithm data, the input and output of the relevant data are adjusted in time in combination with the sampling analysis data, so as to realize the effective operation of the efficient distributed outlier detection algorithm based on big data.
[0104] Big data risk assessment model algorithm based on box plot:
[0105] Five statistics in the data (minimum value, first quartile, median, third quartile and maximum value) are used to observe whether the data has symmetry, the degree of distribution dispersion and other information, especially for the identification of outliers.
[0106] The so-called boxplot, also known as box-and-whisker plot, box-type plot, box-shaped plot or box-shaped plot, is a statistical chart used to show the dispersion of a set of data. The drawing of the boxplot relies on actual data. It does not require the prior assumption that the data obeys a specific distribution form and does not impose any restrictive requirements on the data. It simply truly and intuitively shows the original appearance of the data shape. On the other hand, the boxplot provides a standard for identifying outliers: outliers are defined as values less than Q1-1.5IQR or greater than Q3+1.5IQR, that is, the standard for judging outliers is based on quartiles and interquartile ranges. Quartiles have a certain degree of resistance. As much as 25% of the data can become arbitrarily far without greatly disturbing the quartiles, so outliers cannot affect this standard, and the results of identifying outliers are relatively objective.
[0107] Research on application scenarios of sampling strategy decision support:
[0108] The research objective of the sampling strategy decision support scenario application is to conduct differentiated scenario analysis for sampling control strategy formulation based on historical quality information of suppliers and related material types, as well as the specific properties of the materials themselves. Furthermore, based on the optimized sampling strategy and sampling plan, combined with various application scenarios such as the characteristics of various materials, supply methods, and testing requirements, differentiated sampling analysis items are further refined from aspects such as testing items and testing methods.
[0109] According to actual business, the application scenarios can be classified as follows (partial):
[0110] (1) Confirmation of materials for random inspection (material selection in multi-material scenarios)
[0111] (2) Order supply analysis (setting minimum sampling quantity in the case of massive orders)
[0112] (3) Analysis of suppliers with low qualification rates (sampling strategy adjustment)
[0113] (4) New supplier analysis (confirmation of new supplier's inspection level)
[0114] (5) Order Price Analysis (Deviation Analysis Scenario)
[0115] (6) Order delivery period analysis process (obtaining the latest order delivery value based on historical data)
[0116] (7) Selection of testing institutions (division of testing capabilities, allocation of guidance)
[0117] This study takes the analysis scenario of suppliers with low qualified rates as an example. It requires comprehensive information such as quality problem list data, supplier random inspection qualified rate analysis, random inspection results, inspection item data and inspection item level. Through the big data modeling and analysis of the random inspection strategy decision support system, the inspection level of the corresponding low qualified rate supplier for different material categories is determined, and random inspection plan preparation information is generated to provide guidance for formulating the random inspection ratio of materials for low qualified rate suppliers.
[0118] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
[0119] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method includes only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A method for analyzing detection data and sampling strategy models based on big data technology, characterized in that: The following steps are involved: Constructing a knowledge graph for power material spot inspections: Extracting entities and relationships from multiple dimensions, including material suppliers, equipment information, and operating status, to construct a knowledge graph with an "entity-relationship-entity" triple structure. This knowledge graph is constructed using a top-down approach, including ontology construction, knowledge extraction, fusion, and storage. Data collection and data pool management: Multi-source data from testing equipment is collected through wired Ethernet ports, serial ports, or USB interfaces, and stored in a data pool after cleaning and conversion. The data pool integrates a storage engine, a computing engine, and a data governance module. Establish a material sampling inspection strategy rule base: define inspection items, quantity, and organization sampling inspection strategy elements, and formulate differentiated sampling inspection rules based on the supplier's historical quality data and material characteristics; Sampling inspection strategy decision support system modeling: Based on the knowledge graph and rule base, the sampling inspection strategy is generated using artificial intelligence algorithms, data mining methods and risk assessment models; Scenario application and strategy optimization: Apply sampling strategies in low-qualification supplier analysis and new supplier testing scenarios, and dynamically optimize model parameters based on real-time data.
2. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: In the knowledge graph construction step, structured, semi-structured and unstructured data are preprocessed by ETL tools, uniformly converted into structured data and loaded into the knowledge ontology library.
3. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The data acquisition module supports real-time data collection of transformers and lightning arresters, as well as text / database file parsing of cable protection pipes and line hardware materials.
4. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The differentiated sampling inspection rules include: increasing the sampling inspection volume for suppliers with many historical quality problems, adding inspection items for new suppliers or materials using new technologies, and designing a count-adjusted sampling inspection plan based on the MIL-STD-1916 standard.
5. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The artificial intelligence algorithm includes a binary tree search algorithm, which realizes intelligent matching with the rule base by converting the knowledge graph tree structure into a binary tree.
6. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The data mining method comprises: Decision tree algorithm, used to classify and filter features of material detection data; Canonical correlation coefficient analysis is used to calculate the correlation between test items; The quadratic exponential smoothing method is used to predict the material failure rate.
7. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The risk assessment model is constructed based on the box plot. The minimum, quartile and maximum values of the data are analyzed to divide the data into reasonable areas, warning areas and abnormal areas. The probability density and weight density of the test parameters are introduced to establish the total probability function to achieve risk level division.
8. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The scenario applications include: Analysis scenario for suppliers with low pass rates: Comprehensive quality problem list and sampling pass rate data, dynamically adjusting inspection levels; Testing agency selection scenario: Allocate random inspection tasks based on the capabilities and carrying capacity of the testing center.
9. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The data pool management module uses MySQL and ClickHouse databases and supports data service output.
10. The method for analyzing detection data and sampling strategy models based on big data technology according to claim 1, characterized in that: The outlier detection algorithm adopts efficient distributed computing and identifies abnormal points in the detection data by defining the neighborhood cardinality and distance threshold.