Special equipment quality control digital governance data demand identification method based on rooting theory
Through the combination of grounded theory and machine learning, the cross-modal fusion and dynamic adaptation problems of quality control data requirements identification of special equipment are solved, and high-precision data requirements identification and dynamic governance are achieved, and the digital transformation of special equipment is supported.
Patent Information
- Application Number
- CN202510474947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing technology has problems such as extensive identification granularity, lagging dynamic response, difficulty in fusion of cross-modal data, and insufficient embedding of domain knowledge in the identification of quality control data of special equipment. It is difficult to adapt to the dynamic changes of multi-source heterogeneous data and new technology application scenarios, resulting in insufficient quality and safety risk warning and decision-making capabilities.
Using a three-level coding technology based on grounded theory, combined with the improved PageRank algorithm and TD3 reinforcement learning, a multimodal correlation map is built, and through semantic clustering and domain knowledge fusion, accurate identification and dynamic governance of special equipment quality control data is achieved.
The data demand identification accuracy was improved by 42.6%, and a dynamic matching mechanism between data demand and governance standards was built, the solution generation efficiency was improved by 3.8 times, the data coverage rate was increased to 89%, and the quality deviation prediction accuracy was increased by 35%, supporting the digital implementation of special equipment safety specifications.
Smart Images

Figure CN120387638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital governance of the quality and safety of special equipment, and particularly to a method for identifying data requirements for digital governance of quality control of special equipment based on grounded theory. Background Art
[0002] As an important infrastructure to support the operation of the national economy and ensure social public safety, the quality and safety of special equipment are directly related to the safety of people's lives and property and the high-quality development of industries. However, in the face of multi-source heterogeneous data generated throughout the life cycle such as design, manufacturing, inspection, and operation and maintenance, the traditional governance model has prominent problems such as coarse granularity in data requirement identification and lag in dynamic response, which seriously restricts the improvement of quality and safety risk warning and precision decision-making capabilities.
[0003] Massive heterogeneous quality control data of special equipment (including equipment life cycle inspection reports, maintenance records, fault diagnosis data, IoT sensor time series data, production and manufacturing process parameters, user unit management files, regulatory department inspection documents, etc.) provide a data basis for building a digital quality governance system. These multi-source heterogeneous data contain both structured numerical information and a large amount of unstructured text and image data, and have characteristics such as strong professionalism, complex data dimensions, and close spatio-temporal correlation. Effectively identifying the core data requirements in the digital governance process and establishing an association network between data elements are of great value for breaking through the quality control data islands, improving the accuracy of risk warning, and realizing precise quality governance.
[0004] Currently, traditional methods such as expert experience method, literature analysis method, questionnaire survey method, statistical analysis method, and system modeling method are mainly used to identify the data requirements of special equipment quality control. In specific practices, data requirements are mostly extracted through knowledge-driven methods such as industry standard interpretation, management regulation sorting, expert interview research, and historical data analysis, resulting in problems such as rough granularity in requirement identification, lag in dynamic update, and lack of element association. The research on theory-driven data requirement identification methods needs to be deepened urgently:
[0005] (1) The data requirement identification method based on expert experience and literature analysis mainly constructs a data index system by summarizing industry standard specifications, summarizing management practice experience, and sorting out academic research results. This method highly depends on the accumulation of domain knowledge and is easily limited by the cognitive boundaries of experts and the coverage of literature, resulting in problems such as strong subjectivity, slow update and iteration, and difficulty in adapting to the development of new equipment technologies in data requirement identification. Especially in the face of new technology application scenarios such as intelligent sensors and digital twins, traditional methods are difficult to effectively identify the requirements of emerging data elements.
[0006] (2) The data requirement identification method based on statistical analysis mainly mines key data features through mathematical statistics methods such as correlation analysis and regression analysis of historical operation and maintenance data, fault records, and inspection reports. Although this method has the advantage of quantitative analysis, it has three limitations: First, it requires data to have complete annotations and clear structures, and it is difficult to process multi-modal heterogeneous data; second, the statistical significance index cannot effectively reflect the causal mechanism between data elements; third, it lacks the ability to analyze the dynamic evolution law of the entire data life cycle.
[0007] (3) Existing data requirement identification methods generally have systematic deficiencies, which are mainly manifested as: no mapping relationship between data requirements and quality control business scenarios is established, the hierarchical association structure between data elements is ignored, and the quantitative evaluation of data quality characteristics (integrity, timeliness, consistency, etc.) is lacking. This results in the constructed data index system being difficult to support precise and intelligent quality governance decisions.
[0008] Lack of dynamicity: It is difficult to adapt to the iteration of standards such as TSG 07-2019.
[0009] In the digital governance paradigm, as a systematic qualitative research method, grounded theory can refine core concepts from massive business data through three-level coding technology and construct a theoretical model to provide methodological support for data requirement identification. The traditional application process usually includes five stages: data collection, open coding, axial coding, selective coding, and theoretical saturation testing. However, existing methods face significant challenges in the field of special equipment: First, the full-chain data of equipment has multi-modal characteristics (including structured inspection reports, semi-structured operation and maintenance logs, and unstructured on-site records), and traditional coding methods are difficult to achieve cross-modal data fusion; second, the dynamic evolution of quality control standards leads to the continuous expansion of data requirement dimensions, and the static coding system cannot adapt to the changes in governance scenarios; third, the lack of domain knowledge embedding causes deviations in the extraction of core categories, affecting the engineering applicability of the data requirement map.
[0010] In the practice of existing technologies, there are three major technical bottlenecks in directly applying the standard grounded theory for data requirement identification: First, the cross-stage data correlation analysis lacks systematicness, resulting in insufficient completeness in the identification of key quality parameters; second, no coding framework that is dynamically connected to the safety technical specifications of special equipment is established, causing the mismatch between data requirements and regulatory requirements; third, there is a lack of a data requirement transformation mechanism for digital governance platforms, making it difficult to form a configurable data governance solution. Therefore, it is urgent to construct an improved grounded theory method system to achieve the accurate identification and dynamic governance of quality control data requirements for special equipment through coding mechanism optimization and domain knowledge integration, providing reliable data element support for digital transformation. Summary of the Invention
[0011] The object of the present invention is to propose a method for identifying data requirements for digital governance of special equipment quality control based on grounded theory, so as to solve the problems existing in the above-mentioned prior art, construct an intelligent coding analysis process for multi-modal data fusion, realize in-depth mining of quality control data throughout the life cycle of special equipment, improve the accuracy of data requirement identification and dynamic adaptation ability, and provide accurate data element support for digital governance.
[0012] To achieve the above object, the present invention provides the following solution:
[0013] A method for identifying data requirements for digital governance of special equipment quality control based on grounded theory, comprising:
[0014] Obtain data related to special equipment quality control; wherein, the data related to special equipment quality control includes: design specifications, manufacturing logs, inspection reports and operation and maintenance records;
[0015] Based on the grounded theory three-level coding technology, perform text analysis on the data related to special equipment quality control to generate a list of feature items;
[0016] For the list of feature items, extract keywords and related vocabulary for special equipment quality control requirements to obtain a list of related vocabulary of keywords;
[0017] Classify the list of related vocabulary of keywords, and aggregate keywords and related vocabulary of different categories to form a list of data requirements for special equipment quality control.
[0018] Optionally, for the list of feature items, extracting keywords and related vocabulary for special equipment quality control requirements to obtain a list of related vocabulary of keywords includes:
[0019] Based on the list of feature items, identify the core points of special equipment quality control through an improved PageRank algorithm to obtain a list of keywords for special equipment quality control requirements;
[0020] Based on the list of keywords for special equipment quality control requirements, construct a multi-modal association graph, and apply TD3 reinforcement learning to optimize the association rules to obtain a list of related vocabulary of keywords.
[0021] Optionally, classifying the list of related vocabulary of keywords and aggregating keywords and related vocabulary of different categories includes:
[0022] Based on a preset four-dimensional evaluation matrix, perform semantic clustering on the list of related vocabulary of keywords to form a semantic analysis list of related vocabulary of demand keywords;
[0023] Based on the semantic analysis list of related vocabulary of demand keywords, condense and aggregate keywords and related vocabulary of different categories to form a list of data requirements for special equipment quality control.
[0024] Optionally, based on the three - level coding technology of grounded theory, the text analysis of the special equipment quality control - related data includes:
[0025] Collecting original materials; among them, the original materials include: guiding documents and expert interview materials based on the data requirement elements of the digital governance of special equipment quality control;
[0026] Through open coding, extracting and developing concepts in the original materials to form the initial categories of the data requirements for the digital governance of special equipment quality control;
[0027] Through axial coding, repeatedly comparing and testing concepts, refining and differentiating the initial categories, and extracting the main categories of the data requirements for the digital governance of special equipment quality control;
[0028] Through selective coding, identifying the core category that covers all other categories;
[0029] Further conducting a theoretical saturation test to construct a theoretical model of the performance element structure dimension.
[0030] Optionally, the keyword list of special equipment quality control requirements includes:
[0031] Calculating the term frequency - inverse document frequency weighted value for the feature item list;
[0032] Based on the term frequency - inverse document frequency weighted value, considering the PageRank algorithm with improved in - link quality, and calculating the node centrality score accordingly;
[0033] Steps of the PageRank algorithm with improved in - link quality:
[0034] (1) In - link quality assessment, assessing the quality of each in - link pointing to the target web page, considering the authority of the source web page of the in - link, the text content of the link, and the context of the link;
[0035] (2) Adjusting weight distribution, adjusting its weight in the PageRank calculation according to the quality of the in - link;
[0036] (3) PageRank calculation, performing iterative PageRank calculation according to the adjusted weight distribution;
[0037] (4) Iterative update, repeatedly iteratively updating the PageRank value of each page until convergence;
[0038] (5) Output result: obtaining the PageRank value after considering the in - link quality.
[0039] Optionally, constructing a multi - modal association graph, setting up a TD3 reinforcement learning policy network, and obtaining a list of related keywords including:
[0040] Based on the multi-modal association graph, construct a GraphSAGE graph neural network model and calculate the edge weights; among them, the node features of the GraphSAGE graph neural network model include: data update frequency, risk level, and regulatory body; the D-S evidence theory synthesis formula is introduced for edge weight calculation;
[0041] Set the TD3 reinforcement learning policy network to include:
[0042] (1) Graph neural network feature extraction, use the GraphSAGE graph neural network model to encode the text data and extract the feature representation of each node;
[0043] (2) Reinforcement learning environment construction, use the output features of the graph neural network model as the state of the reinforcement learning environment;
[0044] (3) TD3 algorithm application, use the TD3 algorithm to train the keyword selection strategy in the reinforcement learning environment; among them, the TD3 algorithm includes an Actor network and a Critic network, the Actor network selects keywords according to the feature output of the graph neural network model, and the Critic network evaluates the quality of the selection strategy;
[0045] (4) Optimization and iteration, through the training process of the TD3 algorithm, continuously optimize the keyword selection strategy until convergence;
[0046] (5) Keyword extraction result, use the trained model to extract keywords from the text and obtain a list of keywords-related vocabulary.
[0047] Optionally, the preset four-dimensional evaluation matrix includes: data completeness, risk coefficient, monitoring cost, and regulatory compliance;
[0048] Based on the preset four-dimensional evaluation matrix, the semantic clustering of the keyword-related vocabulary list includes:
[0049] Use natural language processing technology to convert the text data into feature vectors;
[0050] Select a preset clustering algorithm to cluster the text features; complete the semantic clustering of the keyword-related vocabulary list.
[0051] The beneficial effects of the present invention are:
[0052] Overcome the technical bottleneck of traditional grounded theory in multi-modal data processing, and the recognition accuracy is improved by 42.6%;
[0053] Construct a dynamic matching mechanism for data requirements-governance standards, and the efficiency of solution generation is increased by 3.8 times;
[0054] The developed domain knowledge enhanced coding framework enables the accuracy of core category extraction to reach 91.2%;
[0055] The formed configurable governance solution supports 12 types of special equipment safety technical specifications such as TSG 07-2019;
[0056] Compared with the prior art, the data requirement identification coverage rate has increased from 67% to 89%, supporting a 35% increase in the prediction accuracy of quality deviation in the manufacturing process;
[0057] By designing an innovative technical path of "data fusion - intelligent coding - dynamic governance", the present invention provides core data governance capability support for the digital transformation of the quality and safety of special equipment, effectively promoting the digital implementation and implementation of the Special Equipment Safety Supervision Regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1 It is a schematic flow chart of a method for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to an embodiment of the present invention;
[0060] Figure 2 It is a schematic structural diagram of a system for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to an embodiment of the present invention;
[0061] Figure 3 It is a schematic diagram of the grounded theory three-level coding (open coding, axial coding, and selective coding) analysis according to an embodiment of the present invention;
[0062] Figure 4 It is a schematic diagram of the detailed solution based on the grounded theory three-level coding technology according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0064] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0065] As Figure 1 shown, this embodiment proposes a method for identifying data requirements for digital governance of special equipment quality control based on grounded theory. The overall steps of this embodiment include:
[0066] Obtain data related to special equipment quality control; among them, the data related to special equipment quality control includes: design specifications, manufacturing logs, inspection reports, and operation and maintenance records;
[0067] Based on the three-level coding technology of grounded theory, perform text analysis on the data related to special equipment quality control to generate a list of feature items;
[0068] Extract keywords and related vocabulary for special equipment quality control requirements from the list of feature items to obtain a list of related vocabulary for keywords;
[0069] Classify the list of related vocabulary for keywords, and aggregate keywords and related vocabulary in different categories to form a list of data requirements for special equipment quality control.
[0070] Further, extracting keywords and related vocabulary for special equipment quality control requirements from the list of feature items to obtain a list of related vocabulary for keywords includes:
[0071] Based on the list of feature items, identify core nodes through an improved PageRank algorithm to obtain a list of keywords for special equipment quality control requirements;
[0072] Based on the list of keywords for special equipment quality control requirements, construct a multi-modal association graph, apply TD3 reinforcement learning to optimize the association rules, and obtain a list of related vocabulary for keywords.
[0073] Further, classifying the list of related vocabulary for keywords and aggregating keywords and related vocabulary in different categories includes:
[0074] Based on a preset four-dimensional evaluation matrix, perform semantic clustering on the list of related vocabulary for keywords to form a semantic analysis list of related vocabulary for demand keywords;
[0075] Based on the semantic analysis list of related vocabulary for demand keywords, condense and aggregate keywords and related vocabulary in different categories to form a list of data requirements for special equipment quality control.
[0076] Specifically, in the method of this embodiment, first obtain data related to special equipment quality control: collect multi-source heterogeneous data such as design specifications, manufacturing logs, inspection reports, and operation and maintenance records through industrial Internet of Things protocols;
[0077] Secondly, perform text analysis on the special equipment quality control related data to extract key feature items: adopt the three-level coding technology of grounded theory, execute open coding to construct an initial concept set, and generate a list of special equipment quality control data feature items;
[0078] Then, based on the list of special equipment quality control data feature items, identify and extract the keywords for special equipment quality control requirements: identify core nodes by improving the PageRank algorithm to obtain a list of keywords for special equipment quality control requirements;
[0079] Among them, the list of special equipment quality control data feature items refers to the set of feature words of the data source for quality control requirement analysis;
[0080] The list of keywords for special equipment quality control requirements refers to the set of keywords refined from the set of feature words;
[0081] The list of related words for keywords refers to the set of extended words related to the keywords;
[0082] The list of semantic analysis of related words refers to the relevant semantic analysis results of the set of extended words of the keywords;
[0083] The list of special equipment quality control data requirements refers to the list of quality control data requirements refined based on the relevant semantic analysis results;
[0084] Then, based on the list of keywords for special equipment quality control requirements, extract the related words of the keywords for special equipment quality control requirements: construct a multi-modal association graph, apply TD3 reinforcement learning to optimize the association rules, and obtain a list of related words for keywords;
[0085] Then, classify the list of related words of the keywords for special equipment quality control requirements: design a four-dimensional evaluation matrix for semantic clustering to form a list of semantic analysis of related words of the keywords for special equipment quality control requirements;
[0086] Finally, based on the semantic analysis list, condense and aggregate keywords and related words of different categories: dynamically adjust the weights in combination with the digital twin verification results to form a list of special equipment quality control data requirements.
[0087] This embodiment realizes the accurate identification and digital management of special equipment quality control data requirements.
[0088] In this embodiment, obtaining the special equipment quality control related data includes:
[0089] The requirements are to obtain the structured process parameters of the MES system in real time through the OPCUA protocol;
[0090] Collect unstructured weld image data by deploying multi-spectral industrial cameras;
[0091] The feature extraction is implemented by constructing edge computing nodes to extract the FFT features of vibration signals, and the sampling frequency ≥ 20 kHz.
[0092] Furthermore, based on the grounded theory three-level coding technique, the text analysis of the special equipment quality control related data includes:
[0093] Collecting original materials; among them, the original materials include: guiding documents and expert interview materials based on the data requirement elements of the digital governance of special equipment quality control;
[0094] Through open coding, extracting and developing concepts in the original materials to form the initial categories of the data requirements for the digital governance of special equipment quality control;
[0095] Through axial coding, repeatedly comparing and testing concepts, refining and differentiating the initial categories, and extracting the main categories of the data requirements for the digital governance of special equipment quality control;
[0096] Through selective coding, identifying the core category that covers all other genera;
[0097] Further conducting theoretical saturation testing to construct a theoretical model of the performance element structure dimension.
[0098] Specifically, as Figure 4 shown, the grounded theory three-level coding technique proposed in this embodiment includes:
[0099] (1) Based on the policy texts and interview materials of the data requirement elements of the digital governance of special equipment quality control, following the research paradigm and research path of the grounded theory, through open coding, extracting and developing concepts in the original materials to form the initial categories of the data requirements for the digital governance of special equipment quality control;
[0100] (2) Through axial coding, repeatedly comparing and testing concepts, refining and differentiating the initial categories, and extracting the main categories of the data requirements for the digital governance of special equipment quality control;
[0101] (3) Through selective coding, identifying the core category that covers all other genera.
[0102] (4) Further conducting theoretical saturation testing to construct a theoretical model of the performance element structure dimension.
[0103] Furthermore, the list of keywords for the special equipment quality control requirements includes:
[0104] For the list of feature items, calculating the term frequency-inverse document frequency weighted value of the feature items;
[0105] Based on the term frequency-inverse document frequency weighted value, considering the in-link quality improved PageRank algorithm, and calculating the node centrality score accordingly;
[0106] Steps for improving the PageRank algorithm by considering the quality of incoming links:
[0107] (1) Evaluation of incoming link quality: Evaluate the quality of each incoming link pointing to the target web page, mainly considering the authority of the source web page of the incoming link, the text content of the link, the context of the link, etc.
[0108] (2) Adjustment of weight distribution: Adjust the weight of the incoming link in the PageRank calculation according to its quality. Higher-quality incoming links can be assigned higher weights, while lower-quality incoming links have lower weights.
[0109] (3) PageRank calculation: Perform iterative PageRank calculation according to the adjusted weight distribution.
[0110] (4) Iterative update: Repeatedly iterate and update the PageRank value of each page until convergence.
[0111] (5) Output result: Obtain the PageRank value after considering the quality of incoming links.
[0112] Furthermore, construct a multi-modal association graph, set up a TD3 reinforcement learning policy network, and obtain a list of keywords and related words including:
[0113] Based on the multi-modal association graph, construct a GraphSAGE graph neural network model and calculate the edge weights; among them, the node features of the GraphSAGE graph neural network model include: data update frequency, risk level, and regulatory body; the D-S evidence theory synthesis formula is introduced for edge weight calculation;
[0114] Setting up the TD3 reinforcement learning policy network includes:
[0115] (1) Feature extraction of graph neural network: Use the GraphSAGE graph neural network model to encode text data and extract the feature representation of each node (word or sentence).
[0116] (2) Construction of reinforcement learning environment: Take the output features of the graph neural network as the state of the reinforcement learning environment.
[0117] (3) Application of TD3 algorithm: Use the TD3 algorithm to train the keyword selection strategy in the reinforcement learning environment. The Actor network selects keywords according to the features output by the graph neural network, and the Critic network evaluates the quality of the selection strategy.
[0118] (4) Optimization and iteration: Continuously optimize the keyword selection strategy through the training process of the TD3 algorithm until convergence.
[0119] (5) Keyword extraction result: Use the trained model to extract keywords from the text.
[0120] Furthermore, the preset four-dimensional evaluation matrix includes: data completeness, risk coefficient, monitoring cost, and regulatory compliance.
[0121] Based on the preset four-dimensional evaluation matrix, semantic clustering of the keyword-related vocabulary list includes the following steps:
[0122] (1) Construct a four-dimensional evaluation matrix.
[0123] 1) Determine the dimensions: The four-dimensional evaluation matrix includes four key dimensions, such as clustering tightness, separability, semantic relevance, and external consistency. These dimensions can be adjusted according to specific tasks.
[0124] 2) Data preprocessing: Preprocess the text data, including word segmentation, stop word removal, part-of-speech tagging, etc., to extract semantic features.
[0125] (2) Semantic clustering.
[0126] 1) Feature extraction: Use natural language processing techniques (such as TF-IDF, Word2Vec, etc.) to convert the text data into feature vectors.
[0127] 2) Clustering algorithm selection: Select a suitable clustering algorithm (such as K-Means, DBSCAN, etc.) to cluster the text features.
[0128] 3) Clustering execution: According to the selected clustering algorithm, divide the text data into several clusters.
[0129] (3) Application of the four-dimensional evaluation matrix.
[0130] 1) Tightness evaluation: Calculate the similarity of samples within each cluster, for example, use the Silhouette Coefficient to measure the tightness of the cluster.
[0131] 2) Separability evaluation: Measure the separation degree between different clusters, for example, calculate the distance between clusters or use the Adjusted Rand Index (ARI).
[0132] 3) Semantic relevance evaluation: Evaluate the semantic consistency of the text within the cluster through semantic analysis tools (such as WordNet, BERT, etc.).
[0133] 4) External consistency evaluation: If there are external labels, external evaluation metrics (such as ARI) can be used to evaluate the consistency between the clustering results and the true labels.
[0134] (4) Optimize the clustering results.
[0135] 1) Adjust clustering parameters: According to the results of the four-dimensional evaluation matrix, adjust the parameters of the clustering algorithm (such as the number of clusters, distance metric, etc.) to optimize the clustering effect.
[0136] 2) Iterative optimization: Repeat the semantic clustering and evaluation steps until the indicators of the four-dimensional evaluation matrix reach a satisfactory level.
[0137] (5) Output the final result.
[0138] 1) Determine the final clusters: Based on the optimized clustering results, determine the final semantic clusters.
[0139] 2) Result analysis: Analyze each cluster, and extract the key semantic information within the cluster, such as through keyword extraction or topic modeling.
[0140] Specifically, this embodiment proposes a data requirement identification method based on the improved grounded theory: Through three-level coding technology, theoretical sampling and continuous comparison are carried out on multi-source heterogeneous data to construct a trinity requirement identification framework of "data elements - business scenarios - quality indicators". The specific implementation process includes:
[0141] Data acquisition layer: Establish a multi-source data acquisition system covering the entire life cycle of "design - manufacturing - installation - use - inspection - scrapping";
[0142] Theoretical construction layer: Use open coding to extract the initial concepts of data elements, establish element associations through axial coding, and form core categories using selective coding;
[0143] Requirement verification layer: Combine digital twin technology to construct a virtual verification environment to achieve dynamic optimization of the data requirement model;
[0144] System implementation layer: Develop an intelligent data requirement identification platform with independent intellectual property rights, integrating technical modules such as natural language processing, knowledge graph, and machine learning.
[0145] The innovation of the present invention lies in:
[0146] (1) Propose a hybrid research method that combines grounded theory and machine learning, breaking through the efficiency bottleneck of traditional qualitative research;
[0147] (2) Construct a hierarchical model of data requirements, revealing the transformation path of "raw data - characteristic indicators - decision-making knowledge";
[0148] (3) Develop an adaptive data requirement identification system to achieve dynamic evolution and scenario adaptation of the requirement model;
[0149] Through the implementation of this solution, the completeness, accuracy, and timeliness of the quality control data system for special equipment can be effectively improved, providing core data support for building an intelligent quality governance platform.
[0150] For data collection in the design phase based on the full - life - cycle data collection system, key parameters (wall thickness, weld position, stress concentration factor) in the 3D model are automatically extracted through the AutoCAD API interface;
[0151] Parse the BOM table to generate a structured bill of materials, and associate with the process parameter library in the PLM system (such as heat treatment temperature threshold, welding process qualification data);
[0152] For data collection in the manufacturing phase, device status data (spindle vibration value of CNC machine tools, current fluctuation of welding robots) in the MES system is obtained in real - time through the OPC UA protocol;
[0153] Associate with ERP work order data to establish the mapping relationship between production batch numbers and quality traceability codes;
[0154] For data collection in the installation phase, a multi - parameter environmental recorder is deployed (accuracy: temperature ±0.5°C, humidity ±3%RH), and the installation site environmental data is recorded every 5 minutes;
[0155] Read the debugging parameters in the PLC (such as the popping pressure of safety valves, response time of interlock devices) through the Modbus TCP protocol;
[0156] For data collection in the usage phase, three - axis vibration sensors (range ±50g, sampling rate 1kHz) are deployed at key parts of the device, and data is uploaded in real - time through the NB - IoT module;
[0157] Build an edge computing node to perform FFT transformation on the original signal to extract characteristic frequencies (such as the bearing fault characteristic frequency band 2 - 5kHz);
[0158] For data collection in the inspection phase, develop an intelligent parsing engine for inspection reports to automatically extract key fields from PDF / scanned documents:
[0159] Structured fields: next inspection date, allowable operating pressure;
[0160] Semi - structured fields: defect description (such as "there are 3 surface cracks in the head transition zone");
[0161] The full - life - cycle data collection system constructs an intelligent data network throughout the entire product process by deeply integrating industrial Internet of Things and digital twin technologies.
[0162] In the design stage, the intelligent parsing engine developed based on the AutoCAD Mechanical API not only automatically extracts the geometric parameters of the 3D model, but also realizes the dynamic mapping of the stress nephogram and the solid model by integrating the finite element analysis module. At the same time, a material - process knowledge graph is constructed to intelligently match the extracted parameters such as wall thickness and weld seams with the heat treatment curves and welding process libraries in the PLM system to ensure design compliance.
[0163] In the manufacturing process, relying on the real - time data channel constructed based on the OPC UA protocol, millisecond - level acquisition of the machine tool vibration spectrum and welding current waveform is achieved. And through blockchain technology, the ERP work order information and quality traceability codes are linked in an immutable chain - like manner to form a digitally traceable digital mainline.
[0164] In the installation stage, the deployed environmental monitoring network uses self - organizing sensor nodes and combines the PLC debugging parameters obtained through the Modbus TCP protocol to construct an installation quality assessment model that meets the ASME standard. When the safety valve test data deviates from the preset threshold, the correction mechanism is automatically triggered.
[0165] During the operation of the equipment, the three - axis vibration sensor group conducts joint time - frequency domain analysis through edge computing nodes, uses wavelet packet decomposition technology to extract the bearing fault characteristic frequency band, and combines the low - power wide - area transmission characteristics of the NB - IoT network to achieve cloud - based collaboration for predictive maintenance decision - making.
[0166] In the inspection stage, through the intelligent parsing engine with multi - modal fusion, OCR and natural language processing technologies are used to semantically deconstruct the unstructured defect descriptions, automatically associate historical inspection data to generate a 3D defect evolution map, and synchronize and update the extracted structured information such as inspection date and pressure threshold to the asset management system to form a closed - loop quality control.
[0167] Data acquisition technology is realized through structured data acquisition and semi - structured data acquisition, and integrity checks are carried out by constructing an Apache NiFi data processing pipeline, an XML parser, and a JSON converter.
[0168] Text data processing based on standardized data format processing is achieved through encoding conversion: batch - converting GB2312 / GBK - encoded files in historical data to UTF - 8, noise filtering: constructing a regular expression library to remove irrelevant characters, metadata injection: adding the device unique identification code and data acquisition timestamp at the head of the text.
[0169] Time - series data processing uses resampling algorithms. For high - frequency data (1kHz), median filtering is used to downsample to 10Hz, and for low - frequency data (5 - minute interval), cubic spline interpolation is used for filling. Outliers are screened and processed.
[0170] The image data processing adopts a standardized process for size adjustment, color normalization, quality inspection, size normalization, feature enhancement, and noise filtering.
[0171] The implementation method of the theory construction layer is realized through a three-level coding technology system:
[0172] Multimodal data adaptation: Unify text, time series, and image data into the coding analysis framework;
[0173] Dynamic weight mechanism: Adjust the coding weight according to the data confidence (such as the sensor accuracy level);
[0174] Domain knowledge integration: Embed special equipment safety technical specifications (such as TSG 07-2019) as coding constraint conditions.
[0175] The open coding technology solution first performs data preprocessing;
[0176] Text data: Use the BERT model to extract semantic vectors and combine the TF-IDF algorithm to screen high-frequency quality control terms;
[0177] Time series data: Use the time domain feature extraction tool (TSFRESH) to generate statistical features (mean, variance, zero-crossing rate) and frequency domain features (FFT main frequency amplitude);
[0178] Image data: Extract deep features (2048-dimensional feature vectors) through the ResNet-50 pre-trained model.
[0179] Generate concepts for the preprocessed data, construct a concept clustering matrix, and set the cosine similarity threshold for concept merging.
[0180] The data governance system constructs an intelligent processing paradigm for multimodal data fusion, realizes dynamic routing of heterogeneous data through the distributed data pipeline built by Apache NiFi, where the XML parser uses XSD Schema validation to achieve the integrity verification of structured data, and the JSON converter standardizes the semi-structured data format through the JOLT specification.
[0181] Implement a multi-level cleaning strategy for historical text data: Eliminate the character set gap through encoding conversion based on the iconv library, construct a noise pattern library containing more than 500 regular expressions to filter non-standard characters, and form a data lineage traceability chain through metadata injection technology.
[0182] The time series data processing introduces an adaptive signal conditioning mechanism, uses sliding window median filtering for high-frequency vibration data to retain the characteristic peak and valley values, performs interpolation compensation based on time series decomposition for low-frequency environmental data, and combines the Grubbs criterion and the isolation forest algorithm to achieve multi-dimensional anomaly detection.
[0183] The image processing pipeline integrates OpenCV and deep learning frameworks, realizes illumination normalization through gamma correction, adopts bilateral filtering to retain edge features for noise suppression, and constructs a quality assessment model based on SSIM to automatically eliminate blurred images.
[0184] The three-level coding technology system innovatively integrates the domain knowledge graph, transforms the TSG specification clauses into coding constraint rules, and realizes the credibility fusion of multi-source data through the confidence weighted algorithm.
[0185] In the open coding stage, a multi-modal feature joint learning method is adopted. In the text dimension, quality defect entities are extracted through the BERT-CRF model. For time series feature engineering, a time-frequency domain hybrid feature pool is constructed. For image feature extraction, a transfer learning strategy is used to fine-tune the fully connected layer of ResNet. Finally, a concept hypergraph is constructed through spectral clustering algorithm, and a dynamic similarity threshold is set to achieve cross-modal concept alignment, forming an interpretable quality knowledge ontology library.
[0186] The main axis coding technical solution is constructed through an association network:
[0187] GraphX is used to construct a multi-modal association graph, with the nodes being the open coding concepts. The calculation method of the edge weights is as follows:
[0188] W_ij = α * Pearson(X_i,X_j) + β * Jaccard(Text_i,Text_j)
[0189] (α = 0.6, β = 0.4, determined through experimental optimization)
[0190] Key path discovery: The PageRank algorithm is used to identify core nodes.
[0191] Association rule mining: The FP-Growth algorithm is applied to discover strong association rules:
[0192]
[0193] The selective coding technical solution establishes a four-dimensional evaluation matrix through core category screening; as shown in Table 1.
[0194] Table 1: Four-dimensional evaluation matrix
[0195] Dimension Weight Evaluation Index Data Support Degree 30% Number of Data Sources Involved / Total Number of Stages Risk Severity 25% Proportion of This Factor in Historical Accidents Monitoring Feasibility 20% Sensor Coverage + Detection Cost Regulatory Compliance 25% Number of Technical Specification Violation Clauses
[0196] The implementation of open coding is through the input of raw data and concept extraction, and finally realized by the main axis coding technology.
[0197] The selective coding decision model adopts a dynamic weight adjustment mechanism:
[0198] Define the weight update formula:
[0199] W_t = W_{t - 1}+η*(Current cycle fault correlation degree)
[0200] (η = 0.05 is the learning rate, updated quarterly)
[0201] The implementation of digital twin technology first constructs a digital twin model.
[0202] The technical solution uses the Unity3D engine to construct a three - dimensional visualization model to achieve the following functions:
[0203] Material parameters: 12 mechanical properties such as elastic modulus (E = 210GPa), Poisson's ratio (ν = 0.3), etc.;
[0204] Kinematics parameters: elevator speed curve (S - shaped acceleration and deceleration, maximum 2.5m / s 2 ); as shown in Table 2;
[0205] Table 2: Achieving multi - scale simulation
[0206]
[0207] Physical property mapping: Inject the actual device material parameters (elastic modulus, Poisson's ratio) into the virtual model;
[0208] Real - time data synchronization: Update sensor data every 200ms through the OPC UA protocol (error < 0.1%);
[0209] The multi - scale simulation system constructs a cross - dimensional collaborative verification platform. At the macroscopic level, the overall operation of the device is restored through multi - body dynamics simulation. At the microscopic level, ANSYS Workbench is used for multi - physical - field coupling analysis. The two - way transfer of the lattice - scale stress nephogram and macroscopic strain data is realized through APDL scripts, and a material damage evolution model is constructed.
[0210] The virtual verification environment constructs a hierarchical scenario library based on the MBSE methodology. Among them, the extreme condition simulation introduces a multi - field coupling loading mechanism: the high - temperature test uses a non - steady - state thermodynamics model to simulate the thermal expansion effect of materials, the alternating load test evaluates the cumulative plastic deformation through the Chaboche model, and the fault injection module integrates the PHM fault tree to construct the propagation paths of 20 failure modes.
[0211] The data acquisition optimization framework innovatively introduces deep reinforcement learning into the digital twin closed - loop. The state space construction uses t - SNE dimensionality reduction technology to extract the essential manifold of 300 - dimensional feature vectors. The action space design includes an adaptive sampling strategy engine. The exploration - exploitation balance is achieved through the Actor - Critic architecture of the Twin - Delayed Deep Deterministic Policy Gradient algorithm (TD3). Among them, the Critic network uses a prioritized experience replay mechanism to accelerate convergence.
[0212] The reward function integrates the information entropy theory. On the premise of ensuring a 98% fault detection coverage rate, an 80% data compression rate is achieved through feature importance ranking, and the optimal value of the learning rate parameter is determined by Bayesian optimization.
[0213] Build a multi-scale simulation model:
[0214] Macro scale: The operating state of the whole machine (such as the elevator running speed curve);
[0215] Micro scale: The stress distribution of key components (based on ANSYS finite element analysis);
[0216] The virtual verification environment design first constructs a scenario library and simulates extreme working conditions;
[0217] High temperature and high pressure test: The temperature gradient rises from 20°C to 300°C (rate 5°C / min);
[0218] Alternating load test: The pressure cycle range is 0 - 35 MPa (frequency 0.5 Hz);
[0219] Fault injection test: Preset 20 types of typical fault modes (seal ring aging, weld cracking);
[0220] The dynamic optimization of the data requirement model is achieved by establishing a reinforcement learning framework;
[0221] State space: A 300-dimensional feature vector output by the digital twin;
[0222] Action space: Adjust the data acquisition frequency (1 Hz - 10 Hz), feature selection weight;
[0223] Reward function: R = 0.7A + 0.3B (A: Fault detection rate, B: Data flow saving rate);
[0224] The TD3 algorithm is used for policy optimization, and the learning rate is set to 3e-4; The hyperparameters are set as shown in Table 3.
[0225] Table 3: Hyperparameter settings
[0226]
[0227] The system implementation layer designs a refined microservice architecture through the intelligent recognition platform architecture; The technical solution services are shown in Table 4.
[0228] Table 4: Technical solution services
[0229] Service Name Function Description Technology Stack DataCollector Multi-source Data Collection ApacheCamel FeatureEngine Feature Extraction and Encoding Python + PyTorch DigitalTwinCore Digital Twin Simulation Unity3D + C#
[0230] High-availability guarantee is achieved by Kubernetes to realize automatic scaling (scaling out is triggered when the CPU utilization rate > 70%), and the Ceph distributed storage is used for the data persistence layer.
[0231] The interaction of the core modules first goes through the data flow design:
[0232] graph LRA [IoT device] --> |MQTT protocol| B (Data acquisition module)
[0233] B --> |Kafka| C {Feature extraction module}
[0234] C --> |gRPC| D [Coding engine]
[0235] D --> |WebSocket| E [Digital twin verification]
[0236] E --> |Feedback data|
[0237] F [Dynamic optimization module]
[0238] The multi-source heterogeneous data acquisition realizes the unified access of multi-modal data through database docking; as shown in Table 5.
[0239] Table 5: Data access
[0240] Data Type Access Method Performance Index Relational Database JDBC Connection Pool (Maximum 100 Concurrencies) Throughput ≥ 5000 Records per Second Industrial Real-time Database OPC DA / UA Protocol Reading latency < 50ms Document Data Elasticsearch Full-text Search Query response < 200ms
[0241] The access of IoT devices adopts the development of an edge computing gateway:
[0242] Supported protocols: Modbus TCP / RTU, MQTT, CoAP.
[0243] Data preprocessing: Outlier filtering (3σ principle) and data compression (zlib algorithm) are completed at the gateway side.
[0244] For the text semantic analysis of intelligent data feature extraction, the RoBERTa-wwm model is used for domain adaptation training.
[0245] The extraction of time series features constructs a multi-scale feature pool; as shown in Table 6.
[0246] Table 6: Feature pool
[0247] Feature Type Extraction Method Dimension Time-domain Feature Mean, Variance, Kurtosis 15 Dimensions Frequency-domain Feature FFT Main Frequency Amplitude, Wavelet Packet Energy 20 Dimensions Nonlinear Feature Approximate Entropy, Lyapunov Exponent 8 Dimensions
[0248] Semantic encoding based on BERT-wwm: Generate 384-dimensional semantic vectors for professional terms such as "bearing wear", and calculate the cosine similarity to discover potential associations (such as the similarity with "lubrication failure" reaches 0.87).
[0249] Causal Inference Rule Base: Pre - define more than 50 industrial fault logic rules (such as "Peak value of vibration spectrum at 1 kHz → Probability of bearing raceway damage + 35%"), supporting the calculation of probabilistic association strength.
[0250] As Figure 3 shown, the three - level coding automated pipeline is divided into;
[0251] Open Coding Layer: Use the BiLSTM - CRF model for named entity recognition to accurately extract entities such as equipment components (96.2% F1) and fault phenomena (92.7% F1).
[0252] Spindle Coding Layer: Apply the FP - Growth algorithm to mine frequent item sets and discover strong association rules such as {"Temperature > 80°C", "Oil viscosity < 46 cSt"} → "Gearbox overheating" (support > 0.3, confidence > 0.85).
[0253] Selective Coding Layer: Calculate the importance of concept nodes based on the PageRank algorithm, construct a fault propagation network graph, and the node size reflects the weight of the fault influence range.
[0254] System implementation feature version difference visualization: Use the ForceAtlas2 algorithm to dynamically layout the version comparison graph, highlight new nodes / edges in red, mark deleted items in blue, and support cross - version impact analysis. Interpretability report generation: Automatically annotate the confidence, support, and case evidence of association rules (such as when associating "Axial displacement exceeds the standard" with "Bearing wear", attach the time - frequency diagram of vibration signals and maintenance records).
[0255] The graphical interface function uses concept drag - and - drop association: Supports manually establishing an association between "Bearing wear" and "Vibration anomaly", and automatically generates a coding report: Outputs the three - level coding results in PDF format (including the association network graph).
[0256] The online learning mechanism of the demand dynamic optimization module processes data streams: Uses Apache Flink to implement real - time feature calculation for model update strategies:
[0257] The digital twin verification module conducts differential analysis through virtual - real interaction design;
[0258] Virtual control panel: Can adjust the working parameters of the digital twin.
[0259] The demand dynamic optimization module constructs a stream-batch integrated processing architecture, and realizes a real-time feature pipeline with millisecond-level latency through the Apache Flink engine: a sliding time window (Window Size = 5s, Slide = 1s) is adopted to calculate the statistical quantities of working condition features, and the CEP complex event processing technology is combined to identify abnormal patterns; the online learning mechanism designs a double-buffer queue, and asynchronously pushes the 300-dimensional feature vectors calculated in real time to the TD3 agent through the Kafka message queue to realize the incremental update of the policy network parameters.
[0260] The digital twin verification module constructs a two-way data mirroring channel, realizes data synchronization between physical entities and virtual models through the OPC UA protocol, and the difference analysis engine uses the dynamic time warping (DTW) algorithm to calculate the similarity of vibration signal waveforms, and triggers the model self-correction process when the error threshold exceeds 5%.
[0261] The virtual control panel integrates a parameter perturbation generator, supports boundary exploration tests for pressure values within the rated range of ±10% with a step size of 0.5%, generates multi-dimensional parameter combinations by combining Sobol sequence sampling, and constructs a pressure-stress response surface model through sensitivity analysis to form an intelligent verification closed loop with self-evolution ability.
[0262] Theoretical credibility evaluation:
[0263] Use Cohen's Kappa coefficient (κ≥0.75) to verify the inter-coder reliability;
[0264] Adopt the member checking method to conduct follow-up visits to 12 domain experts;
[0265] Calculate the theoretical generalization index:
[0266] This embodiment also proposes a special equipment quality control digital governance data demand identification system based on grounded theory, and the system includes the following core modules:
[0267] Multi-source heterogeneous data acquisition module: supports structured database docking, unstructured document parsing, and Internet of Things device access;
[0268] Data feature intelligent extraction module: applies deep learning algorithms to realize text semantic analysis, time series feature extraction, and image pattern recognition;
[0269] Grounded theory coding engine: provides a visual coding workbench, supports concept generation, category classification, and relationship modeling;
[0270] Demand dynamic optimization module: continuously updates the data demand model based on the online learning mechanism;
[0271] Digital Twin Verification Module: Build a virtual quality control scenario for requirement verification.
[0272] The system of this embodiment can be specifically set as follows:
[0273] Data Sensing Subsystem, including a multi-protocol adapter, supporting Modbus TCP / OPC UA / MQTT access;
[0274] Edge Intelligent Gateway, integrating an INT8 quantization inference engine;
[0275] Coding Analysis Subsystem, including a multi-modal analysis platform, deploying an improved PageRank algorithm and GraphSAGE model;
[0276] Dynamic Optimizer, implementing the update of the policy network based on the TD3 algorithm;
[0277] Virtual Verification Subsystem, including a physical field simulation module, supporting fault mode injection testing;
[0278] Self-Calibration Module, automatically updating model parameters when the data drift > 5%;
[0279] Solution Generation Subsystem, including: Standard Mapper, with a knowledge graph of TSG specification clauses built-in;
[0280] Deployment Optimizer, outputting a Kubernetes cluster configuration list.
[0281] The Coding Analysis Subsystem integrates a three-stage processing pipeline: the first stage performs feature selection based on the improved TF-IDF, the second stage runs the Gephi visualization tool to generate a concept association graph, and the third stage applies the fuzzy Delphi method for expert verification;
[0282] The Dynamic Optimizer has a built-in conflict resolution module, which automatically switches to the Murphy averaging method for evidence fusion when a paradox occurs in the Dempster combination rule.
[0283] In this embodiment, the Edge Intelligent Gateway integrates a lightweight STFT analysis module, with a window shift interval ≤ 5ms;
[0284] The Virtual Verification Subsystem supports WebGL 3D visualization, realizing dynamic rendering of defect cloud maps;
[0285] The Deployment Optimizer defines an automatic scaling policy: add 2 nodes when CPU > 75%, and trigger data sharding when memory > 80%.
[0286] The WebGL 3D visualization module supports LOD hierarchical loading: display the device distribution heat map at a viewing distance of 50m, render the stress cloud map at a viewing distance of 10m, and display the weld microstructure at a viewing distance of 1m;
[0287] The data sharding strategy follows the "Administrative Rules for Informatization Work of Special Equipment" TSG Z0002-2009, and independent data shards are established according to equipment categories (boilers / pressure vessels / elevators).
[0288] This embodiment Figure 1 demonstrates the entire process from raw data collection to the generation of a data requirement list, covering core steps such as data corpus construction, three-level coding analysis, theoretical saturation testing, and data requirement refinement.
[0289] Figure 2 describes the modular architecture of the system, including a data collection and preprocessing module, a grounded coding analysis module, a theoretical model construction module, and a structured output module for data requirements, reflecting the technical implementation path of the method.
[0290] Figure 3 Intuitively presents the three-level coding process of grounded theory through examples, including extracting concepts from original statements, associating categories to form main categories, and finally refining core categories and identifying the key logical chains of data requirements.
[0291] This embodiment establishes the first computable grounded theory framework in the field of special equipment.
[0292] Develop a quantitative evaluation index system for the coding process (including 9 first-level indicators and 27 second-level indicators).
[0293] Achieve the paradigm integration of qualitative research and digital technology (the efficiency of theory construction increases by 68%).
[0294] Verified by the National Special Equipment Safety Technology Committee, this solution reduces the false alarm rate of data requirements to 2.3%, and the theoretical model maintains a fitness of over 89% in the promotion and application in 12 enterprises.
[0295] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for identifying data requirements in the digital governance of quality control of special equipment based on grounded theory, characterized in that, Including: Obtain special equipment quality control related data; among them, the special equipment quality control related data includes: design specifications, manufacturing logs, inspection reports, and operation and maintenance records; Based on the grounded theory three-level coding technology, conduct text analysis on the special equipment quality control related data to generate a list of feature items; For the list of feature items, extract keywords and related vocabulary of special equipment quality control requirements to obtain a list of related vocabulary of keywords; Classify the list of related vocabulary of keywords, and aggregate keywords and related vocabulary of different categories to form a list of data requirements for special equipment quality control.
2. The method for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to claim 1, wherein For the list of feature items, extract keywords and related vocabulary of special equipment quality control requirements to obtain a list of related vocabulary of keywords, including: Based on the list of feature items, identify the core points of special equipment quality control through an improved PageRank algorithm to obtain a list of keywords for special equipment quality control requirements; Based on the list of keywords for special equipment quality control requirements, construct a multi-modal association graph, apply TD3 reinforcement learning to optimize the association rules, and obtain a list of related vocabulary of keywords.
3. The method for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to claim 1, characterized in that, Classify the list of related vocabulary of keywords, and aggregate keywords and related vocabulary of different categories, including: Based on a preset four-dimensional evaluation matrix, conduct semantic clustering on the list of related vocabulary of keywords to form a semantic analysis list of related vocabulary of demand keywords; Based on the semantic analysis list of related vocabulary of demand keywords, condense and aggregate keywords and related vocabulary of different categories to form a list of data requirements for special equipment quality control.
4. The method for identifying data requirements for digital governance of quality control of special equipment based on grounded theory according to claim 1, wherein Based on the grounded theory three-level coding technology, conducting text analysis on the special equipment quality control related data includes: Collect original materials; among them, the original materials include: guiding documents and expert interview materials based on the data requirement elements of special equipment quality control digital governance; Through open coding, extract and develop concepts in the original materials to form the initial categories of data requirements for special equipment quality control digital governance; Through axial coding, repeatedly compare and test concepts, refine and distinguish the initial categories, and extract the main categories of data requirements for special equipment quality control digital governance; Through selective coding, identify the core category that covers all other genera; Further conduct a theoretical saturation test to construct a theoretical model of the performance element structure dimension.
5. The method for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to claim 2, wherein The list of keywords for special equipment quality control requirements includes: For the list of feature items, calculate the weighted value of term frequency-inverse document frequency of the feature items; Based on the weighted value of term frequency-inverse document frequency, consider improving the PageRank algorithm by taking into account the in-link quality, and calculate the node centrality score accordingly; Steps for improving the PageRank algorithm by taking into account the in-link quality: (1) In-link quality assessment, conduct quality assessment on each in-link pointing to the target web page, considering the authority of the source web page of the in-link, the text content of the link, and the context of the link; (2) Adjust weight distribution, adjust its weight in the PageRank calculation according to the quality of the in-link; (3) PageRank calculation, perform iterative calculation of PageRank according to the adjusted weight distribution; (4) Iterative update, repeatedly iterate and update the PageRank value of each page until convergence; (5) Output result: Obtain the PageRank value after considering the in-link quality.
6. The method for identifying data requirements for digital governance of quality control of special equipment based on grounded theory according to claim 2, wherein Construct a multi-modal association graph, set up a TD3 reinforcement learning policy network, and obtain a list of keywords-related vocabulary including: Based on the multi-modal association graph, construct a GraphSAGE graph neural network model and calculate the edge weights; among them, the node features of the GraphSAGE graph neural network model include: data update frequency, risk level, and regulatory body; the D-S evidence theory synthesis formula is introduced in the calculation of edge weights. Set up the TD3 reinforcement learning policy network including: (1) Feature extraction of the graph neural network, use the GraphSAGE graph neural network model to encode the text data and extract the feature representation of each node; (2) Construction of the reinforcement learning environment, use the output features of the graph neural network model as the state of the reinforcement learning environment; (3) Application of the TD3 algorithm, use the TD3 algorithm to train the keyword selection strategy in the reinforcement learning environment; among them, the TD3 algorithm includes an Actor network and a Critic network. The Actor network selects keywords according to the feature output of the graph neural network model, and the Critic network evaluates the quality of the selection strategy; (4) Optimization and iteration, through the training process of the TD3 algorithm, continuously optimize the keyword selection strategy until convergence; (5) Keyword extraction result, use the trained model to extract keywords from the text and obtain a list of keywords-related vocabulary.
7. The method for identifying data requirements for digital governance of special equipment quality control based on grounded theory according to claim 3, wherein The preset four-dimensional evaluation matrix includes: data completeness, risk coefficient, monitoring cost, and regulatory compliance; Based on the preset four-dimensional evaluation matrix, perform semantic clustering on the list of keywords-related vocabulary including: Use natural language processing technology to convert the text data into feature vectors; Select a preset clustering algorithm to cluster the text features; complete the semantic clustering of the list of keywords-related vocabulary.
Citation Information
Patent Citations
Text analysis-based hidden danger identification method for pressure-bearing special equipment and terminal
CN115186778A
Data demand value analysis method and related equipment
CN117196149A
Clinical examination result auditing method and system based on artificial intelligence and big data
CN118629571A
Cited By
Space-time data interpolation method based on space-time decoupling correction flow framework
CN121233916A
Shadow API governance method based on self-supervised comparison and map reasoning
CN122263125A