Intelligent knowledge collection and integrated storage method and device for industrial data standard, medium and product
By employing intelligent knowledge acquisition and integrated storage methods, the problem of automated parsing of industrial data standard documents has been solved, enabling digital conversion and efficient management of data, and improving data quality and consistency detection efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-10
AI Technical Summary
Current industrial data standards are typically stored in paper documents, HTML web pages, or PDF documents, which hinders the automated parsing of standard data by industrial software and makes it difficult to guarantee data quality.
This paper provides an intelligent knowledge acquisition and integrated storage method for industrial data standards, including format conversion, online indexing, entity relationship extraction, data fusion and distributed storage, generating structured standard documents and constructing a knowledge graph storage architecture to realize the digitization and integrated storage of data.
It has enabled the digital conversion of industrial data standard documents, improved data quality and consistency testing efficiency, and supported efficient and intelligent data query and management.
Smart Images

Figure CN121636639A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of industrial data standardized storage, in particular to an intelligent knowledge collection and integrated storage method and device for industrial data standards, a medium and a product. BACKGROUND
[0002] Industrial data is the basis for realizing intelligent manufacturing. In the context of the world's sustained low economic growth, major countries in the world actively seize the opportunity of the new round of technological and industrial reform, take intelligent manufacturing as the basis for enhancing the core competitiveness of the country, and pay more and more attention to the supporting and guaranteeing role of standardization in promoting technological innovation and industrial reform. The advent of the digital age has improved the data capabilities of manufacturing enterprises, and data elements have empowered traditional elements such as capital, labor and technology, enhancing the ability of enterprises in the use and configuration of elements and innovation, achieving cost reduction and new value creation, and enhancing market competitiveness. Industrial data is the core element of intelligent manufacturing and the "blood" of modern manufacturing. China's intelligent manufacturing development has gone through the stages of off-plot engineering, manufacturing informatization, integration of informatization and industrialization, and intelligent manufacturing. Industrial data has always been in a core position.
[0003] High-quality industrial data is throughout the design, production, management, service and other manufacturing activities, and is the prerequisite for guaranteeing the self-awareness, self-decision, self-execution, self-adaptation and self-learning of intelligent manufacturing systems based on industrial internet and industrial big data. It is also the basis for reliable operation of industrial automation systems.
[0004] Standardized digital data is the benchmark for achieving high-quality industrial data. In order to promote the needs of industrial software for automatic representation, analysis and exchange of data, a large number of industrial data standards have been developed by standardization experts from various countries, including a large number of formats, protocols, access interfaces, information codes, data elements, data dictionaries, data models, ontologies and other technical requirements for data representation and exchange, which are used to guarantee the quality of industrial data and cover product data at various stages of enterprise design, manufacturing, sales, support and retirement, as well as manufacturing management data such as procurement orders, production plans and financial management.
[0005] Current industrial data standards are usually stored in paper files, Html web pages or PDF documents, which are not conducive to the automatic analysis needs of industrial software for standard data. Standard data extracted in a digital form from standard files is the benchmark data for guaranteeing the high-quality operation of industrial data. Industrial data standards expressed in a digital form can promote industrial data to comply with industrial data standards, thereby guaranteeing reliable data quality. SUMMARY
[0006] The purpose of this application is to provide an intelligent knowledge acquisition and integrated storage method, device, medium and product for industrial data standards, which can realize the digitization of industrial data standard documents.
[0007] To achieve the above objectives, this application provides the following solution: Firstly, this application provides an intelligent knowledge acquisition and integrated storage method for industrial data standards, including: Acquire multi-source heterogeneous industrial data standard files, and perform format conversion and acquisition configuration on the industrial data standard files to generate an editable preprocessed standard document with clearly defined extraction rules; Based on the preprocessed standard document, online indexing is performed to mark key management information and data block boundary markers, forming a structured standard document with indexing marks; Based on the structured standard document, entities, attributes, and relationships between entities in the industrial data are extracted to obtain an initial extracted dataset containing entity-attribute-relationship data. The initial extracted dataset is subjected to entity fusion, relation fusion, and attribute fusion operations to eliminate data redundancy and logical errors, and generate a standardized industrial data knowledge set. The standardized industrial data knowledge set is imported into a knowledge graph storage architecture to construct an industrial standard benchmark library. A distributed storage scheme is used to integrate the data in the industrial standard benchmark library to obtain a distributed industrial knowledge repository.
[0008] Optionally, the industrial data standard files include: paper standard scans, HTML format standard web pages, and PDF format standard documents.
[0009] Optionally, the data collection configuration includes defining document types, setting the names of fields to be extracted, and the mapping relationships between fields; The document types include: standard documents and technical specifications; the field names to be extracted include: standard number and ICS classification number.
[0010] Optionally, the key management information includes: standard name, standard number, publication date, implementation date, list of drafters and drafting unit; The data block boundary markers include: metadata start marker and metadata end marker.
[0011] Optionally, a distributed storage scheme is used to integrate the data in the industry standard benchmark library to obtain a distributed industrial knowledge repository, specifically including: By utilizing a distributed architecture with redundant backups and dynamic node expansion, the data in the industrial standard benchmark library is stored in an integrated manner to obtain a distributed industrial knowledge repository.
[0012] Optionally, the industrial data standard file includes: standard application scenarios, standard clause content, data dictionary entries, and interface parameter specifications.
[0013] Optionally, after using a distributed storage scheme to integrate the data in the industry standard benchmark library to obtain a distributed industry knowledge repository, the method further includes: A data logic rule base is constructed based on the data in the distributed industrial knowledge repository. The data logic rule base is used to perform integrity and consistency checks on the industrial standard benchmark data, and multi-mode data query services are provided based on the distributed industrial knowledge repository.
[0014] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the intelligent knowledge acquisition and integrated storage method for the industrial data standard described in any one of the above.
[0015] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent knowledge acquisition and integrated storage method for the industrial data standard described above.
[0016] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the intelligent knowledge acquisition and integrated storage method for the industrial data standard described above.
[0017] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides an intelligent knowledge acquisition and integrated storage method, device, medium, and product for industrial data standards. The method includes: acquiring multi-source heterogeneous industrial data standard files, performing format conversion and acquisition configuration on the industrial data standard files to generate an editable preprocessed standard document with clearly defined extraction rules; performing online indexing processing on the preprocessed standard document, marking key management information and data block boundary markers to form a structured standard document with indexing marks; extracting entities, attributes, and relationships between entities from the structured standard document to obtain an initial extraction dataset containing entity-attribute-relationship elements; performing entity fusion, relationship fusion, and attribute fusion operations on the initial extraction dataset to eliminate data redundancy and logical errors, generating a standardized industrial data knowledge set; importing the standardized industrial data knowledge set into a knowledge graph storage architecture to construct an industrial standard benchmark library; and using a distributed storage scheme to integrate the data in the industrial standard benchmark library to obtain a distributed industrial knowledge repository. This application automatically acquires knowledge from industrial data standards in electronic and paper formats, performs format conversion, and distributes and integrates different types of digital standards in a knowledge base system, realizing the digitization of industrial data standard files. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is an application environment diagram of an intelligent knowledge acquisition and integrated storage method for industrial data standards according to an embodiment of this application.
[0020] Figure 2 This is a flowchart illustrating an intelligent knowledge acquisition and integrated storage method for industrial data standards, provided in Embodiment 1 of this application.
[0021] Figure 3 This is a schematic diagram of the standard upload and conversion process provided in Embodiment 1 of this application.
[0022] Figure 4 This is a schematic diagram of the data acquisition configuration management process provided in Embodiment 1 of this application.
[0023] Figure 5 This is a schematic diagram of data indexing and processing provided for Embodiment 1 of this application.
[0024] Figure 6 This is a schematic diagram of the content recognition and extraction process provided in Embodiment 1 of this application.
[0025] Figure 7 This is a schematic diagram of the basic system for digital transformation research provided in Embodiment 1 of this application.
[0026] Figure 8 This is a flowchart illustrating an intelligent knowledge acquisition and integrated storage method for industrial data standards, as provided in Embodiment 2 of this application.
[0027] Figure 9 This is a schematic diagram of the data collection item configuration provided in Embodiment 2 of this application.
[0028] Figure 10 This is a schematic diagram illustrating a tabular metadata example provided in Embodiment 2 of this application. Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] The intelligent knowledge acquisition and integrated storage method for industrial data standards provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server.
[0032] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.
[0033] Example 1: In one exemplary embodiment, such as Figure 2As shown, an intelligent knowledge acquisition and integrated storage method for industrial data standards is provided. This method is executed by computer equipment, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps S1 to S6. Wherein: S1. Obtain multi-source heterogeneous industrial data standard files, and perform format conversion and acquisition configuration on the industrial data standard files to generate an editable preprocessed standard document with clear extraction rules.
[0034] The industrial data standard files include: paper standard scans, HTML format standard web pages, and PDF format standard documents.
[0035] The contents of the industrial data standard document include: standard application scenarios, standard clauses, data dictionary entries, and interface parameter specifications.
[0036] The data collection configuration includes defining document types, setting the names of fields to be extracted, and mapping relationships between fields.
[0037] The document types include: standard documents and technical specifications; the field names to be extracted include: standard number and ICS classification number.
[0038] S2. Based on the preprocessed standard document, perform online indexing processing, mark key management information and data block boundary markers, and form a structured standard document with indexing marks.
[0039] Key management information includes: standard name, standard number, publication date, implementation date, list of drafters and drafting organization.
[0040] Data block boundary markers include: metadata start marker and metadata end marker.
[0041] S3. Based on the structured standard document, extract the entities, attributes, and relationships between entities from the industrial data to obtain an initial extracted dataset containing entity-attribute-relationship information.
[0042] S4. Perform entity fusion, relationship fusion, and attribute fusion operations on the initial extracted dataset to eliminate data redundancy and logical errors, and generate a standardized industrial data knowledge set.
[0043] S5. Import the standardized industrial data knowledge set into a knowledge graph storage architecture to build an industrial standard benchmark library.
[0044] S6. A distributed storage scheme is used to integrate the data in the industrial standard benchmark library to obtain a distributed industrial knowledge repository.
[0045] Specifically, this embodiment utilizes a distributed architecture with redundant backup and dynamic node expansion to integrate the data in the industrial standard benchmark library, thereby obtaining a distributed industrial knowledge repository.
[0046] After using a distributed storage scheme to integrate the data in the industrial standard benchmark library to obtain a distributed industrial knowledge repository, the following steps are also included: A data logic rule base is constructed based on the data in the distributed industrial knowledge repository. The data logic rule base is used to perform integrity and consistency checks on the industrial standard benchmark data, and multi-mode data query services are provided based on the distributed industrial knowledge repository.
[0047] Specifically, this embodiment can be implemented using the following specific methods: 1. Data Acquisition and Processing. The system is used to process multi-source heterogeneous standard documents, including standard documents from various sources and formats such as paper standards, HTML web pages, and PDF scans, and supports data acquisition, indexing, and processing.
[0048] 2. Intelligent Data Identification, Extraction, and Fusion. The system aims to improve the quality and accuracy of industrial data, supporting digital transformation and standardization in the industrial sector. This process involves not only data collection and processing but also emphasizes the intelligent identification and extraction of data, including standard documents, drafters, drafting units, terminology, clauses, classification codes, metadata, data dictionaries, interfaces, abbreviations, and scope of application, to build a comprehensive, accurate, and reliable industrial data benchmark library.
[0049] 3. Establish a standard benchmark library. Referring to patent [ZL 2022 1 0593280.6] "A Standard Digital Management and Maintenance System and Method Based on Knowledge Graph," a standard benchmark library can be constructed by extracting, processing, and storing standard data such as classification codes, metadata, and data dictionaries. This library is used to test the consistency and compliance between industrial data and standard data. This library serves as the benchmark or standard for industrial data quality testing. It includes: 3.1 Data Collection and Extraction: 3.1.1 Standard Upload and Conversion: like Figure 3 As shown, users can upload standard documents in Word or PDF format to the system. PDF documents can be converted into editable Word documents using the document conversion function.
[0050] 3.1.2 Data Acquisition Configuration Management: like Figure 4As shown, the system supports format configuration for data extraction from standard documents of different formats, including document type, field name, field correspondence, etc.
[0051] 3.1.3 Data Indexing and Processing: like Figure 5 As shown, the system supports online data indexing and processing of standard documents, supports online Word editing functions, and allows setting document table of contents format, content format, and marking of drafters, drafting units, publication dates, etc. It also supports data block markers, such as metadata start markers and end markers.
[0052] 3.1.4 Content Recognition and Extraction: like Figure 6 As shown, the system supports automatic or semi-manual content extraction from standard document data. The extracted content includes standard documents, drafters, drafting units, application scenarios, scope, clauses, terms, abbreviations, data elements, classification codes, data fields, and other content.
[0053] The system supports functions such as data extraction, editing, saving, data import, and export.
[0054] 3.1.5 Data import: After data extraction is complete, the system supports importing the extracted and edited data into the knowledge graph storage system for querying, analyzing, and checking standard data.
[0055] 3.2 Distributed Knowledge Base Integrated Storage: The system can store data and knowledge extracted from industrial data standards. This data is organized in the form of a knowledge graph and supports multi-source heterogeneous industrial data, such as product component data, manufacturing management data, transaction data, master data, and OWL ontology. The system adopts a distributed design to ensure high availability and scalability of data. Through an integrated storage solution, different types of data can be seamlessly integrated and managed, achieving unified data management and efficient access.
[0056] 3.2.1 Data quality benchmark database storage: As a storage system for a data quality benchmark database, this system stores and manages data quality benchmark data to facilitate standard conformity comparison and testing of industrial data quality. The system supports precise storage of standard data, ensuring data integrity and consistency. By storing this data, it provides essential foundational support for nationally prioritized data quality research and testing.
[0057] 3.2.2, Basic System for Digital Transformation Research: Please see Figure 7As a foundational system for digital transformation research, the system supports the storage and research of high-maturity standard data in the intelligent standards model proposed by IEC / ISO. The system can store and process high-maturity standard data, such as machine-readable and executable content and machine-controllable documents. By providing these advanced functions, the system can promote research on the digital transformation of industrial data standards and improve machines' ability to automatically parse, understand, and execute standard data.
[0058] 3.2.3 Establishment of Knowledge Rule Base: By establishing a data logic rule base, data quality functions such as integrity and consistency can be tested based on the rule base.
[0059] To implement the above method, this embodiment also provides a corresponding device implementation scheme, as shown below: The intelligent industrial data knowledge acquisition and integrated storage system includes a data acquisition module, a data indexing module, a data identification and extraction module, a data fusion module, a data storage module, a data management module, a data query module, and a result display module.
[0060] The data acquisition module is used to acquire multi-source heterogeneous standard documents, including standard documents from various sources and formats such as paper standards, HTML web pages, and PDF scans.
[0061] The data indexing module is used to index and process various types of collected data; PDF documents can be converted into editable Word documents through the document conversion function; standard documents can be indexed and processed online, and online Word editing functions are supported. Document table of contents and content formats can be set, and the drafter, drafting unit, and publication date can be marked. Data block markers, such as metadata start markers and end markers, are also supported.
[0062] The data identification and extraction module is used to intelligently identify entities and attributes (features) related to industrial data, including product names, component names, and feature / characteristic names.
[0063] It also extracts data relationships, including integration relationships, instance relationships, aggregation relationships, composition relationships, inclusion relationships, and descriptive relationships.
[0064] The standard document data can be extracted automatically or by semi-manual annotation. The extracted content includes standard documents, drafters, drafting units, application scenarios, scope, clauses, terms, abbreviations, data elements, classification codes, data fields, etc.
[0065] The data fusion module is used to merge industrial data, eliminate redundancy and errors, and carry out knowledge fusion such as entity fusion, relationship fusion, and attribute fusion.
[0066] The data storage module is used to import the extracted and edited data into the knowledge graph storage system after the data extraction is completed, so that standard data can be queried, analyzed and checked.
[0067] The data management module is used to manage standard document data that is being extracted and has already been extracted and stored, including data deletion; binding and unbinding entity relationships; and adding, editing, and importing relationships in standard documents.
[0068] The data query module includes intelligent question answering, voice search, and Cypher query; among them, the intelligent question answering uses natural language processing technology and knowledge graph association to enable intelligent dialogue between users and the system, providing accurate answers.
[0069] The voice search function calls an internet voice recognition interface, allowing users to search using voice input. The system can recognize and parse the voice, convert it into text, and then retrieve it from the knowledge graph.
[0070] The results display module is used to display query results.
[0071] Its display methods include categorized display, constructing a classification structure tree based on the attributes of standard data, such as standard classification number (ICS, or CCS), and the hierarchical structure of standard documents (standard name, level of standard bar); it supports classifying entities and relationships in the knowledge graph into different structure tree categories; and it supports display by classification structure tree.
[0072] Compared with the prior art, this embodiment has at least the following advantages: 1. This embodiment can assist enterprises in the digital transformation of paper or electronic standard documents, realizing the conversion of standard documents into standard data and forming a digital standard knowledge base.
[0073] This embodiment provides a solution for the automated conversion of unstructured standard documents into structured standard data. A structured standard knowledge base can facilitate precise knowledge retrieval, question answering, and development. Through natural language processing technology and knowledge graph association, it enables intelligent dialogue between the user and the system, providing accurate answers; voice search facilitates user searches using voice input, and the system can recognize and parse the speech, converting it into text and retrieving it from the knowledge graph. This patented technology is used to convert PDFs to Word documents, extract and form tabular data, offering strong interpretability, high accuracy and precision, and reliable extracted standard data suitable for large-scale, batch applications.
[0074] 2. This embodiment can be used to build a digital standard library, providing standardized benchmark data to ensure high-quality industrial data.
[0075] For the manufacturing industry, the quality of industrial data directly determines the level of intelligence in product design, processing, assembly, logistics, warehousing, inspection and testing, use, and after-sales support. A standard benchmark database collected from national industry standards can be used to test the consistency / compliance of industrial data with standard data, assisting enterprises in achieving automated and intelligent data quality testing or online quality monitoring of data during the production process. This patented technology solves the problem of traditional data quality testing relying on standard file formats, transforming it into testing based on standard data as a benchmark and basis, thus improving the efficiency and intelligence of data quality testing.
[0076] Example 2: In one exemplary embodiment, such as Figure 8 As shown, an intelligent knowledge acquisition and integrated storage method for industrial data standards is provided, including the following steps: A1. Read standard document files uploaded by users from a specific directory on the server.
[0077] A2. Determine the type of the standard file. If the file is in PDF format, convert it to Word format.
[0078] A3. Read and parse the directory structure and content of the Word document, and then start parsing various entity data from the read directory structure.
[0079] A4. Read the annotations made during the creation of the data collection task, and parse the basic information of the document from the annotations, including: ICS, CCS, standard number, standard name, publication date, implementation date, drafter, and drafting unit. For example, if the annotation is "ICS", then the text circled in the annotation is the corresponding value of the annotation; other attributes are similar.
[0080] A5. Parse the referenced chapter content, traverse the chapter and paragraph content. If the content starts with GB or ISO, it is considered a reference standard. Use spaces to separate the current paragraph text. The first part of the separation is the reference standard number, and the remaining parts are merged into the reference standard name.
[0081] A6. Read the content of the relevant chapters and paragraphs, and parse out the main content and scope of application. If a paragraph contains the keywords "specify" or "define," it is considered the main content; if it contains the keyword "applies to," it is considered the scope of application.
[0082] A7. Read the terminology section and extract the Chinese name, English name, definition, source code, and reference clause number. The specific process is as follows: First, try the first method.
[0083] A711. Read the content of each paragraph one by one and determine the content outline level. If the content outline level is the table of contents, use the current paragraph text to obtain the "clause number" of the current content from the chapter number map parsed in step 3. At the same time, use a space " " to separate the current paragraph text content. The first part of the separated content is the "Chinese name" of the term entity, and the second part is the "English name" of the term.
[0084] A712. If the outline level of the text content is not a table of contents but the main text, then determine whether it starts with "note". If so, the current paragraph content is identified as a "note" of the term entity.
[0085] A713. If the main text begins with "[", it is considered source information. After removing the square brackets before and after, use "," to separate the two parts. The first part is the "Source Document Number", and the second part is the "Source Document Clause Number".
[0086] A714. If the text does not begin with "Note" or "[", it is considered a "definition" of the term entity.
[0087] Repeat steps A711-A714 until the terminology section is resolved.
[0088] If the first method fails to parse the terminology information, try the second method: A721. Read all terminology chapter and paragraph content and group the terminology content by serial number. Each group contains a set of terminology definitions.
[0089] A722. Analyze the content of each group one by one. The first item in a single group is the "clause number".
[0090] A723. The second item consists of the Chinese name and the English name, separated by a space. The first part of the separated text is the "Chinese name" of the term entity, and the second part is the "English name" of the term.
[0091] A724. For a detailed analysis of the logic starting from the third item, please refer to steps 2 to 4 of the first method.
[0092] Repeat steps A721 to A724 above until all groups have been parsed.
[0093] A8. Determine if there are abbreviation chapters by checking the chapter titles. If there are abbreviation chapters, parse the abbreviation content. The specific steps are to traverse all paragraphs, separating the paragraph text with commas ",", the second part of the separation is "Chinese vocabulary", and the first part is separated by "-", the first part is "abbreviations" and the second part is "English vocabulary".
[0094] A9. Read the configuration information entered on the page and determine the metadata style of the current standard document, such as... Figure 9 As shown.
[0095] A10. If it's a table format, read the table content and parse identifiers, Chinese names, English names, Chinese abbreviations, English abbreviations, etc. If it's a data block format, parse the data block content according to the page configuration information. For example... Figure 10 As shown. The process of reading data in tabular form: A1011. Read all cells in the first row of the table, and iterate through all cells to get the content in each cell.
[0096] A1012. Read the next row and all cells in the current row, retrieve all cells and read their contents, identifying the cell attributes based on the cell index (starting from 0) and the header content read in step one. For example... Figure 3 The table header corresponding to "002" with index 0 is "Identifier", the table header corresponding to "Forging Product Metadata" with index 1 is "Chinese Name", and the table header corresponding to "Metadata" with index 2 is "English Name", etc.
[0097] A1013. Repeat step A1012 until all lines have been read.
[0098] Data block reading process: A1021. Traverse the metadata chapters and paragraphs, and simultaneously traverse the list of metadata attribute values entered in the second step of the data collection task. Figure 9 The values entered in the document are the corresponding fields, such as "identifier" and "Chinese name".
[0099] A1022. If the current paragraph text begins with the current attribute value, then the current attribute value is parsed. For example, if the text content is "Chinese Name: Work Name", then the second part separated by ":" is the value of the attribute "Chinese Name"; if the text content is "Definition: Name of a Musical Work.", then the second part separated by ":" is the value of the attribute "Definition".
[0100] A1023. Repeat step A1022 until all fields have been parsed.
[0101] A1024. Then repeat steps A1021 to A1023 to parse all the metadata blocks.
[0102] A11. Traverse all category code chapter paragraphs and tables, and read the category code table content. The specific process is similar to the process of reading table-format metadata in step 10.
[0103] A12. Traverse the entire document directory structure and read the clause content.
[0104] A13. Convert all the data read in the above steps into corresponding entities and attributes, and assign a unique ID to each entity.
[0105] A14. Establish the association between entities and standards by using the unique id value of the entity and the unique value of the current standard document to establish the association.
[0106] A15. Save all entities, attributes, and relationships to the library.
[0107] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 11 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements an intelligent knowledge acquisition and integrated storage method based on industrial data standards.
[0108] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0109] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0110] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0111] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0113] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0114] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0115] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0116] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An intelligent knowledge acquisition and integrated storage method of an industrial data standard, characterized in that, The method comprises the following steps: acquiring a multi-source heterogeneous industrial data standard file, and performing format conversion and collection configuration on the industrial data standard file to generate a preprocessed standard document that is editable and has clear extraction rules; performing online indexing processing according to the preprocessed standard document, marking key management information and data block boundary markers to form a structured standard document with indexing markers; extracting entities, attributes and inter-entity relationships of industrial data from the structured standard document to obtain an initial extraction data set containing entities-attributes-relationships; performing entity fusion, relationship fusion and attribute fusion operations on the initial extraction data set to eliminate data redundancy and logical errors, and generating a standardized industrial data knowledge set; importing the standardized industrial data knowledge set into a knowledge graph storage architecture to construct an industrial standard benchmark library; integrally storing data in the industrial standard benchmark library by using a distributed storage scheme to obtain a distributed industrial knowledge storage library.
2. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, The industrial data standard file includes a paper standard scan, an HTML format standard webpage and a PDF format standard document.
3. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, The collection configuration includes defining a document type, setting a field name to be extracted and a mapping relationship between fields; The document type includes a standard literature and a technical specification; and the field name to be extracted includes a standard number and an ICS classification number.
4. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, The key management information includes a standard name, a standard number, an issue date, an implementation date, a list of drafters and a drafting unit; The data block boundary markers include a metadata start marker and a metadata end marker.
5. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, The distributed storage scheme is used to integrally store data in the industrial standard benchmark library to obtain a distributed industrial knowledge storage library, which specifically includes: The data in the industrial standard benchmark library is integrally stored by using a distributed architecture with redundant backup and node dynamic expansion to obtain a distributed industrial knowledge storage library.
6. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, The industrial data standard file includes a standard application scenario, a standard clause content, a data dictionary entry and an interface parameter specification.
7. The intelligent knowledge acquisition and integrated storage method of industrial data standards according to claim 1, characterized in that, After the data in the industrial standard benchmark library is integrally stored by using the distributed storage scheme to obtain the distributed industrial knowledge storage library, the following steps are further included: Based on the data in the distributed industrial knowledge storage library, a data logic rule library is constructed; The data in the industrial standard benchmark library is detected for integrity and consistency by using the data logic rule library, and a multi-mode data query service is provided based on the distributed industrial knowledge storage library.
8. A computer device comprising: A memory, a processor and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the intelligent knowledge acquisition and integrated storage method of the industrial data standard according to any one of claims 1-7.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the intelligent knowledge acquisition and integrated storage method of the industrial data standard according to any one of claims 1-7.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the intelligent knowledge acquisition and integrated storage method of the industrial data standard according to any one of claims 1-7.
Citation Information
Patent Citations
Standard digital management and maintenance system and method based on knowledge graph
CN114792145A