A method and system for semantic-enhanced conversion of heterogeneous industrial data to PDV format

CN122817170APending Publication Date: 2026-09-25BEIJING YUANHUI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611250602.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种异构工业数据到PDV格式的语义增强转换方法和系统,以解决现有技术中因架构孤立且仅作语法转换,导致转换结果语义浅层、无法识别隐含关系的问题

Benefits of technology

在本申请中,通过插件化的适配器架构加载对应适配器插件,将各异构工业数据源文件统一解析为包含结构层数据、几何层数据、语义层数据、属性层数据及关系层数据的目标数据集,打破了不同设计工具间的格式壁垒,可快速扩展支持新的数据格式,降低多源异构数据的集成复杂度与维护成本;基于与数据格式对应的预设映射规则集,将目标数据集转换为初级语义图并同步生成关联的目标数据包,使语义知识与轻量化几何模型在转换源头即建立结构化关联,确保输出结果同时具备可查询的知识结构和可交互的三维呈现,为企业构建统一数字孪生体扫清数据障碍;通过语义识别模型对初级语义图进行特征分析与语义识别,得到隐含语义信息,能够从看似孤立的几何与属性信息中挖掘出未显式声明的功能单元与设计意图,使输出的PDV标准语义文件蕴含远超原始数据的深层知识,提升知识图谱的价值;基于隐含语义信息对初级语义图进行语义增强处理,生成语义增强三元组,并将语义增强三元组与初级语义图进行融合和一致性校验生成目标语义图,使得智能发现的语义洞见被正式编码为可推理的知识图谱节点与边,确保增强后的知识体系逻辑自洽;对目标语义图进行多维度质量评估,基于质量评估结果,将目标语义图转换为PDV标准语义文件,并与目标数据包整合成标准PDV模型包,为下游智能问答、设计辅助与工艺生成等应用提供格式规范、语义丰富且可视可交互的高质量产品知识包。同时,质量评估结果可持续反馈至语义识别模型进行迭代优化,使框架的转换质量与语义识别能力随使用持续提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817170A_ABST
    Figure CN122817170A_ABST
Patent Text Reader

Abstract

The application discloses a heterogeneous industrial data to PDV format semantic enhancement conversion method and system; by acquiring a plurality of heterogeneous industrial data source files, converting the heterogeneous industrial data in each heterogeneous industrial data source file into a target data set in a preset format; based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, converting the target data set into a primary semantic graph, and generating a target data packet associated with the primary semantic graph; through a semantic recognition model, performing feature analysis and semantic recognition on the primary semantic graph to obtain implicit semantic information, performing semantic enhancement processing on the primary semantic graph, generating a semantic enhancement triple, fusing the semantic enhancement triple with the primary semantic graph, and performing consistency checking to generate a target semantic graph; based on a quality evaluation result of the target semantic graph, converting the target semantic graph into a PDV standard semantic file, combining the target data packet, and generating a standard PDV model package.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic conversion technology, and more specifically, to a semantically enhanced conversion method and system for heterogeneous industrial data to PDV format. Background Technology

[0002] Against the backdrop of the rapid development of intelligent manufacturing and digital twins, the industrial sector generates a massive amount of design data every day from various tools such as mechanical design, building information modeling, and electronic design automation. These heterogeneous data use their own different data formats and semantic standards, forming information silos.

[0003] Currently, the mainstream solution for processing heterogeneous industrial data in the industry is to develop independent conversion tools for each data format. These tools parse the source files of different formats separately and then convert them to the target format using customized scripts or hard-coded rules. The independent and isolated architecture of these tools forces enterprises to deploy and maintain these different conversion tools separately when facing multiple data sources, forming vertical "conversion silos." However, this approach has significant drawbacks. For example, the siloed architecture leads to high system maintenance costs, difficulty in expansion, and slow adaptation to new data formats. More importantly, existing tools can only perform format conversion at the syntactic level, lacking the ability to understand the deep semantics of the data. They cannot identify the implicit functional relationships and design intentions between parts, resulting in conversion results that remain at a shallow, structured level of "geometry + labels," rather than semantic knowledge that machines can understand. Summary of the Invention

[0004] The main objective of this application is to provide a semantically enhanced conversion method and system for heterogeneous industrial data to PDV format, in order to solve the problem that the conversion results are semantically shallow and unable to identify implicit relationships due to the isolated architecture and only syntax conversion in the prior art.

[0005] To achieve the above objectives, the first aspect of this application proposes a semantically enhanced conversion method for heterogeneous industrial data to PDV format, comprising:

[0006] Multiple heterogeneous industrial data source files are acquired, and the heterogeneous industrial data in each heterogeneous industrial data source file is converted into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data. Based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, the target dataset is converted into a primary semantic graph, and a target data package associated with the primary semantic graph is generated. The implicit semantic information is obtained by performing feature analysis and semantic recognition on the primary semantic graph using a semantic recognition model. Based on the implicit semantic information, the primary semantic graph is semantically enhanced to generate semantically enhanced triples. The semantically enhanced triples are then fused with the primary semantic graph and a consistency check is performed to generate the target semantic graph. The target semantic graph is subjected to quality assessment to obtain quality assessment results. Based on the quality assessment results, the target semantic graph is converted into a PDV standard semantic file. The target data packet and the PDV standard semantic file are integrated into a standard PDV model packet.

[0007] Secondly, this application provides a semantically enhanced conversion system for heterogeneous industrial data to PDV format, comprising: The acquisition module is used to acquire multiple heterogeneous industrial data source files and convert the heterogeneous industrial data in each heterogeneous industrial data source file into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data. The conversion module is used to convert the target dataset into a primary semantic graph based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, and generate a target data package associated with the primary semantic graph. The recognition module is used to perform feature analysis and semantic recognition on the primary semantic graph through a semantic recognition model to obtain implicit semantic information. The generation module is used to perform semantic enhancement processing on the primary semantic graph based on the implicit semantic information, generate semantic enhancement triples, fuse the semantic enhancement triples with the primary semantic graph and perform consistency verification to generate the target semantic graph. The evaluation module is used to perform quality evaluation on the target semantic graph, obtain quality evaluation results, convert the target semantic graph into a PDV standard semantic file based on the quality evaluation results, and integrate the target data packet and the PDV standard semantic file into a standard PDV model packet.

[0008] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a semantically enhanced conversion method for heterogeneous industrial data to PDV format.

[0009] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a computer, implements a semantically enhanced conversion method for heterogeneous industrial data to PDV format.

[0010] The technical solutions provided by the embodiments of this application may include the following beneficial effects: In this application, a plug-in-based adapter architecture loads corresponding adapter plugins to uniformly parse various heterogeneous industrial data source files into a target dataset containing structural, geometric, semantic, attribute, and relational layer data. This breaks down format barriers between different design tools, allows for rapid expansion to support new data formats, and reduces the integration complexity and maintenance costs of multi-source heterogeneous data. Based on a preset mapping rule set corresponding to the data format, the target dataset is converted into a primary semantic graph and associated target data packages are generated simultaneously. This establishes a structured association between semantic knowledge and lightweight geometric models at the source of conversion, ensuring that the output results possess both a queryable knowledge structure and an interactive 3D presentation, thus removing data obstacles for enterprises to build a unified digital twin. Furthermore, a semantic recognition model performs feature analysis and semantic recognition on the primary semantic graph to obtain implicit semantic information, enabling the extraction of... By extracting undeclared functional units and design intents from seemingly isolated geometric and attribute information, the output PDV standard semantic file contains deeper knowledge than the original data, enhancing the value of the knowledge graph. Based on implicit semantic information, the primary semantic graph undergoes semantic enhancement processing to generate semantically enhanced triples. These triples are then fused with the primary semantic graph and undergo consistency verification to generate the target semantic graph. This formally encodes the semantic insights discovered through intelligent discovery into reasonable knowledge graph nodes and edges, ensuring the logical consistency of the enhanced knowledge system. A multi-dimensional quality assessment of the target semantic graph is performed. Based on the assessment results, the target semantic graph is converted into a PDV standard semantic file and integrated with the target data package to form a standard PDV model package. This provides downstream applications such as intelligent question answering, design assistance, and process generation with a high-quality product knowledge package that is formatted correctly, semantically rich, and visually interactive. Simultaneously, the quality assessment results are continuously fed back to the semantic recognition model for iterative optimization, ensuring that the framework's conversion quality and semantic recognition capabilities continuously improve with use. Attached Figure Description

[0011] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 A flowchart illustrating a semantic enhancement conversion method for heterogeneous industrial data to PDV format provided in this application; Figure 2 This is a schematic diagram of a semantic enhancement conversion system for heterogeneous industrial data to PDV format provided in this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0013] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0014] In this application, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing this application and its embodiments, and are not intended to limit the indicated device, element, or component to having a specific orientation, or to be constructed and operated in a specific orientation.

[0015] Furthermore, in addition to indicating location or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in some cases to indicate a certain dependency or connection relationship. Those skilled in the art can understand the specific meaning of these terms in this application based on the specific circumstances.

[0016] Furthermore, the terms "installation," "setup," "equipped with," "connection," "linked," and "socketing" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0017] The core of this application is to provide a semantic enhancement conversion method for heterogeneous industrial data to PDV format, and a flowchart of one specific implementation is shown below. Figure 1 As shown, the method includes: Step 101: Obtain multiple heterogeneous industrial data source files, and convert the heterogeneous industrial data in each heterogeneous industrial data source file into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data.

[0018] Step 102: Based on the preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, convert the target dataset into a primary semantic graph and generate a target data package associated with the primary semantic graph.

[0019] Step 103: Perform feature analysis and semantic recognition on the primary semantic graph using a semantic recognition model to obtain implicit semantic information.

[0020] Step 104: Based on the implicit semantic information, perform semantic enhancement processing on the primary semantic graph to generate semantic enhancement triples. Then, fuse the semantic enhancement triples with the primary semantic graph and perform consistency verification to generate the target semantic graph.

[0021] Step 105: Perform a quality assessment on the target semantic graph to obtain the quality assessment results. Based on the quality assessment results, convert the target semantic graph into a PDV standard semantic file and integrate the target data packet and the PDV standard semantic file into a standard PDV model packet.

[0022] In step 101 above, multiple heterogeneous industrial data source files are obtained. Heterogeneous industrial data source files (hereinafter referred to as source files) refer to product data files that come from different engineering fields, follow different data standards, or use different software native formats. These include data files that use neutral standard formats for geometry and attribute exchange, dedicated format files directly output by native design software from different fields, and non-standard structured files exported by old design software that has been phased out or is no longer maintained, which do not follow the current industry data exchange standards. These files have essential differences in syntax, semantic definition, and domain professional concepts.

[0023] Then, the heterogeneous industrial data in each heterogeneous industrial data source file is converted into a target dataset in a preset format. This step specifically includes the following steps: Step 111: Load the corresponding adapter plugin according to the file extension of the heterogeneous industrial data source file or the data format specified by the user.

[0024] Step 112: Parse the heterogeneous industrial data source file through the adapter plugin to obtain assembly tree hierarchical relationship data, geometric expression data of each part, geometric semantic feature information, business attribute key-value pair data, and assembly constraint relationship data between each part.

[0025] Step 113: According to the preset format, convert the assembly tree hierarchical relationship data, geometric expression data, geometric semantic feature information, business attribute key-value pair data and assembly constraint relationship data into the structural layer data, geometric layer data, semantic layer data, attribute layer data and relation layer data in the target dataset.

[0026] In step 111 above, the file suffix refers to the string after the last period in the file name of each heterogeneous industrial data source, which is used to identify the data format type to which the file belongs. For example, the suffix of STEP (Standard for the Exchange of Product model data) format files is .stp or .step, and the suffix of IGES (Initial Graphics Exchange Specification) format files is .igs or .iges.

[0027] The user-specified data format refers to the data format type that the user manually selects or enters through the interactive interface. When the file extension is missing, cannot be recognized, or the user believes that the file format automatically determined by the system cannot correctly reflect the actual data organization structure of the file, the user-specified data format shall prevail.

[0028] An adapter plugin is an independent software component developed in accordance with a unified interface specification, used to parse a specific data format and map the parsing results to a preset internal format. Each adapter plugin corresponds to a data format type. The unified interface specification requires each adapter plugin to implement at least geometric data extraction, attribute data extraction, and structural data extraction functions.

[0029] In this embodiment, based on the file extensions of the heterogeneous industrial data source files, the corresponding data format type is searched in a pre-built data format mapping table, or the user-specified data format type is directly read. Then, an adapter plugin uniquely corresponding to that data format type is loaded from the adapter plugin library, and this adapter plugin is connected to the data processing pipeline for subsequent parsing operations. The data format mapping table refers to a pre-stored configuration table showing the correspondence between file extension strings and data format types, which can be constructed based on known data format standards and technical documents.

[0030] In step 112 above, the assembly tree hierarchy data refers to the data describing which sub-assemblies and parts the product is composed of according to which parent-child hierarchy structure. It includes the identifier of each node, the type label of the node type as assembly or part, the reference identifier pointing to the direct parent node of the node, and the number of instances of the node referenced in the current level. The number of instances refers to the number of times a node in the assembly tree is repeatedly referenced under its direct parent node.

[0031] The geometric representation data of each part refers to the data used to describe the three-dimensional geometry of each part, including boundary representation data, lightweight triangular mesh data, and also the minimum coordinates, maximum coordinates, and center point coordinates of the bounding box calculated based on the vertex coordinate sequence. The boundary representation data consists of mathematical surface equations, and the lightweight triangular mesh data consists of vertex coordinate sequences and triangle index sequences.

[0032] Geometric semantic feature information refers to the process of analyzing the mathematical equations of each geometric surface or reading the fields used to identify the type of each geometric surface in the source file to determine the geometric surface type and corresponding key parameters. Geometric surface types include planes, cylinders, spheres, and cones. The key parameter of a plane is the normal vector, the key parameters of a cylinder are the axis direction vector and the radius value, and the key parameters of a sphere are the coordinates of the sphere's center and the radius value.

[0033] Business attribute key-value pair data refers to pairs of data consisting of attribute names and corresponding attribute values ​​extracted from the attribute set associated with each part or assembly in the source file. Attribute names include material, part number, weight, color, etc.

[0034] Assembly constraint relationship data between parts refers to data describing the relative positions and connection methods between parts in an assembly. It includes the parts that constitute the assembly constraint relationship, the corresponding constraint type, and the corresponding part identifier. The constraint types include fitting constraints, alignment constraints, distance constraints, and angle constraints, etc.

[0035] Part identifiers are used to uniquely identify parts in the geometry layer data and attribute layer data, that is, to uniquely identify each part that already exists in the source file.

[0036] In this embodiment, the adapter plugin is used to read the internal data organization structure of the source file segment by segment, and the original information stored in the source file with different syntax structures is transformed into assembly tree hierarchical relationship data, geometric expression data, geometric semantic feature information, business attribute key-value pair data and assembly constraint relationship data with unified semantic expression.

[0037] In step 113 above, the preset format refers to a five-layer key-value pair data template predefined according to the unified input requirements of the data structure. The field names, data types, and nesting levels in each layer of data are fixed. For example, the preset format can be a five-layer data template organized using a JSON (JavaScript Object Notation) object structure. The structure layer data template includes a node ID field, a node type field, a parent node reference field, and an instance count field. The geometry layer data template includes a part identifier field, a boundary representation field, a triangular mesh field, a bounding box field, and a center point coordinate field. The semantic layer data template includes a geometric surface identifier field, a geometric type field, and a parameter list field. The attribute layer data template includes a part identifier field and an attribute mapping table field. The relationship layer data template includes a constraint identifier field, a constraint type field, and a part identifier field for the parts that constitute the assembly constraint relationship.

[0038] The target dataset refers to a unified internal data set formed by organizing five types of data obtained from parsing multiple heterogeneous industrial data source files through an adapter plugin, according to a preset format. This data set summarizes all extractable structural, geometric, semantic, attribute, and relational information from all source files involved in the transformation. Among them, the structural layer data refers to a dataset indexed by node ID, recording the type, parent node reference, and instance count of each node in the assembly tree parsed from all source files. This layer of data unifies the assembly hierarchy relationships of various heterogeneous industrial data source files into a single product structure tree data.

[0039] Geometric layer data refers to a dataset containing boundary representation data, triangular mesh data, bounding box coordinates, and center point coordinates of each part parsed from all source files, indexed by part identifiers. This layer of data summarizes the geometric representation information of each part from various heterogeneous industrial data source files.

[0040] Semantic layer data refers to a dataset indexed by geometric surface identifiers, containing the geometric surface type and corresponding key parameters of each surface parsed from all source files. This semantic layer data is organized according to a preset format using geometric semantic feature information extracted from files from various heterogeneous industrial data sources. Attribute layer data refers to a dataset indexed by part identifiers, recording the mapping table of attribute names and attribute values ​​for each part parsed from all source files. This layer stores the business attribute key-value pairs of each part from various heterogeneous industrial data source files. Relationship layer data refers to a dataset indexed by constraint type identifiers, recording the constraint type of each assembly constraint parsed from all source files and the part identifiers of the parts constituting the assembly constraint relationship. This layer summarizes all assembly constraint relationships between parts from various heterogeneous industrial data source files.

[0041] In this embodiment, the identifiers, type tags, parent node references, and instance counts of each node in the assembly tree hierarchy relationship data are filled into the corresponding fields of the structure layer data template to obtain structure layer data; the boundary representation data, triangular mesh data, bounding box coordinates, and center point coordinates of each part are filled into the corresponding fields of the geometry layer data template to obtain geometry layer data; the geometric surface type and key parameters of each geometric surface in the geometric semantic feature information are filled into the corresponding fields of the semantic layer data template to obtain semantic layer data; the attribute names and attribute values ​​of each part in the business attribute key-value pair data are filled into the corresponding fields of the attribute layer data template to obtain attribute layer data; and the constraint type and part identifier of each constraint in the assembly constraint relationship data are filled into the corresponding fields of the relationship layer data template to obtain relationship layer data.

[0042] In step 102 above, the step "converting the target dataset into a primary semantic graph based on a preset mapping rule set corresponding to the data format of heterogeneous industrial data source files" specifically includes the following steps: Step 201: Obtain the type field value of each node in the structure layer data, and extract at least one structure mapping rule, relation mapping rule, and attribute mapping rule that matches the type field value of each node from the preset mapping rule set.

[0043] Step 202: Based on the structure mapping rules corresponding to each node, map each node to an RDF node that conforms to the PDV format.

[0044] Step 203: Based on the attribute mapping rules corresponding to each node, extract the attribute values ​​corresponding to each RDF node from the attribute layer data, and generate attribute triples based on each RDF node and its corresponding attribute values.

[0045] Step 204: Based on the relation mapping rules corresponding to each node, extract the assembly constraint relationship corresponding to each RDF node from the relation layer data, and generate relation triples based on each RDF node and its corresponding assembly constraint relationship.

[0046] Step 205: Construct a primary semantic graph based on the node IDs, RDF nodes, attribute triples, and relation triples corresponding to each part in the structural layer data.

[0047] In step 201 above, the type field value refers to the field value recorded in each node of the structure layer data, which is used to identify the data entity category corresponding to that node. This field value comes from the type declaration of the data entity in the source file. For example, in the STEP format file, the type field value corresponding to the ADVANCED_FACE entity is ADVANCED_FACE.

[0048] A preset mapping rule set refers to a predefined set of rules used to map entity types, attributes, and relationships in the source data format to corresponding elements in the PDV standard format. The preset mapping rule set includes three types of mapping rules: structure mapping rules, relationship mapping rules, and attribute mapping rules. Each mapping rule contains source pattern matching conditions, target pattern generation logic, triggering conditions, and priority information. The source pattern matching conditions include data format type fields and entity type fields, used to define which entity type in which data format the rule applies to. The target pattern generation logic includes rules for generating PDV standard classes or PDV standard attributes, used to specify the PDV standard elements output after a successful match. The triggering conditions define the data context in which the rule takes effect. The priority information determines the execution order when multiple rules match simultaneously; rules with lower priority values ​​are executed first, indicating higher execution priority, stricter triggering conditions, and a narrower applicable data context. These rules must be executed first to ensure that specialized processing takes effect before generalized processing.

[0049] Structure mapping rules refer to the rules in the preset mapping rule set used to map the type of source data entities to the corresponding category in the PDV standard ontology. For example, mapping the ADVANCED_FACE entity type in the STEP format to the pdv:Face category.

[0050] Relationship mapping rules refer to the rules in the preset mapping rule set used to map the relationships between entities in the source data to the corresponding object attributes in the PDV standard ontology. For example, mapping the RelContainedInSpatialStructure relationship in the IFC (Industry Foundation Classes) format to the pdv:isContainedIn object attribute.

[0051] Attribute mapping rules refer to the rules in the preset mapping rule set used to map the attribute fields of the source data entity to the corresponding data attributes in the PDV standard ontology. For example, mapping the OverallHeight attribute field in the IFC format to the pdv:hasHeight data attribute.

[0052] In this embodiment, for each node of the structural layer data in the target dataset, its type field value is obtained. Using this type field value as a search condition in a preset mapping rule set, mapping rules whose entity type matches the type field value in the source pattern matching conditions are found. If multiple matching mapping rules exist, they are extracted sequentially according to their priority values ​​in ascending order. The extracted mapping rules are then categorized into structural mapping rules, relational mapping rules, and attribute mapping rules.

[0053] In step 202 above, an RDF node conforming to the PDV format refers to a semantic graph node that conforms to the PDV standard ontology definition and is expressed in the form of RDF (Resource Description Framework) triples. Each RDF node contains a Uniform Resource Identifier (URI) and a category declaration. The URI is a string identifier used to uniquely identify an RDF node in the RDF semantic graph. This identifier is generated by concatenating the PDV standard class name specified by the target schema generation logic in the corresponding structure mapping rule of the node and the node ID of the node in the structure layer data. The value of the category declaration is determined by the PDV standard class specified by the target schema generation logic in the structure mapping rule.

[0054] In this embodiment of the application, for each node, a PDV standard class is generated according to the target pattern generation logic in the corresponding structure mapping rule, a Uniform Resource Identifier is generated for the node, and a type declaration triple is generated with the Uniform Resource Identifier as the subject and the PDV standard class specified by the structure mapping rule as the category. The semantic graph node corresponding to the type declaration triple is used as the PDV format RDF node mapped by the node.

[0055] For example, for a node in the structure layer data with a type field value of IfcWall, its corresponding structure mapping rule specifies the PDV standard class as pdv:Wall. Then, a Uniform Resource Identifier (URI) is generated for this node: Wall_001, and a type declaration triple is generated: Wall_001 a pdv:Wall. The semantic graph node corresponding to this type declaration triple is used as the RDF node in PDV format after the node is mapped: Wall_001.

[0056] In step 203 above, an attribute triple refers to a triple that uses the Uniform Resource Identifier corresponding to the RDF node as the subject, the PDV standard data attribute specified by the attribute mapping rule as the predicate, and the attribute value extracted from the attribute layer data as the object, and is used to describe the data attribute information of the RDF node.

[0057] In this embodiment, based on the node ID of each RDF node in the structural layer data, the attribute name and corresponding attribute value associated with the node ID are found in the attribute layer data. According to the attribute mapping rules, each attribute name is mapped to a PDV standard data attribute. An attribute triple is generated with the Uniform Resource Identifier of the RDF node as the subject, the mapped PDV standard data attribute as the predicate, and the corresponding attribute value as the object.

[0058] For example, for the :Wall_001 node, the attribute name OverallHeight and its attribute value 3.0 are obtained from the attribute layer data. According to the attribute mapping rules, OverallHeight is mapped to pdv:hasHeight, generating the attribute triple:Wall_001 pdv:hasHeight "3.0"^^xsd:double.

[0059] In step 204 above, a relation triple refers to a triple with the Uniform Resource Identifier corresponding to the RDF node as the subject, the PDV standard object attribute specified by the relation mapping rule as the predicate, and the Uniform Resource Identifier of another RDF node with assembly constraint relationship as the object, which is used to describe the relationship information between RDF nodes.

[0060] In this embodiment, the assembly constraint relationship associated with the node ID of each RDF node is found from the relation layer data. The constraint type and the identifiers of each part constituting the assembly constraint relationship are obtained. The constraint type is mapped to the PDV standard object attribute according to the relation mapping rule. The relation triple is generated with the Uniform Resource Identifier of the RDF node as the subject, the PDV standard object attribute as the predicate, and the Uniform Resource Identifier of another RDF node constituting the assembly constraint relationship as the object.

[0061] For example, for the node :Wall_001, the assembly constraint relationship associated with the node ID of :Wall_001 is found in the relation layer data. A fitting constraint is found between :Wall_001 and :Slab_001. According to the relation mapping rules, the fitting constraint is mapped to pdv:hasConnection. The relation triple :Wall_001pdv:hasConnection :Slab_001 is generated with the Unicode of :Wall_001 as the subject, pdv:hasConnection as the predicate, and the Unicode of :Slab_001 as the object.

[0062] In step 205 above, the node ID refers to a string or numerical value in the structure layer data used to uniquely identify each node in the assembly tree, and the node IDs of different nodes are different from each other.

[0063] A primary semantic graph refers to an RDF graph structure formed by associating type declaration triples, attribute triples, and relation triples of each RDF node according to node ID. This graph structure uses each RDF node as a graph node and the predicates in the attribute triples and relation triples as edges connecting the graph nodes. Each RDF node carries PDV standard class information through type declaration triples, data attribute information through attribute triples, and association information with other RDF nodes through relation triples.

[0064] In this embodiment, the node ID corresponding to each part is used as the association key. RDF nodes, attribute triples, and relation triples are aggregated according to their node IDs, so that the RDF node corresponding to the same node ID, along with its attribute triples and relation triples, constitutes a complete node description. All node descriptions are organized according to the RDF graph structure to obtain a primary semantic graph. In the primary semantic graph, RDF nodes are connected through relation triples, and each RDF node itself carries category and attribute information through type declaration triples and attribute triples.

[0065] Regarding the step "generating the target data package associated with the primary semantic graph", one optional implementation method is to extract the triangular mesh data, bounding box coordinates, and center point coordinates of each part from the geometric layer data in the target dataset, construct the visual geometric model corresponding to each part, extract the geometric surface type and key parameters of each geometric surface from the semantic layer data, bind the node ID corresponding to each part to the corresponding visual geometric model with the node ID as the association key, so that the extracted geometric data establishes a reference relationship with the corresponding RDF node in the primary semantic graph, and generates the target data package.

[0066] The target data package can be understood as a lightweight 3D visualization data package associated with each RDF node in the primary semantic graph. It contains triangular mesh data, bounding box coordinates and center point coordinates of each part, as well as the geometric surface type and key parameters of each geometric surface. It is used to present the 3D geometric shape of the product in the visualization interface and establish interactive associations with the semantic graph nodes.

[0067] This application embodiment realizes the conversion of a target dataset in a unified format into a primary semantic graph conforming to the PDV standard and the synchronous generation of associated target data packages, so that semantic knowledge and visualization data are established in a structured relationship at the source of conversion, breaking down the format barriers between different design tools.

[0068] In step 103 above, the step "to perform feature analysis and semantic recognition on the primary semantic graph using a semantic recognition model to obtain implicit semantic information" specifically includes the following steps: Step 301: Based on the primary semantic graph, construct the input graph structure and map each graph node in the input graph structure to a node embedding vector.

[0069] Step 302: Aggregate the node embedding vectors of each graph node and the node embedding vectors of their corresponding neighboring nodes to generate the context semantic vectors of each graph node.

[0070] Step 303: Based on the context semantic vector, perform graph analysis tasks on each graph node to obtain implicit semantic information. The graph analysis tasks include at least one of node classification, link prediction, or subgraph recognition.

[0071] In step 301 above, the input graph structure refers to a graph data structure with RDF nodes in the primary semantic graph as graph nodes and relation triples in the primary semantic graph as graph edges, which is used as input to the semantic recognition model.

[0072] A node embedding vector refers to mapping the type declaration and attribute information of a graph node to a fixed-dimensional real number vector, which is used to represent the features of the graph node in vector space.

[0073] In this embodiment, each RDF node in the primary semantic graph is converted into a graph node in the input graph structure, and each relation triple in the primary semantic graph is converted into a graph edge connecting two graph nodes in the input graph structure, thereby constructing the input graph structure. The type declaration triple and attribute triple of each graph node in the input graph structure are encoded into a fixed-dimensional node embedding vector. Specifically, the type declaration triple of the graph node is mapped to a type embedding vector, and the attribute values ​​in each attribute triple of the graph node are mapped to attribute embedding vectors. The type embedding vector and each attribute embedding vector are concatenated to obtain the node embedding vector of the graph node. The above encoding process can be implemented using graph embedding methods or other representation learning methods; this application does not limit this.

[0074] In step 302 above, neighboring nodes refer to other graph nodes in the input graph structure that are directly connected to the current graph node through graph edges. The number of neighboring nodes for each graph node is determined by the number of relation triples corresponding to that RDF node in the primary semantic graph.

[0075] The context semantic vector is a vector obtained by aggregating the node embedding vector of the current graph node with the node embedding vectors of its neighbors. It is used to represent the local topological structure and neighborhood semantic information of the current graph node in the graph structure.

[0076] In this embodiment, for each graph node, its neighboring nodes directly connected by graph edges in the input graph structure are determined. A graph convolution aggregation method can be used to perform a weighted summation of the node embedding vector of the current graph node and the node embedding vectors of its neighboring nodes to obtain an aggregated vector. This aggregated vector is then processed using a nonlinear activation function to obtain the context semantic vector of the graph node. The above graph convolution aggregation method can be implemented using other aggregation methods.

[0077] In step 303 above, the graph analysis task refers to the semantic analysis task performed on graph nodes based on contextual semantic vectors, including three tasks: node classification, link prediction, and subgraph identification. Node classification can be understood as determining whether a graph node belongs to a specific semantic category based on its contextual semantic vector, such as determining whether a group of RDF nodes with cylindrical surface features jointly correspond to a functional component. Link prediction can be understood as determining whether there is an undeclared semantic relationship between two graph nodes based on their contextual semantic vectors. Subgraph identification can be understood as determining whether a subgraph composed of multiple graph nodes matches a specific semantic pattern based on their contextual semantic vectors. A specific semantic pattern refers to a semantic pattern that repeatedly appears in engineering semantics and is composed of multiple RDF nodes according to a fixed topology, such as a motor mounting bracket pattern or a pipe connection component pattern.

[0078] Implicit semantic information refers to the output of a graph analysis task, which includes the semantic pattern type, the set of entity identifiers involved, the confidence level, and the suggested semantic type field. The semantic pattern type identifies the category of the identified semantic pattern, the set of entity identifiers involved records the RDF node IDs corresponding to each graph node participating in the semantic pattern, the confidence level is the degree of credibility of the implicit semantic information, and the suggested semantic type field is the PDV standard semantic type that is suggested to be added to the group of entities.

[0079] In this embodiment, the contextual semantic vectors of each graph node are input into the classification prediction layer of the semantic recognition model. A fully connected layer and a normalization function output the probability distribution of each graph node in each PDV standard semantic category. The graph analysis task is then performed based on the probability distribution. For example, one implementation for generating implicit semantic information could be: analyzing the primary semantic graph of a mechanical assembly using a semantic recognition model, identifying multiple RDF nodes with cylindrical surface features and one RDF node with planar features that spatially constitute a motor mounting base subgraph pattern, generating implicit semantic information with the semantic pattern type of functional unit, the set of entity identifiers involved being the identifiers of the aforementioned RDF nodes, a confidence level of 0.92, and a suggested semantic type field of motor mounting base functional unit. The process of generating implicit semantic information using the classification prediction layer described above can be implemented using other methods, and this application does not limit this approach.

[0080] This application's embodiments enable the extraction of implicit semantic information from primary semantic graphs that cannot be obtained through rule transformation, thereby elevating the transformation results from shallow structured data to deep semantic knowledge containing functional relationships and design intentions. This solves the problem that existing tools can only complete grammatical format transformations and lack the ability to understand the deep semantics of data.

[0081] In step 104 above, the step "based on implicit semantic information, perform semantic enhancement processing on the primary semantic graph to generate semantically enhanced triples" specifically includes the following steps: Step 401: Assign corresponding entity identifiers to the values ​​of suggested semantic type fields in implicit semantic information according to the preset identifier generation rules.

[0082] Step 402: Extract the enhanced rule template that matches the implicit semantic information from the pre-stored enhanced rule template library, fill the entity identifier and the values ​​of each field in the implicit semantic information into the corresponding variable placeholders in the enhanced rule template, and generate a semantic enhancement triplet.

[0083] In step 401 above, the identifier generation rule refers to the predefined naming rules used to assign Uniform Resource Identifiers to newly generated semantic entities, including the generation method of the prefix string and the sequence number. The prefix string can be determined according to the PDV standard semantic type name corresponding to the suggested semantic type field, and the sequence number can be generated by incrementing an integer.

[0084] A semantic entity refers to a newly added semantic object corresponding to a set of RDF nodes with a specific semantic pattern type identified from the primary semantic graph through a semantic recognition model. This semantic object does not exist in the primary semantic graph. It is a PDV standard semantic type instance created during the semantic enhancement process based on implicit semantic information. For example, after identifying multiple cylindrical hole feature nodes and a planar feature node that together constitute a motor mounting base functional unit, a new semantic entity representing this functional unit is created.

[0085] The suggested semantic type field refers to the PDV standard semantic type name recorded in the implicit semantic information that is suggested to be added to the identified entity set. The value of the suggested semantic type field can be, for example, a motor mounting base functional unit, anodizing process features, etc.

[0086] An entity identifier is a Uniform Resource Identifier (URI) generated according to identifier generation rules for a new semantic entity corresponding to implicit semantic information. It is used to uniquely identify the new semantic entity in the target semantic graph.

[0087] In this embodiment, the suggested semantic type field in the implicit semantic information is read, the category name of the corresponding PDV standard semantic type is determined according to the value of the suggested semantic type field, and the entity identifier is generated according to the identifier generation rules. One possible way to generate the entity identifier is to use the category name as a prefix string, assign an incrementing sequence number as a suffix string, and concatenate the prefix string and the suffix string to generate the entity identifier.

[0088] For example, when the value of the suggested semantic type field is a motor mounting base functional unit, the category name MotorMount is used as the prefix string according to the identifier generation rules, and the incrementing number 001 is assigned as the suffix string. The prefix string and the suffix string are concatenated to generate the entity identifier: MotorMount_001.

[0089] In step 402 above, the enhanced rule template library refers to a pre-stored set of rules containing multiple enhanced rule templates. Each enhanced rule template corresponds to a semantic pattern type. The enhanced rule template contains variable placeholders for binding specific entity identifiers and attribute values ​​at runtime.

[0090] An enhanced rule template refers to a template in the enhanced rule template library that corresponds to a specific semantic pattern type. The enhanced rule template defines the structure of the subject, predicate, and object of the semantic enhancement triple, where the positions of the subject, predicate, and object are represented by variable placeholders.

[0091] The fields in implicit semantic information refer to the semantic pattern type, the set of entity identifiers involved, the confidence level, and the data items of the suggested semantic type field contained in the implicit semantic information.

[0092] Variable placeholders refer to variable names in the enhanced rule template that begin with a question mark and are used to be replaced by specific entity identifiers or attribute values ​​when the rule is instantiated.

[0093] Semantic augmentation triples refer to RDF triples generated by replacing variable placeholders in augmentation rule templates with specific values. They are used to formally encode implicit semantic information into reasonable knowledge nodes and edges in the target semantic graph.

[0094] In this embodiment, based on the semantic pattern type in the implicit semantic information, an enhanced rule template matching the semantic pattern type is extracted from the enhanced rule template library. Entity identifiers are filled into the variable placeholders at the subject position in the enhanced rule template; the PDV standard class corresponding to the suggested semantic type field is filled into the variable placeholders at the type declaration position in the enhanced rule template; each entity identifier in the relevant entity identifier set is sequentially filled into the variable placeholders at the object position in the enhanced rule template; and the confidence level is filled into the variable placeholders for the confidence attribute in the enhanced rule template, ultimately generating a semantically enhanced triplet.

[0095] The step of "fusing and verifying the consistency between the semantically enhanced triples and the primary semantic graph to generate the target semantic graph" specifically includes the following steps: Step 411: Add the semantically enhanced triples to the primary semantic graph to obtain the fused semantic graph.

[0096] Step 412: Based on the class hierarchy axiom, attribute axiom, cardinality axiom, disjointness axiom, and functional axiom in the preset PDV standard ontology, perform consistency verification on each RDF node in the fused semantic graph to identify conflicting triples, remove the conflicting triples in the fused semantic graph, and obtain the target semantic graph.

[0097] In step 411 above, the fused semantic graph refers to the merged graph structure that contains RDF nodes and enhanced semantic nodes after adding semantic enhancement triples to the primary semantic graph. Enhanced semantic nodes refer to the newly added RDF nodes representing semantic entities in the semantic enhancement triples.

[0098] Conflicting triples refer to triples in a fused semantic graph that violate the predefined PDV standard ontology axioms and contain logical contradictions. They include three types: type conflict triples, attribute conflict triples, and relation conflict triples. Type conflict triples refer to triples in which the class declaration of an RDF node contradicts the class hierarchy axiom or disjointness axiom in the PDV standard ontology. Attribute conflict triples refer to triples in which the attribute value of an RDF node contradicts the attribute axiom or functional axiom in the PDV standard ontology. Relation conflict triples refer to triples in which the relationship between RDF nodes contradicts the cardinality axiom in the PDV standard ontology.

[0099] In this embodiment, semantically enhanced triples are added to the primary semantic graph, establishing connections between the semantically enhanced triples and existing RDF nodes and relation triples in the primary semantic graph, thus obtaining a fused semantic graph.

[0100] In step 412 above, the preset PDV standard ontology refers to the set of rules in the predefined PDV format standard used to constrain the logical relationships between RDF nodes, including five types of axioms: class hierarchy axioms, attribute axioms, cardinality axioms, disjointness axioms, and functional axioms.

[0101] The class hierarchy axiom refers to the rules defining the inheritance and hierarchical relationships between PDV standard classes in the PDV standard ontology. It is used to constrain the PDV standard class to which an RDF node belongs to to conform to the parent-child relationship between PDV standard classes. For example, if the pdv:Wall class is a subclass of the pdv:BuildingElement class, then an RDF node declared as type pdv:Wall belongs to the pdv:BuildingElement type.

[0102] Attribute axioms refer to the rules that define the value range, data type, and constraints of each attribute in the PDV standard ontology. They are used to constrain the attribute values ​​of RDF nodes to conform to the value range and data type defined for that attribute. For example, if the value range of the pdv:hasHeight attribute is positive real numbers, then a negative value would violate the attribute axiom.

[0103] The cardinality axiom refers to the rules in the PDV standard ontology that define the minimum and maximum number of occurrences of each object attribute. It is used to constrain the number of values ​​of the same object attribute between RDF nodes to be within a preset range. For example, if the minimum cardinality of the pdv:hasComponent attribute is 1, then an assembly RDF node must contain at least one component.

[0104] The disjoint axiom refers to the rule defined in the PDV standard ontology that two PDV standard classes cannot be declared as the class of the same RDF node at the same time. It is used to constrain an RDF node to belong to mutually exclusive PDV standard classes at the same time. For example, if the pdv:Wall class and the pdv:Slab class are disjoint, then the same RDF node cannot be declared as both of these PDV standard classes at the same time.

[0105] The functional axiom refers to the rule in the PDV standard ontology that an object attribute can only have one object for a given subject. It is used to constrain the uniqueness of the values ​​of certain object attributes of RDF nodes. For example, if the pdv:hasMaterial attribute is a functional attribute, then a part RDF node can only have one material attribute value.

[0106] In this embodiment of the application, based on the class hierarchy axiom in the preset PDV standard ontology, the type declaration triples of each RDF node in the fusion semantic graph are traversed, and it is checked whether the PDV standard classes declared by the RDF node conform to the parent-child class relationship between PDV standard classes. If the RDF node declares multiple PDV standard classes at the same time and there is no parent-child class relationship between these PDV standard classes, then the type declaration triples corresponding to the RDF node are determined to be conflicting triples.

[0107] Based on the disjoint axiom in the predefined PDV standard ontology, each RDF node in the fused semantic graph is traversed and checked to see if it is simultaneously declared as a mutually exclusive PDV standard class as defined in the disjoint axiom. If an RDF node is simultaneously declared as a mutually exclusive PDV standard class, then the type declaration triple corresponding to that RDF node is determined to be a conflicting triple. For example, if an RDF node is simultaneously declared as both pdv:Wall and pdv:Slab, and pdv:Wall and pdv:Slab are defined as mutually exclusive PDV standard classes in the disjoint axiom, then the type declaration triples :Wall_001 a pdv:Wall and :Wall_001 a pdv:Slab corresponding to that RDF node are determined to be conflicting triples.

[0108] Based on the attribute axioms in the predefined PDV standard ontology, the attribute values ​​of each RDF node in the fused semantic graph are traversed and checked to see if they satisfy the value range and data type constraints defined in the attribute axioms. If the attribute value of an RDF node is not within the value range defined in the attribute axioms or does not conform to the defined data type, the attribute triple corresponding to that RDF node is determined to be a conflicting triple. For example, if the pdv:hasHeight attribute value of an RDF node is -3.0, while the attribute axioms define the value range of the pdv:hasHeight attribute as positive real numbers, then the attribute triple: Wall_001 pdv:hasHeight -3.0 is determined to be a conflicting triple.

[0109] According to the cardinality axiom in the predefined PDV standard ontology, the occurrence count of the same object attribute in each RDF node is checked to see if it satisfies the minimum and maximum occurrence count constraints defined in the cardinality axiom. If the occurrence count of the same object attribute in an RDF node is less than the minimum cardinality or greater than the maximum cardinality, the relation triple corresponding to that object attribute is determined to be a conflicting triple. For example, if the pdv:hasComponent attribute of an assembly RDF node appears only 0 times, while the minimum cardinality of the pdv:hasComponent attribute is defined as 1 according to the cardinality axiom, then the assembly RDF node is determined to lack a necessary component relation, and the relation triple corresponding to the assembly RDF node is a conflicting triple.

[0110] Based on the functional axioms in the predefined PDV standard ontology, each RDF node's functional object attribute is traversed and checked to see if it has multiple different objects for the same subject. If an RDF node's functional object attribute has multiple different objects for the same subject, then all relation triples except the first relation triplet corresponding to that functional object attribute are considered conflicting triples. For example, if an RDF node's pdv:hasMaterial attribute points to two different objects, Material_A and Material_B, and the functional axioms define pdv:hasMaterial as a functional attribute, then the relation triplet corresponding to that RDF node, Part_001 pdv:hasMaterial :Material_B, is considered a conflicting triple.

[0111] The conflicting triples mentioned above are removed from the fused semantic graph, while the remaining triples are retained to obtain the target semantic graph. The target semantic graph refers to the semantic graph obtained after removing the conflicting triples from the fused semantic graph, which conforms to the predefined PDV standard ontology axioms. There are no logical contradictions between the RDF nodes and triples in this semantic graph, and it can be used for subsequent quality assessment and output.

[0112] This application embodiment formally encodes the implicit semantic insights of intelligent recognition into reasonable knowledge graph nodes and edges, and ensures the logical consistency of the enhanced knowledge system through consistency verification, avoiding the impact of logical contradictions introduced during the semantic enhancement process on the usability of the output PDV model.

[0113] In step 105 above, the step "perform quality assessment on the target semantic map and obtain quality assessment results" specifically includes the following steps: Step 501: Calculate semantic richness based on the number of RDF nodes, attribute triples, relation triples, and PDV classes corresponding to each RDF node in the target semantic graph.

[0114] Step 502: Perform semantic contradiction detection on the target semantic graph to obtain the semantic contradiction types and the corresponding number of contradiction types, so as to calculate the consistency score.

[0115] Step 503: Calculate the coverage based on the number of attribute triples and relation triples in the target semantic graph, and the total amount of information in the target dataset.

[0116] Step 504: Generate quality assessment results based on semantic richness, consistency score, and coverage.

[0117] In step 501 above, the number of PDV classes refers to the number of PDV standard classes used in the type declaration triples of each RDF node in the target semantic graph. This number is obtained by deduplicating the objects of the type declaration triples of all RDF nodes in the target semantic graph. Semantic richness is a quantitative indicator used to measure the knowledge density contained in the target semantic graph.

[0118] In this embodiment, the number of RDF nodes, attribute triples, relation triples, and PDV classes corresponding to each RDF node in the target semantic graph are statistically analyzed. The number of PDV classes is obtained by extracting the object of the type declaration triple of each RDF node in the target semantic graph (the PDV standard class declared by the RDF node) and deduplicating all PDV standard classes. Type diversity, relation density, and attribute fill rate are calculated respectively. Type diversity is the ratio of the number of PDV classes actually used to the total number of classes in the PDV ontology. Relation density is the ratio of the number of relation triples to the product of the number of RDF nodes and the average expected number of relations. Attribute fill rate is the ratio of the number of attribute triples to the product of the number of RDF nodes and the preset number of available attribute types. The average expected number of relations can be set according to the total number of object attribute types defined in the PDV standard ontology, and the preset number of available attribute types can be set according to the total number of data attribute types defined in the PDV standard ontology or the number of attribute key-value pairs contained in the attribute layer data in the target dataset.

[0119] The semantic richness is obtained by weighting and summing the above-mentioned type diversity, relation density and attribute filling rate according to the preset weight coefficients. The preset weight coefficients corresponding to the above calculation of semantic richness can be set according to the importance of each PDV standard class in the PDV standard ontology and the attention priority of each sub-indicator in the actual application scenario.

[0120] In step 502 above, semantic contradiction type refers to the types of semantic conflicts existing in the target semantic graph, including three types: type ambiguity, relation redundancy, and attribute contradiction. Type ambiguity refers to the situation where the same RDF node is declared as multiple similar categories; relation redundancy refers to the situation where the same semantic relation is repeatedly expressed through multiple different paths; attribute contradiction refers to the situation where the same attribute of the same RDF node has inconsistent values ​​in different positions; the number of contradiction types refers to the number of times each semantic contradiction type is detected in the target semantic graph; consistency score refers to a quantitative indicator used to measure the degree of logical harmony within the target semantic graph.

[0121] In this embodiment, all RDF nodes and triples in the target semantic graph are traversed to detect the existence of three semantic contradiction types: type ambiguity, relation redundancy, and attribute contradiction. The occurrence frequency of each semantic contradiction type is counted as the number of contradiction types. The number of contradiction types corresponding to each semantic contradiction type is weighted and summed with a preset weight to obtain the total conflict weight. Then, the ratio of 1 minus the total conflict weight to the maximum possible conflict weight is calculated to obtain the consistency score. The preset weights corresponding to each semantic contradiction type can be set according to the degree of influence of each semantic contradiction type on logical consistency. The maximum possible conflict weight refers to the conflict weight value calculated when all RDF nodes in the target semantic graph have type ambiguity contradictions, all relation triples have relation redundancy contradictions, and all attribute triples have attribute contradictions. It is calculated by weighting the number of RDF nodes, the number of relation triples, and the number of attribute triples according to the preset weights corresponding to each semantic contradiction type.

[0122] In step 503 above, the total amount of information in the target dataset refers to the total amount of information contained in each layer of data in the target dataset, including the number of attribute key-value pairs corresponding to each part in the attribute layer data and the number of assembly constraint relationships between each part in the relation layer data.

[0123] Coverage refers to a quantitative metric used to measure the proportion of attribute and relation information in the target dataset that has been successfully transformed and retained in the target semantic graph.

[0124] In this embodiment, the total number of attribute key-value pairs corresponding to each part in the attribute layer data and the total number of assembly constraint relationships between each part in the relation layer data are counted separately from the target dataset. The total number of these two types of totals is added together to obtain the total amount of information in the target dataset. The sum of the number of attribute triples and the number of relation triples in the target semantic graph is taken as the amount of information retained. The ratio of the amount of information retained to the total amount of information in the target dataset is calculated to obtain the coverage rate.

[0125] In step 504 above, the quality assessment result refers to the comprehensive quality score and corresponding processing method determination result obtained by weighting three indicators: semantic richness, consistency score, and coverage. The processing method determination result includes three types: automatic approval, alarm triggering, and conversion failure. Among them, automatic approval means that the target semantic graph can be directly processed for subsequent serialization output; alarm triggering requires pausing the automatic processing process and notifying manual review and correction; conversion failure requires reconfiguring the mapping rules or retraining the semantic recognition model.

[0126] In this embodiment, semantic richness, consistency score, and coverage are weighted and summed according to preset weight coefficients to obtain a comprehensive quality score. The preset weight coefficients corresponding to the calculation of the comprehensive quality score can be set according to the priority of semantic richness, consistency score, and coverage in actual application scenarios. This application does not limit the value of the preset weight coefficients. The comprehensive quality score is compared with preset judgment thresholds: when the comprehensive quality score reaches the preset pass threshold, it is judged as automatic pass; when the comprehensive quality score is in the preset alarm range, it is judged as triggering an alarm and entering the manual review and correction process; when the comprehensive quality score is lower than the preset failure threshold, it is judged as conversion failure. The preset judgment thresholds include three graded thresholds: preset pass threshold, preset alarm range, and preset failure threshold. Each graded threshold can be set according to the required level of conversion quality in the actual application scenario. The above comprehensive quality score and corresponding processing method judgment results are integrated into a quality assessment result.

[0127] This application implements a multi-dimensional quality quantification evaluation of the enhanced semantic graph. It comprehensively measures the knowledge density, logical harmony, and information retention of the output PDV model through three indicators: semantic richness, consistency score, and coverage, providing an objective basis for the hierarchical processing of the output results.

[0128] Regarding the step "Based on the quality assessment results, convert the target semantic graph into a PDV standard semantic file, and integrate the target data package and the PDV standard semantic file into a standard PDV model package": The processing is graded according to the processing method determination result in the quality assessment results. Specifically: when the processing method determination result is automatic pass, all RDF triples in the target semantic graph are output according to the serialization format specified by the PDV standard to generate a PDV standard semantic file; when the processing method determination result is alarm trigger, the target semantic graph and quality assessment results are sent to the manual review and correction process. After receiving the manually corrected semantic graph, it is re-serialized and output. The process involves generating a PDV standard semantic file. During this process, manual correction operations are recorded, including adding missing triples, deleting incorrect triples, and correcting type declarations. The target semantic map before correction is used as a negative sample, and the target semantic map after manual correction is used as a positive sample to form paired training data. When the number of accumulated correction records reaches a preset number, the semantic recognition model is fine-tuned and updated using the paired training data. After each model update, the new model is compared and tested in an isolated environment. Once the quality of the new model is confirmed to be no lower than that of the currently used model, it is deployed online to achieve a closed-loop iteration in which the conversion quality and semantic recognition capability continuously improve with use.

[0129] When the processing method determines that the conversion has failed, it is marked as needing to reconfigure the preset mapping rule set or retrain the semantic recognition model. After reconfiguration or retraining, steps 102-104 are executed again until a target semantic graph that has passed the quality assessment is obtained. The RDF triples in the target semantic graph that has passed the quality assessment are output according to the serialization format specified by the PDV standard to generate a PDV standard semantic file.

[0130] The aforementioned PDV standard semantic files and target data packages are organized according to a preset directory structure, so that the RDF nodes in the PDV standard semantic files and the geometric data in the target data packages establish a reference relationship through node IDs, generating a standard PDV model package. This achieves hierarchical output and standardized encapsulation of the enhanced semantic graph, ensuring that the final PDV model package is semantically rich, logically consistent, and visually interactive. It provides a high-quality product knowledge package with standardized format for downstream applications such as intelligent question answering, design assistance, and process generation. At the same time, the quality assessment results can be continuously fed back to the semantic recognition model for iterative optimization.

[0131] The PDV standard semantic file refers to a semantic data file that conforms to the PDV standard specification, generated by outputting the RDF triples in the target semantic graph according to the serialization format specified by the PDV standard. The file is stored in the RDF triple serialization format and contains all RDF nodes in the target semantic graph and their attribute triples and relation triples.

[0132] A standard PDV model package refers to a standardized product digital model package containing semantic data and visualization data, generated by organizing PDV standard semantic files and target data packages according to a preset directory structure. In this model package, RDF nodes in the PDV standard semantic files and geometric data in the target data packages establish a reference relationship through node IDs, enabling the semantic information and visualization information to be associated at the data level. The aforementioned preset directory structure can be configured according to the interface specifications of downstream applications; this application does not limit the specific form of the preset directory structure.

[0133] Figure 2 This is a schematic diagram illustrating a specific implementation of a semantic enhancement conversion system for heterogeneous industrial data to PDV format provided in this application embodiment. (Refer to...) Figure 2 The system may include: The acquisition module 21 is used to acquire multiple heterogeneous industrial data source files and convert the heterogeneous industrial data in each heterogeneous industrial data source file into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data. The conversion module 22 is used to convert the target dataset into a primary semantic graph based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, and generate a target data package associated with the primary semantic graph. The identification module 23 is used to perform feature analysis and semantic recognition on the primary semantic graph through a semantic recognition model to obtain implicit semantic information; The generation module 24 is used to perform semantic enhancement processing on the primary semantic graph based on the implicit semantic information, generate semantic enhancement triples, fuse the semantic enhancement triples with the primary semantic graph and perform consistency verification to generate the target semantic graph. Evaluation module 25 is used to perform quality evaluation on the target semantic graph, obtain quality evaluation results, convert the target semantic graph into a PDV standard semantic file based on the quality evaluation results, and integrate the target data packet and the PDV standard semantic file into a standard PDV model packet.

[0134] This application provides a semantic enhancement conversion system for heterogeneous industrial data to PDV format, which is used to implement the aforementioned semantic enhancement conversion method for heterogeneous industrial data to PDV format. Therefore, the specific implementation of the semantic enhancement conversion system for heterogeneous industrial data to PDV format can be found in the embodiment section of the semantic enhancement conversion method for heterogeneous industrial data to PDV format above. The specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.

[0135] This application also provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the semantic enhancement conversion method for heterogeneous industrial data to PDV format described above.

[0136] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, implements the steps of any of the above-described methods for semantically enhanced conversion of heterogeneous industrial data to PDV format.

[0137] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.

[0138] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements any of the above-described semantic enhancement conversion methods for heterogeneous industrial data to PDV format.

[0139] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0140] Obviously, those skilled in the art should understand that the various units or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps into a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0141] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A semantically enhanced conversion method for heterogeneous industrial data to PDV format, characterized in that, include: Multiple heterogeneous industrial data source files are acquired, and the heterogeneous industrial data in each heterogeneous industrial data source file is converted into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data. Based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, the target dataset is converted into a primary semantic graph, and a target data package associated with the primary semantic graph is generated. The implicit semantic information is obtained by performing feature analysis and semantic recognition on the primary semantic graph using a semantic recognition model. Based on the implicit semantic information, the primary semantic graph is semantically enhanced to generate semantically enhanced triples. The semantically enhanced triples are then fused with the primary semantic graph and a consistency check is performed to generate the target semantic graph. The target semantic graph is subjected to quality assessment to obtain quality assessment results. Based on the quality assessment results, the target semantic graph is converted into a PDV standard semantic file. The target data packet and the PDV standard semantic file are integrated into a standard PDV model packet.

2. The method according to claim 1, characterized in that, Convert heterogeneous industrial data from various heterogeneous industrial data source files into a target dataset in a preset format, including: Load the corresponding adapter plugin based on the file extension of the heterogeneous industrial data source file or the data format specified by the user. The adapter plugin parses the heterogeneous industrial data source file to obtain assembly tree hierarchical relationship data, geometric expression data of each part, geometric semantic feature information, business attribute key-value pair data, and assembly constraint relationship data between each part. According to the preset format, the assembly tree hierarchical relationship data, the geometric expression data, the geometric semantic feature information, the business attribute key-value pair data, and the assembly constraint relationship data are respectively converted into structural layer data, geometric layer data, semantic layer data, attribute layer data, and relationship layer data in the target dataset.

3. The method according to claim 1, characterized in that, Based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, the target dataset is converted into a primary semantic graph, including: Obtain the type field value of each node in the structure layer data, and extract at least one structure mapping rule, relationship mapping rule, and attribute mapping rule that match the type field value of each node from the preset mapping rule set; Based on the structure mapping rules corresponding to each node, each node is mapped to an RDF node conforming to the PDV format; Based on the attribute mapping rules corresponding to each node, the attribute values ​​corresponding to each RDF node are extracted from the attribute layer data, and attribute triples are generated based on each RDF node and its corresponding attribute values. Based on the relation mapping rules corresponding to each node, the assembly constraint relationship corresponding to each RDF node is extracted from the relation layer data, and a relation triplet is generated based on each RDF node and its corresponding assembly constraint relationship. Based on the node IDs of each part, each RDF node, the attribute triples, and the relation triples in the structural layer data, a primary semantic graph is constructed.

4. The method according to claim 1, characterized in that, The primary semantic graph is subjected to feature analysis and semantic recognition using a semantic recognition model to obtain implicit semantic information, including: Based on the primary semantic graph, an input graph structure is constructed, and each graph node in the input graph structure is mapped to a node embedding vector. The node embedding vectors of each graph node and the node embedding vectors of their corresponding neighbor nodes are aggregated to generate the context semantic vectors of each graph node. Based on the context semantic vector, a graph analysis task is performed on each graph node to obtain implicit semantic information. The graph analysis task includes at least one of node classification, link prediction, or subgraph identification.

5. The method according to claim 1, characterized in that, Based on the implicit semantic information, the primary semantic graph is semantically enhanced to generate semantically enhanced triples, including: According to the preset identifier generation rules, assign corresponding entity identifiers to the values ​​of the suggested semantic type fields in the implicit semantic information; Extract the enhanced rule template that matches the implicit semantic information from the pre-stored enhanced rule template library, and fill the corresponding variable placeholders in the enhanced rule template with the values ​​of the entity identifier and each field in the implicit semantic information to generate a semantically enhanced triplet.

6. The method according to claim 1, characterized in that, The semantically enhanced triples are fused with the primary semantic graph and a consistency check is performed to generate the target semantic graph, including: The semantically enhanced triples are added to the primary semantic graph to obtain the fused semantic graph; Based on the class hierarchy axiom, attribute axiom, cardinality axiom, disjointness axiom, and functional axiom in the preset PDV standard ontology, consistency checks are performed on each RDF node in the fused semantic graph to identify conflicting triples. Conflicting triples in the fused semantic graph are then removed to obtain the target semantic graph.

7. The method according to claim 1, characterized in that, The target semantic graph is subjected to quality assessment, and the quality assessment results are obtained, including: The semantic richness is calculated based on the number of RDF nodes, attribute triples, relation triples, and PDV classes corresponding to each RDF node in the target semantic graph. Semantic contradiction detection is performed on the target semantic graph to obtain the semantic contradiction types and the corresponding number of contradiction types, so as to calculate the consistency score; The coverage rate is calculated based on the number of attribute triples and relation triples in the target semantic graph and the total amount of information in the target dataset. Based on the semantic richness, the consistency score, and the coverage, a quality assessment result is generated.

8. A semantically enhanced conversion system for heterogeneous industrial data to PDV format, characterized in that, include: The acquisition module is used to acquire multiple heterogeneous industrial data source files and convert the heterogeneous industrial data in each heterogeneous industrial data source file into a target dataset in a preset format. The target dataset includes structural layer data, geometric layer data, semantic layer data, attribute layer data, and relation layer data. The conversion module is used to convert the target dataset into a primary semantic graph based on a preset mapping rule set corresponding to the data format of the heterogeneous industrial data source file, and generate a target data package associated with the primary semantic graph. The recognition module is used to perform feature analysis and semantic recognition on the primary semantic graph through a semantic recognition model to obtain implicit semantic information. The generation module is used to perform semantic enhancement processing on the primary semantic graph based on the implicit semantic information, generate semantic enhancement triples, fuse the semantic enhancement triples with the primary semantic graph and perform consistency verification to generate the target semantic graph. The evaluation module is used to perform quality evaluation on the target semantic graph, obtain quality evaluation results, convert the target semantic graph into a PDV standard semantic file based on the quality evaluation results, and integrate the target data packet and the PDV standard semantic file into a standard PDV model packet.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the semantically enhanced conversion method for heterogeneous industrial data to PDV format as described in any one of claims 1-7.

10. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform a semantic enhancement conversion method for heterogeneous industrial data to PDV format as described in any one of claims 1-7.