Method, apparatus, device, and medium for identifying building block

The use of Large Language Models for semantic alignment and building block identification in information models addresses inefficiencies in existing methods, enabling efficient integration and development of domain-specific sub-models for improved interoperability.

WO2026064987A1PCT designated stage Publication Date: 2026-04-02SIEMENS AG +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for developing information models across different domains are inefficient and require significant manual effort, leading to scalability issues and difficulty in extending and tailoring heterogeneous information models to fit individual business needs.

Method used

Utilizing Large Language Models (LLM) for semantic alignment and identifying building blocks across multiple information models, translating terms with the same semantics into unified terms, and converting these models into a common structure to facilitate interoperability and sub-model generation.

Benefits of technology

This approach enables efficient and time-saving integration of heterogeneous information models, accelerating new business startups and easing engineering processes by providing a more efficient method for data integration and model development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024121132_02042026_PF_FP_ABST
    Figure CN2024121132_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, device, and medium for identifying building block. The method comprising: obtaining multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term; translating terms with the same semantics in the multiple information models into a predetermined unified term; and identifying a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models. Embodiments of the present disclosure provide an easier and more efficient way to identify building blocks which can be converted into sub-models from heterogeneous information models for domain-specific usage. This can accelerate new business startup for the customers and ease the engineering process for data integration.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device, and medium for identifying building blockFIELD

[0001] The present disclosure relates to the technical field of information technology, in particular to a method, apparatus, device, and medium for identifying building block.BACKGROUND

[0002] Information model, also known as conceptual data model, is configured to model data and information according to user's point of view. It is an abstraction from the real world to the information world. It emphasizes its semantic expression function and is easy for users to understand. It is also a language for communication between users and database designers and can be used for database design.

[0003] Information model is a method used to define the general representation of information. It is the basis of object-oriented analysis. The basic idea of information model is to describe three contents: objects, object attributes, and the relationship between objects. There are certain relationships between objects, which are expressed in the form of attributes. The information model is described in two basic forms: one is text description, including the description and explanation of all objects and relationships in the system; The other is graphical representation, which provides a global perspective, considering the coherence, completeness and consistency of the system. By using the information model, we can use different applications to reuse, change and share the managed data. The significance of using information models lies not only in the modeling of objects, but also in the description of the correlation between objects.SUMMARY

[0004] Embodiments of the present disclosure propose a method, apparatus, device, and medium for identifying building block.

[0005] In a first aspect, a method for identifying building block is provided. The method includes:

[0006] obtaining multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term;

[0007] translating terms with the same semantics in the multiple information models into a predetermined unified term; and

[0008] identifying a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.

[0009] In a second aspect, an apparatus for identifying building block is provided. The apparatus includes:

[0010] an obtaining module, configured to obtain multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term;

[0011] a translating module, configured to translate terms with the same semantics in the multiple information models into a predetermined unified term; and

[0012] an identifying module, configured to identify a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.

[0013] In a third aspect, an electronic device is provided. The electronic device comprising a processor and a memory, wherein an application program executable by the processor is stored in the memory for causing the processor to execute a method for identifying building block as described in any of the above.

[0014] In a fourth aspect, a computer-readable medium comprising computer-readable instructions stored thereon is provided, wherein the computer-readable instructions for executing a method for identifying building block as described in any of the above.

[0015] In a fifth aspect, a computer program product comprising a computer program, when the computer program is executed by a processor for executing a method for identifying building block as described in any of the above.

[0016] According to the above technical solutions, obtaining multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term; translating terms with the same semantics in the multiple information models, into a predetermined unified term; and identifying a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the multiple information models. Therefore, embodiments of the present disclosure provide an easier and more efficient way to identify building blocks which can be converted into sub-models for domain-specific business usage from heterogeneous information models. Much higher efficiency and time saving for development of information models to integrate heterogeneous information models including standard or non-standard ones. This can accelerate new business startup for the customers and ease the engineering process for data integration.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To make technical solutions of examples of the present disclosure clearer, accompanying drawings to be used in description of the examples will be simply introduced hereinafter. Obviously, the accompanying drawings to be  described hereinafter are only some examples of the present disclosure. Those skilled in the art may obtain other drawings according to these accompanying drawings without creative labor.

[0018] Fig. 1 is an exemplary flowchart of a method for identifying building block according to an embodiment of the present disclosure.

[0019] Fig. 2 is an exemplary schematic diagram of a process for identifying building block according to an embodiment of the present disclosure.

[0020] Fig. 3 is an exemplary schematic diagram of the process of generating sub-model based on LLM according to embodiments of the present disclosure.

[0021] Fig. 4 is an exemplary structural diagram of an apparatus for identifying building block according to an embodiment of the present disclosure.

[0022] Fig. 5 is an exemplary structural diagram of an electronic device according to an embodiment of the present disclosure.

[0023] List of reference numbers: DETAILED DESCRIPTION

[0024] To make the purpose, technical scheme, and advantages of the disclosure clearer, the following examples are given to further explain the disclosure in detail. Nouns and pronouns related to people in this patent application are not limited to specific gender.

[0025] To be concise and intuitive in description, the scheme of the disclosure is described below by describing several representative embodiments. Many details in the embodiments are only used to help understand the scheme of the disclosure. However, it is obvious that the technical scheme of the disclosure can be realized without being limited to these details. To avoid unnecessarily blurring the scheme of the disclosure, some embodiments are not described in detail, but only the framework is given. Hereinafter, "including" refers to "including but not limited to" , "according to. . . " refers to "at least according to. . ., but not limited to. . . " . When the number of an element is not specifically indicated below, it means that the element can be one or more, or can be understood as at least one.

[0026] Standard information model like Asset Administration Shell (AAS) enables interoperability across companies, physical entities (equipment, devices, sensors, etc. ) , across product lifecycles, across engineering lifecycles, across business lifecycles, etc. However, developing a new standard information model requires extraction of semantic information from broad information models in the market like industrial standard models, national or international models, non-standard models, etc. Secondly, industry consists of countless domains which have overlap and differences among each other. An information model will not be sustainable and scalable if huge manual work and long audit procedure are required for each domain. Therefore, a more efficient method is expected to ease the development of information model for different domains. Moreover, an information model as a template will not fit all kinds of situations. Therefore, fast extension and tailoring of the information model will be necessary. However, uncountable terms from different information models are individually named and defined with different semantic  definitions by different stakeholders. This makes it difficult to make extension and tailoring of heterogeneous information models to fit individual business.

[0027] In the prior art, there are no efficient method to construct information models for different domains. The models are usually manually created and published by investigation of market demands with long audit procedure.

[0028] With the emerging of LLM (Large Language Model) , machine can understand human natural language much better than ever. Embodiment of the present disclosure take the advantage of LLM for semantical alignment of heterogeneous information models and common Sub-model generation from the heterogeneous information models.

[0029] Fig. 1 is an exemplary flowchart of a method for identifying building block according to an embodiment of the present disclosure. As shown in Figure, the method includes:

[0030] Step 101: obtaining multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term.

[0031] Here, multiple information models can be implemented as multiple heterogeneous information models containing multiple model types, such as standard information models or non-standard information models. For example, multiple heterogeneous information models may include industrial standard information models, national or international information models, non-standard information models, etc. Each information model may contain one or more pairs, and each pair contains respective terms and respective semantic descriptions of the respective terms.

[0032] Step 102: translating terms with the same semantics in the multiple information models into a predetermined unified term.

[0033] Here, terms with the same semantics in the information models are translated into a unified term.

[0034] For example, term "automatic robotic arm" in Model A, term "robot" in Model B, and term "robotic equipment" in Model C. Based on semantic description of the term "automatic robotic arm" in Model A, semantic description of the term "robot" in Model B, and semantic description of the term "robot equipment" in Model C, it is determined that the terms "robotic arm" , "robot" , and "robot equipment" have the same semantics. Therefore, translate term "robotic arm" in Model A, term "robot" in Model B, and term "robotic equipment" in Model C into a unified term, such as "robotic arm" .

[0035] For example, term "PLC" in Model A, term "PLC equipment" in Model B, and term "PLC node" in Model C. Based on semantic description of the term "PLC" in Model A, semantic description of the term "PLC equipment" in Model B, and semantic description of the term "PLC node" in Model C, it is determined that the terms "PLC" , "PLC equipment" , and "PLC node" have the same semantics. Therefore, translate term "PLC" in Model A, term "PLC equipment" in Model B, and term "PLC node" in Model C into a unified term, such as "PLC" .

[0036] In one embodiment, the translating terms with the same semantics in the multiple information models into a predetermined unified term includes: receiving a user vocabulary comprising the unified term and a semantic definition of the unified term; calculating respective similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term; determining respective target semantic definitions with respective similarities greater than a predetermined threshold; translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.

[0037] The calculation methods for semantic similarity mainly include traditional methods and deep learning methods. Traditional methods mainly focus on the lexical level, using algorithms such as TF-IDF to solve the matching problem at the lexical level. Deep learning methods focus more on semantic matching by constructing deep text matching models. These models can be divided into two types: representational and interactive. The representational model focuses more on building the representation layer and adopts a twin network structure, allowing the twin towers to share parameters and map two sentences to a vector space, thereby achieving overall matching at the sentence level. This method can better capture the semantic relationships between sentences and provide more accurate semantic similarity calculation. In deep learning methods, DSSM (Deep Structured Semantic Models) is an important model that uses massive click exposure logs of queries and titles in search engines, expresses queries and titles as low latitude semantic vectors using DNN, and calculates the distance between the two semantic vectors through cosine distance to ultimately train a semantic similarity model. DSSM can not only predict the semantic similarity between two sentences, but also obtain the low dimensional semantic vector expression of a certain sentence. For example, similarity can be determined through Euclidean distance (L2 norm) , Manhattan distance (L1 norm) , and Minkowski distance.

[0038] In one embodiment, the calculating similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term includes: inputting the multiple information models into a first LLM; inputting the user vocabulary into the first LLM; enabling the first LLM to calculate respective similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term; the determining respective target semantic definitions with respective similarities greater than a predetermined threshold comprises: enabling the first LLM to determine respective target semantic definitions with respective similarities greater than the threshold; the method comprises: receiving a term translation list from the first LLM, and the term translation list comprises the unified term and respective terms in respective pairs which comprise the respective target semantic definitions; the translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term comprises: translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term, based on  the term translation list.

[0039] LLM is a deep learning model trained on a large amount of textual data, capable of generating natural language text or understanding the meaning of language text. Large language models are typically based on Transformer architecture, using pre trained objectives such as Language Modeling to improve performance by increasing model size, training data, and computational resources. Large language models can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and are an important pathway towards artificial intelligence. LLM may include: Claude 2; LLaMA; GPT-3; ChatGPT, and so on.

[0040] Step 103: identifying a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.

[0041] A building block is a repeatable data structure in multiple information models that contains a unified term and common attributes related to the unified term in the multiple information models. Repeatable data structures of all information models can be found out, as all the information models are translated to the same set of terms, they can be compared with each other.

[0042] To implement the building block identification, all information models can be transformed to graph data to provide extra structural information with nodes representing all terms and edges representing all relationships between all terms.

[0043] For example, building block {,, Robot Device “: ,, DataTypeRobotType “} , {,, Mission “: [,, grasping “, ,, loading “, ,, unloading “] } can be identified with compare of model 1 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} and model 2 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} .

[0044] In one embodiment, the method includes: confirming correctness of the term translation list in response a confirmation instruction triggered by a user, before the translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.

[0045] In one embodiment, the method includes: determining a target schema; converting the building block into a sub-model which conforms to the target schema.

[0046] Schema is a structured information model that describes a collection of information and the relationships between this information. This abstract representation can help developers manage, query, and modify data more conveniently. A schema consists of the following elements: data type, column names, and constraints.

[0047] For example, the target schema can be JSON schema, XML schema, and so on.

[0048] JSON Schema is a specification used to describe JSON data, which can be used to define the structure, format, and constraints of JSON data objects. Through JSON Schema, JSON data can be validated, verified, and documented to ensure its correctness and integrity.

[0049] XML Scherma is an XML language used to describe and approximate XML documents. In terms of functionality, it is very similar to DTD and is used to define and describe the structure and content schema of XML documents. An XML Schema document is a text file with an extension of ". xsd" that follows the syntax rules of XML. W3C regulations require that the root element of an XML Schema document must be "schema" and the namespace must be: “https:  / / www. w3. org / 2001 / XMLSchema” . All content is added to the root tag<schema>, where 'xsd'is the prefix of the namespace and can be defined arbitrarily, usually set as'xsd 'or'xs'. There are two attributes in the<schema>declaration: the name attribute and the XMLNS attribute. The name attribute specifies the name of Schema, which can be omitted. The XMLNS attribute specifies the namespace of the schema document. XSD documents can define the following content of XML: defining elements that can appear in the document; Define the attributes that can appear in the document; Define which element is a child element; Define the order of sub elements; Define the number of sub elements; Define whether an element is empty or can contain text; Define the data types of elements and attributes; Define default and fixed values for elements and attributes.

[0050] The main components of XML Schema comprise:

[0051] (1) Type

[0052] Simple type: an element that does not contain any child elements or attributes, but only contains textual content.

[0053] Complex type: an element that contains child elements or attributes.

[0054] (2) Elements

[0055] <element name= "element name" type= "data type" minOccurs= "int" maxOccurs= "int"  / >

[0056] The name attribute indicates the name of the XML element. The type attribute indicates the data type of the XML element, which can be selected from built-in XML data types or user-defined data types. The minIncidents attribute indicates the minimum number of occurrences of an XML element, with a minimum value of 0, and is an optional attribute. The maxEvents attribute indicates the maximum number of occurrences of an XML element, with a minimum value of 1 and a maximum value of unbounded, indicating infinite occurrences. It is an optional attribute.

[0057] (3) Attributes

[0058] Only elements of complex types can have attributes; Elements can have simple or complex types, while attributes can only have simple types.

[0059] <element name= "element_name" type= "dataType"  / >

[0060] <xsd: complexType name= "dataType" >

[0061] <xsd: attribute name= "attribute_name" type= "simple_type" use= "use_method"

[0062] default= "value" fixed= "value" >

[0063] < / xsd: attribute>

[0064] < / xsd: complexType>

[0065] Element_name refers to the name of the corresponding element in the XML file. Attributename refers to the name of an attribute. Simple_Type refers to the data type of an attribute, which can be a built-in data type or a custom data type defined by the simple type element. Use_sthod indicates the actual value requirements for attributes in XMD elements, which can be optional required、prohibited. Among them, optional means that the attribute value is optional and is the default value: ; Required means that the attribute value must exist, and this attribute value must appear at least once; Prohibited means that the attribute value cannot appear and is used to restrict the use of the attribute in the restriction element. Default refers to the default value of a property. Fixed means that if a property exists, its content can only be the value specified by this property and cannot be changed.

[0066] (4) Group definition

[0067] Element group is the grouping of several elements together. The element group must be a direct child element of the schema root element. If other types of elements need to have element groups as child elements, they must be implemented by referencing ref form.

[0068] (5) Annotation

[0069] To facilitate reading and understanding of XML Schema documents, it is necessary to add comment statements to explain the relevant content. XML Schema supports the use of<! -------> Annotation method. In addition, another dedicated<annotation / >element is provided to add comments, which has better readability and can also be read by other applications. The<annotation / >element contains two child elements: <documentation / >: This element mainly stores information suitable for reading. <appinfo / >: This element mainly stores information for other applications. Any number of<documentation / >and<appinfo / >child elements can appear in the<annotation / >element without any order requirement.

[0070] In one embodiment, the converting the building block into a sub-model which conforms to the target schema includes: inputting the target schema and the building block into a second LLM; enabling the second LLM to convert the building block into a sub-model which conforms to the target schema.

[0071] In one embodiment, the method includes: replacing respective structures of the unified term in the multiple information models with the sub-model.

[0072] For example, building block {,, Robot Device “: ,, DataTypeRobotType “} , {,, Mission “: [,, grasping “, ,, loading “, ,, unloading “] } can be identified with compare of model 1 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “:  [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} and model 2 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} . After the building block is converted into a sub-model with a XML Scherma, data structure of {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} of model 1 is replaced with the sub-model and data structure of {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} of model 2 is replaced with the sub-model. Model 1 and model 2 can achieve interoperability after replacing respective data structures in Model 1 and Model 2 with the sub-model converted by a common building block of model 1 and model 2.

[0073] In one embodiment, the method includes: obtaining a new information model that does not belong to the multiple information models, the new information model comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term; determining a target term in the new information model that has the same semantics as the unified term; replacing structure of the unified term in the new information model with the sub-model.

[0074] For example, building block {,, Robot Device “: ,, DataTypeRobotType “} , {,, Mission “: [,, grasping “, ,, loading “, ,, unloading “] } can be identified with compare of model 1 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} and model 2 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} .

[0075] After the building block is converted into a sub-model with a XML Scherma, data structure of {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} of model 1 is replaced with the sub-model and data structure of {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} of model 2 is replaced with the sub-model.

[0076] In a new model 3, data structure {,, Robot node ": ,, DataTypeRobotType" , ,, Mission ": [,, grinding" , ,, detection ", ,, loading" , ,, unloading "] are included. It is determined that “Robot node” is a target term that has the same semantics as the unified term “Robot Device” . Data structure {,, Robot node " : ,, DataTypeRobotType" , ,, Mission ": [,, grinding", ,, detection ", ,, loading" , ,, unloading " ] is replaced in model 3 with the sub-model.

[0077] Therefore, after replacing respective data structures in model 1, model 2 and model 3 with the sub-model converted by a common building block of model 1, model 2 and model 3 can achieve interoperability.

[0078] Fig. 2 is an exemplary schematic diagram of a process for identifying building block according to an  embodiment of the present disclosure. The process includes: user input information models for the construction of the common sub-model. With user defined prompts and vocabulary, LLM translates each original information model with common terms. For each translation, the user can verify the output and make LLM rework in case the translation is not fully correctly done. Then the user gets all information models aligned semantically to the same vocabulary. The aligned information models are compared with each other to identify common building blocks. With user defined prompts and schema for common sub-model, LLM constructs the common building blocks in the user-desired way. The user can verify the output and make LLM rework in case the construction is not fully correctly done. Then the user gets the common sub-model from the heterogeneous information models.

[0079] As shown in Figure 2, the process includes:

[0080] Step 20: Input multiple information models. Next, execute the process of semantic alignment of the models 30.

[0081] The process of semantic alignment of the models 30 includes:

[0082] Step 21: Based on a user vocabulary that includes respective unified terms and respective semantic descriptions of the respective unified terms, translate terms with the same semantics in the multiple information models into a corresponding unified term in the user vocabulary.

[0083] Step 22: Verify translation results by user.

[0084] Step 23: Check if the translation results are accurate. If so (corresponding to the "Y" branch) , proceed to Step 24 and its subsequent steps; Otherwise (corresponding to the "N" branch) , return to Step 21.

[0085] Then, complete the process of semantic alignment of the models 30 and proceed to Step 24.

[0086] Step 24: Identify common building blocks from translated multiple information models. Next, execute common sub-model generation 40.

[0087] Common sub-model generation 40 includes:

[0088] Step 25: LLM converts each building block into its own sub-modules or a unified sub-model that includes all building blocks, based on inputted schema.

[0089] Step 26: User validates the sub-models or unified sub-model.

[0090] Step 27: Check if the sub-models or unified sub-model are accurate. If so (corresponding to the "Y" branch) , proceed to Step 28 and its subsequent steps; Otherwise (corresponding to the "N" branch) , return to step 25.

[0091] Then, complete common sub-model generation 40 and proceed to step 28.

[0092] Step 28: Output the sub-models or unified sub-model.

[0093] Figure 3 is an exemplary schematic diagram of the process of generating sub-model based on LLM according to embodiments of the present disclosure.

[0094] As shown in Figure 3, the available information models 50 are outputted as they are as ,,Raw Models “to be  translated. The information models 50 can be composed in different formats like JSON, XML, etc. For example, the information models 50 includes non-standard models 51, industrial standard models 52, national standard models 53 and international standard models 54.

[0095] The available information models 50 are also outputted as a list of pairs of { “term” : “semantic definition” } . Each term and its semantic definition are organized in { “key (term) ” : “value (definition) ” } style, e.g., { “Robot” : “machine with components of robot or robot arm” } .

[0096] A term is a short textual identifier in natural language or formatted human-understandable code name to represent the semantic definition that is defined in the documentations, e.g., ,, PLC “has the meaning of programmable logic controller, and ,, rotation speed “has the meaning of the speed of rotation of an object.

[0097] With vocabulary 56 defined by user 70 and prompts 57 for translation, The first LLM 58 outputs a list of the pair of each term and its translation, namely ,, Term-Translation List “59. User vocabulary 56 is user-defined list of terms and their semantic definition to be used. For example, {,, Robotic Device “: ,, Automated machine with components of robot or robot arm “} .

[0098] Translation prompts 57 are designed structural language to make first LLM 58 understand the task to be done.

[0099] For example:

[0100] {Term-Definition-List} is what we get from available information models 50, e.g., { “PLC” : “programmable logic controller” } . {User-Defined-Vocabulary} is for example {,, Robotic Device “: ,, Automated machine with components of robot or robot arm “} . {Input-Information-Model} is for example {\ "Robot\" : \ "DataTypeRobotType\" } . {Target-Format} is for example ,, JSON “. User 70 check if the Term-Definition-List is accurate.

[0101] As output translated model 61, for example, ,, Robot “and ,, Robot Arm “are translated to user defined,, Robotic Device “in form {,, Robot “: ,, Robotic Device (Vocabulary) “, ,, Robot Arm “: ,, Robotic Device “} . With the text swapper 60, the terms in raw models are all replaced with translated terms. For example, raw model {,, Robot “: DataTypeRobotType} and {,, Robot Arm “: DataTypeRobotType} are all translated to {,, Robotic Device “: DataTypeRobotType} .

[0102] With the building block identifier 62, the repeatable structures, namely the building blocks, can be found out, as all the information models are translated to the same set of terms, they can be compared with each other. To implement the building block identifier 62, all information models can be transformed to graph data to provide extra structural information with nodes representing all terms and edges representing all relationships between all terms.

[0103] For example, building block {,, Robot Device “: ,, DataTypeRobotType “} , {,, Mission “: [,, grasping “, ,, loading “, ,, unloading “] } can be identified with compare of model 1 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, detection “, ,, loading “, ,, unloading “] , ,, Feature “: ,, transportation “} and model 2 {,, Robot Device “: ,, DataTypeRobotType “, ,, Mission “: [,, grasping “, ,, loading “, ,, unloading “, ,, weighting “] , ,, Configuration “: ,, parallel “} .

[0104] With user defined schema 64 with prompts 65 for model construction, The second LLM 66 outputs common sub-model 67. User 70 check if the sub-model 67 is accurate. The schema 64 is the data format to organize metamodel information, it can be JSON-Schema, XML-Schema, etc., For example:

[0105] Model construction prompts are designed structural language to make LLM understand the task to be done, for example:

[0106] Common sub-model 67 is the result of the system. It follows the user-defined format, for example:

[0107] Embodiments of the present disclosure provide an easier and more efficient way to find common sub-models from heterogeneous information models for domain-specific business usage. This can be a supplement for Anchor Model. This may accelerate the integration of data models from different product  / production  / application lifecycle and for different engineering domains. This also works for Products like Jupiter OIE, AX, TIA Portal, Teamcenter, Tecnomatix, SINUMERIK, etc. Much higher efficiency and time saving for development of information models to integrate heterogeneous information models including standard or non-standard ones. This can accelerate new business startup for the customers and ease the engineering process for data integration. The usage of LLM can also work for other customization purposes.

[0108] Fig. 4 is an exemplary structural diagram of an apparatus for identifying building block according to an embodiment of the present disclosure. As shown in Figure 4, apparatus 400 for identifying building block comprises: an obtaining module 401, configured to obtain multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term; a translating module 402, configured to translate terms with the same semantics in the multiple information models into a predetermined unified term; and an identifying module 403, configured to identify a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.

[0109] In one embodiment, the translating module 402 is configured to receive a user vocabulary comprising the unified term and a semantic definition of the unified term, calculate respective similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term, determine respective target semantic definitions with respective similarities greater than a predetermined threshold, and translate respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.

[0110] In one embodiment, the apparatus 400 comprising: a converting module 404, configured to determine a target schema, and convert the building block into a sub-model which conforms to the target schema.

[0111] In one embodiment, the apparatus 400 comprising: a replacing module 405, configured to replace respective structures of the unified term in the multiple information models with the sub-model.

[0112] In one embodiment, the apparatus 400 comprising: a replacing module 405, configured to obtain a new information model that does not belong to the multiple information models, the new information model comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term, determine a target term in the new information model that has the same semantics as the unified term, and replace structure of the unified term in the new information model with the sub-model.

[0113] Embodiments of the present disclosure also propose an electronic device with a processor memory architecture. Fig. 5 is an exemplary structural diagram of an electronic device according to an embodiment of the present disclosure. As shown in Figure 5, electronic device 500 includes a processor 501, a memory 502, and a computer program stored on memory 502 that can run on processor 501. When the computer program is executed by processor 501, the method for identifying building block as described in either of the above is implemented. Among them, memory 502 can be implemented as various storage media such as electrically erasable programmable read-only memory (EEPROM) , flash memory, programmable program read-only memory (PROM) , etc. Processor 501 can be implemented to include one or more central processors or one or more field programmable gate arrays, wherein the field programmable gate array integrates one or more central processor cores. Specifically, the central processing unit or core can be implemented as a CPU, MCU, DSP, and so on.

[0114] It should be noted that not all steps and modules in the above processes and structural diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution sequence of each step is not fixed and can be adjusted as needed. The division of each module is only for the convenience of describing the functional division used. In actual implementation, a module can be divided into multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be in the same device or different devices.

[0115] The hardware modules in each implementation can be implemented mechanically or electronically. For example, a hardware module can include specially designed permanent circuits or logic devices (such as dedicated processors, such as FPGA or ASIC) to complete specific operations. Hardware modules can also include programmable logic devices or circuits temporarily configured by software (such as general-purpose processors or other programmable processors) for performing specific operations. As for the specific use of mechanical methods, either dedicated permanent circuits or temporarily configured circuits (such as software configuration) to implement hardware modules, it can be determined based on cost and time considerations.

[0116] The above is only a preferred embodiment of the present disclosure and is not intended to limit the scope of  protection of the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

[0117] Independent of the grammatical term usage, individuals with male, female or other gender identities are included within the term.

Claims

1.A method for identifying building block, comprising:obtaining (101) multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term;translating (102) terms with the same semantics in the multiple information models into a predetermined unified term; andidentifying (103) a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.2.The method according to claim 1, wherein the translating (102) terms with the same semantics in the multiple information models into a predetermined unified term comprises:receiving a user vocabulary comprising the unified term and a semantic definition of the unified term;calculating respective similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term;determining respective target semantic definitions with respective similarities greater than a predetermined threshold;translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.3.The method according to claim 2,wherein the calculating similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term comprises: inputting the multiple information models into a first LLM; inputting the user vocabulary into the first LLM; enabling the first LLM to calculate respective similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term;wherein the determining respective target semantic definitions with respective similarities greater than a predetermined threshold comprises: enabling the first LLM to determine respective target semantic definitions with respective similarities greater than the threshold; the method comprises:receiving a term translation list from the first LLM, and the term translation list comprises the unified term and respective terms in respective pairs which comprise the respective target semantic definitions;wherein the translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term comprises:translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term, based on the term translation list.4.The method according to claim 3, comprising:confirming correctness of the term translation list in response a confirmation instruction triggered by a user, before the translating respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.5.The method according to claim 1, comprising:determining a target schema;converting the building block into a sub-model which conforms to the target schema.6.The method according to claim 5, wherein the converting the building block into a sub-model which conforms to the target schema comprises:inputting the target schema and the building block into a second LLM;enabling the second LLM to convert the building block into a sub-model which conforms to the target schema.7.The method according to claim 6, comprises:replacing respective structures of the unified term in the multiple information models with the sub-model.8.The method according to claim 6, comprises:obtaining a new information model that does not belong to the multiple information models, the new information model comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term;determining a target term in the new information model that has the same semantics as the unified term;replacing structure of the unified term in the new information model with the sub-model.9.An apparatus for identifying building block, comprising:an obtaining module (401) , configured to obtain multiple information models, each of the multiple information models comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term;a translating module (402) , configured to translate terms with the same semantics in the multiple information models into a predetermined unified term; andan identifying module (403) , configured to identify a building block from the translated multiple information models, wherein the building block comprises the unified term and a common structure of the unified term in the translated multiple information models.10.The apparatus according to claim 9, wherein the translating module (402) is configured to receive a user vocabulary comprising the unified term and a semantic definition of the unified term, calculate respective  similarities between respective semantic definitions in the multiple information models and the semantic definition of the unified term, determine respective target semantic definitions with respective similarities greater than a predetermined threshold, and translate respective terms in respective pairs which comprise the respective target semantic definitions into the unified term.11.The apparatus according to claim 9, comprising:a converting module (404) , configured to determine a target schema, and convert the building block into a sub-model which conforms to the target schema.12.The apparatus according to claim 11, comprising:a replacing module (405) , configured to replace respective structures of the unified term in the multiple information models with the sub-model.13.The apparatus according to claim 11, comprising:a replacing module (405) , configured to obtain a new information model that does not belong to the multiple information models, the new information model comprises multiple pairs, and each of the multiple pairs comprises a term and a semantic definition of the term, determine a target term in the new information model that has the same semantics as the unified term, and replace structure of the unified term in the new information model with the sub-model.14.An electronic device, comprising a processor (501) and a memory (502) , wherein an application program executable by the processor (501) is stored in the memory (502) for causing the processor (501) to execute a method for identifying building block according to any one of claims 1-8.15.A computer-readable medium comprising computer-readable instructions stored thereon, wherein the computer-readable instructions for executing a method for identifying building block according to any one of claims 1-8.16.A computer program product comprising a computer program, upon the computer program is executed by a processor for executing a method for identifying building block according to any one of claims 1-8.

Citation Information

Patent Citations

  • Word, expression, and sentence translation management tool

    US20030101044A1

  • Systems, methods, and apparatus for automated mapping and integrated workflow of a controlled medical vocabulary

    US20110066425A1

  • Methods and systems for multi-engine machine translation

    US20130262080A1

  • Translation Protocol for Large Discovery Projects

    US20140358518A1