Industrial Internet information modeling optimization method based on natural language processing

Through automated natural language processing methods, modeling element clusters are extracted and divided to construct information models, which solves the problems of insufficient data quality and integrity in traditional methods and improves the effectiveness and reliability of information models.

CN119669794BActive Publication Date: 2025-09-23BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411740368.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-09-23
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Traditional industrial Internet information modeling methods based on natural language processing rely on manual data collection and processing, resulting in incomplete and inaccurate model construction, poor data quality and integrity, and affecting the effectiveness and reliability of the information model.

Method used

A method based on natural language processing is adopted. The preset natural language processing model is input through the preset data source and model prompt information to automatically extract modeling elements. The modeling element clusters are divided based on the target cluster number and semantic similarity threshold to construct an information model, including a combination of tree and network structure layers.

Benefits of technology

It improves the effectiveness and reliability of information models, reduces data quality and integrity issues, enhances the compatibility and accuracy of information models, and optimizes model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669794B_ABST
    Figure CN119669794B_ABST
Patent Text Reader

Abstract

The present invention provides an industrial Internet information model modeling optimization method based on natural language processing, which relates to the field of computer technology; the method comprises: extracting modeling elements by inputting a preset data source and model prompt information into a preset natural language processing model; the model prompt information comprises first-stage prompt information and second-stage prompt information; dividing the modeling elements into a corresponding number of modeling element clusters based on the number of target clusters; constructing an information model based on the modeling element clusters; the information model is used to indicate the relationship between each modeling element; the information model comprises a tree structure layer, or a combination of a tree structure layer and a mesh structure layer; wherein the tree structure layer comprises an object class layer corresponding to the information model, a cluster category layer corresponding to each modeling element cluster, and a representative element layer corresponding to each modeling element cluster; the method can solve the problem of poor effect and reliability of the information model; and improve the effect and reliability of the information model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an industrial Internet information modeling optimization method based on natural language processing. Background Art

[0002] An information model is an abstract and standardized description of real-world information and data, expressing the structure, relationships, and flow of information through unified standards. In industrial production, information models are used to describe the interrelationships between elements such as production processes, equipment, techniques, and materials. This ensures efficient and accurate information exchange and sharing between different systems and equipment, supporting production decision-making and control. As the foundation for the standardized representation and flow of production information, information models are an effective solution for improving the integration and interoperability of production information.

[0003] Traditional industrial Internet information modeling and optimization methods based on natural language processing include: manually collecting and processing data and building information models.

[0004] However, due to the difficulty in obtaining data and difficulty in standardization, traditional modeling methods rely on manual collection and processing of data and lack sufficient automated tools and technologies, which is time-consuming and labor-intensive, resulting in incomplete and inaccurate model construction, poor data quality and integrity, and thus poor effectiveness and reliability of information models. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides an industrial Internet information modeling optimization method based on natural language processing to eliminate or improve one or more defects in the existing technology. It can solve the problem of poor effectiveness and reliability of information models.

[0006] One aspect of the present invention provides an industrial Internet information modeling optimization method based on natural language processing, the method comprising the following steps:

[0007] Inputting the preset data source and model prompt information into the preset natural language processing model to extract modeling elements; the model prompt information includes first-stage prompt information and second-stage prompt information; the first-stage prompt information is used to prompt the preset natural language processing model to extract entity types from the preset data source according to the preset entity type list; the second-stage prompt information is used to prompt the preset natural language processing model to extract modeling elements based on the extracted entity types;

[0008] Divide the modeling elements into a corresponding number of modeling element clusters based on the number of target clusters;

[0009] An information model is constructed based on the modeling element cluster; the information model is used to indicate the relationship between the various modeling elements, including the internal relationship corresponding to the modeling elements within each modeling element cluster, and the cross-cluster relationship corresponding to the modeling elements between different modeling element clusters. The information model includes a tree structure layer, or a combination of a tree structure layer and a network structure layer.

[0010] In some embodiments of the present invention, before dividing the modeling elements into a corresponding number of modeling element clusters based on the target number of clusters, the method further includes:

[0011] Perform cluster analysis on modeling elements based on a preset number of clusters, and calculate the silhouette coefficients corresponding to clustering of modeling elements under different numbers of clusters;

[0012] The target number of clusters is determined from a preset number of clusters based on the silhouette coefficient.

[0013] In some embodiments of the present invention, before dividing the modeling elements into a corresponding number of modeling element clusters based on the target number of clusters, the method further includes:

[0014] Determine the target semantic similarity threshold corresponding to each modeling element cluster;

[0015] Compare the modeling elements in each modeling element cluster pairwise to determine the semantic similarity corresponding to each pair of modeling elements;

[0016] In each modeling element cluster, modeling element pairs whose semantic similarity is greater than or equal to the corresponding target semantic similarity threshold are aggregated.

[0017] In some embodiments of the present invention, determining a target semantic similarity threshold corresponding to each modeling element cluster includes:

[0018] Within a preset semantic similarity threshold range, analyzing and obtaining a first relationship between the silhouette coefficient corresponding to each modeling element cluster and the semantic similarity, and a second relationship between the number of modeling elements corresponding to each modeling element cluster and the semantic similarity;

[0019] Based on the first relationship and the second relationship, a target semantic similarity threshold corresponding to each modeling element cluster is determined within a preset semantic similarity threshold range.

[0020] In some embodiments of the present invention, before determining the target semantic similarity threshold corresponding to each modeling element cluster, the method further includes:

[0021] determining the number of modeling elements in each modeling element cluster;

[0022] For modeling element clusters with more than 2 modeling elements, a step of determining a target semantic similarity threshold corresponding to each modeling element cluster is performed.

[0023] In some embodiments of the present invention, after inputting the preset data source and model prompt information into the preset natural language processing model and extracting the modeling elements, the method further includes: converting the modeling elements into corresponding word vector forms.

[0024] In some embodiments of the present invention, the tree structure layer includes an object class layer corresponding to the information model, a cluster category layer corresponding to each modeling element cluster, and a representative element layer corresponding to each modeling element cluster; the object class layer is used to indicate the modeling object corresponding to the information model; the cluster category layer is used to indicate the category corresponding to each modeling element cluster; and the representative element layer is used to indicate the representative element corresponding to each modeling element cluster.

[0025] In some embodiments of the present invention, constructing an information model based on a cluster of modeling elements includes:

[0026] Based on the modeling element clusters and association relationships, a network structure layer is constructed;

[0027] Determine the average word vector of the modeling element word vectors in each modeling element cluster respectively, and compare the semantic similarity with the preset modeling element cluster category to obtain a comparison result; the comparison result is used to indicate the similarity between each modeling element cluster and each preset modeling element cluster category;

[0028] Based on the comparison results, the category corresponding to each modeling element cluster is determined, and the cluster category layer is constructed;

[0029] Based on the internal relationships, the representative elements corresponding to each modeling element cluster are determined and the representative element layer is constructed.

[0030] In some embodiments of the present invention, constructing an information model based on a cluster of modeling elements further includes:

[0031] At least two information sub-models are constructed based on the cluster of modeling elements;

[0032] An information model is obtained by combining at least two information sub-models.

[0033] In some embodiments of the present invention, the preset data source includes a text data source and / or a knowledge base corresponding to a preset natural language processing model; when the preset data source includes a text data source, the model prompt information is input into the preset natural language processing model, and before the modeling elements are extracted, it also includes: inputting the text data source as an attachment into the preset natural language processing model.

[0034] The industrial Internet information model modeling optimization method based on natural language processing of the present invention can solve the problem of poor effectiveness and reliability of information models; by presetting a natural language processing model, according to the preset entity type list indicated by the first-stage model prompt information, the entity type is automatically obtained from the preset data source, and then through the second-stage model prompt information and entity type, the entity information is automatically extracted from the preset data source without relying on manual collection and processing, avoiding the problem of poor data quality and integrity, thereby improving the effectiveness and reliability of the information model; at the same time, using a unified preset entity type list to obtain entity types can overcome the compatibility barriers between different standards and improve the compatibility of information models.

[0035] In addition, after determining the target semantic similarity threshold, the modeling elements in each modeling element cluster are optimized. That is, according to the target semantic similarity threshold, it is determined whether the modeling elements can be aggregated to reduce the number of modeling elements, improve model accuracy, and thus optimize model performance.

[0036] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.

[0037] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. In the drawings:

[0039] Figure 1 A flowchart of a method for optimizing an industrial Internet information model based on natural language processing provided by one embodiment of the present invention;

[0040] Figure 2 A schematic diagram of text-driven prompt learning provided by an embodiment of the present invention;

[0041] Figure 3 A schematic diagram of knowledge-driven prompt learning provided by an embodiment of the present invention;

[0042] Figure 4 A graph showing the relationship between the silhouette coefficient and the number of clusters provided in one embodiment of the present invention;

[0043] Figure 5A schematic diagram of the division of modeling element clusters provided by an embodiment of the present invention;

[0044] Figure 6 A schematic diagram of a first relationship and a second relationship provided in one embodiment of the present invention;

[0045] Figure 7 Another schematic diagram of a first relationship and a second relationship provided in accordance with an embodiment of the present invention;

[0046] Figure 8 Another schematic diagram of a first relationship and a second relationship provided in accordance with an embodiment of the present invention;

[0047] Figure 9 Another schematic diagram of a first relationship and a second relationship provided in accordance with an embodiment of the present invention;

[0048] Figure 10 Another schematic diagram of a first relationship and a second relationship provided in accordance with an embodiment of the present invention;

[0049] Figure 11 A semantic similarity heat map provided by an embodiment of the present invention;

[0050] Figure 12 Another semantic similarity heat map provided by an embodiment of the present invention;

[0051] Figure 13 Another semantic similarity heat map provided by an embodiment of the present invention;

[0052] Figure 14 Another semantic similarity heat map provided by an embodiment of the present invention;

[0053] Figure 15 Another semantic similarity heat map provided by an embodiment of the present invention;

[0054] Figure 16 A schematic diagram of the framework of an information model provided by an embodiment of the present invention;

[0055] Figure 17 A schematic diagram of another information model framework provided by an embodiment of the present invention;

[0056] Figure 18 A hierarchical clustering dendrogram provided by an embodiment of the present invention;

[0057] Figure 19 A schematic diagram of another information model framework provided by an embodiment of the present invention;

[0058] Figure 20 A radar chart showing the category similarity of modeling element clusters provided by one embodiment of the present invention;

[0059] Figure 21 A schematic diagram of another information model framework provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0061] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0062] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0063] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0064] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0065] The following is a detailed introduction to the industrial Internet information modeling optimization method based on natural language processing provided by this application.

[0066] like Figure 1 As shown, the embodiment of the present application provides an industrial Internet information model modeling optimization method based on natural language processing. The implementation of the method can rely on a computer program, which can be run on a computer device such as a smart phone, a tablet computer, a personal computer, or a server. This embodiment does not limit the operating subject of the method. The method includes at least the following steps:

[0067] Step S101: input model prompt information into a preset natural language processing model to extract modeling elements.

[0068] In some embodiments of the present invention, extracting modeling elements refers to the process of extracting searchable and editable entity information related to the modeling purpose of the modeling object from a preset data source using a preset natural language processing model. The extracted modeling elements are used to construct an information model.

[0069] The preset natural language processing model refers to a pre-trained language model, including but not limited to the Chat Generative Pre-trained Transformer (ChatGPT) or the Language Model for Dialogue Applications (LaMDA). Entity information refers to specific data describing actual objects, concepts, or events, including their attributes, characteristics, and status. For example, in the industrial field, entity information refers to independently existing things or concepts that are identified and defined within the industry, such as manufacturer, equipment name, and production date.

[0070] Before extracting modeling elements, you must first clarify the modeling purpose of the modeled object. For example, if the modeling purpose is to manage equipment operation, the modeling elements extracted should focus on the equipment's operating status. If the modeling purpose is to assemble manufacturing equipment, the modeling elements should include information about the equipment's functions and appearance.

[0071] In some embodiments of the present invention, the preset data source includes a text data source and / or a knowledge base corresponding to a preset natural language processing model; wherein the text data source includes unstructured text related to the modeling object, such as technical documents, operation manuals, white papers, etc.; the knowledge base corresponding to the preset natural language processing model includes the language knowledge and background knowledge learned during the model training process, as well as the knowledge acquired from the Internet through search engines.

[0072] In the case where the preset data source is a text data source, before the model prompt information is input into the preset natural language processing model and the modeling elements are extracted, the method further includes: inputting the text data source as an attachment into the preset natural language processing model.

[0073] The text data source is uploaded to the pre-set natural language processing model in a nearby format. Information is extracted from the pre-set data source using text-driven prompt learning. The pre-set data source and model prompt information are combined as input to the pre-set natural language processing model. The pre-set natural language processing model is used to obtain more complete and accurate modeling elements, filtering out irrelevant elements in the text data source to reduce the search space and computational complexity. The rich semantic information accumulated by the pre-set natural language processing model after training on large-scale text data captures subtle semantic relationships between sentences or paragraphs, thereby understanding complex context. Extracting precise modeling elements from the pre-set data source makes the results more relevant and accurate.

[0074] In some embodiments of the present invention, the process of obtaining modeling elements from a preset data source is designed as a sequential, multi-round question-answering process, comprising two phases, each phase comprising multiple rounds of question-answering. Accordingly, the model prompt information includes first-phase prompt information and second-phase prompt information; the first-phase prompt information prompts the preset natural language processing model to extract entity types from the preset data source according to a preset list of entity types; the second-phase prompt information prompts the preset natural language processing model to extract modeling elements based on the extracted entity types.

[0075] The entity type list refers to a list of categories or types used to classify and identify different entities in the text, including but not limited to identifiers, attributes, events, services, parameters, relations, classes, etc.

[0076] In practice, during each conversational turn, prompts can be constructed based on previously extracted information, serving as input to the pre-set natural language processing model. Meanwhile, in the second phase, relevant information is further extracted and classified based on the entity types extracted in the first phase, combined with specific constraints such as modeling objectives and key modeling elements. Finally, the information extracted from each conversational turn is combined into the final structured data, which is the required modeling element.

[0077] For example: Reference Figure 2 Taking the construction of an information model in the industrial field as an example, the preset data source is text-based instructions, white papers, technical documents and other materials related to the modeling object, and the preset data source is uploaded as an attachment; in the first stage, the first-stage prompt information is used to instruct the preset natural language processing model to extract entity types (such as identifiers, attributes) from the preset data source according to a given entity type list; in the second stage, the second-stage prompt information is used to instruct the preset natural language processing model to extract entities from the preset data source according to the entity type as modeling elements (such as equipment number, part number, cutting speed, feed speed).

[0078] In actual implementation, access to text data sources is limited; in addition, the quality of the text data sources obtained can be low, resulting in poor extraction of modeling elements. Therefore, this embodiment also proposes a knowledge-driven prompt learning approach for modeling element extraction, where the preset data source also includes a knowledge base corresponding to a preset natural language processing model.

[0079] Compared to text-driven prompt learning, knowledge-driven prompt learning requires no additional input. Instead, it generates relevant entities directly from a knowledge base based on a pre-defined natural language processing model through the design of targeted prompts. This enhances the versatility and adaptability of the information model, effectively extracting entity information with high accuracy even when data sources are missing or incomplete.

[0080] For example: Reference Figure 3 , in order to build an information model for the industrial field; in the first stage, the first stage prompt information is used to instruct the preset natural language processing model to extract entity types (such as identifiers, events) from the knowledge base of the preset natural language processing model according to a given entity type list; in the second stage, the second stage prompt information is used to instruct the preset natural language processing model to extract entities from the knowledge base of the preset natural language processing model according to the entity type as modeling elements (such as version numbers, artifact loading, etc.).

[0081] Since text-driven hint learning methods may be limited by the quality and coverage of the text, and knowledge-driven hint learning methods may be limited by the known information in the knowledge base, in actual implementation, it is possible to combine the two to complement each other's shortcomings, maximize coverage of relevant information, reduce omissions, and improve the comprehensiveness and accuracy of information extraction.

[0082] Taking the automatic extraction of modeling elements of CNC machine tools as an example; combining the text-driven method and the knowledge-driven method, in the text-driven method, CNC equipment series standards, CNC machine tool operation manuals and CNC machine tool technical documents are selected as the preset data sources to input the preset natural language processing model, and entity types and entity information are extracted based on the seven key elements of the 3IM framework (Identifier, Attribute, Event, Service, Parameter, Relationship, Class). In the first stage, the prompt message is: "Given the text is the content of the attachment, given the entity type list: [Identifier, Attribute, Event, Service, Parameter, Relationship, Class], what entity types may be included in this text?"; in the second stage, the prompt message is "Based on the above text, please output the entities with the entity type of "Identifier" in the form of a list, similar to: ["Entity Type", "Entity Name"], ["Entity Type", "Entity Name 2"]...", and "Among the entities with the entity type of "Identifier" just given, which entities are related to "CNC Machine Tools"?", etc. In the knowledge-driven approach, entity types and entity information are extracted from the knowledge base of the preset natural language processing model itself. In the first stage, the prompt message is "Given a list of entity types: [identifier, attribute, event, service, parameter, relationship], what entity types may the modeling object "CNC machine tool" contain?" In the second stage, the prompt message is "Please output entities with entity type "identifier" in the form of a list, similar to: ["entity type", "entity name"], ["entity type", "entity name 2"]..."

[0083] Step S102 : dividing the modeling elements into a corresponding number of modeling element clusters based on the target number of clusters.

[0084] In some embodiments of the present invention, in order to reasonably classify modeling elements so that the constructed information model is clearer, more orderly, and easier to maintain and expand, it is necessary to divide the modeling elements according to the target number of clusters to obtain a corresponding number of modeling element clusters.

[0085] The target number of clusters is determined from the preset number of clusters set based on the analysis results obtained after cluster analysis is performed on the modeling elements according to the preset number of clusters set for global optimization of the modeling elements.

[0086] Specifically, a set of preset cluster numbers is first set, which includes several cluster numbers. Then, within each cluster number, the modeling elements are clustered to obtain clusters of modeling elements, and the corresponding silhouette coefficient is calculated. The silhouette coefficients for different cluster numbers are then compared to comprehensively evaluate the compactness of modeling elements within clusters and the degree of separation between elements in clusters. This determines the target number of clusters, and the modeling elements are then divided into the corresponding number of clusters based on the target number of clusters.

[0087] Specifically, before dividing the modeling elements into a corresponding number of modeling element clusters based on the target number of clusters, it also includes: performing cluster analysis on the modeling elements based on a preset cluster number set, calculating the silhouette coefficients corresponding to the clustering of the modeling elements under different cluster numbers; and determining the target number of clusters in the preset cluster number set based on the silhouette coefficients.

[0088] In some embodiments of the present invention, the target number of clusters refers to the number of clusters corresponding to the largest silhouette coefficient among the silhouette coefficients.

[0089] refer to Figure 4 , taking the automatic extraction of modeling elements of CNC machine tools as an example, the preset cluster number set is set to [2,10], and then all modeling elements are hierarchically clustered within the range of the set, the corresponding silhouette coefficient under each cluster number is calculated, and the relationship between the cluster number and the silhouette coefficient is plotted, as shown in Figure 4 As shown. Figure 4Analysis shows that initially, as the number of clusters increases, the number of modeling elements within each cluster decreases, the compactness within the cluster increases, and the differences between clusters increase, resulting in a larger silhouette coefficient. When the number of clusters continues to increase to a certain level, the number of modeling elements contained in each cluster becomes too small, the cluster's stability and representativeness deteriorate, and the silhouette coefficient deteriorates. Although further increasing the number of clusters may lead to a better clustering solution and a slight increase in the silhouette coefficient, it does not reach its initial peak. Therefore, the silhouette coefficient reaches its maximum when the number of clusters is 6, indicating that both the compactness within the cluster and the separation between clusters are optimal, resulting in a good overall effect. Therefore, the target number of clusters is set at 6.

[0090] After determining the target number of clusters, the modeling elements of the CNC machine tool are divided into a corresponding number of modeling element clusters. In order to provide an intuitive reference for further analysis, all modeling elements are clustered and the corresponding cluster scatter plots are drawn, as shown in Figure 2. Figure 5 shown.

[0091] Figure 5 There are six clusters displayed, named Cluster0, Cluster1, Cluster2, Cluster3, Cluster4 and Cluster5. Among them, Cluster0 has 5 modeling elements, Cluster0 = {cutting speed, spindle speed, feed adjustment rate, feed override, feed speed}; Cluster1 has 4 modeling elements, Cluster1 = {operating status, working status, alarm status, operating time}; Cluster2 has 5 modeling elements, Cluster2 = {equipment type, factory batch, manufacturer, serial number, production date}; Cluster3 has 8 modeling elements, Cluster3 = {spindle servo system, spindle drive system, main transmission system, servo system, position follow-up system, feed servo system, feed transmission system, feed drive system}; Cluster4 has 12 modeling elements, Cluster4 = {CNC system, programmable controller, auxiliary system, sensor, cooling system, lubrication system, protection system, hydraulic system, oil pressure system, fixture system, chip removal system, clamping system}; Cluster5 has 2 modeling elements, Cluster5 = {material inbound, material outbound}.

[0092] However, even though a limited number of modeling elements are captured, these elements may contain unnecessary redundant information, complicating the construction of information model relationships, increasing system development and maintenance costs, and reducing responsiveness. Furthermore, if too few modeling elements are retained, the information model may lose essential details and key semantic information, resulting in information loss and insufficient model accuracy. Therefore, further screening and optimization of modeling elements is necessary to strike a balance between model complexity and information completeness.

[0093] In some embodiments of the present invention, the modeling elements within each cluster are locally optimized according to a target semantic similarity threshold to reduce the number of modeling elements. The target semantic similarity threshold is an indicator used to determine whether modeling elements can be clustered and is used to assess the similarity of modeling elements within a cluster. Within the same cluster, if the semantic similarity of two modeling elements exceeds the target semantic similarity threshold, one of the modeling elements is retained and the other is deleted until the similarity of all modeling elements within the cluster falls below the target semantic similarity threshold.

[0094] Specifically, after dividing the modeling elements into a corresponding number of modeling element clusters based on the target number of clusters, the method also includes: determining a target semantic similarity threshold corresponding to each modeling element cluster; comparing the modeling elements in each modeling element cluster pairwise to determine the semantic similarity corresponding to each pair of modeling elements; and in each modeling element cluster, aggregating the modeling element pairs whose semantic similarity is greater than or equal to the corresponding target semantic similarity threshold.

[0095] Specifically, within the pre-set semantic similarity threshold range, the changes between the silhouette coefficient of each modeling element cluster and the number of retained modeling elements are analyzed, and the target semantic similarity threshold corresponding to each modeling element cluster is selected based on the analysis results; in this process, by analyzing the silhouette coefficients corresponding to different semantic similarity thresholds, it is determined which semantic similarity threshold makes the clustering effect better; at the same time, by observing the number of retained modeling elements, the modeling elements can be locally optimized as much as possible while ensuring the clustering effect, balancing the clustering quality and data volume while improving the integrity and practicality of the model.

[0096] Specifically, determining the target semantic similarity threshold corresponding to each modeling element cluster includes: analyzing within a preset semantic similarity threshold to obtain a first relationship between the silhouette coefficient corresponding to each modeling element cluster and the semantic similarity, and a second relationship between the number of modeling elements corresponding to each modeling element cluster and the semantic similarity; based on the first relationship and the second relationship, determining the target semantic similarity threshold corresponding to each modeling element cluster within the preset semantic similarity threshold.

[0097] Take the automatic extraction of modeling elements of CNC machine tools as an example to illustrate. Figure 5When locally optimizing the modeling element clusters (Cluster0, Cluster1, Cluster2, Cluster3, and Cluster4) in , the semantic similarity threshold range is set to [0.75, 1.05] with a step size of 0.05. Too low a threshold may remove a large number of modeling elements, resulting in incomplete information expression. Therefore, setting the lower threshold to 0.75 can strike a balance between data integrity and simplicity and ensure data quality. At the same time, although the theoretical maximum value of semantic similarity is 1, the upper threshold reaches 1.05. This is to ensure that the algorithm can still perform well when approaching the limit value, and to ensure its reliability in practical applications. In actual implementation, the semantic similarity threshold range can be adjusted according to actual conditions. This embodiment does not limit the value and step size of the semantic similarity threshold range.

[0098] In order to intuitively demonstrate the process of selecting the semantic similarity threshold, Cluster0 to Cluster4 are analyzed in detail. Figures 6 to 10 The results are presented in . Specifically, each modeling element cluster has two corresponding relationship subgraphs: a first relationship subgraph corresponding to the semantic similarity threshold-silhouette coefficient and a second relationship subgraph corresponding to the semantic similarity threshold-number of retained modeling elements. The first relationship subgraph depicts the impact of different semantic similarity thresholds on the silhouette coefficient, while the second relationship subgraph shows the relationship between the semantic similarity threshold and the number of retained modeling elements.

[0099] Depend on Figure 6 As can be seen, when the semantic similarity threshold for Cluster 0 is greater than or equal to 0.80, the silhouette coefficient remains at its maximum, but the number of modeling elements does not change. In contrast, when the semantic similarity threshold is 0.75, the silhouette coefficient only slightly decreases from the maximum value, while the number of modeling elements decreases by 2. This shows that a threshold of 0.75 maintains high clustering quality while effectively reducing redundant data and improving overall data quality. Therefore, 0.75 is selected as the semantic similarity threshold for Cluster 0.

[0100] Depend on Figure 7As can be seen, when the semantic similarity threshold is greater than or equal to 0.90, the silhouette coefficient of Cluster 1 decreases significantly, while the number of modeling elements remains unchanged. When the semantic similarity threshold is within the specified range, the silhouette coefficient is larger and the number of modeling elements is optimized. This indicates that choosing an appropriate semantic similarity threshold can lead to more compact clusters, reflecting the effectiveness of the optimization. Although the silhouette coefficient and the number of retained modeling elements are the same at thresholds of 0.75, 0.80, and 0.85, 0.85 is a turning point, likely reflecting a key change in data characteristics. Here, considering both the clustering effect and the need for data redundancy removal, 0.85 is selected as the semantic similarity threshold for Cluster 1.

[0101] Depend on Figure 8 As can be seen, when the semantic similarity threshold is greater than or equal to 0.80, the silhouette coefficient decreases significantly, while the number of modeling elements is not optimized. This indicates that at these higher thresholds, the clustering effect deteriorates and the data distribution loses its well-structured quality. In contrast, at a threshold of 0.75, not only is the silhouette coefficient larger, improving clustering quality, but the number of modeling elements is also effectively streamlined, reducing data redundancy and thus improving the overall efficiency of the model. Therefore, 0.75 is selected as the semantic similarity threshold for Cluster2.

[0102] Depend on Figure 9 It can be seen that as the threshold of semantic similarity gradually increases, the silhouette coefficient shows a fluctuating trend of first increasing, then decreasing, and then increasing again. Specifically, when the threshold increases from 0.80 to 0.85, the silhouette coefficient increases significantly. However, as the threshold continues to increase to 0.90, the silhouette coefficient decreases slightly. But when the threshold continues to increase to 0.90 and above, the silhouette coefficient increases again and reaches its highest value between the thresholds of 0.95 and 1.05. Figure 9 It can be seen that as the threshold increases, the number of retained modeling elements also gradually increases. When the threshold is greater than or equal to 0.95, the number of modeling elements remains unchanged at 8. Figure 9 From this perspective, for Cluster 3, selecting a higher threshold (such as 0.95) can achieve better clustering results, but it also results in poor optimization of the modeling elements, and therefore is not considered an appropriate semantic similarity threshold. Lower thresholds (0.75 or 0.80) correspond to slightly lower silhouette coefficients than those at higher thresholds, but this difference is not significant, and the clustering effect is not significantly affected. Furthermore, these lower thresholds can effectively reduce the number of modeling elements, remove redundant elements, and improve cluster compactness. Although the silhouette coefficient and the number of retained modeling elements are the same for thresholds of 0.75 and 0.80, 0.80, as a turning point, better reflects the subtle changes in the optimization process. Therefore, 0.80 was selected as the semantic similarity threshold for Cluster 3.

[0103] Depend on Figure 10 It can be seen that as the semantic similarity threshold gradually increases, the silhouette coefficient generally shows an increasing trend. Figure 10 It can be seen that as the threshold increases, the number of retained modeling elements also gradually increases. When the threshold is greater than or equal to 0.95, the number of modeling elements remains unchanged at 12. Overall, although a higher threshold range can make the silhouette coefficient reach a larger value, it also significantly increases the number of retained modeling elements, resulting in more redundancy. In contrast, when the threshold is 0.75 and 0.80, the silhouette coefficient is slightly lower than the maximum value, but still remains at a high level, and the number of retained modeling elements is smaller, and the redundancy removal effect is better. Figure 10 It can be seen from the figure that the threshold of 0.80 is a turning point. Further increasing the threshold will lead to increased redundancy, and further lowering the threshold does not significantly improve the clustering effect. Therefore, choosing 0.80 as the semantic similarity threshold can effectively reduce redundant elements while ensuring the clustering effect, achieving the best balance.

[0104] Based on the above analysis, considering the clustering effect and the need for redundancy removal, 0.75, 0.85, 0.75, 0.80 and 0.80 were selected as the target semantic similarity thresholds for Cluster0 to Cluster4 respectively.

[0105] After determining the target semantic similarity threshold, the modeling elements in each modeling element cluster are optimized. That is, according to the target semantic similarity threshold, it is determined whether the modeling elements can be aggregated to reduce the number of modeling elements, improve model accuracy, and thus optimize model performance.

[0106] Specifically, the similarity between different modeling elements in the same modeling element cluster is calculated, and a semantic similarity heat map is drawn to intuitively present the relationship between the elements. When the semantic similarity of two modeling elements exceeds the target semantic similarity threshold corresponding to the modeling element cluster, one of the modeling elements is retained and the other is deleted until the similarity of the modeling elements in the modeling element cluster is lower than the corresponding target semantic similarity threshold.

[0107] like Figures 11 to 15 As shown in the figure, the automatic extraction of modeling elements of CNC machine tools is used as an example to illustrate. Figure 11 It can be seen that the semantic similarity between cutting speed and spindle speed is 0.78, so we choose to keep cutting speed and delete spindle speed. Similarly, the semantic similarity between feed rate and feed adjustment rate is 0.76, so we choose to keep feed rate and delete feed adjustment rate. By optimizing the modeling elements in Cluster0, the elements that are finally retained include cutting speed, feed rate, and feed speed. Figure 12It can be seen that the semantic similarity between the running state and the working state is 0.88. We choose to keep the running state and delete the working state. After optimizing the modeling elements in Cluster1, the elements that are finally retained include alarm state, running state, and running time. Figure 13 It can be seen that the semantic similarity between factory batch and serial number is 0.84. We choose to keep the factory batch and delete the serial number. By optimizing the modeling elements in Cluster2, the elements that are finally retained include manufacturer, factory batch, equipment type, and production date. Figure 14 It can be seen that the semantic similarity between the servo system and the position follower system is 0.89. The servo system is retained and the position follower system is deleted. The semantic similarities between the feed servo system and the feed drive system and the feed transmission system are 0.93 and 0.84 respectively. The feed servo system is retained and the feed drive system and the feed transmission system are deleted. The semantic similarities between the spindle servo system and the main transmission system and the spindle drive system are 0.81 and 0.91 respectively. The spindle servo system is retained and the main transmission system and the spindle drive system are deleted. Here, the elements that have been deleted in the previous round are not discussed. After optimizing the modeling elements in Cluster3, the elements that are finally retained include the servo system, the feed servo system, and the spindle servo system. Figure 15 As can be seen, the semantic similarity between the clamping system and the fixture system is 0.83, so the clamping system is retained and the fixture system is deleted. The semantic similarity between the hydraulic system and the oil pressure system is 0.92, so the hydraulic system is retained and the oil pressure system is deleted. After optimization of the modeling elements of Cluster 4, the remaining elements are the CNC system, chip removal system, clamping system, lubrication system, hydraulic system, auxiliary system, cooling system, protection system, programmable controller, and sensor.

[0108] In some embodiments of the present invention, a method is used that combines global optimization based on cluster analysis with local optimization in combination with a semantic similarity threshold. Global optimization ensures the rational classification of the overall data and finds the most appropriate clustering structure, thereby optimizing the grouping effect of modeling elements and improving the accuracy and efficiency of subsequent analysis and modeling. Local optimization optimizes information density by removing redundant or repeated information, thereby improving the accuracy of classification. This not only makes the information model more accurate and efficient, but also reveals more detailed hierarchical relationships between modeling elements, helping to build a clearer and more reasonable hierarchical structure. By combining global optimization and local optimization, not only the accuracy and efficiency of the information model are ensured, but also a systematic and hierarchical organizational structure is provided for the modeling elements.

[0109] In addition, when the number of modeling elements in a modeling element cluster is small, there is no need to perform local optimization on the modeling element cluster. Figure 5Cluster5 in ,Since the number of modeling elements in Cluster5 is small and it covers the core functions, it is directly retained without further local optimization.

[0110] Specifically, before determining the target semantic similarity threshold corresponding to each modeling element cluster, the method further includes: determining the number of modeling elements in each modeling element cluster; and for modeling element clusters with a number of modeling elements greater than 2, executing the step of determining the target semantic similarity threshold corresponding to each modeling element cluster.

[0111] In addition, after the modeling elements are extracted, in order to facilitate clustering processing and semantic similarity analysis in the subsequent modeling optimization process, each extracted modeling element needs to be converted into a corresponding word vector.

[0112] Specifically, the preset data source and model prompt information are input into the preset natural language processing model, and after the modeling elements are extracted, the method further includes: converting the modeling elements into corresponding word vector forms.

[0113] In some embodiments of the present invention, each modeling element is converted into a corresponding word vector by using Bidirectional Encoder Representations from Transformers (BERT) to capture complex semantic relationships.

[0114] Step S103 : constructing an information model based on the modeling element clusters; wherein the information model is used to indicate the relationships between the modeling elements, including internal relationships corresponding to the modeling elements within each modeling element cluster, and cross-cluster relationships corresponding to the modeling elements between different modeling element clusters.

[0115] In some embodiments of the present invention, taking into account the complex associations between model elements, a model representation framework of "Underground Root Structure" (URS) is proposed, including a tree structure, or a combination of a tree structure and a mesh structure. A reasonable model structure helps to clearly express the relationship between elements within and across models, improve the comprehensibility and maintainability of the model, and facilitate subsequent updates and expansions, thereby reducing maintenance costs and time. Unlike the hierarchical associations in traditional tree structures, the last layer of the URS model representation framework develops a large number of roots to better express the complex, unstructured, and even intertwined extended attribute relationships in the information model. This not only retains the hierarchical organizational advantages of the tree structure, but also combines the association of the mesh structure to flexibly adapt to complex extended attributes, providing a more comprehensive model representation method.

[0116] Specifically, the information model includes a tree structure layer, or a combination of a tree structure layer and a network structure layer. The tree structure layer includes an object class layer corresponding to the information model, a cluster category layer corresponding to each modeling element cluster, and a representative element layer corresponding to each modeling element cluster. The object class layer is used to indicate the modeling objects corresponding to the information model; the cluster category layer is used to indicate the categories corresponding to each modeling element cluster; and the representative element layer is used to indicate the representative elements corresponding to each modeling element cluster.

[0117] For example, reference Figure 16 Taking the four-layer URS model representation framework as an example, the first three layers adopt a tree structure, leveraging its clear hierarchy and structure to make information classification and management more intuitive. The fourth layer's structure is transformed into a horizontal branch network with a mesh structure and associated relationships to enhance the model's ability to express relationships. This enables the URS model representation framework to support cross-hierarchical and horizontal connections and the establishment of associated relationships, allowing for flexible adjustment of the model structure and connection methods according to different needs.

[0118] At the same time, in order to more accurately represent the complexity and dynamic nature of the information model, the information model should be able to effectively reflect the relationships between elements within and across the model. This embodiment refers to the Unified Modeling Language (UML) and defines relationships as inheritance, realization, dependency, association, aggregation, and composition.

[0119] Among them, the composition relationship refers to a whole-part relationship, which is stronger than the aggregation relationship; the aggregation relationship refers to a whole-part relationship, which is weaker than the composition relationship; the association relationship refers to a class knowing the properties of another class; the inheritance relationship refers to a class inheriting the functions of another class and can add additional functions; the implementation relationship refers to a class implementing the functions of one or more interfaces; the dependency relationship refers to a class depending on another class, usually through parameters, variables, etc. at runtime.

[0120] After the introduction of relations, the URS model representation framework is as follows Figure 17 As shown in the figure, the URS model representation framework defines a basic information model framework.

[0121] exist Figure 17 In [1], the first layer is the object class layer, i.e., the modeling object; the second layer is the cluster category layer, which is used to represent the category corresponding to each modeling element cluster.

[0122] For example, taking industrial systems as an example, based on the different information types and functional requirements commonly found in industrial systems, they can be further divided into six categories: specification, process, control, management, business, and custom. The specification category includes modeling elements related to standards, technical parameters, and feature descriptions. The process category covers modeling elements related to operations, processing flows, and process steps. The control category involves modeling elements related to system control, operating instructions, and control logic. The management category includes modeling elements related to management activities, resource allocation, and organizational structure. The business category involves modeling elements related to business activities, transaction processes, and business rules. Custom categories refer to special modeling elements that cannot be categorized into the aforementioned categories.

[0123] In some embodiments of the present invention, modeling element classification at the cluster level is performed by calculating the average word vector of each modeling element within each cluster and comparing it with the semantic similarity of the pre-set modeling element cluster category. This more accurately reflects the overall characteristics of the cluster, avoids the bias that may be caused by classifying individual elements, and improves classification accuracy.

[0124] The third layer is the representative element layer. This layer is constructed by selecting appropriate representative elements from the fourth layer, taking into account the structure of the fourth layer and the actual structural characteristics of the modeled object, the degree of association between modeled elements, and other factors. If a suitable representative element cannot be selected from a modeling element cluster as a common parent for the remaining elements in that cluster, all modeling elements in that cluster are elevated to the representative element layer. In actual implementation, it is possible that all modeling element clusters in the fourth layer are elevated to the representative element layer. In this case, the information model constructed using the URS model representation framework can include only the tree structure layer.

[0125] The fourth layer is the network structure layer. In some embodiments of the present invention, the relationship between modeling elements is pre-constructed in an artificial form based on domain knowledge and professional judgment. By manually constructing the relationship between modeling elements, it is ensured that the information model accurately reflects the actual business logic and operation process, thereby enhancing the accuracy and practicality of the model.

[0126] Specifically, an information model is constructed based on modeling element clusters, including: constructing a mesh structure layer based on modeling element clusters and association relationships; determining the average word vector of the modeling element word vectors within each modeling element cluster, and performing semantic similarity comparison with preset modeling element cluster categories to obtain a comparison result; the comparison result is used to indicate the similarity between each modeling element cluster and each preset modeling element cluster category; based on the comparison result, determining the category corresponding to each modeling element cluster, and constructing a cluster category layer; based on the internal relationship, determining the representative element corresponding to each modeling element cluster, and constructing a representative element layer.

[0127] Taking the automated extraction of modeling elements for CNC machine tools as an example, in order to more clearly demonstrate the hierarchical relationships of modeling elements within modeling element clusters after global and local optimization, hierarchical clustering dendrograms corresponding to the modeling element clusters are drawn. By combining the dendrogram structure with the actual structural characteristics of the modeling object and other factors for analysis and judgment, the modeling element mapping at the mesh structure layer is completed.

[0128] by Figure 5 Take the modeling element clusters Cluster3 and Cluster5 in the example, refer to Figure 18 The feed servo system and the spindle servo system are relatively close and are initially clustered together. This cluster is then merged with the servo system to form a higher-level cluster. Ultimately, the feed servo system and the spindle servo system are placed side by side, belonging to the same hierarchical structure as the servo system. The two modeling elements of Cluster 5 are directly mapped into the URS model representation framework.

[0129] After the mesh structure layer modeling element mapping is completed, the initial form of the information model is obtained as follows Figure 19 As shown in the figure, "*" indicates the part that needs to be mapped but has not been mapped yet. After that, a cluster category layer is constructed, and the modeling elements of each cluster are classified into the preset modeling element cluster category. The average word vector of the modeling element word vectors retained in the cluster is calculated and compared with the preset modeling element cluster category. The classification results are displayed through a radar chart, as shown in the figure. Figure 20 As shown in Figure 2, we can intuitively see the similarity of each modeling element cluster in the six categories. Figure 20 It can be seen that the modeling elements of Cluster0 to Cluster5 belong to the control class, management class, specification class, control class, control class and business class respectively.

[0130] Based on the above classification, the modeling elements of CNC machine tools are mapped to the URS model representation framework for representation, and the corresponding relationships are constructed in combination with domain knowledge. The model architecture is composed of Figure 19 The initial form shown develops into Figure 21 The finished form is shown. The optimized architecture shows that the number of CNC machine tool modeling elements has been reduced by 13 compared to the initial form, accounting for 36.11% of the total. The automated modeling optimization method proposed in this article not only optimizes the number of modeling elements and streamlines the model structure, but also ensures information integrity and retains essential information and relationships, thereby replacing traditional manual modeling methods.

[0131] In practical implementation, due to the complexity and diversity of industrial systems, which typically involve multiple functional modules, disciplines, and hierarchical structures, the modeling process often requires breaking the entire system into multiple information sub-models. These sub-models can describe the characteristics of different parts or layers of the system, such as the device layer, control layer, and application layer, each with its own independent function or task. By rationally combining these sub-models, a complete information model can be formed, which not only simplifies the modeling process but also improves the system's flexibility and scalability, ultimately achieving accurate modeling that meets diverse needs.

[0132] Specifically, constructing the information model based on the modeling element cluster further includes: constructing at least two information sub-models based on the modeling element cluster; and combining the at least two information sub-models to obtain the information model.

[0133] Furthermore, to provide a more accurate and comprehensive description of the system, multiple information models can be combined to create more complex information modeling. After optimizing the information model, it is stored. This ensures that when the same object needs to be modeled in the future, the stored model and relationships can be directly called upon without re-modeling, improving work efficiency and model reusability. At the same time, model consistency is guaranteed.

[0134] In summary, the industrial Internet information model modeling optimization method based on natural language processing provided by this embodiment extracts modeling elements by inputting a preset data source and model prompt information into a preset natural language processing model; the model prompt information includes first-stage prompt information and second-stage prompt information; the first-stage prompt information is used to prompt the preset natural language processing model to extract entities from the preset data source according to a preset entity type list; the second-stage prompt information is used to prompt the preset natural language processing model to sort out the extracted modeling elements based on the extracted entities; the modeling elements are divided into a corresponding number of modeling element clusters based on the number of target clusters; an information model is constructed based on the modeling element clusters; the information model includes a tree structure layer and a mesh structure layer; the tree structure layer includes an object class layer corresponding to the information model, a cluster category layer corresponding to each modeling element cluster, and a cluster class layer corresponding to each modeling element cluster. The representative element layer corresponding to the model element cluster; the object class layer is used to indicate the modeling objects corresponding to the information model; the cluster category layer is used to indicate the categories corresponding to each modeling element cluster; the representative element layer is used to indicate the representative elements corresponding to each modeling element cluster; it can solve the problem of poor effect and reliability of the information model; through the preset natural language processing model, according to the preset entity type list indicated by the first-stage model prompt information, the entity type is automatically obtained from the preset data source, and then the entity information is automatically extracted from the preset data source through the second-stage model prompt information and entity type, without relying on manual collection and processing, avoiding the problem of poor data quality and integrity, thereby improving the effect and reliability of the information model; at the same time, using a unified preset entity type list to obtain entity types can overcome the compatibility barriers between different standards and improve the compatibility of the information model.

[0135] In addition, after determining the target semantic similarity threshold, the modeling elements in each modeling element cluster are optimized. That is, according to the target semantic similarity threshold, it is determined whether the modeling elements can be aggregated to reduce the number of modeling elements, improve model accuracy, and thus optimize model performance.

[0136] Corresponding to the above method, the present invention also provides an industrial Internet information model modeling optimization method device based on natural language processing, including a processor, a memory and a computer program / instructions stored on the memory, the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the aforementioned industrial Internet information model modeling optimization method based on natural language processing.

[0137] The present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the aforementioned industrial Internet information model modeling optimization method based on natural language processing.

[0138] The present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned industrial Internet information model modeling optimization method based on natural language processing.

[0139] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0140] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0141] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0142] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for optimizing industrial Internet information modeling based on natural language processing, characterized in that: The method comprises the following steps: Inputting model prompt information into a preset natural language processing model to extract modeling elements; the model prompt information includes first-stage prompt information and second-stage prompt information; the first-stage prompt information is used to prompt the preset natural language processing model to extract entity types from a preset data source according to a preset entity type list; the second-stage prompt information is used to prompt the preset natural language processing model to extract modeling elements based on the extracted entity types; Dividing the modeling elements into a corresponding number of modeling element clusters based on a target number of clusters; the target number of clusters refers to the number of clusters corresponding to the largest silhouette coefficient among the silhouette coefficients; before dividing the modeling elements into the corresponding number of modeling element clusters based on the target number of clusters, the method further includes: performing cluster analysis on the modeling elements based on a preset number of clusters set, calculating the silhouette coefficients corresponding to the clustering of the modeling elements under different numbers of clusters; and determining the target number of clusters in the preset number of clusters set based on the silhouette coefficients; An information model is constructed based on the modeling element cluster; the information model is used to indicate the relationship between the various modeling elements, including internal relationships corresponding to the modeling elements within each modeling element cluster and cross-cluster relationships corresponding to the modeling elements between different modeling element clusters; the information model includes a tree structure layer, or a combination of the tree structure layer and a network structure layer; Before dividing the modeling elements into a corresponding number of modeling element clusters based on the target cluster number, the method further includes: Determining a target semantic similarity threshold corresponding to each modeling element cluster; comparing the modeling elements in each modeling element cluster pairwise to determine the semantic similarity corresponding to each pair of modeling elements; and aggregating, within each modeling element cluster, modeling element pairs whose semantic similarity is greater than or equal to the corresponding target semantic similarity threshold; The determining of the target semantic similarity threshold corresponding to each modeling element cluster includes: analyzing, within a preset semantic similarity threshold range, to obtain a first relationship between the silhouette coefficient corresponding to each modeling element cluster and the semantic similarity, and a second relationship between the number of modeling elements corresponding to each modeling element cluster and the semantic similarity; and determining, based on the first relationship and the second relationship, the target semantic similarity threshold corresponding to each modeling element cluster within the preset semantic similarity threshold range.

2. The industrial Internet information modeling optimization method based on natural language processing according to claim 1 is characterized in that: Before determining the target semantic similarity threshold corresponding to each modeling element cluster, the method further includes: determining the number of modeling elements in each of the modeling element clusters; For the modeling element clusters having more than 2 modeling elements, the step of determining the target semantic similarity threshold corresponding to each modeling element cluster is performed.

3. The industrial Internet information modeling optimization method based on natural language processing according to claim 1 is characterized in that: The method further includes inputting the model prompt information into a preset natural language processing model to extract the modeling elements, and converting the modeling elements into corresponding word vector forms.

4. The industrial Internet information modeling optimization method based on natural language processing according to claim 1 is characterized in that: The tree structure layer includes an object class layer corresponding to the information model, a cluster category layer corresponding to each modeling element cluster, and a representative element layer corresponding to each modeling element cluster; the object class layer is used to indicate the modeling object corresponding to the information model; The cluster category layer is used to indicate the category corresponding to each modeling element cluster; the representative element layer is used to indicate the representative element corresponding to each modeling element cluster.

5. The method for optimizing the industrial Internet information model based on natural language processing according to claim 4 is characterized in that: The constructing of the information model based on the modeling element cluster includes: Constructing the network structure layer based on the modeling element clusters and association relationships; Determine the average word vector of the modeling element word vectors in each modeling element cluster respectively, and compare the semantic similarity with the preset modeling element cluster category to obtain a comparison result; the comparison result is used to indicate the similarity between each modeling element cluster and each preset modeling element cluster category; Based on the comparison results, determining the category corresponding to each modeling element cluster and constructing the cluster category layer; Based on the internal relationship, the representative element corresponding to each modeling element cluster is determined, and the representative element layer is constructed.

6. The industrial Internet information modeling optimization method based on natural language processing according to claim 1 is characterized in that: The constructing of the information model based on the modeling element cluster further includes: Constructing at least two information sub-models based on the modeling element cluster; The at least two information sub-models are combined to obtain the information model.

7. The industrial Internet information modeling optimization method based on natural language processing according to claim 1 is characterized in that: The preset data source includes a text data source and / or a knowledge base corresponding to the preset natural language processing model; In the case where the preset data source includes the text data source, before inputting the model prompt information into the preset natural language processing model and extracting the modeling elements, it also includes: inputting the text data source as an attachment into the preset natural language processing model.

Citation Information

Patent Citations

  • Natural language processing model construction method based on big data

    CN114328939A

  • Text classification method and device based on reinforcement learning, computer equipment and medium

    CN114780727A