Automated Engineering Learning Framework for Cognitive Engineering

By applying machine learning and artificial intelligence in the engineering stage of industrial automation, the problem of underdeveloped application of AI at this stage is solved, and the effect of improving the efficiency and accuracy of automation engineering tasks is achieved.

CN114730169BActive Publication Date: 2025-05-06SIEMENS AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080072079.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-16
Filing Date
2020-08-11
Publication Date
2025-05-06
Estimated Expiration
2040-08-11

AI Technical Summary

Technical Problem

Due to the various technical problems of application of artificial intelligence (AI) in the engineering stage of industrial automation, such as scarce engineering data, short duration of engineering stage, and difficulty in acquiring human intentions and knowledge, the application of AI in the engineering stage has not been fully developed.

Method used

By providing methods, systems, and devices to complete automated engineering tasks using machine learning or artificial intelligence. Specific implementations include training a neural network based on automated source code, generating code embeddings to perform code classification and semantic code search, and generating probability distributions associated with hardware components based on some hardware configurations.

Benefits of technology

The application of AI in the engineering stage of industrial automation is achieved, which improves the efficiency and accuracy of automation engineering tasks, reduces the cost of engineering stage, and improves the degree of automation of code reusability and hardware configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114730169B_ABST
    Figure CN114730169B_ABST
Patent Text Reader

Abstract

Due to the large amount of data available from sensors, the application of artificial intelligence (AI) in industrial automation is mainly concentrated in the runtime stage. The present invention describes methods, systems and devices that can use machine learning or artificial intelligence (AI) to complete automation engineering tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 887,822, filed on August 16, 2019, the disclosure of which is incorporated herein by reference in its entirety. Background Art

[0003] Industrial automation is undergoing a technological revolution in smart production, thanks to the latest breakthroughs in smart robots, sensors, big data, advanced materials, advanced supercomputing, the Internet of Things, cyber-physical systems, and artificial intelligence. These systems are currently being integrated into factories, power grids, transportation systems, buildings, homes, and consumer devices through software. The life cycle of an industrial automation system can be divided into two phases: engineering and runtime. The engineering phase refers to the activities that occur before the system is operational. These engineering activities can include hardware selection, hardware configuration, automation code development, testing, and simulation. On the other hand, the runtime phase refers to the activities that occur during the operation of the system. Example runtime activities include control, signal processing, monitoring, prediction, etc.

[0004] Due to the availability of large amounts of data from sensors, the application of artificial intelligence (AI) in industrial automation is mainly focused on the runtime phase. For example, time series prediction algorithms have been very successful in signal processing. Planning and constraint satisfaction are used for control and control code generation. Anomaly detection algorithms have become very popular in system monitoring against cyber attacks. Probabilistic graphical models and neural networks are used for prediction and health management of complex cyber-physical systems such as wind and gas turbines.

[0005] However, this paper recognizes that the use of AI in the engineering phase has not yet been fully developed due to various technical issues associated with applying AI in the engineering phase. Summary of the invention

[0006] Embodiments of the present invention address and overcome one or more of the shortcomings described herein by providing methods, systems, and apparatus that can use machine learning or artificial intelligence (AI) to complete automated engineering tasks.

[0007] In example aspects, a computing system, such as an automation engineering system, can use machine learning to perform various automation engineering tasks. The system can train a neural network based on an automation source code, for example, a neural network of a learning module can be configured to learn a programmable logic controller (PLC) source code of a programmable logic controller (PLC) and an automation source code of a manufacturing system. The system can generate code embeddings according to training. For example, a learning module can be configured to generate a code embedding that defines a PLC source code snippet and an automation source code snippet as corresponding vectors in a space based on learning of the PLC source code and the automation source code. Based on the code embedding, the system can determine a category associated with a particular code snippet. Alternatively or additionally, based on the code embedding, the system can determine that different code snippets are similar to a given source code snippet. In particular, in some cases, a semantic code search module can generate a score associated with a neighbor compared to a vector in a space defined by the source code to determine that different code snippets are similar to a given source code (e.g., a PLC source code or an automation source code) in terms of code syntax or code functionality. In some cases, the semantic code search module can also be configured to score neighbors compared to vectors in the space to determine that different code snippets are similar to source code snippets (e.g., PLC source code or automation source code) in terms of code functionality but different in code syntax.

[0008] In another example aspect, the system can generate a probability distribution associated with hardware components based on a partial hardware configuration. For example, the hardware recommendation module can be configured to learn a hardware configuration for completing an automation engineering task. The hardware recommendation module can also be configured to receive a partial hardware configuration. Based on the learned hardware configuration, the hardware recommendation module can determine a plurality of hardware components and corresponding probabilities associated with the plurality of hardware components. The corresponding probabilities can define a hardware component that completes a partial hardware configuration of the plurality of hardware components, thereby defining a complete hardware configuration that completes the automation engineering task. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The foregoing and other aspects of the present invention can be best understood from the following detailed description when read in conjunction with the accompanying drawings. For the purpose of illustrating the invention, presently preferred embodiments are shown in the accompanying drawings, however, it should be understood that the invention is not limited to the specific means disclosed. The drawings include the following drawings:

[0010] Figure 1 is a block diagram of an example automation engineering system configured to learn automation code and hardware to automate automation engineering tasks according to various embodiments described herein.

[0011] Figure 2 is a graph depicting example performance results of an automated engineering system predicting code with human annotations.

[0012] Figure 3 is a graph depicting example performance results of different classifiers for an automated engineering system.

[0013] Figure 4 is a graph depicting example performance results of an automated engineering system using code to predict code.

[0014] Figure 5 Example code snippets that may be input into an automation engineering system are shown so that the automation engineering system may identify similarities between the code snippets.

[0015] Figure 6 is a flow chart illustrating example operations that may be performed by an automation engineering system according to example embodiments.

[0016] Figure 7 An example of a computing environment is shown in which embodiments of the invention may be implemented. DETAILED DESCRIPTION

[0017] First, the present invention recognizes that, to date, various issues have limited the ability to apply artificial intelligence to the engineering phase of industrial automation. For example, engineering data is generally scarce due to its proprietary nature. As another example, the duration of the engineering phase is typically shorter than the runtime phase. For example, some industrial automation systems have been operating for more than 30 years. Therefore, the engineering phase is generally considered less important than the operational phase. As another example, capturing human intent and knowledge is often a daunting technical challenge. Capturing engineering expertise in expert systems can be time-consuming and expensive. However, this article recognizes that manufacturing is shifting from centralized mass production to distributed batch production. Even if the duration of the engineering phase is shorter relative to the runtime phase, this transition increases the cost of the engineering phase relative to the total cost of automation.

[0018] However, the present invention further recognizes that automation engineering involves various technical challenges that distinguish it from general-purpose software. For example, software development for automation engineering is typically done by automation engineers rather than software experts, which can result in non-reusable and imprecise code. In addition, the interaction of automation engineering software (AES) with the physical world typically requires engineers to understand the hardware configuration that defines how sensors, actuators, and other hardware are connected to the digital and analog inputs and outputs of a given system. Therefore, defining the hardware configuration for the development of engineering tools may require multiple iterations between hardware and software development. This tight coupling between hardware and software products may result in longer development cycles for AES and other effects.

[0019] According to various embodiments described in the present invention, automated engineering tasks are performed using machine learning. In particular, for example, machine learning can be used according to embodiments described in the present invention to perform code classification, semantic code search, and hardware recommendation. Code classification generally refers to organizing existing and new code into functional libraries, such as code functions with different categories. Example categories include, but are not limited to, signal processing, signal generation, and robotic motion control. In some cases, as production requirements change, new functions may need to be efficiently integrated into production. For another example, frequent reconfiguration of production systems may require a higher degree of code reusability. The present invention recognizes that semantic code search can, for example, help engineers with productivity by allowing them to find functionally equivalent code. In some cases, similar code can inform their decision-making when writing software to automate hardware they have never experienced before. With regard to the task of hardware recommendation, automated engineering includes the task of integrating various hardware components with software to achieve production goals. Therefore, selecting a good hardware configuration may be a key activity. In an example embodiment, the system can recommend hardware to help engineers perform hardware configuration automation. For example, in some cases, given a partial hardware configuration, the system can predict the complete hardware configuration.

[0020] Now refer to Figure 1 According to various embodiments described herein, an example architecture or automated engineering system 100 can perform automated engineering tasks. The system 100 includes a code classification module 102, a semantic code search module 104, and a hardware recommendation module 106, configured to perform various automated engineering tasks, as described herein. The system 100 also includes a learning module 108, which can include one or more neural networks.

[0021] The system 100 may include one or more processors and memory having stored applications, agents, and computer program modules to implement embodiments of the present disclosure, including a learning module 108, a code classification module 102, a semantic code search module 104, and a hardware recommendation module 106. A module may refer to a software component that performs one or more functions. Each module may be a discrete unit, or the functions of multiple modules may be combined into one or more units that form part of a larger program.

[0022] Various data may be collected for training the learning module 108. In an example, the learning module 108 may be trained on an industrial data set 110 and a manufacturer data set 112. The industrial data set 110 may include source code 114 used by an industrial automation system. In particular, for example, the source code 114 may include a program for a programmable logic controller (PLC). The manufacturer data set 112 may include source code 116 and hardware configuration data 118 from a manufacturing automation system. In an example use case, the source code 116 and the hardware configuration data 118 may include data collected from a public source, such as the Arduino Project Hub at create.arduino.cc. Example test data sets are referenced throughout this disclosure to illustrate example results generated by the system 100, but it should be understood that the embodiments are not limited to example data sets and example results from example data sets. In particular, the example data set includes 2,927 Arduino projects and 683 PLC projects. Various source code 116 and / or hardware configuration data 118 may be associated with the respective projects. In addition, the manufacturer data set 112 may include various metadata indicating various information related to the corresponding code or data, such as, but not limited to, the project category, project title, summary, various tags or descriptions of the project, and hardware configuration (e.g., components and supplies) associated with a particular code. In the example, the manufacturer data set 112 includes data associated with multiple projects, and each of the multiple projects is classified so that each project is associated with a category. In the example, the code classification module 102 uses the category of the project as a label. As further described in the present invention, the code classification module 102 can apply a label to a given code to predict the classification of the code. In some cases, the title, summary, label, and description metadata of a given project can provide an upper baseline for label classification. That is, for example, providing additional features can improve the predictive performance of the machine learning model.

[0023] Similarly, the hardware configuration data 118 can be associated with multiple projects. In an example, the hardware configuration data 118 includes a list of components required for a particular project. In some examples, the hardware configuration data 118 is managed at the data management module 120 to clean the data. The data management module 120 can perform automated operations to clean the data. In addition or alternatively, an automation expert can clean the data at the data management module 120. For example, in an example test case, the hardware configuration data 118 includes more than 6,500 unique components for 2,927 projects. Continuing with this example, at the data management module 120, it is determined that some of the 6,500 components are not actually unique. For example, in some cases, the same component has different names in different lists (e.g., "resister 10k" vs. "Resistor 10k ohm"). Such naming differences can be removed at the data management module 120. In some examples, components can be renamed according to their functions in order to manage the hardware configuration data 118. In particular, for example, an abstract functional level can be defined for hardware in order to correctly name components. Example categories in the first level of abstraction (e.g., level 1) may include, but are not limited to, actuators, Arduino, communications, electronics, human-machine interface, materials, memory, power supplies, and sensors. Example categories in the second or finer level of abstraction (e.g., level 2) may include, but are not limited to: actuators {acoustic, air, flow, motor}, Arduino {large, medium, other, small}, communications {ethernet, fiber, radio, serial, wifi}, electronics {capacitors, diodes, relays, resistors, transistors}, human-machine interface {buttons, displays, inputs, LEDs}, materials {adapters, terminal blocks, screws, solder, wiring}, memory {solid state}, power supplies {batteries, regulators, shifters, power supplies, transformers}, and sensors {acceleration, acoustics, cameras, encoders, fluids, gps, miscellaneous, optical, imaging, pv, rfid, temperature}.

[0024] Although two levels of abstraction for classifying hardware components are presented for purposes of example, it should be understood that hardware may be abstracted to additional or alternative levels to manage data, and all such data management is considered to be within the scope of the present invention. In addition, although specific classifications are presented for purposes of example, it should be understood that hardware components may be classified according to alternative, fewer, and / or additional classifications, and all such classifications are considered to be within the scope of the present invention.

[0025] Similar to the manufacturer data set 112, the industrial data set 110 may include data collected from public sources. For example, PLC code may be collected from the OS CAT library at www.oscat.de, which is a publicly available PLC program library that is independent of vendors. Different categories of reusable code functions may be found in the OSCAT library, such as signal processing (SIGPRO), geometric calculations (GEOMETRY), and string operations (STRINGS). In the example, the collected source code 114 may include its associated categories, such as in the file annotation section. These categories may be extracted and may be used as tags by the code classification module 102. In the example, the source code 114 is written in the SCL language, but it should be understood that the embodiments are not limited thereto.

[0026] Continue to refer to Figure 1 , given a code snippet, the code classification module 102 can predict a label associated with the code. In some cases, before the code is input to the code classification module 102, it is preprocessed at one or more preprocessing modules (e.g., a first or PLC code preprocessor 122 or a second code preprocessor 124). For example, the PLC code preprocessor 122 can process the source code 114 to extract various features from the source code 114. Similarly, the second code preprocessor 124 can process the source code 116 to extract various features. For example, the preprocessors 122 and 124 can discard comments and / or references to hardware so that the learning module 108 is input with high-quality data. Table 1 below illustrates example features that can be extracted at the PLC code preprocessor 122 and the second code preprocessor 124. In the example, the manufacturer data set 112 contains more features than the industrial data set 110, and therefore, there are some features in the manufacturer data set 112 that are not available in the industrial data set 110. In particular, the example source code 114 does not contain includes functions and project data, such as tags, titles, descriptions, and components. However, it should be understood that the features in the table are presented as examples and that alternative or additional features may be extracted according to other example embodiments.

[0027]

[0028]

[0029] Table 1

[0030] Continue to refer to Figure 1, the system 100 may also include a feature selection module 126, which is configured to select features from the features extracted by the PLC code preprocessor 122 and the second code preprocessor 124. The feature selection module 126 can select different features to combine different feature sets with each other, thereby adjusting the learning module 108 to perform feature space research. In addition, the feature selection module 126 can format various features in a manner suitable for machine learning. For example, the code can be represented by different feature combinations, such as includes functions, functions, annotations, symbols, and keywords. For another example, the code document can be represented by a combination of tags, titles, and descriptions. Alternatively, in some cases, code representation and code document features can be combined.

[0031] Based on the selected features or feature combinations from the feature selection module 126, the learning module 108 can generate a code embedding 128. The feature selection module 126 can generate text representations of the selected features, and the learning module 108 can embed those text representations with vectors so that the source code 114 and the source code 116 are associated with multiple vectors. Therefore, the learning module 108 can map the code to vectors in the space. In addition, for example, the learning module 108 can map code segments that are similar to each other to points that are close to each other in the space. The learning module 108 can perform various algorithms or calculations to generate the code embedding 128. For example, the learning module 108 can include a first or document to vector (doc2vec) processor 130 configured to perform a doc2vec algorithm and a second or term frequency-inverse document frequency (tf-idf) processor 132 configured to perform a tf-idf algorithm, but it should be understood that alternative processors and methods can be implemented to generate the code embedding 128 as needed.

[0032] In the example, code embeddings 128 are generated by a first or doc2vec processor 130 and a second or tf-idf processor 132, and these embeddings are compared to each other. The learning module 108 can execute the doc2vec processor 130 to generate hyperparameters of interest, which may include embedding dimensions and training algorithms (e.g., distributed memory and distributed bag of words). In the example further described herein, the negative sample is 5.

[0033] In some examples, after the code embeddings 128 are generated, the code classification module 102 can be trained. In particular, the code classification module 102 can include a supervised neural network or model. The code embeddings 128 can be input into the code classification module 102 such that the code embeddings 128 define the input samples. The encoding labels can be output by the encoding classification module 102 such that the encoding labels define target values ​​for the encoding classification module 102. The target values ​​or encoding labels can correspond to the categories of the original industrial dataset 110 and the manufacturer dataset 112. The code classification module 102 can include various classifiers, such as a logistic regression classifier 134, a random forest classifier 136, etc. In some cases, the classifiers can be compared, such as using an F 1 - Rating indicator. F 1 - Scoring usually considers precision (p) and recall (r) to measure the accuracy of the test. Mathematically, F 1 -score can be defined as the harmonic mean of precision (p) and recall (r). In various examples described in this invention, the F1-score can be calculated as

[0034] In an example, the code classification module 102 uses a lower bound and an upper bound for the code label classification. In some cases, the lower bound can be defined by training the code label classifier using random embeddings, and the upper bound can be defined by training the code label classifier using human annotations. In particular, for example, the manufacturer dataset 112 can include annotations, such as labels and descriptions, which can be used in combination or independently of each other. In an example, referring to Figure 2 , annotation configurations (e.g., tags, descriptions, tags, and descriptions) can be embedded into code embedding 128 using doc2vec processor 130 and tf-idf processor 132, and the tag classifications produced by each of processors 130 and 132 can be processed using the respective F 1 Ratings for comparison.

[0035] Also refer to Figure 2 , showing an example of a doc2vec processor 130 and a tf-idf processor 132 1 In particular, the F of the tag annotation configuration 202, the description annotation configuration 204, and the description and tag annotation configuration 206 are shown. 1 Rating 200. Figure 2As shown, according to the example, the doc2vec processor 130 has better performance than the tf-idf processor 132. In the example, the embedding dimension of the doc2vec processor 130 is set to 50, and the tf-idf processor 132 generates an embedding dimension of 1,469 for the tag annotation configuration 202; an embedding dimension of 66,310 for the description annotation configuration 204; and an embedding dimension of 66,634 for the description and tag annotation configuration 206. As shown in the example, the description and tag annotation configuration 206 provides F 1 The upper limit of the score of 200 is 0.8213.

[0036] Also refer to Figure 3 , shows an example F 1 The score 300 illustrates an example comparison of the performance of a logistic regression classifier 134 and a random forest classifier 136 using an example 50-dimensional code embedding 128 from the doc2vec processor 130. In this example, the logistic regression classifier 134 performs better than the random forest classifier 136, but it should be understood that performance may vary, for example, based on the input data and other factors. In this example, a lower bound can be established by generating a 50-dimensional random code embedding 128 and predicting labels using the logistic regression classifier 134. Also refer to Figure 4 , Example F 1 The score of 400 comes from the code classification module 102 using the code to predict the label. In particular, example F 1 The score 400 indicates that the lower bound of 0.3538 is defined by the tf-idf processor 132 and the doc2vec processor 130. The feature selection module 126 may select the example feature 402 for predicting the label. In some cases, after establishing the upper and lower bounds, different code features 402 may be used to predict the label. Figure 4 In the example shown, the code embedding 128 including the features "includes function" and "function" produces slightly better performance than the random baseline because, for example, the information contained in these features is limited. In contrast, according to the example, the code embedding 128 including other code features significantly improves the classification accuracy. For example, the code embedding 128 including the code features "token" and "code" respectively produces similar F 1 Ratings 400, 0.63 and 0.67. Figure 4 The example results of also illustrate that the annotation feature with a score of 0.67 can contain valuable information that can be used to predict code labels. For example, the code embedding 128 including the combination of the feature code and annotation and the code embedding 128 including the combination of the feature code and title produce F 1 A score of 400 is approximately 0.71, which is the highest score in the example. Thus, according to various embodiments, without being bound by theory, the prediction performance with code feature embeddings can be better than human annotation embeddings.

[0037] Reference again Figure 1 , the system 100 may also include a semantic code search module 104, which is configured to find a program in response to a code or code snippet, wherein the found program is similar to the code or code snippet. In the context of automation engineering, according to various embodiments, similarity can be defined in terms of syntax and functionality. For example, syntax similarity 504 can help engineers find useful functions in a given context, and functional similarity can let engineers know that other automation solutions have been designed. In some cases, the doc2vec processor 130 can bring similar documents close to each other in the embedding space (e.g., code embedding 128). For example, for a given code embedding 128 associated with a source code snippet 116, the nearest neighbor closest to the given code embedding 128 can represent code similar to the code snippet. Therefore, the semantic code search module 104 can identify one or more nearest neighbors 138 associated with the code snippet based on the code embedding 128 of the code snippet. In addition, according to various embodiments, the function structure can be captured in the code embedding 128, so the semantic code search module 104 can not only find the nearest neighbors 138 of documents with similar syntax.

[0038] In the example, the quality of the code embedding 128 is verified by randomly sampling 50 code snippets from the source code 116. In this example, the sample code snippets can be scored based on the similarity of each code snippet to its first 3 nearest neighbors. In some cases, for each code snippet pairing, a similarity rating of code syntax and code structure is given. For example, a rating of 1 can represent similarity, and a rating of 0 can represent a lack of similarity. Code syntax can refer to the use of similar variable and function names. Code structure can refer to the use of similar code arrangements, such as, for example, if-then-else and for loops. In the example further described in the present invention, software experts can provide ratings and associated confidence scores. As an example and not a limitation, the confidence score can range from 1 (lowest confidence) to 5 (highest confidence). Regardless of how the score is calculated, the confidence score can represent the confidence level of a given expert in the evaluation process. Continuing with the example described in the present invention, 5 of the 50 samples are eliminated during the expert evaluation. For example, if at least one of the first 3 nearest neighbors is an empty file or contains code in a programming language different from the corresponding code snippet or example, the example can be eliminated.

[0039] Continuing with the example introduced above, Table 2 below shows the average code syntax and code structure similarity scores given by experts. Table 2 includes high confidence levels (e.g., average confidence ≥ 4.5) to eliminate the impact of uncertain answers. The agreement between experts can also be measured by Fleiss Kappa (k). As shown in the example shown in Table 2, the syntax and structure similarity scores of the first 1 neighbor are very high (0.68 and 0.61, respectively), but the similarity scores of the first 2 and first 3 neighbors are significantly reduced (below 0.50). In addition, in this example, the experts' syntax similarity scores are basically consistent (0.61≤k≤0.80) and the structural similarity scores are consistent (0.41≤k≤0.60). Therefore, based on the example results, the code embedding 128 from the doc2vec processor 130 can capture both syntax similarity and structural similarity.

[0040]

[0041]

[0042] Table 2

[0043] For another example, referring to Table 3 below, in another example use case, three similar code snippets are selected, and three different code snippets are selected. As shown in Table 3, the cosine similarity associated with the code embedding 128 of the corresponding code snippets can be measured. The example code snippets shown in Table 30 show that the agreement between experts is very strong, and the confidence in the similarity and lack of similarity of the first 3 nearest neighbors is very high. The results of this example illustrate that the code snippets that the experts believe are most similar to each other are also defined as similar by the semantic code search module 104. In particular, the code snippets that the experts believe are most similar to each other are also close to each other in the code embedding space defined by the learning module 108. In addition, the code snippets that the experts believe are most different from each other are far apart in the code embedding space defined by the learning module 108.

[0044]

[0045] Table 3

[0046] Now refer to Figure 5, a first example code snippet 500 and a second example code snippet 502 are shown. The example code snippets 500 and 502 define similar Arduino code snippets generated by ArduCode of the source code 116. The first example code snippet 500 and the second example code snippet 502 define different levels of similarity. For example, the Arduino program has setup() and loop() functions to initialize the program and specify the control logic executed each cycle. From a syntax perspective, the two example programs use the same example standard functions: 0pinMode() configures the hardware-connected Arduino board pin as input or output; analogRead() reads an analog value from a pin; Serial.print() prints ASCII characters via the serial port; delay() pauses the program for an amount of time specified by a parameter (in ms); and analogWrite() writes an analog value to a pin.

[0047] Still refer to Figure 5 Semantically, the example program reads sensor values ​​(e.g., 1 value in example code snippet 500 and 3 values ​​in example code snippet 502); scales the sensor values ​​to a range (e.g., from 300-1024 to 0-255 using map() in example code snippet 50054 and (x+100) / 4 in example code snippet 502); prints the scaled sensor values ​​via the serial port; writes analog values ​​to LEDs (e.g., a single LED in example code snippet 500 and three LEDs in example code snippet 502); and pauses the program (e.g., 10ms in example code snippet 500 and 100ms in example code snippet 502). In the example program including the first and second example code snippets 500 and 502, the order in which the above operations are scheduled is different. Functionally, example programs 500 and 502 perform the same task, namely, creating a heat map for sensor values ​​using LEDs. Although in Figure 5 While there are some syntactic similarities in the examples shown, according to various embodiments, the semantic code search module 104 can capture semantic and functional similarities between programs or code snippets.

[0048] Refer again Figure 1, the example system 100 may also include a hardware recommendation module 106, which may be configured to predict hardware components for completing a task given a partial list of hardware components. For example, in some cases, given a partial list of hardware components, the hardware recommendation module 106 may identify other hardware components that are often used in conjunction with the partial list. In some examples, the hardware configuration data 118 may be messy or inconsistent. In such examples, the hardware configuration data 118 may be cleaned in the data management module 120 to define silver standard or clean hardware configuration data 140. In some cases, the automation expert may use the data management module 120 to generate silver standard data. The hardware recommendation module 106 may use the clean hardware configuration data 140 to learn the joint probability distribution of the hardware components represented by the hardware configuration data 118. Then, given a partial list of hardware components, the hardware recommendation module 106 may generate a conditional probability associated with a corresponding hardware component for completing the partial list. To perform its tasks, the hardware recommendation module 106 may include various machine learning or neural networks. For example, the hardware recommendation module 106 may include a Bayesian network module 142 and an autoencoder module 144 .

[0049] In an example implementation, the Bayesian network module 142 and the autoencoder module 144 generate predictions or recommendations for hardware based on random hardware configurations defined in the clean hardware configuration data 140. In this example, the categories of the clean hardware configuration data 140 define random variables of the Bayesian network learned by the Bayesian network module 142. In this example implementation, the Bayesian network module 142 can use Pomegrate to learn the structure of the Bayesian network to fit the model with 70% of the clean hardware configuration data 140. In this example, the Bayesian network for the level 1 component consists of 9 nodes and the network for the level 2 component consists of 45 nodes. However, the present invention recognizes that with the number of variables (45) in the level 2 configuration, the initialization of the Bayesian network requires a large amount of time. Therefore, in this example, the autoencoder module 144 is implemented as Keras to learn a low-dimensional representation of the clean hardware configuration data 140. The decoder of the autoencoder module 144 can learn to reconstruct the original input from the low-dimensional representation. To avoid overfitting, the autoencoder module 144 may use L1 and L2 regularization terms.

[0050] In an example, the hardware recommendation module 106 can recommend a predetermined number (represented by k in the present invention) of hardware components. Therefore, the hardware recommendation module 106 can recommend the top k hardware components, and the resulting model can be evaluated in terms of precision@k. Precision(p)@k can represent the relevant portion of the recommended hardware components in the top k sets. In the example described in the present invention represented in Table 4 below, for each hardware configuration in the test data, one hardware component is omitted and its precision@k is measured. Table 4 shows example results of a random baseline, a Bayesian network module 142, and an autoencoder module 144. As shown, for level 1 hardware prediction, the performance of the random baseline improves linearly from p@1=0.1, p@3=0.32, and p@5=0.54 to p@9=1. The Bayesian network also improves linearly from p@1=0.32, p@3=0.59, and p@5=0.79. In this example, the autoencoder provides the best performance and the best improvement at p@1 = 0.36, p@3 = 0.79, and p@5 = 0.95. As shown in the figure, the autoencoder's p@3 performs the same as the Bayesian network's p@5, which is 0.79. In addition, as shown in the figure, the autoencoder achieves an accuracy greater than 0.95 at p@5 in the example.

[0051]

[0052] Table 4

[0053] 42 Continuing with this example, as described above, learning a Bayesian network for the hardware components of Level 2 may be computationally impractical. Therefore, in this example, the autoencoder module 144 learns the hardware components of Level 2 to define the example p@k results shown in Table 5 below. Referring to Table 5, according to the example, the overall p@k of the autoencoder module 144 for Level 2 is lower than Level 1 because the example Level 2 hardware configuration is sparser than Level 1. However, the improvement of this example over the random baseline is 10 times that of p@1, 5 times that of p@3, 4 times that of p@5, and 3 times that of p@10.

[0054]

[0055] Table 5

[0056] Now refer to Figure 6, a computing system, such as the automation engineering system 100, may perform example operations 600 to perform various automation engineering tasks. At 602, the system may train a neural network based on automation source code. For example, the neural network of the learning module 108 may be configured to learn a programmable logic controller (PLC) source code for a programmable logic controller (PLC) and an automation source code for a manufacturing system. At 604, the system may generate a code embedding based on the training. For example, the learning module 108 may be configured to generate a code embedding that defines a PLC source code snippet and an automation source code snippet as corresponding vectors in a space based on learning the PLC source code and the automation source code. At 606, based on the code embedding, the system may determine a category associated with a particular code snippet.

[0057] For example, the code classification module 102 may receive a PLC code embedding from the learning module 108, wherein the PLC code embedding defines a segment of the PLC source code as a vector in space. Based on the PLC code embedding, the code classification module 102 may determine a category associated with the PLC source code segment. Alternatively or additionally, the code classification module 102 may be further configured to receive a manufacturing code embedding from the learning module 108, wherein the manufacturing code embedding defines a segment of the automation source code as a vector in space. Based on the manufacturing code embedding, the code classification module 108 may determine a category associated with the automation source code segment. In some cases, the system may extract multiple features from the PLC source code and the automation source code. The system may include a feature selection module 126, which may be configured to select a specific feature or a feature combination from a plurality of features extracted from the PLC source code and the automation source code. The neural network may be adjusted based on one or more selected features or a combination of selected features. For example, the feature selection module 126 may be configured to provide the features selected from the PLC source code and the automation source code to the learning module in one or more combinations so as to adjust the learning module based on one or more combinations of selected features.

[0058] Still refer to Figure 6, at 608, the system may determine that different code snippets are similar to a given source code snippet. For example, the system may include a semantic code search module 104, which may be configured to receive a PLC code embedding from a learning module. Based on the PLC code embedding, the semantic code search module 104 may determine that different code snippets define neighbors near a vector in a space of the PLC source code snippet, thereby determining that different code snippets are similar to the PLC source code snippet. Alternatively or additionally, the system may generate a specific manufacturing code embedding that defines a specific snippet of the automation source code as a vector in a space. Based on the manufacturing code embedding, the semantic code search module 104 may determine that different code snippets define neighbors near a vector in a space of a specific automation source code snippet, thereby determining that different code snippets are similar to specific automation source code snippets. In particular, the semantic code search module 104 may generate a score associated with a neighbor compared to a vector in a space to determine that different code snippets are similar to a given source code (e.g., a PLC source code or an automation source code) in terms of code syntax or code functionality. In some cases, the semantic code search module 104 may also be configured to score neighbors compared to vectors in space to determine that different code snippets are similar to source code snippets (e.g., PLC source code or automation source code) in code functionality but different in code syntax.

[0059] Continue to refer to Figure 6 At 610, the system may generate a probability distribution associated with the hardware components based on the partial hardware configuration. For example, the hardware recommendation module 106 may be configured to learn a hardware configuration for completing the automation engineering task. The hardware recommendation module may also be configured to receive the partial hardware configuration. Based on learning the hardware configuration, the hardware recommendation module 106 may determine a plurality of hardware components and corresponding probabilities associated with the plurality of hardware components. The corresponding probabilities may define a selection of a hardware component that completes the partial hardware configuration from the plurality of hardware components, thereby defining a complete hardware configuration for completing the automation engineering task.

[0060] Figure 7 An example of a computing environment in which embodiments of the present invention may be implemented is shown. The computing environment or automation engineering system 700 includes a computer system 510, which may include a communication mechanism such as a system bus 521 or other communication mechanism for transferring information within the computer system 510. The computer system 510 also includes one or more processors 520 coupled to the system bus 521 for processing information. For example, the code classification module 102, the semantic code search module 104, the hardware recommendation module 106, and the learning module 108 may include or be coupled to one or more processors 520.

[0061] The processor 520 may include one or more central processing units (CPUs), graphics processing units (CPUs), or any other processors known in the art. And generally, the processor as described in the present invention is a device for executing machine-readable instructions stored on a computer-readable medium for performing tasks, and may include any one or a combination of hardware and firmware. The processor may also include a memory storing machine-readable instructions, which can be executed for performing tasks. The processor acts on information by manipulating, analyzing, modifying, converting, or transmitting information for use by an executable program or information device, and / or by routing the information to an output device. For example, the processor may use or include the capabilities of, for example, a computer, a controller, or a microprocessor, and may be adjusted using executable instructions to perform special functions that are not performed by a general-purpose computer. The processor may include any type of suitable processing unit, including but not limited to a central processing unit, a microprocessor, a reduced instruction set computer (RISC) microprocessor, a complex instruction set computer (CISC) microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on a chip (SoC), a digital signal processor (DSP), etc. In addition, the processor 520 can have any suitable micro-architecture design, including any number of constituent components, such as registers, multiplexers, arithmetic logic units, a cache controller for controlling read / write operations to cache memory, a branch predictor, etc. The micro-architecture design of the processor may be able to support any of a variety of instruction sets. The processor can be coupled (electrically coupled and / or as an executable component) to any other processor that enables interaction and / or communication between them. The user interface processor or generator is a known element that includes electronic circuits or software or a combination of both for generating display images or portions thereof. The user interface includes one or more display images that enable a user to interact with the processor or other devices. The system bus 521 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may allow information (e.g., data (including computer executable code), signaling, etc.) to be exchanged between various components of the computer system 510,

[0062] The system bus 521 may include, but is not limited to, a memory bus or memory controller, a peripheral bus, an accelerated graphics port, etc. The system bus 521 may be associated with any suitable bus architecture, including, but not limited to, Industry Standard Architecture (ISA), Micro Channel Architecture (MCA), Enhanced ISA (EISA), Video Electronics Standards Association (VESA) architecture, Accelerated Graphics Port (AGP) architecture, Peripheral Component Interconnect (PCI) architecture, PCI-Express architecture, Personal Computer Memory Card International Association (PCMCIA) architecture, Universal Serial Bus (USB) architecture, etc.

[0063] Continue to refer to Figure 7 , the computer system 510 may also include a system memory 530 coupled to the system bus 521 for storing information and instructions to be executed by the processor 520. The system memory 530 may include computer-readable storage media memory in volatile and / or non-volatile form, such as read-only memory (ROM) 531 and / or random access memory (RAM) 532. RAM 532 may include other dynamic storage devices (e.g., dynamic RAM, static RAM, and synchronous DRAM). ROM 531 may include other static storage devices (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). In addition, the system memory 530 may be used to store temporary variables or other intermediate information during the execution of instructions by the processor 520. The basic input / output system 533 (BIOS) contains basic routines that help transfer information between internal elements of the computer system 510, such as during startup, and may be stored in ROM 531. RAM 532 may contain data and / or program modules that the processor 520 can access immediately and / or is currently operating on by the processor. System memory 530 may additionally include, for example, an operating system 534, application programs 535, and other program modules 536. Application programs 535 may also include a user portal for developing applications to allow parameters to be entered and modified as needed.

[0064] The operating system 534 may be loaded into the memory 530 and may provide an interface between other application software executing on the computer system 510 and the hardware resources of the computer system 510. More specifically, the operating system 534 may include a set of computer executable instructions for managing the hardware resources of the computer system 510 and providing common services to other applications (e.g., managing memory allocation between various applications). In certain example embodiments, the operating system 534 may control the execution of one or more program modules described as being stored in the data memory 540. The operating system 534 may include any operating system now known or that may be developed in the future, including but not limited to any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.

[0065] The computer system 510 may also include a disk / media controller 543 coupled to the system bus 521 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 541 and / or a removable media drive 542 (e.g., a floppy disk drive, an optical drive, a tape drive, a flash drive, and / or a solid-state drive). The storage device 540 may be added to the computer system 510 using an appropriate device interface (e.g., a small computer system interface (SCSI), an integrated device electronics (IDE), a universal serial bus (USB), or a firewire). The storage devices 541, 542 may be external to the computer system 510.

[0066] The computer system 510 may also include a field device interface 565 coupled to the system bus 521 to control field devices 566, such as devices used in a production line. The computer system 510 may include a user input interface or GUI 561, which may include one or more input devices, such as a keyboard, touch screen, tablet, and / or pointing device, for interacting with a computer user and providing information to the processor 520.

[0067] In response to the processor 520 executing one or more sequences of one or more instructions contained in a memory such as system memory 530, the computer system 510 may perform some or all of the processing steps of an embodiment of the present invention. Such instructions may be read into the system memory 530 from another computer-readable medium of the memory 540, such as a magnetic hard disk 541 or a removable media drive 542. The magnetic hard disk 541 and / or the removable media drive 542 may contain one or more data storage devices and data files used by embodiments of the present invention. The data storage device 540 may include, but is not limited to, a database (e.g., relational, object-oriented, etc.), a file system, a flat file, a distributed data storage device in which data is stored on more than one node of a computer network, a peer-to-peer network data storage device, etc. The data storage device may store various types of data, such as skill data, sensor data, or any other data generated according to embodiments of the present invention. The data storage device contents and data files may be encrypted to increase security. The processor 520 may also be used in a multi-processing arrangement to execute one or more instruction sequences contained in the system memory 530. In alternative embodiments, hard-wired circuits may be used instead of software instructions or in combination with software instructions. Therefore, the embodiments are not limited to any specific combination of hardware circuits and software.

[0068] As described above, the computer system 510 may include at least one computer-readable medium or memory for storing instructions programmed according to embodiments of the present invention and for containing data structures, tables, records, or other data described in the present invention. As used in the present invention, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processor 520 for execution. Computer-readable media can take a variety of forms, including but not limited to non-transient, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical disks, solid-state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disks 541 or removable media drives 542. Non-limiting examples of volatile media include dynamic memory, such as system memory 530. Non-limiting examples of transmission media include coaxial cables, copper wires, and optical fibers, including lines that constitute the system bus 521. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communications.

[0069] The computer-readable medium instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and traditional process programming languages, such as "C" programming language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, as a separate software package part on the user's computer, part on the user's computer and part on a remote computer or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In some embodiments, electronic circuits including, for example, programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs) can be personalized electronic circuits to execute computer-readable program instructions by utilizing the state information of computer-readable program instructions to perform various aspects of the present invention.

[0070] The present invention describes various aspects of the present invention with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer-readable medium instructions.

[0071] The computing environment 1300 may also include a computer system 510 operating in a networked environment using logical connections to one or more remote computers, such as remote computing devices 580. The network interface 570 may communicate with other remote devices 580 or systems and / or storage devices 541, 542, for example, via a network 571. The remote computing device 580 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device, or other public network node, and typically includes many or all of the elements described above with respect to the computer system 510. When used in a network environment, the computer system 510 may include a modem 572 for establishing communications over a network 571 (e.g., the Internet). The modem 572 may be connected to the system bus 521 via the user network interface 570 or via another appropriate mechanism.

[0072] The network 571 can be any network or system known in the art, including the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection or a series of connections, a cellular telephone network, or any other network or medium that can facilitate the communication between the computer system 510 and other computers (e.g., a remote computing device 580). The network 571 can be wired, wireless, or a combination thereof. The wired connection can be implemented using Ethernet, a universal serial bus (USB), RJ-6, or any other wired connection known in the art. The wireless connection can be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular networks, satellites, or any other wireless connection methods generally known in the art. In addition, multiple networks can work alone or communicate with each other to facilitate communication in the network 571.

[0073] It should be understood that the system memory 530, Figure 7 The depicted program modules, applications, computer-executable instructions, codes, etc. are merely illustrative and not exhaustive, and the processing described as being supported by any particular module may alternatively be distributed among multiple modules or performed by different modules. In addition, various program modules, scripts, plug-ins, application programming interfaces (APIs), or any other suitable computer-executable code hosted locally on the computer system 510, the remote device 580, and / or hosted on other computing devices accessible via one or more networks 571 may be provided to support the processing by the computer system 510, the remote device 580, and / or other computing devices accessible via one or more networks 571. Figure 1 The functions and / or additional or alternative functions provided by the program modules, applications or computer executable codes shown. In addition, the functions can be modularized in different ways, so that the functions described as being composed of Figure 1The processing jointly supported by the collection of program modules depicted in the present invention may be performed by a fewer or greater number of modules, or the functionality described as supported by any particular module may be supported at least in part by another module. In addition, the program modules supporting the functionality described in the present invention may form part of one or more applications executable on any number of systems or devices according to any suitable computing model, such as a client-server model, a peer-to-peer model, etc. In addition, the processing described as being supported by Figure 1 Any functionality supported by any program module depicted in the drawings may be implemented, at least in part, in hardware and / or firmware on any number of devices.

[0074] It should also be understood that the computer system 510 may include alternative and / or additional hardware, software or firmware components other than those described or depicted without departing from the scope of the present invention. More specifically, it should be understood that the software, firmware or hardware components described as forming part of the computer system 510 are merely illustrative, and in various embodiments, some components may not exist, or additional components may be provided. Although various illustrative program modules have been depicted and described as software modules stored in the system memory 530, it should be understood that the functions described as supported by the program modules can be implemented by any combination of hardware, software and / or firmware. It should also be understood that in various embodiments, each of the above modules may represent a logical partition of the supported functions. The logical partition is depicted for the convenience of explaining the functions and may not represent the structure of the software, hardware and / or firmware used to implement the functions. Therefore, it should be understood that in various embodiments, the functions described as provided by a specific module may be provided at least in part by one or more other modules. In addition, in some embodiments, one or more of the depicted modules may not exist, while in other embodiments, additional modules that are not depicted may exist and may support at least a portion of the described functions and / or additional functions. Furthermore, while certain modules may be depicted and described as sub-modules of another module, in certain embodiments these modules may be provided as stand-alone modules or sub-modules of other modules.

[0075] Although specific embodiments of the present invention have been described, those skilled in the art will recognize that many other modifications and alternative embodiments are within the scope of the present disclosure. For example, any function and / or processing capability described with respect to a particular device or component can be performed by any other device or component. In addition, although various illustrative implementations and architectures have been described according to embodiments of the present disclosure, those skilled in the art will appreciate that many other modifications to the illustrative implementations and architectures described in the present invention are also within the scope of the present disclosure. In addition, it should be understood that any operation, element, component, data, etc. described in the present invention as being based on another operation, element, component, data, etc., can be additionally based on one or more other operations, elements, components, data, etc. Therefore, the phrase "based on" or its variants should be interpreted as "based at least in part on".

[0076] Although the embodiments have been described in language specific to structural features and / or methodological behaviors, it should be understood that the invention is not necessarily limited to the specific features or behaviors described. Instead, specific features and actions are disclosed as illustrative forms of implementing the embodiments. Conditional language, such as "may," "might," and the like, unless otherwise expressly stated or otherwise understood in the context in which they are used, is generally intended to convey that certain embodiments may include, while other embodiments do not include, certain features, elements, and / or steps. Therefore, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or steps in any way, or that one or more embodiments must include logic for determining whether such features, elements, and / or steps are included or will be performed in any particular embodiment with or without user input or prompting.

[0077] The flowchart and block diagram in the figure illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a part of a module, segment or instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the box may not appear in the order marked in the figure. For example, two boxes displayed in succession can actually be executed substantially at the same time, or these boxes can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each box of the block diagram and / or flowchart illustration, as well as the combination of boxes in the block diagram and / or flowchart illustration, can be implemented by a dedicated hardware-based system that performs a specified function or behavior or performs a combination of dedicated hardware and computer instructions.

Claims

1. An automated engineering system, comprising: one or more modules; A processor, configured to execute the one or more modules; and A memory, configured to store the one or more modules, wherein the one or more modules include: a learning module configured to learn a programmable logic controller PLC source code of a programmable logic controller PLC and an automation source code of a manufacturing system, wherein the learning module comprises a neural network and the learning module is further configured to be trained on an industrial dataset comprising PLC source code used by an industrial automation system and a manufacturer dataset comprising source code from a manufacturing automation system; a feature selection module configured to select features extracted from the automation source code of the industrial dataset and the manufacturer dataset and provide the selected features to the learning module; and Based on learning the PLC source code and the automation source code, a code embedding is generated that defines the PLC source code snippet and the automation source code snippet as corresponding vectors in a space, wherein the corresponding vectors are represented by mapping the code to a vector in a vector space, so that the PLC source code and the source code from the manufacturing automation system are associated with multiple vectors in the vector space, wherein, in order to generate the code embedding, the learning module is also configured to execute a document to vector processor and / or a term frequency-inverse document frequency processor.

2. The automation engineering system according to claim 1, wherein the one or more modules of the automation engineering system further comprises a code classification module, wherein the code classification module is configured to: receiving a PLC code embedding from the learning module, the PLC code embedding defining a PLC source code snippet as a vector in a space; and Based on the PLC code embedding, a category associated with the PLC source code snippet is determined.

3. The automation engineering system according to claim 2, wherein the code classification module is further configured to: receiving a manufacturing code embedding from the learning module, the manufacturing code embedding defining an automation source code snippet as a vector in a space; and Based on the manufacturing code embedding, a category associated with the automation source code snippet is determined.

4. The automation engineering system according to claim 1, wherein the one or more modules further comprise a feature selection module, wherein the feature selection module is configured to: selecting features extracted from the PLC source code and the automation source code; and Selected features from the PLC source code and the automation source code are provided to the learning module in one or more combinations to adjust the learning module based on the one or more combinations of selected features.

5. The automation engineering system according to claim 1, wherein the one or more modules further comprise a semantic code search module, wherein the semantic code search module is configured to: receiving a PLC code embedding from the learning module, the PLC code embedding defining a PLC source code snippet as a vector in a space; and Based on the PLC code embedding, it is determined that different code snippets define neighbors near a vector in a space of the PLC source code snippets, thereby determining that the different code snippets are similar to the PLC source code snippet.

6. The automation engineering system according to claim 1, wherein the one or more modules further comprise a semantic code search module, wherein the semantic code search module is configured to: receiving a manufacturing code embedding from the learning module, the manufacturing code embedding defining an automation source code snippet as a vector in a space; and Based on the manufacturing code embedding, it is determined that a different code snippet defines a neighbor near a vector in a space of the automation source code snippet, thereby determining that the different code snippet is similar to the automation source code snippet.

7. The automated engineering system according to claim 5, wherein: The semantic code search module is further configured to score the neighbors compared to vectors in space to determine whether the different code snippet is similar to the PLC source code snippet in code syntax or code functionality.

8. The automated engineering system according to claim 5, wherein: The semantic code search module is further configured to score the neighbors compared to vectors in space to determine that the different code snippet is similar to the PLC source code snippet in code functionality but different in code syntax.

9. The automation engineering system according to claim 1, wherein the one or more modules further comprise a hardware recommendation module, wherein the hardware recommendation module is configured to: Receive partial hardware configuration; Based on the partial hardware configuration, generating a probability distribution associated with hardware components; and Based on the probability distribution, a predetermined number of hardware components are identified to complete the partial hardware configuration.

10. A method performed by a computing system, the method comprising: training a neural network on an industrial dataset containing programmable logic controller (PLC) source code and a manufacturer dataset containing source code from manufacturing automation systems; selecting features extracted from the automation source code of the industrial dataset and the manufacturer dataset and providing the selected features to the neural network; Based on the training, a code embedding is generated that defines the PLC source code snippet and the automation source code snippet as corresponding vectors in a space, wherein the corresponding vectors are represented by mapping the code to vectors in the vector space, so that the PLC source code and the source code from the manufacturing automation system are associated with multiple vectors in the vector space, wherein in order to generate the code embedding, a document to vector processor and / or a term frequency-inverse document frequency processor is performed.

11. The method according to claim 10, further comprising: generating a specific PLC code embedding that defines a specific PLC source code snippet as a vector in a space; and Based on the particular PLC code embedding, a category associated with the particular PLC source code snippet is determined.

12. The method according to claim 10, further comprising: generating a manufacturing-specific code embedding that defines a specific automation source code snippet as a vector in a space; and Based on the specific manufacturing code embedding, a category associated with the automation source code snippet is determined.

13. The method according to claim 10, further comprising: extracting a plurality of features from the PLC source code and the automation source code; selecting a specific feature or a combination of features from the plurality of features extracted from the PLC source code and the automation source code; and The neural network is adjusted based on one or more selected features or a combination of selected features.

14. The method according to claim 10, further comprising: generating a specific PLC code embedding, wherein the specific PLC code embedding defines a specific PLC source code fragment as a vector in a space; and Based on the specific PLC code embedding, it is determined that different code snippets define neighbors near a vector in a space of the specific PLC code snippet, thereby determining that the different code snippet is similar to the specific PLC source code snippet.

15. The method according to claim 10, further comprising: generating a specific manufacturing code embedding that defines a specific automation source code snippet as a vector in a space; and Based on the specific manufacturing code embedding, it is determined that different code snippets define neighbors near a vector in a space of the specific automation source code snippet, thereby determining that the different code snippet is similar to the specific automation source code snippet.

16. The method according to claim 14, further comprising: A score associated with the neighbor compared to the vector in the space is generated to determine whether the different code snippet is similar to the PLC source code snippet in code syntax or code functionality based on the score.

17. The method according to claim 14, wherein: The method further includes generating a score associated with the neighbor compared to a vector in space, and determining, based on the score, that the different code snippet is similar to the PLC source code snippet in code functionality and different in code syntax.

18. The method according to claim 10, further comprising: Receive partial hardware configuration; Based on the partial hardware configuration, generating a probability distribution associated with the hardware components; and Based on the probability distribution, a predetermined number of hardware components are identified to complete the partial hardware configuration.

19. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 10 to 18 are implemented.

Citation Information

Patent Citations

  • Method for autonomously learning source codes by computer

    CN109977205A

  • Knowledge-based programmable logic controller with flexible in-field knowledge management and analytics

    US20170017221A1