A Deep Learning-Based Industrial Service Component Recommendation Method
Through deep learning technology, the service component recommendation method is built on edge nodes, which solves the problem of lack of computing resources at edge nodes, achieves high-precision dynamic update recommendations, and improves the utilization rate of service components.
Patent Information
- Application Number
- CN202211580842.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-09
AI Technical Summary
In distributed industrial systems, the prior art cannot effectively solve the problem of service components recommendation of edge nodes, especially due to the lack of computing resources, the recommendation method cannot be applied.
Using a deep learning-based method, the historical service component sequence and single service component description information are obtained from edge nodes, vector representation is generated through encoding, and semantic vectors are used to build a semantic library to achieve global dynamic update recommendations.
It improves the accuracy and adaptability of recommendations, can realize dynamic updates at edge nodes, solves the problem of computing resource limitations, and improves the utilization rate of service components.
Smart Images

Figure CN115858929B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an industrial service component recommendation method based on deep learning. Background Art
[0002] With the full use of the cyber-physical system as the mark of Industry 4.0, the industrial system architecture has gradually evolved from a centralized system to a distributed system architecture such as cloud-edge collaboration. Therefore, industrial software has shifted from tight coupling to loose coupling. Facing complex industrial application scenarios, interoperability and reusability need to be achieved. In the service-oriented architecture system in the software field, its flexibility, composability, and independence make it considered a potential solution. Therefore, the development of industrial software will encapsulate each functional block in a service-oriented style to form service components, so as to meet the requirements of interoperability and reusability and effectively reduce the development cost. In this context, with the gradual increase of industrial distributed applications, the corresponding number of service components has grown rapidly. How to recommend service components to improve the utilization rate of components and build value-added applications of multiple components has become a new technical problem.
[0003] In response to this technical problem, related technical solutions have proposed a service component recommendation method that clusters related service components according to attributes such as function, theme, category, domain, and service quality. However, from the perspective of service component composition, the call association between components can be regarded as the actual call situation when service components are executed. Service component recommendation based on call association is more important in the actual application of component composition. In some other related technical solutions, deep learning technology is combined to solve this problem. However, the dynamic changes of service components are often not considered. In a distributed environment, the component calls of edge nodes are constantly changing. In addition, its recommendation method may not be applicable to edge nodes due to the lack of computing resources at the edge. Summary of the Invention
[0004] In view of this, in order to at least partially solve one of the above technical problems or defects, the purpose of the embodiments of the present invention is to provide an industrial service component recommendation method based on deep learning, achieving a global dynamic update recommendation effect, and solving the problem of limited computing resources where service component recommendation technology cannot be applied to edge nodes in a centralized environment.
[0005] Specifically, the technical solution of the present application provides an industrial service component recommendation method based on deep learning, including the following steps:
[0006] Obtain the historical service component sequence and the description information of a single service component from the edge nodes;
[0007] Encode the preprocessed historical service component sequence to obtain a first vector representation, and encode the preprocessed description information to obtain a second vector representation;
[0008] Generate a first masked sequence after masking the first vector representation, and generate a second masked sequence after masking the second vector representation;
[0009] Predict the sequence adjacent to the first vector representation, combine the prediction result with the first masked sequence to obtain a first semantic vector, predict the sequence adjacent to the second vector representation, and combine the prediction result with the second masked sequence to obtain a second semantic vector;
[0010] Construct a first semantic library of the service scheduling sequence according to the first semantic vector, and construct a second semantic library of a single service component according to the second semantic vector;
[0011] Obtain a target service request, obtain a target semantic vector by matching with the vectors in the second semantic library, and obtain a target component scheduling sequence by matching according to the sequence of the target semantic vector in the first semantic library.
[0012] In a feasible embodiment of the solution of the present application, the encoding the preprocessed historical service component sequence to obtain a first vector representation, and encoding the preprocessed description information to obtain a second vector representation includes:
[0013] Obtain an input sequence after preprocessing to remove punctuation marks and whitespace characters, where the input sequence includes the preprocessed historical service component sequence and the preprocessed description information;
[0014] Add a first token to the beginning of the sequence of the input sequence to obtain a token embedding sequence, and split the token embedding sequence through a delimiter to obtain a plurality of embedding vector sequences;
[0015] Encode according to the embedding vector sequences to obtain the first vector representation and the second vector representation.
[0016] In a feasible embodiment of the solution of the present application, the encoding according to the embedding vector sequences to obtain the first vector representation and the second vector representation includes:
[0017] Input the embedding vector sequences into an encoding model, and calculate the similarity between the query vector and the keyword vector in the embedding vector sequences through the attention function in the encoding model;
[0018] Normalize the vector elements in the embedding vector sequence according to the similarity, and perform a fully connected process on the normalization result through a feed-forward network of positions to obtain an embedding vector representation; the embedding vector representation includes the first vector representation and the second vector representation.
[0019] In a feasible embodiment of the solution of the present application, after the step of inputting the embedding vector sequence into the encoding model and calculating the similarity between the query vector and the keyword vector in the embedding vector sequence through the attention function in the encoding model, it further includes:
[0020] Perform low-dimensional projection on the query vector, keyword vector, and data item calculated by the attention function;
[0021] Perform linear projection on the output vector obtained after low-dimensional projection to determine the parameter weights of the attention function.
[0022] In a feasible embodiment of the solution of the present application, generating the first masked sequence after masking the first vector representation and generating the second masked sequence after masking the second vector representation includes:
[0023] Input the token vector representation and the token position into the masked language model, and predict the masked position through the masked language model; the token vector representation includes the first vector representation and the second vector representation;
[0024] Replace the tokens in the token vector representation according to the masked positions to obtain the first masked sequence and the second masked sequence.
[0025] In a feasible embodiment of the solution of the present application, the step of replacing the tokens in the token vector representation according to the masked positions to obtain the first masked sequence and the second masked sequence includes:
[0026] Randomly arrange the masked positions of the tokens in the token vector representation, and replace the tokens in the vector representation after random arrangement to obtain an initial sequence;
[0027] Pad the sequence length of the initial sequence and pad the number of replaced tokens in the padded sequence to obtain the first masked sequence and the second masked sequence.
[0028] In a feasible embodiment of the solution of the present application, predicting the sequence adjacent to the first vector representation, combining the prediction result with the first masked sequence to obtain the first semantic vector, predicting the sequence adjacent to the second vector representation, and combining the prediction result with the second masked sequence to obtain the second semantic vector includes:
[0029] Input the embedding vector representation containing the first token into a multi-layer perceptron with a single hidden layer, and output a binary classification result through the multi-layer perceptron;
[0030] The binary classification result is used to characterize the probability that the embedding vector sequence is an adjacent sequence.
[0031] In a feasible embodiment of the solution of the present application, the obtaining the target service request, obtaining the target semantic vector by matching with the vectors in the second semantic library, and obtaining the target component scheduling sequence according to the sequence of the target semantic vector in the first semantic library includes:
[0032] Determine the first quantity and the centroid vector of the clustering clusters in the first corpus, where the first quantity is determined according to the number of dimensions of the first semantic vector;
[0033] Determine the Euclidean distance between the first semantic vector and the centers of several of the clustering clusters according to the attribute dimensions of the first semantic vector and the centroid vector;
[0034] Divide the first semantic vectors in the first corpus into several clustering clusters according to the minimum value among several of the Euclidean distances;
[0035] Determine the second distance between the target semantic vector and the clustering clusters, and determine the target component scheduling sequence according to the first semantic vectors in the clustering cluster corresponding to the minimum value of the second distance.
[0036] In a feasible embodiment of the solution of the present application, after the step of dividing the first semantic vectors in the first corpus into several clustering clusters according to the minimum value among several of the Euclidean distances, it further includes:
[0037] Determine the intra-cluster similarity according to the average distance between the first semantic vector and other semantic vectors in the current clustering cluster;
[0038] Determine the inter-cluster dissimilarity according to the average distance between the first semantic vector and semantic vectors in other clustering clusters except the current clustering cluster;
[0039] Determine the silhouette coefficient according to the intra-cluster similarity and the inter-cluster dissimilarity, and adjust the first quantity of the clustering clusters according to the silhouette coefficient.
[0040] In a feasible embodiment of the solution of the present application, the industrial service component recommendation method further includes:
[0041] Obtain a linear combination of the masked language model loss function and the adjacent sequence prediction loss function, and adjust the parameters of the masked language model and the multi-layer perceptron with a single hidden layer according to the result of the linear combination.
[0042] The advantages and beneficial effects of the present invention will be partially given in the following description, and the other parts can be understood through the specific implementation manners of the present invention:
[0043] The technical solution of this application provides a method for recommending industrial service components based on deep learning. On the one hand, it solves the problem of recommendation accuracy. That is, traditional recommendation methods mostly train and obtain semantic information based on the attributes within a single component, etc. In this solution, historical component request sequences and single service component description statements are used as training data for two semantic trainings respectively. And for the semantic vectors of a single service, the semantics generated at the sequence level and the semantics generated from the single service description are combined. Such a consideration is closer to the actual application scenario and the accuracy is improved. On the other hand, since service sequences and single service component description statements are used as training data, the real-time generated historical service sequence data and new service description statements are periodically summarized to the cloud center for training, and the results are transmitted to the edge nodes, so as to achieve the global dynamic update recommendation effect, and solve the problem that the service component recommendation technology in a centralized environment cannot be applied to the computing resource limitations of edge nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 It is a flowchart of the steps of a method for recommending industrial service components based on deep learning provided in the technical solution of this application;
[0046] Figure 2 It is a model training architecture diagram when the training data in the technical solution of this application is a service scheduling sequence;
[0047] Figure 3 It is a schematic diagram of the implementation environment of the technical solution of this application;
[0048] Figure 4 It is a flowchart of the steps of another method for recommending industrial service components based on deep learning provided in the technical solution of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0050] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0052] As pointed out in the content of the background art, on the one hand, in related technical solutions, clustering related service components according to attributes such as function, theme, category, field, and quality of service is an effective service component recommendation method, and a QoS prediction method based on clustering is proposed for service component recommendation based on the above attributes of service components, or recommendation is performed based on component function clustering; however, most of these recommendation methods focus on the attributes of service components themselves for service component recommendation. From the perspective of service component combination, the call association between components can be regarded as the actual call situation when service components are executed, and service component recommendation based on call association is more important in the actual application of component combination. On the other hand, most service component recommendation methods are solved by combining deep learning techniques, such as service recommendation QoS prediction based on deep feature learning proposed in the application environment, or service component recommendation based on deep feature collaborative filtering; however, such solutions often do not consider the dynamic changes of service components. In a distributed environment, the component calls of edge nodes are constantly changing. On the other hand, traditional recommendation methods may not be applicable to edge nodes due to the lack of computing resources at the edge.
[0053] Regarding the deficiencies of the prior art pointed out in the above content, on the first hand, as Figure 1 shown, the technical solution of this application provides a deep learning-based industrial service component recommendation method, and the method includes steps S01 - S06:
[0054] S01. Obtain the historical service component sequence and the description information of a single service component from the edge nodes;
[0055] Specifically, in the embodiment, the edge nodes collect the historical service component sequence and the description of a single service component from the edge devices, and then perform necessary preprocessing operations.
[0056] S02. Encode the preprocessed historical service component sequence to obtain a first vector representation, and encode the preprocessed description information to obtain a second vector representation;
[0057] Specifically, in the embodiment, the edge nodes transmit the preprocessed component scheduling sequence and the description of a single service component to the cloud center, and the cloud center transmits the obtained embedding vectors into the Encoder for two rounds of training respectively to obtain two vector representations, namely the vector representation at the service component scheduling sequence level, i.e., the first vector representation, and the vector representation of the description of a single service component, i.e., the second vector representation.
[0058] S03. Generate a first masked sequence after masking the first vector representation, and generate a second masked sequence after masking the second vector representation;
[0059] S04. Predict the sequence adjacent to the first vector representation, combine the prediction result with the first masked sequence to obtain a first semantic vector, predict the sequence adjacent to the second vector representation, and combine the prediction result with the second masked sequence to obtain a second semantic vector;
[0060] Specifically, in the embodiment, steps S03 and S04 transmit the output vectors of the two Encoders into the pre-trained model respectively, mainly performing two task operations. Taking the incoming data as the service scheduling sequence as an example, the first task is the masked language model, and the second is to predict whether the incoming component sequence pair is an adjacent sequence pair. The embedding vectors after passing through the pre-trained model are more accurate and generate a semantic library according to subsequent operations. The semantic library consists of two major parts, namely the semantic vector of the service scheduling sequence and the semantic vector of a single service component. Among them, the first masked sequence is the masked sequence at the service component scheduling sequence level output by the masked language model, and the second masked sequence is the masked sequence of the description of a single service component output by the masked language model; the first semantic vector and the second semantic vector also satisfy the aforementioned corresponding relationship, which will not be elaborated here.
[0061] S05. Construct a first semantic library for the service scheduling sequence according to the first semantic vector, and construct a second semantic library for a single service component according to the second semantic vector;
[0062] S06. Obtain a target service request, obtain a target semantic vector by matching with vectors in the second semantic library, and obtain a target component scheduling sequence by matching according to the sequence of the target semantic vector in the first semantic library;
[0063] Specifically in the embodiment, steps S05 and S06 are that in the embodiment, the cloud center clusters the semantic vectors of the service sequences in the generated semantic library through the K-means clustering algorithm, and sends the clustering result and the semantic library to the edge node. The edge node receives a single service component requested by the edge device currently, obtains the semantic vector of the single service of the request in the semantic library, calculates the distance from the center vector of each cluster in the clustering result, and selects a suitable category to recommend the service component scheduling sequence therein to the edge device.
[0064] Specifically in the embodiment, the preprocessing process in the embodiment, that is, step S02 in the method, may include steps S021 - S023:
[0065] S021. Obtain an input sequence after preprocessing and removing punctuation marks and space characters, where the input sequence includes a preprocessed historical service component sequence and preprocessed description information;
[0066] S022. Add a first token at the beginning of the sequence of the input sequence to obtain a token embedding sequence, and split the token embedding sequence through a delimiter to obtain a number of embedding vector sequences;
[0067] S023. Encode according to the embedding vector sequences to obtain the first vector representation and the second vector representation.
[0068] Specifically in the embodiment, when the edge node receives the historical service component scheduling sequence of the edge device, first add the [cls] token at the beginning of the sequence to represent the information at the sequence level, and the subsequent output result is the semantic vector information of the sequence. Then, insert the delimiter [sep] in the sequence to split the service component sequence into a sequence pair for training and predicting the next sentence task in the pre-trained model, and each component name in the service component sequence is a token.
[0069] Particularly, in the process of model training in the embodiment, the mathematical description of the historical service component scheduling sequence in the training data is as follows:
[0070] C seq ={C1 C2 … C n-1 C n}(n≥2)
[0071] Where, Q seq represents the historical service component sequence; C nRepresents the name of a single component that makes up a component sequence; n represents the number of service components called in the component sequence.
[0072] In addition, the description statements for a single service component in the embodiments also need to be preprocessed by removing punctuation marks and whitespace characters. Similar to the foregoing implementation process, the [cls] token is added, and a separator [sep] is inserted to divide the description statements into statement pairs. During the process of model training in the embodiments, the mathematical description of the description statements of a single service component in the training data is as follows:
[0073] L sen ={C ∪ M rest ∪ E ∪ D}
[0074] Where L sen represents the description statement of a single service component; C represents the name of the service component; M rest represents the Restful style type of the service component (GET, POST, DELETE, PUT); E represents the category to which the service component belongs; D represents the description of the implementation function for the component.
[0075] More specifically, the data modeling of the description statements of a single service component in the embodiments is designed into four parts. Three parts consider the characteristics of the key attributes in a single component, and the last part is the description of the implementation function of the component. Together, they form the description statement of a single service component, and the component semantic vectors generated in subsequent training are more accurate than traditional methods.
[0076] In subsequent embodiments, the operation of performing data embedding initialization can be mainly divided into three parts. Taking the incoming data as a service scheduling sequence as an example, first, the embedding of tokens is generated according to the processed service sequence. Second, since the sequence is segmented by [sep] and the model needs to distinguish it, fragment embeddings need to be generated, that is, the first subsequence is represented by the same vector, and the second subsequence is represented by another same vector. Finally, position embeddings are generated. The input is the position information of each token in the sequence starting from 0, so as to obtain the vectors corresponding to each position. Generally, random initialization is used to let the model learn it by itself. The three parts of the preprocessed embeddings are linearly added to obtain the value of the preprocessed service sequence embedding.
[0077] Furthermore, the embodiment normalizes the output results, that is, normalizes the elements in each sample, and then puts the results into a position-based feed-forward network (equivalent to a fully connected layer) to change the input dimension and thus act on two fully connected layers, which is equivalent to two one-dimensional convolutional layers with a kernel window of 1. Then, the results are normalized to obtain the final embedded vector representation. It should be noted that the number of Encoder blocks can be specified according to the situation. The more the number, the more resources are required for training. Determine the number according to the actual application scenario. After the data is passed in twice, the Encoder outputs the semantic vector of the service component scheduling sequence and the vector representation of a single service component description.
[0078] In some feasible implementation manners, the encoding process in step S023 in the embodiment can be further divided into steps S0231 - S0232:
[0079] S0231. Input the embedded vector sequence into the encoding model, and calculate the similarity between the query vector and the keyword vector in the embedded vector sequence through the attention function in the encoding model;
[0080] S0232. Normalize the vector elements in the embedded vector sequence according to the similarity, and perform a fully connected process on the normalization result through a position-based feed-forward network to obtain an embedded vector representation; wherein, the embedded vector representation includes the first vector representation and the second vector representation.
[0081] Specifically in the embodiment, when the cloud center receives the embedded vector generated by the edge node, it inputs it into the Encoder model, that is, the encoding model. First, the input embedded information is input into the multi-head attention mechanism; in this mechanism, the query and key in the attention function used in the embodiment are of equal length and both equal to dk, the length of value is dv, and the output is also dv (consistent with the length of value). The calculation process of the similarity is to take the inner product of the query and the key as the similarity. The larger the inner product value of two vectors, the higher the similarity between the two vectors. Then, divide it by the square root of dk and put it into the softmax to obtain the output result. For example, in the embodiment, there are n k-v pairs, so there will be n non-negative weights that add up to 0. The formula is as follows:
[0082]
[0083] In some feasible implementation manners, the parameter weight W in the embodiment has no learning process, so Q, K, and V are projected into a low dimension, projected h times, and the attention function is performed h times. The output vectors of the h times are combined together and linearly projected once to obtain the final output, so that the parameter weight of the projection can be learned. The formula description is as follows:
[0084] MultHead(Q, K, V) = Concat(head1, …, headh)W o
[0085]
[0086] In some possible embodiments, step S03 in the method may further include steps S031 - S032:
[0087] S031. Input the token vector representation and the token position into a masked language model, and predict the masked positions through the masked language model; wherein, the token vector representation includes the first vector representation and the second vector representation;
[0088] S032. Replace the tokens in the token vector representation according to the masked positions to obtain a first masked sequence and a second masked sequence.
[0089] Specifically, in the embodiment, as Figure 2 shown, the vector representation output by the Encoder through step S02 is applied to two tasks in the pre - trained model. One of them is the training of the Mask LM language model, that is, the masked language model. Each time, with a probability of 15%, some components in the service sequence are randomly replaced with artificial special tokens <mask>, since the subsequent fine-tuning tasks will not occur <mask>, so in the selected subsequence, with an 80% probability, the selected token is changed to <mask>, replace with a random token with a 10% probability and keep the original token with a 10% probability. The example is as follows (taking the input data as the service scheduling sequence):
[0090] Input: getTemperature robotGetpose setPosition[MASK]
[0091] Label: [MASK]=robotJogging
[0092] While performing the masking operation on the input service sequence, obtain the label of the masked position, and then use the sample to train the model. By predicting the masked position, learn these service sequences. The purpose of training this task is to predict the probability of the correct word masked by mask, which is described as follows:
[0093] P(getTemperature robotGetpose setPosition robotJogging|getTemperaturerobotGetpose setPosition[MASK]) = P{mask = robotJogging|getTemperaturerobotGetpose setPosition}
[0094] In some feasible ways, step S032 in the method may further include steps S0321 - S0322:
[0095] S0321. Randomly shuffle the masked positions of the tokens in the token vector representation, and replace the tokens in the shuffled vector representation to obtain the initial sequence;
[0096] S0322. Pad the sequence length of the initial sequence, and pad the number of replaced tokens in the padded sequence to obtain the first masked sequence and the second masked sequence.
[0097] Specifically in the embodiment, in the code for randomly masking some tokens, the n - pred variable represents the number of tokens to be masked, and cand_maked_pos represents which positions are candidates and can be masked (for example, <sep>, [CLS], and other tokens are not masked and have no meaning. The sequence is then shuffled and replaced with [MASK] based on the value of random(). Two zero padding operations are performed: the first to pad the sequence length so that all sentences in a batch are the same length, and the second to pad the number of masks. Different sentence lengths result in different numbers of words being masked, and we need to ensure that the number of masks in the same batch is the same. Therefore, a meaningless [0] is added at the end.
[0098] In this embodiment, the vector representation output by the encoder in step S02 is used to perform two tasks in the pre-trained model: the other is to predict whether two sequences in a sequence pair generated by the previously input pre-processed sequence are adjacent. Therefore, in some feasible implementations, step S04 in this embodiment may include step S041:
[0099] S041. Input the embedding vector representation containing the first word into a multilayer perceptron with a single hidden layer, and output a binary classification result through the multilayer perceptron; wherein the binary classification result is used to represent the probability that the embedding vector sequence is an adjacent sequence.
[0100] Specifically in the embodiment, in the training sample (taking the input data as the service scheduling sequence getTemperature robotGetpose setPosition robotJogging as an example):
[0101] 50% probability of selecting adjacent sequence pairs (positive class):
[0102] <cls>getTemperature robotGetpose <sep>setPosition robotJogging <sep>
[0103] Select a random sequence pair (negative class) with a 50% probability:
[0104] <cls>getTemperature robotGetpose <sep>setPosition robotClose <sep>
[0105] will <cls>The corresponding output is placed in the fully connected layer to predict whether the sequence pair is adjacent.
[0106] More specifically, in the embodiment, a multi-layer perceptron with a single hidden layer can be used to predict whether the second sequence is the next sequence of the first sequence in the model input sequence pair. Due to the self-attention in the encoder, the special token " <cls>"indicates that the two input sequences have been encoded. Therefore, the output layer of the multi-layer perceptron classifier takes X as input, where X is the output of the hidden layer of the multi-layer perceptron, and the input of the MLP hidden layer is the encoded " <cls>”For the token, the loss function is the cross - entropy loss formula for binary classification as follows:
[0107]
[0108] Among them, p i represents the label of sample i, where the positive class is 1 and the negative class is 0; y i represents the probability that sample i is predicted as the positive class. For the positive class, samples of adjacent sequence pairs are selected with a 50% probability, and for the negative class, samples of random sentence pairs are selected with a 50% probability.
[0109] Furthermore, in the embodiment, there are two parameters positive and negative in the code to record the number of positive and negative samples in the NLP task. The ratio is preferably close to 1:1 in a batch because the generation probability of the positive class and the parent class is specified as 50%. The output <cls>The token semantic vectors need to go through a classification layer, and the vector dimension changes from the original output dimension d_model to 2 for a binary classification task to predict whether the two sequences in the sequence pair generated from the previously preprocessed input are adjacent.
[0110] In the embodiment, step S06 is to perform k-means clustering on the embedding vectors of each service component sequence in the obtained semantic library. The value of K, that is, the number of clusters, is specified during the clustering process. The specification of the value of K needs to be determined according to the data size of the service component sequences in the actual application scenario. The cloud center transmits the clustering result and the semantic library to the edge node. The edge node finds the corresponding service component semantic vector in the semantic library according to the requested components of the current edge device, compares the distance with the center point vectors of each cluster in the cluster result, selects the cluster with the closest distance, and recommends the component sequence it contains to the edge device.
[0111] Furthermore, in some feasible embodiments, step S06 in the method may include steps S061 - S064:
[0112] S061. Determine the first number of clustering clusters and the centroid vectors in the first corpus, where the first number is determined according to the dimension number of the first semantic vector;
[0113] S062. Determine the Euclidean distance between the first semantic vector and the center points of several clustering clusters according to the attribute dimension of the first semantic vector and the centroid vectors;
[0114] S063. Divide the first semantic vectors in the first corpus into several clustering clusters according to the minimum value among several Euclidean distances;
[0115] S064. Determine the second distance between the target semantic vector and the clustering clusters, and determine the target component scheduling sequence according to the first semantic vector in the clustering cluster corresponding to the minimum value of the second distance.
[0116] Specifically, in the embodiment, the data used in step S06 (in steps S061 - S064) is the service scheduling sequences in the semantic library generated by the pre-trained model. <cls>The semantic vector of the label (this vector represents the information features at the level of the service component scheduling sequence). First, in the embodiment, the number of clusters K needs to be specified, and the centroid vectors are randomly selected. In terms of the selection of K, it will affect the clustering effect, and it is determined according to the actual number of service component sequence semantics generated. Here, we choose the number of dimensions of the semantic vector as the number of K. In the specific implementation process, it can be determined according to the actual application scenario. X is the input semantic vector of the service component sequence (that is, obtained through the pre-trained model) <cls>(semantic vector representation), each vector has attributes of m dimensions, C is the initialized cluster center, calculate the Euclidean distance from each service semantic vector to the cluster center point, and the formula is as follows:
[0117]
[0118] The service semantic vector with the smallest distance forms a cluster. Then, update the cluster center point within the cluster, that is, calculate the mean of each dimension, and then divide the cluster classes until the mean vector is not updated.
[0119] In some feasible embodiments, the clustering effect can be verified by the silhouette coefficient. That is, step S06 in the embodiment may further include steps S065 - S067:
[0120] S065. Determine the intra - cluster similarity according to the average distance between the first semantic vector and other semantic vectors in the current clustering cluster;
[0121] S066. Determine the inter - cluster dissimilarity according to the average distance between the first semantic vector and semantic vectors in other clustering clusters except the current clustering cluster;
[0122] S067. Determine the silhouette coefficient according to the intra - cluster similarity and the inter - cluster dissimilarity, and adjust the first number of the clustering clusters according to the silhouette coefficient.
[0123] Specifically in the embodiment, first calculate the average distance a(i) from sample i to other samples in the same cluster, which represents the intra - cluster similarity of sample i. Then, calculate the average distance b(i) from sample i to all samples in another cluster, which represents the inter - cluster dissimilarity of the sample. The coefficient formula is as follows:
[0124]
[0125] If s(i) is close to 1, it indicates that the sample clustering is reasonable. If s(i) is close to - 1, it indicates that sample i should be classified into another cluster. If s(i) is approximately 0, it indicates that the sample is on the boundary of two clusters. Verify the clustering effect through this parameter. When the effect is not obvious, the value of K can be appropriately increased.
[0126] In some feasible embodiments, during the training process of two tasks (Mask LM language model and adjacent sequence prediction model), the method of the embodiment may further include step S07: Obtain the linear combination of the masked language model loss function and the adjacent sequence prediction loss function, and adjust the parameters of the masked language model and the single - hidden - layer multi - layer perceptron according to the result of the linear combination.
[0127] Specifically, in the embodiment, the integrated loss of the final pre-trained model is a linear combination of performing the previous two tasks, namely, the masked language model loss function and the next sentence prediction loss function. Through the above pre-trained model, the service component sequence and each token of the service description statement after preprocessing can be finally obtained as semantic expression vectors with a length of 128 bits. The size of the semantic expression vector and the scale of the model can be adjusted by the number of Ecoder blocks, the size of the hidden layer, and the number of heads in the self-attention mechanism.
[0128] The following combines the accompanying drawings of the specification to describe the specific implementation process of the method in the technical solution of the present application more completely and in detail as follows:
[0129] The embodiment recommends service components for edge devices based on a distributed cloud-edge collaborative industrial manufacturing environment. The specific environment is a three-layer architecture, namely, the cloud center, edge nodes, and edge devices, as shown in the attachment Figure 4 As shown, the dataset for training is the service component sequence generated by the sensor component and the robotic arm component in a typical industrial production environment, as well as the semantic description information of a single service component. In the embodiment, the JSON information description of a single service component is as follows (taking the getTemperature service component as an example)
[0130]
[0131]
[0132] As Figure 3 shown, the first step is the preprocessing of the edge node service component sequence and a single service description statement:
[0133] Taking the service component scheduling sequence data as an example, when obtaining the service component sequence, some special tokens are inserted, such as
[0134] <cls>(Indicating the semantics at the sequence level), <sep>(For splitting the component sequence) The specific insertion rule is half (rounded up) of the number of service components in the service scheduling sequence. The example is as follows:
[0135] Original service component scheduling sequence:
[0136] getTemperature robotGetpose setPosition robotJogging
[0137] After preprocessing:
[0138] <cls>getTemperature robotGetpose <sep>setPosition robotJogging <sep>
[0139] Among them, the descriptions of the service components are as follows:
[0140] getTemperature: Obtain the sensor temperature;
[0141] robotGetpose: Obtain the current coordinates of the robotic arm;
[0142] setPosition: Set the target coordinates for the robotic arm to move to;
[0143] robotJogging: Move the robotic arm to the target position;
[0144] During the preprocessing process, the token embedding is the normal word vector, that is, nn.Embedding() in PyTorch. The role of the segment embedding is to use the information of the embedding to let the model distinguish between the upper and lower sequences. We set all the tokens of the upper sequence to 0 and all the tokens of the lower sequence to 1, so that the model can judge the start and end positions of the upper and lower sentences. For example:
[0145] <cls>getTemperature robotGetpose <sep>setPosition robotJogging <sep> 0 0 0 0 1 1 1
[0147] The position embedding values are trained by the model. Therefore, during initialization, the values are generally input sequentially starting from 0, and finally the three embedding values are linearly added together.
[0148] Punctuation marks and whitespace characters are removed from the description statement of a single service component. Subsequently, similar to the above preprocessing operations, the [cls] token is added, and the separator [sep] is inserted to divide the description statement into a statement pair and data initialization operations. Taking the description of the getTemperature service component as an example:
[0149] Mathematical modeling of the input data generates the description statement data of the service component (for the specific mathematical modeling, refer to steps S021 - S023 in the embodiments, which will not be elaborated here):
[0150] getTemperature GET TemperatureSensor Obtain industrial workshop indoor temperature data, outdoor temperature data
[0151] The result after preprocessing is:
[0152] <cls>getTemperature GET TemperatureSensor <sep>Obtain industrial workshop indoor temperature data, outdoor temperature data
[0153] In the above preprocessing <seq>The insertion rule is between data models E and D, that is, between the category to which the service component belongs and the description of the functions implemented by the component. <seq>, because the first three parts are individual attributes in the component while the fourth part is a description statement of the component function. Such a selection is conducive to the execution of the subsequent training task of predicting the next sentence.
[0154] Step 2, the preprocessed service component sequence and the single service component description data are input into the Encoder to generate component vector representations:
[0155] It should be noted that when the preprocessed data is transmitted from the edge node to the cloud center, it is not transmitted in real time, but historical data is transmitted periodically to avoid waste of resources. Since the global data is aggregated to the cloud center, the training semantics and subsequent clustering operations have been executed in the cloud server, avoiding the actual situation of lack of computing resources at the edge node. The cycle time is adjusted according to the actual application scenario.
[0156] The embodiment uses the Encoder in the Transformer model and traverses in sequence according to the number of Encoder blocks by inputting the preprocessed vector embeddings. The position embeddings need to be learned by the model itself, so the numerical input adopts the method of random initialization. Note that when using it, a sufficiently long random position embedding parameter is required for learning.
[0157] The specific data input steps are to add the three previously initialized parameters, namely the token embedding, the segment embedding, and the position embedding vectors, and input them into the previous Encoder block. Define tokens as two input sequences of length 8, where each token is the index of the vocabulary. The forward inference of the Encoder using the input tokens returns the encoded results, where each token is represented by a vector, and its length is defined by the hyperparameter num_hiddens. This hyperparameter is usually the hidden size (number of hidden units) of the encoder. Through this operation, we can obtain the vector representations of each token after passing through the Encoder.
[0158] Step 3, train the masked language model and the next sentence prediction model to generate the service component sequence semantic vector and the single service component description semantic vector:
[0159] 1. What needs to be noted in the task of the masked language model: With a 15% probability, there is a 10% probability of inserting a random component token. This accidental noise makes the model less biased towards the masked token in its bidirectional context encoding (especially when the label token remains unchanged), so the insertion probability of the mask cannot be changed.
[0160] The masked token prediction uses a multi-layer perceptron with a single hidden layer. In the forward inference, it requires two inputs: the encoded results of the previous Encoder and the token positions for prediction. The output is the prediction results for these positions.
[0161] First, initialize the masked language model (MLM). The operation sequence is the fully connected layer, ReLU, LayerNorm, and the fully connected layer. Note that the activation function in the word task uses the Gaussian error linear unit, and the formula is as follows:
[0162]
[0163] The error function is:
[0164]
[0165] where Φ(x) is the cumulative distribution function of the standard normal distribution. When the embodiment performs prediction, it is necessary to input the token vector result of the Encoder and the position of the predicted word, and then perform the prediction operation, which is a classification problem. The forward inference of the masked language model returns the prediction results at all masked positions. For each prediction, the size of the result is equal to the size of the vocabulary. By the true label of the predicted token under the mask, we can calculate the cross-entropy loss of the masked language model task in pre-training.
[0166] In addition, it should be noted that in the code where the embodiment randomly masks some tokens, the variable n-pred represents the number of tokens to be masked, and cand_maked_pos represents which positions are candidates and can be masked (such as <sep>, [CLS] These tokens cannot be masked as they are meaningless. Then, the sequence content is shuffled by the shuffle function, and then replaced with [MASK] according to the value of random(). Next, two Zero Padding operations are performed. The first is to pad the length of the sequence so that the sentences in a batch are of the same length. The second is to pad the number of masks because different sentence lengths will result in different numbers of words being masked. It is necessary to ensure that the number of masks is the same in the same batch, so meaningless [0] also needs to be added at the end.
[0167] 2. In the next sentence prediction model, this task uses a multi-layer perceptron with a single hidden layer to predict whether the second sequence is the next sequence of the first sequence in the model input sequence pair. Due to the self-attention in the encoder, the special token " <cls>"indicates that the two input sequences have been encoded. Therefore, the output layer of the multi-layer perceptron classifier takes X as input, where X is the output of the hidden layer of the multi-layer perceptron, and the input of the MLP hidden layer is the encoded " <cls>”For the token, the loss function is the cross-entropy loss formula for binary classification as follows:
[0168]
[0169] Among them, p i represents the label of sample i, where the positive class is 1 and the negative class is 0; y i represents the probability that sample i is predicted as the positive class. For the positive class, samples of adjacent sequence pairs are selected with a 50% probability, and for the negative class, samples of random sentence pairs are selected with a 50% probability.
[0170] In the specific code of the embodiment, there are two parameters, positive and negative, which are used to record the number of positive and negative samples in the NLP task. The ratio is preferably close to 1:1 in one batch because the generation probability of the positive class and the parent class is specified as 50%. The output <cls>The token semantic vectors need to go through a classification layer, with the vector dimension reduced from the original output dimension d_model to 2, to perform a binary classification task for predicting whether the two sequences in the sequence pair generated from the previously preprocessed input are adjacent.
[0171] The integrated loss of the final pre-trained model is a linear combination of the losses of the previous two tasks, namely the masked language model loss function and the next sentence prediction loss function. Through the above pre-trained model, the semantic expression vectors of each token of the preprocessed service component sequence and service description statement can be finally obtained, with a length of 128 bits. The size of the semantic expression vector and the scale of the model can be adjusted by the number of Ecoder blocks, the size of the hidden layer, and the number of heads in the self-attention mechanism.
[0172] The output tokens after passing through the pre-trained model twice represent different semantic vectors. The first time is for training based on the service component scheduling sequence data. <cls>The vector in it is the semantic vector of the entire service scheduling sequence, generating the semantic library of the service scheduling sequence. Each of the other tokens is the semantic vector of each API at the dynamic level in this scheduling. To better represent the semantic vector of each service at the scheduling level, the quantity is counted, and the operation of calculating the average value of the semantic vectors with a length of 128 for each is as follows:
[0173]
[0174] Among them, E i represents the semantic vector of the i-th service at the global dynamic scheduling sequence level; N is the number of sequences in the historical service scheduling sequence that meet the condition of the existence of this service; L i represents the semantic vector of the i-th service in the current scheduling sequence. Obtained by pre-training for the second time according to the single service description statement <cls>It is the semantic vector of a 128-bit single service at the service description level. Finally, a linear addition operation is performed on the semantic vectors of each service obtained above at the scheduling level, and finally a semantic library represented by a single service vector is obtained.
[0175] Step 4: Cluster the generated semantic vectors of the service sequence through the K-means algorithm:
[0176] The data used in this step is the service scheduling sequence in the semantic library generated by the pre-trained model <cls>The semantic vector of the label (this vector represents the information features at the level of the service component scheduling sequence).
[0177] First, specify the number of clusters K and randomly select the centroid vectors. In terms of the selection of K, it will affect the clustering effect and is determined according to the actual number of service component sequence semantics generated. Here, we choose the dimensionality of the semantic vector as the number of K. The scheme user can decide according to the actual application scenario. X is the input service component sequence semantic vector (i.e., obtained through the pre-trained model) <cls>(semantic vector representation), each vector has attributes of m dimensions, C is the initialized cluster center, calculate the Euclidean distance from each service semantic vector to the cluster center point, the formula is as follows:
[0178]
[0179] The service semantic vector with the smallest distance forms a cluster. Then, update the cluster center point within the cluster, that is, calculate the mean value of each dimension, and then perform cluster classification until the mean vector is not updated.
[0180] To verify the clustering effect, the silhouette coefficient is used in the embodiment:
[0181] First, calculate the average distance a(i) from sample i to other samples in the same cluster, which represents the similarity within the cluster of sample i. Then, calculate the average distance b(i) from sample i to all samples in another cluster, which represents the dissimilarity between clusters of the sample. The coefficient formula is as follows:
[0182]
[0183] If s(i) is close to 1, it indicates that the sample clustering is reasonable. If s(i) is close to -1, it indicates that sample i should be classified into another cluster. If s(i) is approximately 0, it indicates that the sample is on the boundary of two clusters. The clustering effect is verified through this parameter. When the effect is not obvious, the value of K can be appropriately increased.
[0184] Step 5: The cloud center transmits the clustering result and the semantic library to the edge node for service recommendation for edge devices:
[0185] The edge node receives a single service component requested by the edge device currently, matches the service component information with the semantic library to obtain the semantic vector of the single service of the current request, calculates the distance between the current semantic vector and the center vector of each cluster in the clustering result, selects the cluster with the closest distance, and recommends the scheduling sequence in this cluster. Since in the clustering process, the clustering result of the sequence is <cls>(Component sequence semantic vector) for clustering, so the component sequence cluster with the closest distance is selected for recommendation, indicating that using this single service component is most likely to perform the tasks of the component sequences in this class, thereby making dynamic recommendations for edge devices.
[0186] It should be added that when the cloud center transmits the clustering results and the semantic library to the edge nodes in the embodiment, it is transmitted periodically to avoid waste of communication resources and achieve global dynamic updates. The specific transmission period is determined by the actual user of this solution according to the data scale of the application scenario.
[0187] From the above specific implementation process, it can be summarized that the technical solution provided by the present invention has the following advantages or advantages compared with the prior art:
[0188] The technical solution of this application is different from traditional recommendations that generate service semantics based on the attributes of service components themselves. Considering the perspective of service composition, using service component sequences as semantic extraction makes the generation of semantics closer to the actual industrial production environment. Therefore, a combination of the two is adopted to generate the semantic library of service components. In addition, a mathematical expression of the description statement of a single service component and the service component scheduling sequence is proposed. On the other hand, in the recommendation process, it is not a recommendation of a single service component, but a recommendation of a service component scheduling sequence guided by the task objective based on the single service requested by the user, and the recommendation benefit is greater.
[0189] In addition, for the industrial production distributed cloud-edge collaborative environment where the computing resources of edge nodes are scarce, an architecture mode is proposed in which pre-training and clustering operations are performed in the cloud center, and then the results are transmitted to the edge nodes for service component recommendation. This architecture also takes into account the dynamic recommendation of service components. In the solution, the edge nodes transmit service data periodically, and the cloud center pushes the clustering results and the semantic library periodically, solving the problems of service dynamic recommendation and lack of edge node resources in the distributed environment.
[0190] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0191] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device.
[0192] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0193] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0194] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.< / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / sep> < / seq> < / seq> < / sep> < / cls> < / sep> < / sep> < / cls> < / sep> < / sep> < / cls> < / sep> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / cls> < / sep> < / sep> < / cls> < / sep> < / sep> < / cls> < / sep> < / mask> < / mask> < / mask>
Claims
1. An industrial service component recommendation method based on deep learning, characterized in that, Including the following steps: Obtain the historical service component sequence and the description information of a single service component from edge nodes; Encode the preprocessed historical service component sequence to obtain a first vector representation, and encode the preprocessed description information to obtain a second vector representation; Generate a first masked sequence after masking the first vector representation, and generate a second masked sequence after masking the second vector representation; Predict the sequence adjacent to the first vector representation, combine the prediction result with the first masked sequence to obtain a first semantic vector, predict the sequence adjacent to the second vector representation, and combine the prediction result with the second masked sequence to obtain a second semantic vector; Construct a first semantic library of the service scheduling sequence according to the first semantic vector, and construct a second semantic library of a single service component according to the second semantic vector; Obtain a target service request, obtain a target semantic vector by matching with vectors in the second semantic library, and obtain a target component scheduling sequence by matching the sequence in the first semantic library according to the target semantic vector; The encoding the preprocessed historical service component sequence to obtain a first vector representation and encoding the preprocessed description information to obtain a second vector representation includes: Obtain an input sequence after preprocessing to remove punctuation marks and whitespace characters, where the input sequence includes the preprocessed historical service component sequence and the preprocessed description information; Add a first token to the beginning of the sequence of the input sequence to obtain a token embedding sequence, and split the token embedding sequence through a delimiter to obtain a number of embedding vector sequences; Encode according to the embedding vector sequences to obtain the first vector representation and the second vector representation; The encoding according to the embedding vector sequences to obtain the first vector representation and the second vector representation includes: Input the embedding vector sequences into an encoding model, and calculate the similarity between the query vector and the keyword vector in the embedding vector sequences through an attention function in the encoding model; Normalize the vector elements in the embedding vector sequences according to the similarity, and perform a fully connected process on the normalization result through a feed-forward network of positions to obtain an embedding vector representation; the embedding vector representation includes the first vector representation and the second vector representation.
2. The industrial service component recommendation method based on deep learning according to claim 1, wherein After the step of inputting the embedding vector sequences into an encoding model and calculating the similarity between the query vector and the keyword vector in the embedding vector sequences through an attention function in the encoding model, it further includes: Perform a low-dimensional projection on the query vector, keyword vector, and data item calculated by the attention function; Perform a linear projection on the output vector obtained after the low-dimensional projection to determine the parameter weights of the attention function.
3. The industrial service component recommendation method based on deep learning according to claim 1, wherein The generating a first masked sequence after masking the first vector representation and generating a second masked sequence after masking the second vector representation includes: Input the token vector representation and the token position into the masked language model, and predict the masked position through the masked language model; the token vector representation includes the first vector representation and the second vector representation; Replace the tokens in the token vector representation according to the masked position to obtain a first masked sequence and a second masked sequence.
4. The method for recommending industrial service components based on deep learning according to claim 3, wherein, The step of replacing the tokens in the token vector representation according to the masked position to obtain a first masked sequence and a second masked sequence includes: Randomly arrange the masked positions of the tokens in the token vector representation, and replace the tokens in the vector representation after the random arrangement to obtain an initial sequence; Pad the sequence length of the initial sequence, and pad the number of replaced tokens in the padded sequence to obtain the first masked sequence and the second masked sequence.
5. A method for recommending industrial service components based on deep learning according to claim 1, characterized in that, Predict the sequence adjacent to the first vector representation, combine the prediction result with the first masked sequence to obtain a first semantic vector, and predict the sequence adjacent to the second vector representation, combine the prediction result with the second masked sequence to obtain a second semantic vector, including: Input the embedding vector representation containing the first token into a multi-layer perceptron with a single hidden layer, and output a binary classification result through the multi-layer perceptron; The binary classification result is used to characterize the probability that the embedding vector sequence is an adjacent sequence.
6. The industrial service component recommendation method based on deep learning according to claim 1, characterized in that The step of obtaining the target service request, obtaining the target semantic vector by matching with the vectors in the second semantic library, and obtaining the target component scheduling sequence by matching according to the sequence of the target semantic vector in the first semantic library includes: Determine the first number of clustering clusters in the first corpus and the centroid vector, where the first number is determined according to the dimensionality of the first semantic vector; Determine the Euclidean distance between the first semantic vector and the centers of several clustering clusters according to the attribute dimension of the first semantic vector and the centroid vector; Divide the first semantic vectors in the first corpus into several clustering clusters according to the minimum value among several Euclidean distances; Determine the second distance between the target semantic vector and the clustering cluster, and determine the target component scheduling sequence according to the first semantic vector in the clustering cluster corresponding to the minimum value of the second distance.
7. The method for recommending industrial service components based on deep learning according to claim 6, wherein, After the step of dividing the first semantic vectors in the first corpus into several clustering clusters according to the minimum value among several Euclidean distances, it further includes: Determine the intra-cluster similarity according to the average distance between the first semantic vector and other semantic vectors in the current clustering cluster; Determine the inter-cluster dissimilarity according to the average distance between the first semantic vector and semantic vectors in other clustering clusters except the current clustering cluster; Determine the silhouette coefficient according to the intra-cluster similarity and the inter-cluster dissimilarity, and adjust the first number of the clustering clusters according to the silhouette coefficient.
8. A method for recommending industrial service components based on deep learning according to any one of claims 1-7, characterized in that, The industrial service component recommendation method further includes: Obtain a linear combination of the masked language model loss function and the adjacent sequence prediction loss function, and adjust the parameters of the masked language model and the multi-layer perceptron with a single hidden layer according to the result of the linear combination.
Citation Information
Patent Citations
Method and device for reporting scheduling request in narrowband Internet of Things system
CN108633096A
Sequence recommendation method, system and device and medium
CN114897145A