Method and system for evaluating semantic similarity between software code identifier and related natural language description by using bidirectional recurrent neural network

The BiLSTM neural network architecture solves the accuracy of the semantic similarity evaluation of code identifiers and annotations, and realizes efficient and robust semantic similarity evaluation, which improves the automation level of code understanding and maintenance.

CN120578413APending Publication Date: 2025-09-02SICHUAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510676668.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and efficiently evaluate the deep semantic similarity between software code identifiers and natural language descriptions, resulting in difficulty in understanding and maintaining codes.

Method used

Using a bidirectional long and short-term memory (BiLSTM) neural network architecture, the semantic representation of code identifiers and annotations is learned and compared through the twin network structure, and combined with specific preprocessing and output combination strategies, the Logits binary cross entropy loss and Adam optimization algorithm are trained.

Benefits of technology

Improves the accuracy of semantic similarity evaluation between code identifiers and comments, reduces manual review workload, improves the quality and maintainability of code documents, and is robust to wording and abbreviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a method and a system for evaluating semantic similarity between software code identifiers and related natural language descriptions by utilizing a BiLSTM (Bidirectional Recurrent Neural Network). According to the method, a twin BiLSTM network model sharing weight is constructed, code identifiers and descriptions are respectively input and subjected to feature extraction, then extracted feature vectors are combined, and a semantic similarity score is predicted through a full connection layer. The system comprises a data processing unit, a model training unit and a similarity prediction unit, and realizes automatic evaluation of consistency between code identifiers and annotations. According to the method, the limitation of a traditional text similarity evaluation method in processing professional texts in the code field is solved, the quality of code documents is improved, and the maintainability and the understandability of software are enhanced. The effectiveness of the method is verified through experiments on a text similarity data set, and it is expected that targeted high-performance expression can be obtained after training and fine tuning are conducted on specific field data of code identifiers and annotations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the fields of computer-implemented automated text analysis and natural language processing (NLP). More specifically, the present invention relates to methods and systems for evaluating semantic similarity or relatedness between different textual inputs, particularly between structured code elements (e.g., function or variable names, i.e., identifiers) and their corresponding unstructured or semi-structured natural language descriptions (e.g., code comments or documentation). The present invention further relates to the application of deep learning models, particularly recurrent neural networks (RNNs) and long short-term memory (LSTM) networks, to efficiently understand and compare the semantic content of these text pairs.

[0002] Applying natural language processing and neural networks to text content analysis, including similarity analysis, is an active area of ​​research and development. For example, neural networks are used as machine learning models to predict outputs, while long short-term memory (LSTM) neural network layers are used for tasks such as document classification. Patent analysis has also been combined with natural language processing, using embedding techniques to calculate patent similarity. Furthermore, research has employed ensemble approaches combining multiple BERT-based models to enhance the accuracy of semantic similarity in patent documents. This invention addresses this technical area by applying these concepts to a specific type of text pair: code identifiers and comments. This application not only lies at the intersection of software engineering (code analysis) and artificial intelligence (natural language processing, deep learning), but advances in this specific area are expected to drive the development of automated software engineering tools that improve code quality, maintainability, and developer productivity by ensuring semantic consistency between documentation and code. Background Art

[0003] During software development and maintenance, semantic consistency between code identifiers (such as function and variable names) and their associated natural language descriptions (such as code comments and documentation strings) is crucial for code understandability, maintainability, and team collaboration. However, manually checking this consistency is time-consuming and error-prone. Therefore, developing techniques to automatically assess this semantic similarity is of great value.

[0004] Traditional methods for assessing text similarity primarily include string-based metrics, corpus-based methods, and knowledge-based methods. String-based metrics, such as Levenshtein distance or Jaccard similarity based on N-grams, primarily focus on surface-level lexical matching and struggle to capture deeper semantic meaning. They perform poorly for synonyms, paraphrases, or different expressions describing the same concept.

[0005] Corpus-based methods, such as latent semantic analysis (LSA) or vector space models using term frequency-inverse document frequency (TF-IDF) weighting combined with cosine similarity, rely on the statistical co-occurrence of words in large amounts of text. While these methods can capture some semantic connections in some cases, they often struggle with polysemy and synonymy and may not perform well on high-dimensional, sparse data, such as short text or specialized domain vocabulary. Cosine similarity also has limitations, such as ignoring the magnitude of vectors (i.e., the influence of text length) and being constrained by data normalization assumptions.

[0006] Knowledge-based methods, such as those using lexical ontologies like WordNet, calculate similarity by leveraging predefined semantic relationships (such as hyponymy and synonymy). However, the performance of such methods is highly dependent on the coverage and quality of the knowledge base. Existing knowledge bases may not be comprehensive enough for terminology in specific programming languages, project-specific abbreviations, or vocabulary in emerging technology fields.

[0007] The limitations of these traditional methods are particularly evident when applied to assessing the similarity between code identifiers and comments. Code identifiers are often concise and may contain abbreviations or domain-specific terms, while code comments vary in style, length, and level of detail. This leads to a significant "semantic gap" between identifiers and comments, where they may use different vocabularies or levels of abstraction to describe the same concepts. The conciseness of identifiers also provides traditional methods with limited lexical features. Furthermore, the meaning of terms in code is often highly dependent on context.

[0008] Assessing semantic similarity in short texts presents numerous challenges, including ambiguity, polysemy, data sparsity, and the presence of anomalies such as misspellings and non-standard terminology. These challenges can degrade the performance of traditional machine learning algorithms. Even when defining the similarity task itself, unclear conditions or uncommensurable mappings can arise, further increasing the complexity of the task.

[0009] With the development of machine learning technology, some early neural network models (such as feedforward networks with word embeddings) have been applied to text similarity tasks. These models have improved over traditional methods, but may still not fully capture the sequential dependencies and long-range contextual information that are crucial for understanding programming structures and detailed comments.

[0010] In recent years, more advanced deep learning models, such as the Transformer-based BERT model, have achieved remarkable success in many natural language processing tasks. However, these models may also have limitations, such as length restrictions for very long text inputs, or their specific word segmentation strategies (such as WordPiece) may not be applicable to certain highly specialized domains without adjustment.

[0011] Existing technologies still fall short in accurately and robustly evaluating the deep semantic similarity between code identifiers and natural language descriptions. The limitations of traditional methods, particularly when processing specialized domain texts like code, which possess rich implicit semantic structures, have necessitated the need for advanced neural network models capable of learning deeper semantic representations, such as the bidirectional long short-term memory (BiLSTM) network employed in this invention. There has been a clear trend in the evolution of text similarity technology from simple string matching to statistical methods and now to deep learning models capable of learning semantic representations. This invention aligns with this technological evolution and applies it to a novel and practical problem in software engineering. Clearly articulating these limitations of the existing technology helps highlight the technical problems addressed by this invention and the beneficial effects it brings. Summary of the Invention

[0012] This invention aims to overcome the limitations of existing techniques for accurately and efficiently assessing the semantic similarity between software code identifiers (e.g., function names, class names, variable names) and their associated natural language descriptions (e.g., code comments, docstrings). Existing methods often struggle to capture subtle semantic relationships, leading to inaccurate assessments of documentation quality and hindering code comprehension, maintenance, and collaboration. Specifically, there is currently a lack of reliable automated tools to verify that code comments accurately reflect the functionality implied by their corresponding function names.

[0013] To address the above problems, the present invention provides a computer-implemented method and system that utilizes a specialized neural network architecture, specifically a Bidirectional Long Short-Term Memory (BiLSTM)-based model, to learn and compare semantic representations of code identifiers and their descriptions.

[0014] By capturing deeper context and sequence information, the present invention achieves higher accuracy in semantic similarity assessment compared to traditional methods. It provides an automated way to check the consistency between code and comments, reducing the workload of manual review. This helps to improve and maintain high-quality code documentation, thereby improving the maintainability and understandability of the software. Due to the learned semantic representation, the present invention is more robust to changes in wording, syntax, and abbreviations. Although the example code provided may have been trained using a general dataset, the method can be fine-tuned for domain-specific code bases to achieve better performance.

[0015] Specifically, the method of the present invention comprises the following main steps: S1: Acquire data, parse the input identifier-description pairs, tokenize them, build a shared or independent vocabulary, convert tokens into numeric sequences, and pad or truncate the sequences to a fixed length in preparation for further processing. Specifically, S1 includes: S11: Receive input data, which contains pairs of a first text sequence (e.g., a function name) and a second text sequence (e.g., a function comment). The sample code demonstrates reading data from a CSV file (e.g., train.csv), where each row contains question1, question2, and the label is_duplicate.

[0016] S12: Tokenize each text sequence into a list of tokens (e.g., by using a tokenize function: converting to lowercase, removing characters except non-alphanumeric characters, and splitting by spaces).

[0017] S13: Build a vocabulary based on the corpus of all tokens, including special tokens such as <pad>(fill) and <unk>(unknown word) (e.g., via the build_vocab function).

[0018] S14: Convert the tokenized sequence into a numeric sequence based on the vocabulary (e.g., via the text_to_sequence function).

[0019] S15: pad or truncate the digital sequence to a predefined maximum length (MAX_LEN) (for example, implemented by a pad_sequence function).

[0020] S2: Provides a neural network model that allows the processed identifier sequence and description sequence to be input into the model's parallel embedding layer, and then into the model's shared-weight bidirectional LSTM layer (thus forming a twin network structure for sequence encoding), generating a context-aware vector representation, and combining the output representations from the two BiLSTMs. The combined feature vector is passed through one or more fully connected layers, and finally an output layer predicts the similarity score for each input. Further, S2 specifically includes: S21: An embedding layer (nn.Embedding) that converts the numeric input sequence into a dense vector representation (of dimension EMBEDDING_DIM).

[0021] S22: At least one bidirectional LSTM layer (nn.LSTM) with the specified number of hidden units (HIDDEN_DIM) and number of layers (NUM_LAYERS), configured to process the embedding sequences of the first and second text inputs. Crucially, the LSTM layers used to process the two inputs share weights (forming a twin LSTM structure).

[0022] S23: A mechanism for combining the outputs of a BiLSTM layer from two input sequences. The sample code concatenates the forward hidden state and backward hidden state of the last BiLSTM layer for each input sequence (i.e., hidden1_last_fwd, hidden1_last_bwd, hidden2_last_fwd, hidden2_last_bwd), and then further concatenates the four resulting vectors.

[0023] S24: One or more fully connected layers (nn.Linear, contained in self.fc) that receive the combined BiLSTM output and transform it, typically using a nonlinear activation function (such as nn.ReLU()) and dropout (nn.Dropout), to produce the final output logit representing the similarity.

[0024] S3: Using a dataset of identifier-description pairs with known similarity labels (e.g., similar / dissimilar, or hierarchical ratings), the network in this method is trained using a suitable loss function (e.g., binary cross entropy loss with logits) and an optimization algorithm (e.g., Adam). Specifically, S3 includes: S31: Initialize model parameters.

[0025] S32: Iteratively process batches of training data (pairs of sequences and their true similarity labels).

[0026] S33: For each batch, perform the forward propagation of the model to obtain the predicted similarity logit.

[0027] S34: Use a loss function (e.g., nn.BCEWithLogitsLoss) to calculate the loss value, which measures the difference between the predicted logit and the true label.

[0028] S35: Perform backpropagation to calculate the gradient of the loss with respect to the model parameters.

[0029] S36: Update the model parameters using an optimization algorithm (e.g., optim.Adam with a learning rate of LEARNING_RATE).

[0030] S4: Generate similarity scores based on the logit values ​​for inference or prediction. Further, S4 specifically includes: S41: Input new text sequence pairs (e.g., function name and comment) into the trained model.

[0031] S42: Get the output logit from the model.

[0032] S43: Apply the sigmoid function to the logit to convert it into a probability score (between 0 and 1) representing the likelihood of semantic similarity.

[0033] S44: Optionally, apply a threshold (THRE) to the probability score to perform a binary classification (eg, “similar” or “dissimilar”).

[0034] The BiLSTM architecture employed in this paper, particularly its twin-network-like weight-sharing mechanism, has been shown to be effective in learning semantic similarity between sentences. By applying this architecture to the specific and practical software engineering problem of code identifiers and comments, combined with a specific preprocessing process and output combination strategy, this paper provides a concrete, practical technical solution, rather than just an abstract algorithm. Code identifiers and comments are inherently sequential data, where word order and context are crucial. BiLSTM can effectively capture these dependencies in both directions, making this model well-suited for solving the target problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic block diagram of an exemplary system architecture for evaluating semantic similarity between a code identifier and a description according to one embodiment of the present invention.

[0036] Figure 2 4 is a flowchart of a method for training a semantic similarity evaluation model according to one embodiment of the present invention.

[0037] Figure 3 4 is a flowchart of a method for predicting semantic similarity using a trained model according to one embodiment of the present invention.

[0038] Figure 4 is a more detailed schematic diagram of a BiLSTM-based neural network architecture for semantic similarity evaluation according to one embodiment of the present invention. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] 1 System Overview (See Figure 1 ) This invention proposes a method and system for evaluating the semantic similarity between software code identifiers (such as function names) and their associated natural language descriptions (such as code comments). The system primarily comprises a data processing unit, a model training unit, and a similarity prediction unit. The data processing unit receives raw text data and converts it into a format suitable for model processing. The model training unit uses the preprocessed data to train a core semantic similarity model. The similarity prediction unit then uses the trained model to perform similarity scoring on new code identifier-description pairs.

[0041] The system can be deployed on computing devices with appropriate computing power, such as servers or personal computers, which are equipped with processors, memory, storage, and are capable of running corresponding software instructions (for example, a Python environment with libraries such as PyTorch installed). The setting DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu") in the code indicates that the system can automatically choose to use the GPU or CPU for calculations based on hardware conditions, reflecting the adaptability of the hardware.

[0042] 2 Data Acquisition and Preprocessing Module Data acquisition and preprocessing are the basis of the entire process, and their purpose is to convert raw text data into numerical representations that can be effectively learned by the neural network model.

[0043] 2.1 Input Data Format The input data consists of paired text sequences. In this example, the primary data source is a CSV file (e.g., csv_file = r"D:\chorme download\train.csv\train.csv" specified in the code, although the path is exemplary). This CSV file is expected to contain at least three columns: question1 (representing the first text sequence, such as a function name), question2 (representing the second text sequence, such as a function comment), and is_duplicate (a binary label representing whether the two text sequences are similar, with 1 indicating similarity and 0 indicating dissimilarity). The code snippet first reads this CSV file and then constructs positive samples (similar pairs) and negative samples (dissimilar pairs) based on the value of the is_duplicate column.

[0044] While the primary path currently uses CSV files, the code also includes a parse_data(file_path) function designed to parse a text file in a specific format (e.g., result_gpt-4o.txt). Records in this file consist of key-value pairs such as guess_name:, description_guess:, separated by blank lines. This function can be used as an alternative or alternative method for data acquisition, processing data from different sources or formats.

[0045] 2.2 Tokenization After obtaining the text pairs, each text sequence needs to be tokenized. This is achieved using the tokenize(text) function. This function first converts the input text to lowercase ( text.lower() ) to eliminate case differences. Then, the regular expression re.sub(r'[^a-z0-9\s]', '', text) is used to remove non-alphanumeric characters (but retain spaces). This simplifies the vocabulary and removes potential noise. Finally, text.split() is used to split the cleaned text into a list of tokens by space.

[0046] 2.3 Vocabulary Construction After tokenization, a vocabulary needs to be built, mapping tokens to unique integer indices. This is done with the build_vocab(tokenized_texts) function. This function takes a list of tokenized text sequences, first collects all the tokens, and counts the frequency of each token using collections.Counter.

[0047] The vocabulary vocab contains two special tokens when it is initialized: <pad>: Used for filling, the index is usually 0.

[0048] <unk>: This is used to represent unregistered words (words not in the vocabulary), and its index is usually 1. The vocabulary is then updated based on word frequency. Tokens that meet a specific condition (count > 0 in the code, meaning all words that have appeared) are added to the vocabulary and assigned indices starting at 2. The final vocabulary size, vocab_size, is determined by the number of distinct tokens plus the number of special tokens.

[0049] 2.4 Sequence Conversion and Padding After building the vocabulary, each tokenized text sequence needs to be converted into a sequence of numbers. This is achieved through the text_to_sequence(tokens, vocab) function, which traverses the token list tokens and finds the index of each token in vocab. If the token does not exist, it is used <unk>The index of the word.

[0050] Since neural network models usually require fixed-length input, it is necessary to pad or truncate the numeric sequence. The pad_sequence(sequence, max_len, pad_value) function is responsible for this task. It receives a numeric sequence sequence, a preset maximum length MAX_LEN (for example, configured as 100 in the code), and a padding value pad_value (usually <pad>The index of the token, which is 0. If the sequence length is less than MAX_LEN, pad_value is added to the end of the sequence until it reaches MAX_LEN. If the sequence length is greater than MAX_LEN, the sequence is truncated from the end to MAX_LEN.

[0051] 2.5 Dataset Partitioning and Loading After preprocessing, the data is converted into PyTorch tensors. sequences_guess_tensor and sequences_real_tensor store the numerical representations of all first and second text sequences respectively, and labels_tensor stores the corresponding labels.

[0052] The dataset is then split into training and test sets using the sklearn.model_selection.train_test_split function. The split ratio is controlled by test_size (e.g., 0.15), and random_state (e.g., 42) is used to ensure reproducibility. Importantly, the parameter stratify=labels_tensor ensures that the proportion of samples in each class in the training and test sets remains consistent with that in the original dataset, which is crucial for classification tasks, especially when the classes are imbalanced.

[0053] To facilitate model training, we define a custom PyTorch Dataset class, SimilarityDataset. Its instances accept a partitioned sequence tensor and a label tensor. The __len__ method returns the dataset size, and the __getitem__ method returns a pair of sequences and their labels, given an index.

[0054] Finally, torch.utils.data.DataLoader is used to create data loaders (train_loader and test_loader) for the training and test sets. DataLoader can organize the data into batches according to the specified BATCH_SIZE (for example, 32) and shuffle the training data at the beginning of each training round (epoch) (shuffle=True for train_loader), which helps the model generalize.

[0055] 3 BiLSTM-based similarity model architecture (BiLSTMSimilarity) (see Figure 4 ) The core semantic similarity assessment model of this paper is a neural network based on a bidirectional long short-term memory (BiLSTM) network, designed to efficiently extract and compare semantic information from two input text sequences. This model is implemented in PyTorch as a class called BiLSTMSimilarity, which inherits from nn.Module.

[0056] 3.1 Model Initialization (__init__) The model's __init__ constructor accepts the following key parameters to define its structure: vocab_size: The size of the vocabulary, which determines the input dimension of the embedding layer.

[0057] embedding_dim: The dimension of the word embedding vector (for example, EMBEDDING_DIM = 128 in the code).

[0058] hidden_dim: The dimension of the LSTM hidden layer (for example, HIDDEN_DIM = 64 in the code). For a bidirectional LSTM, the actual output dimension at each time step is hidden_dim * 2.

[0059] num_layers: The number of layers in the LSTM stack (for example, NUM_LAYERS = 4 in the code).

[0060] dropout_prob: The dropout ratio (e.g. 0.2) applied to different parts of the model (e.g. after the embedding layer, between LSTM layers, between fully connected layers) for regularization and to prevent overfitting.

[0061] The internal components of the model include: Embedding layer (self.embedding): nn.Embedding(vocab_size, embedding_dim,padding_idx=vocab). This layer converts the input sequence of integer indices (from tokenization and vocabulary mapping) into dense word embedding vectors. The padding_idx parameter specifies the indices of the padding tokens so that these padding tokens are not considered when computing the gradient, which is crucial for processing variable-length sequences.

[0062] BiLSTM layer (self.lstm): nn.LSTM(embedding_dim, hidden_dim, num_layers=num_layers, bidirectional=True, batch_first=True, dropout=dropout_prob ifnum_layers > 1 else 0).

[0063] bidirectional=True is the core feature of this layer, which means that LSTM processes the input sequence from both forward (left to right) and backward (right to left) directions, thereby capturing more comprehensive contextual information for each token.

[0064] batch_first=True specifies that the dimension order of the input and output tensors is (batch, seq_len, features).

[0065] If num_layers > 1, the dropout parameter applies dropout to the output of each LSTM layer except the last one.

[0066] Key Point: In the forward method, the same self.lstm instance is used to process two different input sequences (embedded1 and embedded2). This means that the weights of the BiLSTM layer are shared. This architecture, known as a Siamese network, is particularly well-suited for learning similarities or distances between two inputs. It forces the model to learn a mapping function that maps semantically similar inputs to similar locations in the feature space.

[0067] Dropout layer (self.dropout): nn.Dropout(dropout_prob). In the code, this dropout layer is applied after the embedding operation and before the input to the LSTM.

[0068] Fully connected layer (self.fc): This is a nn.Sequential container containing a series of linear layers and activation functions, which are used to combine and map the features extracted by BiLSTM to the final similarity score.

[0069] The first linear layer: nn.Linear(hidden_dim * 2 * 2, hidden_dim). Its input dimension hidden_dim * 2 * 2 (i.e., hidden_dim * 4) is calculated as follows: For each input sequence, the final BiLSTM layer outputs a forward hidden state and a backward hidden state, each of dimension hidden_dim. Therefore, the final representation of a sequence (by concatenating the forward and backward states) has dimension hidden_dim * 2. Since there are two input sequences and their final representations are concatenated, the total input dimension is (hidden_dim * 2) + (hidden_dim * 2) = hidden_dim * 4.

[0070] Activation function: nn.ReLU(), introducing nonlinearity.

[0071] Dropout layer: nn.Dropout(dropout_prob), further regularization.

[0072] Output linear layer: nn.Linear(hidden_dim, 1). This layer outputs a scalar value (logit) representing the raw similarity score between the two input sequences.

[0073] 3.2 Forward Propagation (forward(self, seq1, seq2)) The forward method defines the flow path of data in the model: Embedding: Two input sequences seq1 and seq2 (of shape (batch_size, seq_len)) are passed through the shared embedding layer self.embedding, and then dropout is applied. The resulting embedded1 and embedded2 (of shape (batch_size, seq_len, embedding_dim)) are obtained.

[0074] LSTM Processing: embedded1 and embedded2 are each passed through a shared BiLSTM layer, self.lstm. lstm_out1, (hidden1, cell1) = self.lstm(embedded1) lstm_out2, (hidden2, cell2) = self.lstm(embedded2) lstm_out contains the output of each time step, while hidden contains the hidden state of the last time step (cell is the cell state). For a bidirectional multi-layer LSTM, the shape of hidden is (num_layers * num_directions, batch_size, hidden_dim).

[0075] Hidden State Extraction: The model focuses on the final semantic representation of each sequence, which is typically represented by the hidden state of the last time step of the BiLSTM. Because it is a bidirectional LSTM, the forward and backward hidden states of the last time step need to be combined. For an N-layer bidirectional LSTM, index -2 in the hidden tensor corresponds to the forward hidden state of the last layer, and index -1 corresponds to the backward hidden state of the last layer. hidden1_last_fwd = hidden1[-2, :, :] hidden1_last_bwd = hidden1[-1, :, :] Similar operations are performed on hidden2 to obtain hidden2_last_fwd and hidden2_last_bwd.

[0076] Combining Hidden States: Concatenate the four extracted hidden state vectors (two from seq1 and two from seq2) along dimension 1 (feature dimension): combined_hidden = torch.cat((hidden1_last_fwd, hidden1_last_bwd, hidden2_last_fwd, hidden2_last_bwd), dim=1) The shape of combined_hidden is (batch_size, hidden_dim * 4).

[0077] Fully Connected Layers and Output: The combined_hidden vector is fed into the sequential fully connected layer defined in self.fc. The final output is of shape (batch_size, 1). Output.squeeze(1) removes the last dimension, resulting in a tensor of shape (batch_size), where each element is the raw similarity logit for the corresponding input pair.

[0078] This strategy of processing two inputs through a shared BiLSTM and then combining their final representations and feeding them into subsequent layers for classification is an effective way to learn pairwise sequence relations such as similarity.

[0079] 4 Model training process (see Figure 2 ) The goal of model training is to adjust the parameters (weights and biases) of the BiLSTMSimilarity model so that it can accurately distinguish between similar and dissimilar text pairs.

[0080] 4.1 Criterion Select nn.BCEWithLogitsLoss() as the loss function. This loss function combines a Sigmoid layer with Binary Cross-Entropy Loss and is suitable for binary classification problems (outputting a logit representing similarity). Using BCEWithLogitsLoss directly is more numerically stable than first passing it through a Sigmoid layer and then using BCELoss. It measures the difference between the model's output logit and the true label (0 or 1).

[0081] 4.2 Optimizer Use optim.Adam(model.parameters(), lr=LEARNING_RATE) as the optimizer. Adam is a widely used adaptive learning rate optimization algorithm that combines the advantages of the AdaGrad and RMSProp algorithms and is generally able to efficiently find good parameter solutions. LEARNING_RATE (for example, set to 0.001 in the code) is a hyperparameter that controls the parameter update step size.

[0082] 4.3 Training Loop The training process is performed in multiple epochs, each of which is a complete pass through the training dataset. The total number of epochs is specified by NUM_EPOCHS (e.g. 20).

[0083] At the beginning of each epoch, call model.train() to put the model into training mode. This enables the training behavior of dropout and batch normalization (if included in the model).

[0084] Then, iterate over each batch of data provided by train_loader: Data preparation: Move the input sequence (batch_guess, batch_real) and label (batch_labels) of the current batch to the specified computing device DEVICE (GPU or CPU).

[0085] Gradient zeroing: optimizer.zero_grad(). Before calculating new gradients, the previously calculated gradients must be cleared.

[0086] Forward propagation: outputs = model(batch_guess, batch_real). Input the batch data into the model and get the predicted logit.

[0087] Calculate the loss: loss = criterion(outputs, batch_labels). Use the loss function to calculate the loss between the predicted logit and the true label.

[0088] Backward propagation: loss.backward(). Based on the loss value, the gradient of the loss with respect to all trainable parameters of the model is automatically calculated.

[0089] Parameter update: optimizer.step(). The optimizer updates the model parameters based on the calculated gradient.

[0090] During training, you can periodically print the loss value of the current batch and the average training loss at the end of each epoch to monitor the training progress.

[0091] 4.4 Validation After each training epoch, you can choose to evaluate the model's performance on the test set (or a separate validation set). This helps monitor whether the model is overfitting and can be used to tune hyperparameters or implement strategies such as early stopping.

[0092] The verification process is as follows: Calling model.eval() sets the model to evaluation mode. This disables dropout and causes the batch normalization layer (if present) to use its running statistics learned during training.

[0093] Using the torch.no_grad(): context manager, computations within this block do not track gradients, saving memory and computational resources.

[0094] Iterate over each batch of data provided by the test_loader (or validation loader).

[0095] Perform forward propagation to get the output logit.

[0096] Convert logit to probability through the Sigmoid function: torch.sigmoid(outputs).

[0097] Convert the probabilities to binary prediction labels (0 or 1) according to a preset threshold THRE (e.g. 0.5): (torch.sigmoid(outputs) > THRE).long().

[0098] Collect all predicted labels and true labels, and then calculate evaluation metrics such as accuracy (accuracy_score). The validation accuracy after each epoch provides a snapshot of the model performance.

[0099] 5 Evaluation Process After training is complete, a final evaluation is performed on the test set to measure the model's generalization ability on unseen data. The evaluation process is similar to the validation process: the model is set to evaluation mode (model.eval()) and the torch.no_grad() environment is used.

[0100] Predict all samples in the test set, collect the predicted labels and true labels. Then calculate a series of evaluation metrics, including: Accuracy: accuracy_score(all_labels, all_preds), the proportion of samples that are correctly classified.

[0101] Precision, Recall, and F1-Score: precision_recall_fscore_support(all_labels, all_preds, average='binary'). These metrics provide a more comprehensive view of performance for binary classification tasks, especially when the classes may be unbalanced. average='binary' means that the metrics are calculated for the positive class (label 1).

[0102] Together, these metrics reflect the model's ability to distinguish the semantic similarity between code identifiers and their descriptions.

[0103] 6 Example prediction / inference (see Figure 3 ) The trained model can be used to predict the semantic similarity between new, unseen text pairs. The code provides the predict_similarity(text1, text2, model, vocab, max_len, device) function to implement this function.

[0104] The function receives two raw text strings text1 and text2, as well as the trained model, vocabulary, maximum sequence length, and computing device.

[0105] The internal steps are as follows: Set the model to evaluation mode: model.eval().

[0106] Perform the same preprocessing steps on the input text1 and text2 as for the training data: Use the tokenize function to perform lemmatization.

[0107] Use the text_to_sequence function to convert a word sequence into a numeric sequence.

[0108] Use the pad_sequence function to pad / truncate a numeric sequence to obtain a sequence of fixed length.

[0109] Convert the processed sequence to a PyTorch tensor and move it to the specified device.

[0110] In the with torch.no_grad(): environment, the tensorized sequence is input into the model for forward propagation to obtain the output logit.

[0111] Convert logit to a probability value through the Sigmoid function: probability = torch.sigmoid(output).item(). This probability value (between 0 and 1) indicates the possibility that the model predicts that text1 and text2 are semantically similar.

[0112] The function returns this probability value. The final "similar" or "dissimilar" judgment can be made based on this probability and the preset threshold THRE.

[0113] The code also includes an example loop that makes predictions for each pair of text in raw_data (which may come from the result_gpt-4o.txt file, or in the current implementation, more likely refers to a subset or raw form of the data originally loaded from CSV and used for training / testing) and calculates the average similarity probability. This shows how the model can be used in a real application.

[0114] 7 Model Saving and Vocabulary Persistence In order to be able to reload and use a trained model in the future, or deploy the model in a different environment, its parameters and associated vocabulary need to be saved.

[0115] Save model parameters: torch.save(model.state_dict(), 'bilstm_similarity_model.pth'). This command saves all the learnable parameters (weights and biases) of the model to a file named bilstm_similarity_model.pth. state_dict is a Python dictionary object that maps each layer to its parameter tensor.

[0116] Vocabulary saving: Python import json with open('vocab.json', 'w') as f: json.dump(vocab, f) The vocabulary vocab (a dictionary mapping tokens to integer indices) is saved as a JSON file vocab.json. When loading the model for inference, the new input text must be preprocessed using the exact same vocabulary as used during training.

[0117] These saving steps ensure the portability and reusability of the model.

[0118] 8 Implementation Details and Hyperparameters The specific implementation of the present invention depends not only on the model architecture, but is also significantly influenced by a series of configuration parameters and hyperparameters. These parameters together define the specific operating state and learning behavior of the model. Table 1 below lists the key configuration parameters and hyperparameters used in this embodiment and their example values, which are derived from the provided code.

[0119] The selection of these parameters is typically based on experience and experimental tuning. For example, the setting of MAX_LEN requires a trade-off between information retention and computational efficiency; EMBEDDING_DIM and HIDDEN_DIM control the capacity and expressiveness of the model; and NUM_EPOCHS and LEARNING_RATE jointly affect the convergence speed of training and the final performance. Although the patent claims may not limit the specific values ​​of these parameters, providing this information in the detailed description helps those skilled in the art reproduce and understand a specific and effective embodiment of the present invention.

[0120] The core innovation of this invention lies not only in its use of the BiLSTM architecture, but also in its specific configuration (e.g., number of layers, hidden dimensions), the combination of the outputs of the two text branches (concatenating the final hidden states), the choice of loss function (BCEWithLogitsLoss), and the application of this entire process to more efficiently solve the unique problem of assessing semantic similarity between code identifiers and comments than large models. These specific implementation details constitute a key difference between this invention and the prior art.

[0121] 9 Model Performance Examples To demonstrate the effectiveness of the methods and systems described in this article, the provided code was used for training and evaluation on a publicly available text similarity dataset (e.g., the Quora Question Pairs dataset, loaded via the train.csv file in the code). Although this dataset is not specifically targeted at code identifiers and comments, it serves as a proxy task for evaluating the model's ability to learn general semantic similarity. Under specific hyperparameter configurations (as shown in Table 1), the model can achieve the performance metrics shown in Table 2 below.

[0122] Note: The values ​​in Table 2 are exemplary results based on a general text similarity dataset and are not actually used in code labeling. When evaluating the similarity between identifiers and annotations, it is recommended to use data from that specific field for training and evaluation. The performance may be Different.

[0123] These exemplary performance metrics demonstrate that the BiLSTM-based system described in this paper is able to effectively learn semantic similarity between text pairs. By training or fine-tuning on datasets from the target application domain (i.e., code identifiers and annotations), targeted high performance is expected to be achieved, enabling reliable automated assessment of consistency between code and its documentation.< / pad> < / unk> < / unk> < / pad> < / unk> < / pad>

Claims

1. A method for evaluating the semantic similarity between a software code identifier and a related natural language description using a bidirectional recurrent neural network, characterized in that: The following steps are involved: S1: Gets data and parses the input identifier-description pairs, tokenizes them, builds a shared or independent vocabulary, converts tokens into numeric sequences, and pads or truncates the sequences to a fixed length in preparation for further processing; S2: Provide a neural network model that can input the processed identifier sequence and description sequence into the model's parallel embedding layer, and then input them into the model's shared-weight bidirectional LSTM layer to generate context-aware vector representations, and combine the output representations from the two BiLSTMs. Pass the combined feature vector through one or more fully connected layers, and finally predict the similarity score of each input through an output layer; S3: Use a dataset of identifier-description pairs containing labels with known similarity, and adopt a suitable loss function and optimization algorithm to train the network in this method; S4: Generate similarity scores based on the logit values ​​for use in inference or prediction.

2. The method according to claim 1, characterized in that Said S1 further comprises: S11: Receive input data, which includes a paired first text sequence and a second text sequence; S12: Tokenize each text sequence into a token list; S13: Build a vocabulary based on the corpus of all tokens, including special tokens such as <pad>and <unk> ;< / unk> < / pad> S14: Convert the lemma sequence into a numerical sequence based on the vocabulary; S15: Pad or truncate the numeric sequence to a predefined maximum length.

3. The method according to claim 1, characterized in that Said S2 further comprises: S21: An embedding layer that converts the digital input sequence into a dense vector representation; S22: at least one bidirectional LSTM layer configured to process the embedding sequence of the first text input and the second text input, and the LSTM layer for processing the two inputs shares weights; S23: A mechanism for combining the outputs of a BiLSTM layer from two input sequences; S24: One or more fully connected layers that receive the combined BiLSTM output and transform it to produce the final output logit representing the similarity.

4. The method according to claim 1, wherein Said S3 further comprises: S31: Initialize model parameters; S32: iteratively process training data batches; S33: For each batch, perform the forward propagation of the model to obtain the predicted similarity logit; S34: Calculate the loss value using the loss function; S35: Perform backpropagation to calculate the gradient of the loss with respect to the model parameters; S36: Update model parameters using an optimization algorithm.

5. The method according to claim 1, characterized in that Said S4 further comprises: S41: Input the new text sequence pair into the trained model; S42: Get the output logit from the model; S43: Apply the sigmoid function to the logit to convert it into a probability score representing the possibility of semantic similarity; S44: Optionally, applying a threshold to the probability score for binary classification.

6. A system for evaluating the semantic similarity between software code identifiers and related natural language descriptions using a bidirectional recurrent neural network, characterized in that: The system includes: A data processing unit, for obtaining and preprocessing input identifier-description pairs, including lemmatization, vocabulary building, and sequence conversion and filling; A model training unit for training a shared-weight bidirectional LSTM-based neural network model to evaluate the semantic similarity between identifiers and descriptions; The similarity prediction unit is used to use the trained model to perform similarity scoring on new code identifier-description pairs.

7. The system according to claim 6, characterized in that The neural network model includes: Embedding layers, which convert numeric input sequences into dense vector representations; A bidirectional LSTM layer with shared weights to generate context-aware vector representations; Fully connected layer, used to combine BiLSTM outputs and predict similarity scores; The output layer is used to output the similarity logit value.

8. The system according to claim 6, wherein: The model training unit uses a data set containing known similarity labels for training, and adopts a loss function and an optimization algorithm to adjust model parameters.

9. The system according to claim 6, wherein: The similarity prediction unit inputs a new text sequence pair into the trained model, obtains the output logit, and converts it into a probability score representing the possibility of semantic similarity.

10. The system according to claim 9, characterized in that The similarity prediction unit optionally applies a threshold to the probability scores for binary classification.