A method of programming a response time prediction for a question within a question and answer community
Patent Information
- Application Number
- CN202311551011.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-11-20
AI Technical Summary
[0005]又如:CN114238621A所公开的“一种基于Transformer的编程问题帖标题自动生成方法”,其针对的的确是编程社区中的问题,但是,其是用于生成问题帖标题,与对问题回复时间预测无关
[0033]1、本发明针对编程社区的特点,将问答模块中问题的文本特征与数字特征进行融合,通过全连接层实现回复时间的预测,一方面实现了预测回复时间的特定功能,另一方面,该预测综合考虑了编程社区中不同版块自身的特点,使得对于不同版块的问题均能够进行准确的预测。
Smart Images

Figure CN117573871B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method for predicting the response time of questions in a programming Q&A community, for predicting the response time period or time of questions in a programming Q&A community. Background Technology
[0002] With the rapid development of the information technology industry, programming Q&A communities have become important platforms for knowledge sharing. Developers can share their experiences, insights, code, and solutions through these communities. This sharing helps other developers learn new technologies, solve problems, and continuously improve their own learning and growth. In programming Q&A communities, developers can ask questions and receive help, suggestions, and solutions from other community members. This collective wisdom and mutual support enables faster and more effective problem-solving. To facilitate discussions among developers, programming Q&A communities include a Q&A module for discussing issues encountered during use.
[0003] However, the diversity of questions and their popularity across different fields lead to varying response times for each question. Even when a question is raised, it's difficult to know whether it will attract attention from others. This makes it challenging for those involved to predict response times, indirectly reducing development efficiency. Traditional prediction methods typically rely on the response times of previous questions in the same field or similar questions in other fields. However, these methods tend to have significant prediction biases when there are few or no relevant questions, and sometimes fail to provide accurate predictions.
[0004] Research on community question-and-answer models is quite common, but the focus is not on predicting question-and-answer time. For example, CN112765326, "A Method, System and Application for Expert Recommendation in Question-and-Answer Communities," analyzes relevant information about questions and recommends expert users to answer new questions, emphasizing the recommendation relationship between questions and responders, but lacking prediction of the response time.
[0005] For example, CN114238621A discloses "A method for automatically generating titles of programming question posts based on Transformer". It does target questions in the programming community, but it is used to generate titles of question posts and has nothing to do with predicting the time to reply to questions.
[0006] For example, CN116226352A discloses "An answer selection method that considers the spatiotemporal dependency of question and answer". It also focuses on the question, but it selects the answer with a high degree of matching from the existing question and answer as the reply. It does not care about and cannot predict the reply time. Summary of the Invention
[0007] The purpose of this invention is to provide a method for predicting the response time of questions in a programming Q&A community, addressing all or part of the aforementioned problems, so as to predict the time it will take for a new question to be answered in the programming community.
[0008] The technical solution adopted in this invention is as follows:
[0009] A method for predicting response time for questions within a programming Q&A community, comprising:
[0010] The trained prediction model is used to predict the response time to the problem to be predicted; wherein, the training of the prediction model includes:
[0011] We collected the question text, response time, number of questions, number of responses, and popularity of the domain from multiple programming Q&A communities.
[0012] Based on the response time, the question texts are divided into different categories, and each category corresponds to a specific response time range;
[0013] The question text is converted into a text vector, the question text features are extracted from the text vector, and the text features are output through an attention layer; the number of questions, the number of replies, and the domain popularity are concatenated into vectors to obtain numerical features; the text features and numerical features are fused to obtain multimodal features;
[0014] A neural network classification model is trained using the multimodal features of each problem and the response time interval to obtain a prediction model.
[0015] Furthermore, the conversion of the question text into a text vector includes:
[0016] The constructed vocabulary is used to encode all words in the question text to obtain the corresponding text vectors, and the length of the encoding for each word is consistent.
[0017] Furthermore, the conversion of the question text into a text vector includes:
[0018] Word2Vec is used to encode each word in the question text to obtain the corresponding text vector.
[0019] Furthermore, a bidirectional long short-term memory network is used to extract text features from text vectors.
[0020] Furthermore, the fusion of the text features and numerical features includes:
[0021] The text features and numerical features are concatenated into vectors.
[0022] Furthermore, the text vector has a dimension of 100, and the text feature has a dimension of 16.
[0023] To address all or some of the aforementioned problems, this invention also provides a method for predicting response time for questions within a programming Q&A community, comprising:
[0024] The trained prediction model is used to predict the response time to the problem to be predicted; wherein, the training of the prediction model includes:
[0025] We collected the question text, response time, number of questions, number of responses, and popularity of the domain from multiple programming Q&A communities.
[0026] The question text is converted into a text vector, the question text features are extracted from the text vector, and the text features are output through an attention layer; the number of questions, the number of replies, and the domain popularity are concatenated into vectors to obtain numerical features; the text features and numerical features are fused to obtain multimodal features;
[0027] The neural network regression model is trained using the multimodal characteristics and response time of each problem to obtain the prediction model.
[0028] Furthermore, the conversion of the question text into a text vector includes:
[0029] Word2Vec is used to encode each word in the question text to obtain the corresponding text vector.
[0030] Furthermore, a bidirectional long short-term memory network is used to extract text features from text vectors.
[0031] Furthermore, the text vector has a dimension of 100, and the text feature has a dimension of 16.
[0032] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0033] 1. This invention, tailored to the characteristics of programming communities, integrates the textual and numerical features of questions in the question-and-answer module and uses a fully connected layer to predict response times. On the one hand, it achieves the specific function of predicting response times; on the other hand, this prediction comprehensively considers the characteristics of different sections within the programming community, enabling accurate prediction of questions in different sections.
[0034] 2. This invention introduces an attention mechanism when extracting features from question text. Compared with traditional sequential learning where the weights of the acquired features are consistent, this is more in line with the fact that feature weights are inconsistent in real-world problems. This makes the output text features more consistent with the characteristics of programming question-and-answer community questions, thereby making the learning effect of the classifier model more consistent with the characteristics of programming question-and-answer community questions.
[0035] 3. Compared with traditional machine learning and BERT, this invention has higher accuracy and better generalization ability in terms of comprehensive indicators. Attached Figure Description
[0036] The present invention will be described by way of example and with reference to the accompanying drawings, wherein:
[0037] Figure 1 This is the flowchart of Example 1.
[0038] Figure 2 This is a diagram of a text feature extraction model. Detailed Implementation
[0039] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.
[0040] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.
[0041] Example 1
[0042] A method for predicting response times for questions in a programming Q&A community, used to predict the time frame within which a new question will receive a response. For example... Figure 1 As shown, the method includes:
[0043] Question text, responses, and corresponding response times are collected from a programming Q&A community (Q&A module) and categorized as text information. The number of questions and responses in each question's domain, as well as domain popularity, are also collected and categorized as numerical features. Based on the response time, the question text is divided into different categories, each representing a specific response time interval. In this embodiment, the response time intervals are divided into five categories: within 1 hour, 1 hour to 12 hours, 12 hours to 24 hours, 24 hours to one week (or half a month) (this category falls under 24 hours but can be estimated), and one week (or half a month) or more (this category falls under the category of no recent response, i.e., difficult to predict). Each response time interval is identified by a label. This collected or categorized data is used as samples for training the prediction model.
[0044] For the question text, all words are encoded using the constructed vocabulary, transforming the question text into a set of text vectors. This involves converting a single text into a matrix, representing each word in the question text with a vector to obtain a data format suitable for subsequent models. This process is typically called word embedding or text embedding, and the corresponding vectors are usually called embedding vectors. Additionally, to standardize the input format, whitespace is added to the vocabulary to normalize the length of the input question text, truncating excessively long text and adding whitespace to meet the uniform length requirement.
[0045] In some embodiments, the words in the question text are encoded using Word2Vec (i.e., word embeddings). Word2Vec represents each word as a vector based on the context of the question text and learns the parameters of these vectors through a prediction objective function. The core of the Word2Vec network is a single-hidden-layer feedforward neural network, where both the input and output of the network are word vectors. After constructing the word embedding model, words can be transformed into vector form while preserving the relationships between words in their context. The output of this model serves as the basis for subsequent text feature extraction.
[0046] In principle, obtaining the question text and corresponding response time is sufficient to train a prediction model. However, testing revealed that for programming Q&A communities, the prediction model trained solely on this information only accurately predicted the response time for new questions in some sections, but differed significantly from reality in others. Even for the same question, responses in popular areas were consistently faster than in general areas. That is, considering only the question text is insufficient for accurate prediction of response times in programming communities due to the differences between different areas. To address this issue, these differences need to be considered. Therefore, this invention, in addition to collecting historical question texts, also collects the digital features of the questions to provide extra information for prediction and mitigate the impact of area differences on response time prediction.
[0047] In some embodiments, the numerical features of the problem are concatenated into a vector-like numerical feature. This vector serves as another input to the prediction model (the text vector is one input), used to jointly train (or jointly predict) the prediction model with the text features. The numerical features of the problem can provide additional information on top of the text features, offering more comprehensive data for the model to predict response time.
[0048] In some embodiments, a bidirectional long short-term memory network (BiLSTM) is used to extract features from the question text, which are then processed by an attention layer to obtain the text features. Long short-term memory network (LSTM) is a type of recurrent neural network (RNN). In practical applications, RNNs suffer from problems such as vanishing gradients, exploding gradients, and poor ability to handle long-range dependencies. Therefore, this invention uses LSTM instead of RNN for question text feature extraction. LSTM is similar to RNN in its main structure, but it incorporates a forget gate, which effectively avoids problems such as vanishing gradients, exploding gradients, and poor ability to handle long-range dependencies. Therefore, in this invention, LSTM is used to extract features from the question text. Using a bidirectional LSTM allows for feature extraction from both forward and backward directions, thus fully capturing the contextual relationships of the question text.
[0049] In this embodiment of the invention, a bidirectional LSTM is used to extract features from the problem text from both the forward and backward directions. Due to the presence of the LSTM forget gate, the problems of gradient explosion or gradient vanishing during the feature extraction process can be effectively avoided. The output of the bidirectional LSTM is a 16-dimensional vector (the dimension can be set by the user), which can be used for subsequent classification work.
[0050] Meanwhile, traditional sequence learning assigns consistent weights to the acquired features. However, in practical problems, feature weights are often inconsistent. Therefore, this invention employs an attention mechanism that allows the model to assign weights to the sequence output. The bidirectional LSTM outputs are then summed after passing through an attention layer to obtain the final model output, i.e., the text features. The model structure for extracting text features is as follows: Figure 2 As shown.
[0051] As for the spliced digital features, they are themselves derived from the splicing of digital vectors and can be used directly as features.
[0052] For the acquired text and numeric features, the prediction model first fuses these two features. The numeric feature is a 3-dimensional vector, and the text feature is a 16-dimensional vector; both have the same norm. Concatenating these two vectors completes feature fusion. The numeric and text features are concatenated first and second to form a 19-dimensional vector, which serves as the input to the classifier in the prediction model. Additionally, the label for the response time interval is also input into the prediction model for learning. This label can be encoded using one-hot encoding or other encoding methods.
[0053] For the classifier, a neural network classification model is used as the base model for learning. In this embodiment, the multimodal feature dimension is 19, therefore, the input dimension of the classifier is 19. Furthermore, this embodiment aims to predict the response time period of the question using the prediction model, so the classifier is constructed with an output dimension of 5 (or similarly configured dimensions), outputting 5 categories of prediction results, corresponding to the time of sample collection: response within 1 hour, response between 1 and 12 hours, response between 12 and 24 hours, response after 24 hours, and no response in the near future (i.e., difficult to predict). The dimension conversion between the input and output dimensions is accomplished through a fully connected layer. For the classifier's output, the original output of the model is converted into a probability distribution using the Softmax function. Each element is exponentially calculated and then normalized to a probability value, as shown in the following formula.
[0054]
[0055] Where Softmax(x) i Let x be the i-th element of the output probability distribution. i For the i-th element of the multimodal feature (vector) after fusing digital and textual features, the Softmax function is used in the last layer of the model output to ensure that the final output of the model is an effective probability distribution.
[0056] In addition, to prevent overfitting of the neural network, random dropout is introduced before the model output during the training phase. This involves setting a portion of the elements of the input tensor to zero. In this invention, the dropout parameter is set to 0.5, which means that a portion of the input elements are ignored with a probability of 0.5. This can effectively reduce the model's dependence on specific input features and enhance the model's generalization ability.
[0057] The prediction model is trained using the multimodal features and labels (i.e., response time interval representation) of the collected samples. The trained model is then used to predict response time intervals for new questions in the programming community. During prediction, the question text and numerical features are still obtained. The question text is then transformed into vectors and its features are extracted to obtain text features, which, along with the numerical features, are input into the prediction model for prediction. The pseudocode for this embodiment is as follows:
[0058]
[0059]
[0060] Example 2
[0061] This embodiment discloses another method for predicting the response time of questions in a programming question-and-answer community. This method is largely similar to the method in Embodiment 1. The difference is that the purpose of this embodiment is to predict the time (rather than a time period) when a new question will receive a response. Accordingly, this embodiment does not divide the response time interval for each collected question. Instead, it directly uses the time difference between the collected response time and the time when the question was raised as the label of the question. In addition, the classifier uses a neural network regression model (instead of a neural network classification model) as the base model for learning.
[0062] Example 3
[0063] This embodiment modifies the two embodiments described above to obtain another method for predicting response time in programming question-and-answer communities. In this embodiment, the bidirectional LSTM is replaced with an RNN or GRU neural network to extract text features.
[0064] Example 4
[0065] This embodiment modifies Embodiment 1 or Embodiment 2 to obtain another method for predicting response time in programming question-and-answer communities. In this embodiment, when extracting text features using bidirectional LSTM, the dimension of the text vector output by Word2Vec is kept unchanged, and a multi-layer perceptron (MLP) is used to increase the dimension of the digital features.
[0066] Example 5
[0067] To evaluate the model's sensitivity to changes in input parameters, this embodiment conducts a sensitivity experiment to assess the model's stability, robustness, and reliability. The model's input parameters are changed to a certain extent, and the experimental results are shown in Table 1. Accuracy is the proportion of correctly predicted samples to the total number of samples. Recall is defined as shown in formula (8). The average recall of different categories is the model's recall. AUC is the area under the ROC curve, where the ROC curve is plotted with TPR and FPR as the axes. The calculation formulas for TPR and FPR are shown in formula (9), where TP is the number of samples that are actually positive and the model predicts are also positive, TN is the number of samples that are actually negative and the model predicts are also negative, FP is the number of samples that are actually negative but the model predicts are positive, and FN is the number of samples that are actually positive but the model predicts are negative.
[0068]
[0069]
[0070] Table 1. Results of the sensitivity experiment
[0071]
[0072]
[0073] Experimental results show that the best performance is achieved when the word embedding dimension is configured to 100 and the text vector dimension is configured to 16.
[0074] In addition, this embodiment also selects the programming community API (https: / / rapidapi.com / hub) as the experimental object to compare the performance of the present invention (the scheme of embodiment 1) with commonly used machine learning models and BERT. Ten-fold cross-validation is used to evaluate the model's performance and generalization ability. This embodiment selects AUC, recall, and accuracy as evaluation metrics. The comparative experimental results are shown in Table 2. As can be seen from Table 2, the method of fully implementing the present invention outperforms commonly used machine learning models and BERT in multiple metrics.
[0075] Table 2 Comparative Experiment Evaluation Results
[0076]
[0077] Except for the second experiment which only used text features, the other experiments used both text features and numerical features in different models. It can be seen that compared with traditional machine learning, the model constructed in this invention has better performance and generalization ability, can make fuller use of data information, and its effect is also better than BERT.
[0078] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.
Claims
1. A method for predicting response time for questions in a programming question-and-answer community, characterized in that, include: The trained prediction model is used to predict the response time of questions to be predicted in a programming Q&A community, which includes different sections; wherein, the training of the prediction model includes: We collected the question text, response time, number of questions, number of responses, and popularity of the domain from multiple programming Q&A communities. Based on the response time, the question texts are divided into different categories, and each category corresponds to a specific response time range; The question text is converted into a text vector, the question text features are extracted from the text vector, and the text features are output through an attention layer; the number of questions, the number of replies, and the domain popularity are concatenated into vectors to obtain numerical features; the text features and numerical features are fused to obtain multimodal features; A neural network classification model is trained using the multimodal features of each problem and the response time interval to obtain a prediction model.
2. The method for predicting response time for questions in a programming Q&A community as described in claim 1, characterized in that, The process of converting the question text into a text vector includes: The constructed vocabulary is used to encode all words in the question text to obtain the corresponding text vectors, and the length of the encoding for each word is consistent.
3. The method for predicting response time for questions in a programming Q&A community as described in claim 2, characterized in that, The process of converting the question text into a text vector includes: Word2Vec is used to encode each word in the question text to obtain the corresponding text vector.
4. The method for predicting response time for questions in a programming Q&A community as described in claim 1, characterized in that, Text features of text vectors are extracted using a bidirectional long short-term memory network.
5. The method for predicting response time for questions in a programming Q&A community as described in claim 1, characterized in that, The fusion of the text features and numerical features includes: The text features and numerical features are concatenated into vectors.
6. The method for predicting response time for questions in a programming Q&A community as described in claim 1, characterized in that, The text vector has a dimension of 100, and the text feature has a dimension of 16.
7. A method for predicting response time for questions in a programming question-and-answer community, characterized in that, include: The trained prediction model is used to predict the response time of questions to be predicted in a programming Q&A community, which includes different sections; wherein, the training of the prediction model includes: We collected the question text, response time, number of questions, number of responses, and popularity of the domain from multiple programming Q&A communities. The question text is converted into a text vector, the question text features are extracted from the text vector, and the text features are output through an attention layer; the number of questions, the number of replies, and the domain popularity are concatenated into vectors to obtain numerical features; the text features and numerical features are fused to obtain multimodal features; The neural network regression model is trained using the multimodal characteristics and response time of each problem to obtain the prediction model.
8. The method for predicting response time for questions in a programming Q&A community as described in claim 7, characterized in that, The process of converting the question text into a text vector includes: Word2Vec is used to encode each word in the question text to obtain the corresponding text vector.
9. The method for predicting response time for questions in a programming Q&A community as described in claim 7, characterized in that, Text features of text vectors are extracted using a bidirectional long short-term memory network.
10. The method for predicting response time for questions in a programming Q&A community as described in claim 7, characterized in that, The text vector has a dimension of 100, and the text feature has a dimension of 16.
Citation Information
Patent Citations
Transform-based programming problem post title automatic generation method
CN114238621A
Answer selection method considering question and answer space-time dependency relationship
CN116226352A