Method and device for aligning standard items and custom questions of users

By introducing diverse expression methods for government affairs data sets, synonyms replacement and large language model generation in the field of government affairs services, and combining attention-enhancing classification networks, the accuracy and efficiency of matching user common problems with standard government affairs matters are solved, and higher matching accuracy and model adaptability are achieved.

CN120124591APending Publication Date: 2025-06-10INSPUR SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266056.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately match users' common problems and standard government affairs in the field of government services, and it is slow to respond to new problems and changes, has poor flexibility, high data acquisition costs, and limited model generalization capabilities.

Method used

Various expression methods for generating large language models are introduced in the government affairs field, synonym replacement and large language model generation, combined with attention-enhancing classification networks, data sets of standard matters and common user problems are constructed, data augmentation and sample construction are carried out, and matching accuracy and efficiency are improved.

Benefits of technology

It significantly improves the accuracy and efficiency of matching matters and problems, enhances the generalization ability and robustness of the model for different expression methods, and improves the adaptability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124591A_ABST
    Figure CN120124591A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for aligning standard items and popular questions of users, and relates to the technical field of artificial intelligence. Comprising the following steps: step 1, constructing a data set of standard items and popular questions of users, and storing each standard item and corresponding questions possibly proposed by a plurality of users through the data set; 2, expanding data; 3, constructing samples: constructing a positive sample and a negative sample according to the data of the data set, 4, creating an attention-enhanced question classification network, training the question classification network by using the positive sample and the negative sample, classifying whether the question of the user and the input standard item describe the same item or not, and if yes, judging whether the question of the user is the same item or not. And judging whether the problem of the user is a certain item in the standard item list or not, and realizing alignment of the standard item and the popular problem of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method and device for aligning standard matters with common user questions, relating to the field of artificial intelligence technology. Background Art

[0002] In the field of government affairs services, when users handle various matters, they often raise various questions related to the matters. It is necessary to accurately find the standard government affairs matters related to the common questions raised by users in order to improve the efficiency of government affairs services.

[0003] In the existing methods, problem matching can be carried out based on the template matching scheme, but the maintenance cost of the template library is high and the templates need to be continuously updated and expanded. In addition, this method is slow to respond to new questions or changes in the expression methods raised by users, with poor flexibility and unable to effectively cope with complex and changeable user needs. It is also possible to train a classification model based on machine learning methods to classify and identify the standard matter categories corresponding to the questions. However, this method requires a large amount of labeled data for model training, and the data acquisition cost is high. In addition, the generalization ability of the model is limited, and when facing unseen questions or expression methods, the classification accuracy may be low, showing great limitations. Summary of the Invention

[0004] In view of the problems of the existing technology, the present invention provides a method and device for aligning standard matters with common user questions, introducing a government affairs domain dataset, a synonym replacement and a large language model to generate diverse expression methods, and an attention-enhanced classification network, significantly improving the accuracy and efficiency of matter-question matching.

[0005] The specific solution proposed by the present invention is as follows:

[0006] The present invention provides a method for aligning standard matters with common user questions, including:

[0007] Step 1: Construct a dataset of standard matters and common user questions, and save each standard matter and multiple questions that users may raise corresponding thereto through the dataset;

[0008] Step 2: Expand the data:

[0009] Perform synonym replacement of questions: perform word segmentation on the questions in the dataset, replace the synonyms in the word segmentation results through a synonym corpus, and combine them into new questions.

[0010] Use the large language model ChatGLM to perform generation processing on the questions in the dataset, simulate the expression methods of different background users, and expand the dataset.

[0011] Step 3: Construct samples: Based on the data in the dataset, construct positive and negative samples. A positive sample refers to a data pair where the user's common question is correctly matched with the standard matter, and a negative sample refers to a data pair where the user's common question does not match the standard matter.

[0012] Step 4: Create an attention-enhanced question classification network, and use positive and negative samples to train the question classification network to classify whether the user's question and the input standard matter describe the same thing, and determine whether the user's question is one of the matters in the standard matter list, so as to align the standard matter with the user's common question.

[0013] Furthermore, in step 1 of the method for aligning standard matters with users' common questions, constructing a dataset of standard matters and users' common questions includes:

[0014] Determine the standard matters that need to be included in the dataset.

[0015] According to the standard matters, collect and sort out the questions asked by users during the handling process.

[0016] Check whether the standard matters meet the standards, and add the standard matters that meet the standards and the users' questions to the dataset of standard matters and users' common questions.

[0017] Furthermore, in step 2 of the method for aligning standard matters with users' common questions, performing synonym replacement of questions includes:

[0018] For each sorted-out user's common question, perform word segmentation to split the question sentence into individual words or phrases.

[0019] According to the words or phrases in the word segmentation result, find synonyms using a thesaurus of near-synonyms.

[0020] Combine the near-synonyms of the words or phrases to obtain a new question.

[0021] Furthermore, in step 2 of the method for aligning standard matters with users' common questions, using the large language model ChatGLM to perform generation processing on the questions in the dataset includes:

[0022] Set prompt words, simulate handling users of different occupations and different age groups, and generate various possible expressions.

[0023] Generate questions according to the prompt words and various possible expressions.

[0024] Furthermore, in step 4 of the method for aligning standard matters with users' common questions, training an attention-enhanced question classification network includes:

[0025] Use BERT pre-trained on Chinese corpora as the backbone network. Take positive and negative samples as inputs, and use the pre-trained BERT model to encode the input samples. Then, capture the key parts of the encoded sample features through the attention module, output the weighted encoded vector, and finally input the weighted encoded vector into the classification layer for classification to determine whether the user's common question matches the standard matter.

[0026] The present invention also provides a device for aligning standard matters with user's common questions, including a dataset management module, an expansion module, a sample management module, and a classification alignment module.

[0027] The dataset management module constructs a dataset of standard matters and user's common questions, and saves each standard matter and multiple questions that users may ask corresponding to it through the dataset.

[0028] The expansion module expands the data:

[0029] Perform synonym replacement of questions: Segment the questions in the dataset, replace the synonyms in the segmented results through a synonym corpus, and combine them into new questions.

[0030] Use the large language model ChatGLM to generate the questions in the dataset, simulate the expression methods of users in different backgrounds, and expand the dataset.

[0031] The sample management module constructs samples: According to the data in the dataset, construct positive samples and negative samples. Positive samples refer to data pairs where the user's common question correctly matches the standard matter, and negative samples refer to data pairs where the user's common question does not match the standard matter.

[0032] The classification alignment module creates an attention-enhanced question classification network, trains the question classification network using positive and negative samples, classifies whether the user's question and the input standard matter describe the same thing, and determines whether the user's question is one of the matters in the standard matter list, so as to achieve the alignment of standard matters with user's common questions.

[0033] Furthermore, the dataset management module of the device for aligning standard matters with user's common questions constructs a dataset of standard matters and user's common questions, including:

[0034] Determine the standard matters that need to be included in the dataset.

[0035] According to the standard matters, collect and sort out the questions asked by users during the handling process.

[0036] Check whether the standard matters meet the standards, and add the standard matters that meet the standards and user questions to the dataset of standard matters and user's common questions.

[0037] Furthermore, the expansion module of the device for aligning standard items with popular questions of users replaces synonyms of questions, including:

[0038] For each common user question sorted out, word segmentation is performed to split the question sentence into individual words or phrases.

[0039] Use the synonym database to find synonyms based on the words or phrases in the segmentation results.

[0040] Combine words or phrases with synonyms to get new questions.

[0041] Furthermore, the expansion module of the device for aligning standard items with user popular questions uses the large language model ChatGLM to generate and process the questions in the data set, including:

[0042] Set prompt words, simulate users of different professions and ages, and generate various possible expressions.

[0043] Generate questions based on the prompt word and various possible expressions.

[0044] Furthermore, the classification alignment module of the apparatus for aligning standard items with user popular questions trains an attention-enhanced question classification network, including:

[0045] BERT, which has been pre-trained on Chinese corpus, is used as the backbone network. Positive and negative samples are taken as input. The pre-trained BERT model is used to encode the input samples. Then, the attention module is used to capture the key parts of the encoded sample features and output the weighted encoding vector. Finally, the weighted encoding vector is input into the classification layer for classification to determine whether the user's popular questions match the standard items.

[0046] The benefits of the present invention are:

[0047] By sorting out the questions that users may ask, we constructed a "Standard Matters - Common Questions of Users" data set, ensuring the professionalism and accuracy of the data and providing a solid basic data source.

[0048] By using synonym replacement and large language model generation methods, the expression of user questions is enriched. Specifically, through the application of word segmentation and synonym corpus, the expression of user questions is diversified and closer to the actual expression of users. This not only improves the diversity of the question library, but also enhances the generalization ability and robustness of the model for different expressions.

[0049] By generating possible question - asking ways of users of different occupations and ages through large - language models, the question bank covers a wider range of user expressions, further enhancing the comprehensiveness and representativeness of the dataset. This generation method simulates the expressions of users with different backgrounds through prompting words, effectively improving the adaptability and accuracy of the model.

[0050] By constructing positive and negative samples, the classification effect of the model is improved. The reasonable construction of positive and negative samples enables the model to better learn to distinguish matching and non - matching questions during the training process, significantly improving the accuracy of question matching.

[0051] The proposed attention - enhanced question classification network combines the self - attention mechanism with the pre - trained language model, which can more effectively capture the key parts of user questions, thereby improving the accuracy and reliability of question classification. The introduction of the attention mechanism not only enhances the interpretability of the model but also improves the classification performance. Brief Description of the Drawings

[0052] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0053] Figure 2 It is a schematic diagram of the construction process of the dataset.

[0054] Figure 3 It is a schematic diagram of the structure of the attention - enhanced question classification network.

[0055] Figure 4 It is a schematic diagram of the business interaction logic process. Detailed Implementation Modes

[0056] ChatGLM: The full name is Chat General Language Model, which is literally translated as Chat General Language Model. It is an open - source large - language model pre - trained in Chinese and English.

[0057] Token: It is the basic unit in natural language processing. It can be a word, character, or sub - word. The process of decomposing text into tokens is called tokenization, which is an important step in text pre - processing. Tokenization helps convert language into a numerical form that the model can process, laying a foundation for further analysis and processing.

[0058] GPT: A generative pre-trained language model developed by OpenAI. It adopts the Transformer architecture and learns the statistical characteristics and context information of language through large-scale unsupervised pre-training. During the pre-training stage, GPT uses a large amount of text data to train the model, enabling it to generate coherent and contextually relevant text. Subsequently, GPT can complete various natural language processing tasks, such as text generation, dialogue systems, translation, and question answering, by fine-tuning data for specific tasks.

[0059] BERT: Short for Bidirectional Encoder Representations from Transformers, is a pre-trained language model in the field of natural language processing. It adopts the Transformer model architecture and is characterized by its ability to capture bidirectional information in sentence contexts.

[0060] SVM: Short for Support Vector Machine, is a supervised learning model used for classification and regression analysis. Its basic idea is to find an optimal hyperplane to separate data points of different classes in a high-dimensional space.

[0061] Large language model: A large language model refers to a natural language processing model with a large number of parameters and powerful learning capabilities. Such models are usually based on deep learning techniques and adopt deep neural network structures, which contain hundreds of millions to hundreds of billions of parameters. This enables them to learn and understand complex patterns, grammar rules, and semantic relationships in natural language. The advantage of large language models is their ability to handle more complex, abstract, and context-rich natural language tasks, and they also greatly promote research and applications in the field of natural language processing.

[0062] Prompt engineering: In the fields of natural language processing and machine learning, prompt engineering refers to the input text provided by users or systems to the model to trigger the model to generate corresponding outputs. In dialogue systems or generative models, the prompt text is usually the question, request, or task description put forward by the user.

[0063] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.

[0064] Embodiment 1

[0065] The present invention provides a method for aligning standard matters with common user questions, including:

[0066] Step 1: Build a dataset of standard matters and users' common questions, and save each standard matter and multiple questions that users may ask corresponding to it through the dataset.

[0067] Among them, building the dataset of standard matters and users' common questions in Step 1 includes:

[0068] Determine the standard matters that need to be included in the dataset,

[0069] According to the standard matters, collect and sort out the questions asked by users during the handling process,

[0070] Check whether the standard matters meet the standards, and add the standard matters that meet the standards and users' questions to the dataset of standard matters and users' common questions.

[0071] For example, determine the standard matter "Permission for adding or reconstructing a plane intersection on a road" that needs to be included in the dataset, collect and sort out questions that users may ask during the handling process, such as "Reconstruct the intersection". Add the matters that meet the standards and users' questions to the "Standard Matter - User Common Question" dataset to provide basic data for subsequent data expansion and model training, and delete those that do not meet the standards from the user question set.

[0072] Step 2: Expand the data:

[0073] Perform synonym replacement of questions: Perform word segmentation on the questions in the dataset, replace the synonyms in the word segmentation results through a synonym corpus, and combine them into new questions.

[0074] Among them, performing synonym replacement of questions includes:

[0075] For each sorted-out user common question, perform word segmentation, split the question sentence into individual words or phrases, such as "Reconstruct the intersection" being split into "Reconstruct" and "Intersection",

[0076] Find synonyms according to the words or phrases in the word segmentation results using a synonym corpus, such as replacing "Reconstruct" with "Reconstruct and build", and "Intersection" with "Crossing",

[0077] Combine the synonyms of the words or phrases to obtain new questions. Such as generating new sentences like "Reconstruct and build the intersection", "Reconstruct the crossing", "Reconstruct and build the crossing", etc.

[0078] Use the large language model ChatGLM to perform generation processing on the questions in the dataset, simulate the expression ways of users in different backgrounds, and expand the dataset.

[0079] Using the large language model ChatGLM to perform generation processing on the questions in the dataset includes:

[0080] Set prompt words, simulate users of different professions and ages, and generate various possible expressions.

[0081] Generate questions based on the prompt word and various possible expressions.

[0082] For example, by using the prompt "You are a user with no experience in handling affairs, and you are accustomed to expressing yourself in a popular and concise manner.", generate 5-10 similar expressions for the following questions. The order should be reversed to enhance generalization and robustness. Repetition is not allowed, and elements in the original sentence cannot be missing. They should be given in separate paragraphs.

[0083] For example, questions such as “How to rebuild the intersection?”, “Where to rebuild the intersection?”, “What procedures are required for intersection reconstruction?” are generated.

[0084] Step 3: Construct samples: construct positive samples and negative samples based on the data in the dataset. Positive samples refer to data pairs that correctly match the user's common questions with standard items, and the label is 1 to form positive samples.

[0085] Negative samples refer to data pairs where the user's common questions do not match the standard items, and the label is 0. By constructing positive and negative samples.

[0086] For example, the standard items in the sample: "Permit to add or modify a level crossing on a highway" correspond to common user questions: "Intersection modification", "How to apply for permission to modify an intersection", and "What procedures do I need to modify an intersection?"

[0087] For example, in the negative sample, the standard item “issuing a one-time financial subsidy for self-employment” and the mismatched user common question “reconstructing the intersection”.

[0088] Step 4: Create an attention-enhanced question classification network, use positive and negative samples to train the question classification network, classify whether the user's question and the input standard item describe the same thing, determine whether the user's question is an item in the list of standard items, and align the standard items with the user's popular questions.

[0089] The attention-enhanced question classification network trained in step 4 includes:

[0090] BERT, which has been pre-trained on Chinese corpus, is used as the backbone network. Positive and negative samples are taken as input. The pre-trained BERT model is used to encode the input samples. Then, the attention module is used to capture the key parts of the encoded sample features and output the weighted encoding vector. Finally, the weighted encoding vector is input into the classification layer for classification to determine whether the user's popular questions match the standard items.

[0091] Example 2

[0092] The present invention also provides a device for aligning standard matters with common user questions, including a dataset management module, an expansion module, a sample management module, and a classification and alignment module.

[0093] The dataset management module constructs a dataset of standard matters and common user questions, and saves each standard matter and multiple questions that users may ask corresponding to it through the dataset.

[0094] The expansion module expands the data:

[0095] Perform synonym replacement of questions: perform word segmentation on the questions in the dataset, replace the synonyms in the word segmentation results through a synonym corpus, and combine them into new questions.

[0096] Use the large language model ChatGLM to perform generation processing on the questions in the dataset, simulate the expression ways of users with different backgrounds, and expand the dataset.

[0097] The sample management module constructs samples: according to the data in the dataset, construct positive samples and negative samples. Positive samples refer to data pairs in which common user questions are correctly matched with standard matters, and negative samples refer to data pairs in which common user questions are not matched with standard matters.

[0098] The classification and alignment module creates an attention-enhanced question classification network, uses positive samples and negative samples to train the question classification network, classifies whether the user's question and the input standard matter describe the same thing, and determines whether the user's question is one of the matters in the standard matter list, so as to achieve the alignment of standard matters and common user questions.

[0099] For the content such as information interaction and execution process among the above-mentioned modules in the device, since it is based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0100] Similarly, the advantages of the device of the present invention are:

[0101] By sorting out the questions that users may ask, a "standard matter - common user question" dataset is constructed, ensuring the professionalism and accuracy of the data, and providing a solid basic data source.

[0102] Using the methods of synonym replacement and large language model generation, the expression ways of user questions are enriched. Specifically, through the application of word segmentation and synonym corpus, the expressions of user questions are diversified and closer to the actual expressions of users. This not only improves the diversity of the question library, but also enhances the generalization ability and robustness of the model to different expression ways.

[0103] By generating possible ways of asking questions for users of different occupations and ages through large language models, the question bank covers a wider range of user expressions, further enhancing the comprehensiveness and representativeness of the dataset. This generation method simulates the expressions of users with different backgrounds through prompting words, effectively improving the adaptability and accuracy of the model.

[0104] By constructing positive and negative samples, the classification effect of the model is improved. The reasonable construction of positive and negative samples enables the model to better learn to distinguish matching and non-matching questions during the training process, significantly improving the accuracy of question matching.

[0105] The proposed attention-enhanced question classification network combines the self-attention mechanism with the pre-trained language model, which can more effectively capture the key parts of user questions, thereby improving the accuracy and reliability of question classification. The introduction of the attention mechanism not only enhances the interpretability of the model but also improves the classification performance.

[0106] It should be noted that not all steps and modules in the above processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities separately, or some components in multiple independent devices may be jointly implemented.

[0107] The device of the present invention can be used for the interaction logic between users and the system. The user inputs a question from the front-end interaction interface and enters the back-end algorithm processing stage. First, the input question is matched with the matters in the matter library, and through the screening of the matching threshold, a list of standard matters related to the user's question is obtained. Subsequently, the deployed attention-enhanced question classification network is used to obtain the matters related to the user's question in the obtained list of matters. Then, the large language model generates an answer based on the relevant matters obtained by the network and the user's question. Finally, the system outputs the answer through the front-end interaction interface, and the user views the answer and makes an evaluation. If the user is satisfied with the answer, they click like; if not, they click dislike.

[0108] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for aligning standard issues with user common questions, characterized by include: Step 1: Construct a dataset of standard items and common user questions, and save each standard item and the corresponding multiple questions that users may ask through the dataset; Step 2: Augment the data: Replace synonyms in questions: Perform word segmentation on questions in the dataset, replace synonyms in the word segmentation results with the synonym corpus, and combine them into new questions. Use the large language model ChatGLM to generate and process questions in the dataset, simulate the expression methods of users with different backgrounds, and expand the dataset; Step 3: Construct samples: According to the data in the data set, construct positive samples and negative samples. Positive samples refer to the data pairs where the user's common questions correctly match the standard items, and negative samples refer to the data pairs where the user's common questions do not match the standard items. Step 4: Create an attention-enhanced question classification network, use positive and negative samples to train the question classification network, classify whether the user's question and the input standard item describe the same thing, determine whether the user's question is an item in the list of standard items, and align the standard items with the user's popular questions.

2. A method for aligning standard items with user common questions according to claim 1, characterized in that In step 1, a dataset of standard issues and common user questions is constructed, including: Identify standard items that need to be included in the dataset, According to standard matters, collect and sort out the questions asked by users during the handling process, Check whether the standard items meet the standards, and add the standard items and user questions that meet the standards to the data set of standard items and user popular questions.

3. A method for aligning standard items with user popular questions according to claim 1, characterized in that In step 2, synonyms of the question are replaced, including: For each common user question sorted out, word segmentation is performed to split the question sentence into individual words or phrases. Use the synonym database to find synonyms based on the words or phrases in the segmentation results. Combine words or phrases with synonyms to get new questions.

4. A method for aligning standard items with user common questions according to claim 1, characterized in that In step 2, the large language model ChatGLM is used to generate and process the questions in the dataset, including: Set prompt words, simulate users of different professions and ages, and generate various possible expressions. Generate questions based on the prompt word and various possible expressions.

5. A method for aligning standard items with user popular questions according to claim 1, characterized in that In step 4, the attention-enhanced question classification network is trained, including: BERT, which has been pre-trained on Chinese corpus, is used as the backbone network. Positive and negative samples are taken as input. The pre-trained BERT model is used to encode the input samples. Then, the attention module is used to capture the key parts of the encoded sample features and output the weighted encoding vector. Finally, the weighted encoding vector is input into the classification layer for classification to determine whether the user's popular questions match the standard items.

6. A device for aligning standard items with user common questions, characterized by It includes data set management module, expansion module, sample management module and classification alignment module. The data set management module constructs a data set of standard items and common user questions, and saves each standard item and the corresponding multiple questions that users may ask through the data set; Extension module to expand data: Replace synonyms in questions: Perform word segmentation on questions in the dataset, replace synonyms in the word segmentation results with the synonym corpus, and combine them into new questions. Use the large language model ChatGLM to generate and process questions in the dataset, simulate the expression methods of users with different backgrounds, and expand the dataset; Sample management module constructs samples: constructs positive samples and negative samples based on the data in the data set. Positive samples refer to data pairs where the user's common questions correctly match the standard items, while negative samples refer to data pairs where the user's common questions do not match the standard items. The classification alignment module creates an attention-enhanced question classification network, uses positive and negative samples to train the question classification network, classifies whether the user's question and the input standard item describe the same thing, determines whether the user's question is an item in the standard item list, and achieves alignment between standard items and users' popular questions.

7. The device for aligning standard items with user popular questions according to claim 6, characterized in that the data The collection management module builds data sets of standard items and common user questions, including: Identify standard items that need to be included in the dataset, According to standard matters, collect and sort out the questions asked by users during the handling process, Check whether the standard items meet the standards, and add the standard items and user questions that meet the standards to the data set of standard items and user popular questions.

8. The device for aligning standard items with user popular questions according to claim 6, characterized in that The expansion module replaces synonyms of questions, including: For each common user question sorted out, word segmentation is performed to split the question sentence into individual words or phrases. Use the synonym database to find synonyms based on the words or phrases in the segmentation results. Combine words or phrases with synonyms to get new questions.

9. The device for aligning standard items with user popular questions according to claim 6, characterized in that The expansion module uses the large language model ChatGLM to generate and process questions in the dataset, including: Set prompt words, simulate users of different professions and ages, and generate various possible expressions. Generate questions based on the prompt word and various possible expressions.

10. The device for aligning standard items with user popular questions according to claim 6, characterized in that The classification alignment module trains an attention-enhanced question classification network, including: BERT, which has been pre-trained on Chinese corpus, is used as the backbone network. Positive and negative samples are taken as input. The pre-trained BERT model is used to encode the input samples. Then, the attention module is used to capture the key parts of the encoded sample features and output the weighted encoding vector. Finally, the weighted encoding vector is input into the classification layer for classification to determine whether the user's popular questions match the standard items.