Test question recommendation filtering method based on text vector
Through the text vector-based test questions recommendation filtering method, the problems of low efficiency and high labor cost in the existing technology are solved, and the test questions with high similarity are quickly found, the recommendation and filtering efficiency is improved, the labor workload is reduced, and it is applicable to different disciplines and fields.
Patent Information
- Application Number
- CN202510098683.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
AI Technical Summary
The existing test questions recommendation methods have problems such as high labor costs, high migration costs, unfriendly handling complex attributes, and inability to quickly adapt to new majors and positions. It is easy to concentrate on recommending similar test questions based on collaborative filtering, and due to data sparseness, it is impossible to quickly adapt to new majors and positions.
The recommended filtering method for test questions based on text vectors is used. By generating the test requirements text based on user needs in advance, inputting it into the Siamese Bert model and converting it into a 768-dimensional vector, calculating the cosine similarity with the preset test point text vector, generating the test question query statement and obtaining the corresponding test questions, filtering and similarity calculation, until any test question similarity of the current complete test question ≤ the similarity threshold, and generating the test paper.
It has achieved rapid finding test questions with high similarity to the target vector, greatly improving the efficiency of recommendation and filtering, reducing the workload of manual screening and matching test questions, and can automatically find test questions that meet the requirements, providing support for paper grouping or recommended learning resources. It is suitable for test questions in different subjects and fields, and is highly versatile.
Smart Images

Figure CN120067463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of question recommendation, and specifically, to a question recommendation filtering method based on text vectors. Background Art
[0002] With the development of generative AI, the explosive growth of information has made it easier for users to obtain information. However, at the same time, due to problems such as AI hallucinations, the cost for users to distinguish information has increased. In the field of talent assessment, there are currently a large number of professional question banks and questions generated by AI in recent years. How to more accurately recommend questions to the test paper compilation teachers corresponding to the relevant majors and subjects is particularly important.
[0003] Currently, the mainstream recommendation methods mainly include rule-based recommendation technology, content-based recommendation technology, and collaborative filtering technology. Among them, rule-based recommendation technology overly relies on language experts in the professional field to define grammar rules, which requires a large amount of time to extract rules, has too high labor costs, and at the same time has huge migration costs. Among them, content-based recommendation is not friendly to the processing of complex attributes and cannot generate good recommendations for new majors and subjects. Among them, recommendation methods such as collaborative filtering and cognitive diagnosis are prone to concentrating on recommending questions with similar test points and also have problems brought about by data sparsity, and cannot quickly adapt to new majors and positions brought about by the current explosive development.
[0004] Therefore, there is an urgent need for a question recommendation filtering method based on text vectors.
[0005] Regarding the problems in the related art, no effective solution has been proposed yet. Summary of the Invention
[0006] Regarding the problems in the related art, the present invention proposes a question recommendation filtering method based on text vectors to overcome the above-mentioned technical problems existing in the existing related technologies.
[0007] The technical solution of the present invention is realized as follows:
[0008] A question recommendation filtering method based on text vectors includes the following steps:
[0009] Pre-join the examination requirements in the form of natural language text according to user needs in a preset format as the basic input content to generate an examination requirement text;
[0010] Input the examination requirement text into the vector model Siamese Bert to convert it into a 768-dimensional vector;
[0011] Calculate the cosine similarity between the 768-dimensional vector and the test point text vector stored in the preset vector database, including: screening all test points where the cosine similarity between the 768-dimensional vector and the test point text vector > the preset similarity threshold;
[0012] Generate a test question query statement and obtain the corresponding test questions according to the generated test question query statement;
[0013] Perform test question filtering, merge the queried test questions into complete test questions, calculate the similarity of the complete test questions, screen and replace the test questions with a similarity > the preset similarity threshold until the similarity of any test question in the current complete test question ≤ the similarity threshold, and generate a test paper.
[0014] Furthermore, calculating the cosine similarity between the 768-dimensional vector and the test point text vector stored in the preset vector database is expressed as:
[0015] Among them, is the 768-dimensional vector, is the test point text vector, and is the inner product of the vectors and ; and are the norms of the vectors and respectively.
[0016] Furthermore, the generation of the test question query statement includes the following steps:
[0017] Pre-extract the test question features in the test requirement text according to the preset formatting template and use the test question features combined with the test points as the input text;
[0018] Input the input text into the preset table generation model for processing to obtain table data including test points, question types, and the number of questions;
[0019] Input the table data into the preset SQL generation model for processing to obtain the test question query statement;
[0020] Input the test question query statement into the ELASTICSEARCH search engine to query the corresponding test questions.
[0021] Furthermore, the test question features include: the total score, the number of questions, and the proportion of subjective and objective questions in the test requirement text.
[0022] Furthermore, the calculation of the similarity of the complete test questions includes the following steps:
[0023] Pre-use the pre-trained word vector model Word2Vec to convert each word in the text of each test question into a vector;
[0024] Average all the word vectors in any test question to obtain the vector representation of the test question, which is expressed as:
[0025]
[0026] Among them, a test question contains word vectors represented as:
[0027] Calculate the similarity between pairwise test question vectors, which is expressed as:
[0028] Among them, the vector of test question A is The vector of test question B is
[0029] Furthermore, it also includes the following steps:
[0030] If the similarity of the current test question > the similarity threshold, it means that there are similar knowledge points in the current two test questions, then replace any one of the current test questions until the similarity of any test question in the current complete test question ≤ the similarity threshold.
[0031] Advantages of the present invention:
[0032] The present invention pre-stitches the examination requirements in the form of natural language text according to user requirements as the basic input content to generate the examination requirement text; inputs the examination requirement text into the vector model Siamese Bert to convert it into a 768-dimensional vector; calculates the cosine similarity between the 768-dimensional vector and the examination point text vector stored in the preset vector database to generate a test question query statement, and obtains the corresponding test questions according to the generated test question query statement; performs test question filtering, merges the queried test questions into complete test questions, and calculates the similarity of the complete test questions to generate a test paper, realizing the conversion of the test requirement text into a 768-dimensional vector, which is relatively fast for subsequent cosine similarity calculation. For a large-scale test question bank, through vector operations, test questions with relatively high similarity to the target vector can be quickly found, greatly improving the efficiency of recommendation and filtering. At the same time, it can automatically calculate the similarity of complete test questions, reducing the workload of manually screening and matching test questions. By setting parameters such as the similarity threshold, test questions that meet the requirements can be automatically found, providing support for test paper compilation or learning resource recommendation. In addition, it can be applied to test questions in different disciplines and fields. Whether it is disciplines such as mathematics, Chinese, physics, chemistry, or different types of examinations, the text vector method can screen and recommend test questions according to the semantics of the text, with strong versatility, aiming to assist the test paper compilation teacher to complete the test paper compilation work efficiently and scientifically, while saving labor costs, improving the quality of the test paper, and reducing subjective biases through relevant technical means. The entire process covers multiple links from generating examination requirements according to user needs to finally completing the test paper compilation. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0034] Figure 1 It is a schematic flowchart of a test question recommendation and filtering method based on text vectors according to an embodiment of the present invention. Specific embodiments
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0036] According to an embodiment of the present invention, a test question recommendation and filtering method based on text vectors is provided.
[0037] As Figure 1 shown, the test question recommendation and filtering method based on text vectors according to an embodiment of the present invention includes the following steps:
[0038] Step S1, pre-construct the exam requirements in the form of natural language text according to user requirements in a preset format as the basic input content to generate an exam requirement text;
[0039] Step S2, input the exam requirement text into the vector model Siamese Bert to convert it into a 768-dimensional vector;
[0040] In this technical solution, the exam requirement text is input into the vector model Siamese Bert to convert it into a 768-dimensional vector, realizing the conversion of the exam requirement text into a vector representation form that is convenient for mathematical calculation and comparison.
[0041] Step S3, calculate the cosine similarity between the 768-dimensional vector and the test point text vectors stored in the preset vector database, including: screening all test points with the cosine similarity between the 768-dimensional vector and the test point text vectors > the preset similarity threshold;
[0042] Among them, calculating the cosine similarity between the 768-dimensional vector and the test point text vectors stored in the preset vector database is expressed as:
[0043] Among them, is the 768-dimensional vector, is the test point text vector, and is a vector and is the inner product of and are respectively the norms of vectors and ;
[0044] In this technical solution, the value of the cosine similarity is between [-1, 1]. The closer the value is to 1, the more similar the 768-dimensional vector is to the test point text vector.
[0045] Specifically, in application, the similarity threshold can be set to 0.75, and all test points with the cosine similarity between the 768-dimensional vector and the test point text vector > 0.75 are screened as the test points that meet the requirements, providing the basis for subsequent question generation.
[0046] Step S4, generate a question query statement and obtain the corresponding questions according to the generated question query statement, which includes the following steps:
[0047] Step S401, extract the question features in the test requirements text in advance according to a preset formatting template, and use the question features combined with the test points as the input text;
[0048] Among them, the question features include: the total score, the number of questions, and the proportion of subjective and objective questions in the test requirements text;
[0049] Step S402, input the input text into a preset table generation model for processing to obtain table data including test points, question types, and the number of questions;
[0050] Step S403, input the table data into a preset SQL generation model for processing to obtain a question query statement;
[0051] Step S404, input the question query statement into the ELASTICSEARCH search engine to query the corresponding questions;
[0052] Step S5, perform question filtering, merge the queried questions into complete questions, calculate the similarity of the complete questions, and screen and replace the questions with similarity > the preset similarity threshold until the similarity of any question in the current complete question ≤ the similarity threshold, and generate a test paper, which includes the following steps:
[0053] Previously, use the pre-trained word vector model Word2Vec to convert each word in the text of each question into a vector.
[0054] Average all the word vectors in any question to obtain the vector representation of the question, which is expressed as:
[0055]
[0056] Among them, a test question contains word vector representations as follows:
[0057] Calculate the similarity between pairwise test question vectors, which is expressed as:
[0058] Among them, the vector of test question A is The vector of test question B is
[0059] And when the similarity > the similarity threshold, it indicates that there are similar knowledge points in the current two test questions. Then, replace any one of the current test questions until the similarity of any test question in the current complete test question ≤ the similarity threshold.
[0060] Specifically, in this technical solution, the query test questions are combined into a complete test question. It combines the test questions and answers into a complete text, and considers the test questions and their answers as a whole, which is convenient for subsequent comparative analysis of similarity knowledge points.
[0061] In addition, during the screening process, pairwise comparison can be adopted. For example: compare the second question with the first question, the third question with the first and second questions, and so on, to comprehensively compare the test questions in the test paper pairwise and find the test questions with similar knowledge points.
[0062] At the same time, perform repeated replacement and comparison. After finding the test question points with similar knowledge points, directly replace them and then compare them with other test questions one by one. Repeat this process until there are no test questions with similar knowledge points in all the test questions in the entire test paper. Gradually purify the test paper content through continuous cyclic operations to avoid problems such as concentrated test points, repeated knowledge points, and even answer hints for the upper and lower questions.
[0063] Step S5, perform test question filtering. Combine the query test questions into a complete test question, calculate the similarity of the complete test question, and screen and replace the test questions with similarity > the preset similarity threshold until the similarity of any test question in the current complete test question ≤ the similarity threshold. Among them, it includes the following steps:
[0064] In summary, by means of the above technical solution of the present invention, the following effects can be achieved:
[0065] The present invention pre - splices the examination requirements into a natural - language text format according to preset formats based on user requirements as basic input content to generate an examination - requirement text; inputs the examination - requirement text into the vector model Siamese Bert to convert it into a 768 - dimensional vector; calculates the cosine similarity between the 768 - dimensional vector and the text vectors of examination points stored in a preset vector database to generate a test - question query statement, and obtains corresponding test questions according to the generated test - question query statement; filters the test questions, combines the queried test questions into complete test questions, and calculates the similarity of the complete test questions to generate a test paper, realizing the conversion of the examination - requirement text into a 768 - dimensional vector. For subsequent cosine - similarity calculation, the calculation speed is relatively fast. For a large - scale test - question bank, through vector operations, test questions with relatively high similarity to the target vector can be quickly found, greatly improving the efficiency of recommendation and filtering. At the same time, it can automatically calculate the similarity of complete test questions, reducing the workload of manual screening and matching of test questions. By setting parameters such as a similarity threshold, test questions that meet the requirements can be automatically found, providing support for test - paper compilation or recommendation of learning resources. In addition, it can be applied to test questions in different disciplines and fields. The text - vector method can screen and recommend test questions according to the semantics of the text, with strong versatility. It aims to assist the test - paper - compiling teacher to complete the test - paper compilation work efficiently and scientifically, while saving labor costs, improving the quality of the test paper, and reducing subjective biases through relevant technical means. The whole process covers multiple links from generating examination requirements according to user needs to finally completing the test - paper compilation.
[0066] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. After considering the disclosure in the specification and the embodiments, those skilled in the art will easily think of other implementation schemes of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include well - known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0067] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A test question recommendation and filtering method based on text vectors, characterized in that: The following steps are involved: The examination requirements are pre-assembled into a natural language text format according to user needs in a preset format, and used as basic input content to generate the examination requirements text; Convert the Siamese Bert test requirement text input vector model to a 768-dimensional vector; Calculating the cosine similarity between the 768-dimensional vector and the test point text vector stored in the preset vector database, which includes: screening all test points whose cosine similarity between the 768-dimensional vector and the test point text vector is greater than a preset similarity threshold; Generate a test question query statement, and obtain corresponding test questions according to the generated test question query statement; Filter the test questions, merge the query test questions into complete test questions, calculate the similarity of the complete test questions, filter and replace the test questions with similarity greater than the preset similarity threshold, until the similarity of any test question in the current complete test question is ≤ the similarity threshold, and generate the test paper.
2. The test question recommendation and filtering method based on text vector according to claim 1, characterized in that: Calculate the cosine similarity between the 768-dimensional vector and the test point text vector stored in the preset vector database, expressed as: in, is a 768-dimensional vector, is the test point text vector, and is a vector and The inner product of and They are vectors and Model.
3. The test question recommendation and filtering method based on text vector according to claim 1, characterized in that: The step of generating a test question query statement comprises the following steps: Extract the test question features in the test requirement text in advance according to the preset formatting template, and combine the test question features with the test points as the input text; Input the input text into the preset table generation model for processing, and obtain table data containing test points, question types and question quantities; Input the table data into the preset SQL generation model for processing to obtain the test question query statement; Enter the test question query statement into the ELASTICSEARCH search engine to query the corresponding test questions.
4. The test question recommendation and filtering method based on text vector according to claim 3, characterized in that: The test question characteristics include: the total score of the test requirement text, the number of questions and the proportion of subjective and objective questions.
5. The test question recommendation and filtering method based on text vector according to claim 1, characterized in that: The method of calculating the similarity of the complete test questions comprises the following steps: Use the pre-trained word vector model Word2Vec to convert each word in each test text into a vector; All word vectors in any test question are averaged to obtain the vector representation of the test question, which is expressed as: Among them, a test question contains word vector representation as follows: Calculate the similarity between two test question vectors, expressed as: Among them, the vector of test question A is The vector of Question B is 6. The test question recommendation and filtering method based on text vector according to claim 5, characterized in that: The following steps are also included: If the similarity of the current test question is greater than the similarity threshold, it means that there are similar knowledge points between the two current test questions, then any of the current test questions will be replaced until the similarity of any test question in the current complete test question is ≤ the similarity threshold.