Examination question duplicate checking method
By setting up a test question database, dividing functional areas and conducting semantic analysis in the test question plagiarism method, the problem of inaccurate test questions in the existing technology is solved, and more scientific and reasonable similarity calculation and higher plagiarism accuracy are achieved.
Patent Information
- Application Number
- CN202510171217.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing test questions plagiarism check method ignores the differences in the structural characteristics and components of the test questions, resulting in the inaccurate results of the test questions and the quality of the test papers.
By setting up a test question database and storing historical test questions in partitions, identifying the information fields of the test questions to be tested, dividing the test questions into multiple functional areas (material area, question stem area, option area, answer area), and using natural language processing technology and semantic analysis, the logical similarity value and text repetition value of each functional area are calculated, and the overall similarity value of the test questions is finally calculated to determine the overlapping test questions.
By considering the structural characteristics and differences in the components of the test questions, the accuracy and flexibility of the test questions are improved, making the similarity calculation more scientific and reasonable.
Smart Images

Figure CN120068845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of test question information processing, and particularly relates to a method for checking the duplication of examination questions. Background Art
[0002] With the continuous expansion of the scale of various examinations, the work of formulating test questions has become increasingly important.
[0003] The prior art CN114238737A discloses a method for determining the duplication of similar test questions, which includes calculating the hash of existing test questions to obtain the hash codes of existing test questions, then calculating the hash of new test questions and obtaining the hash codes of new test questions, and then comparing the hash codes of existing test questions with the hash codes of new test questions. If they are repeated, the new test questions will be deleted. If not, the new test questions will be marked with similarity to the existing test questions for teachers to observe. However, when dealing with the similarity detection of test questions, traditional test question duplication checking methods usually take the entire content of the test question as a whole and compare its similarity with other test questions. However, this method ignores the structural characteristics of the test questions themselves and the differences in the constituent elements, resulting in inaccurate results in checking the assessment content of the test questions, and thus reducing the quality of the test papers. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the background art, and a method for checking the duplication of examination questions is proposed.
[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions: A method for checking the duplication of examination questions, which specifically includes the following steps: Step 1: Set up a test question database, and at the same time collect historical test question information, and store the historical test question information in the test question database in partitions according to the corresponding information fields. Step 2: Mark the test question to be detected as the target detection test question, identify the information field of the target detection test question, and perform information retrieval in the corresponding storage area in the test question database to obtain similar test questions. Then set up a functional area, and divide the target detection test question and the similar test questions according to the functional area, where the functional area includes a material area, a stem area, an option area, and an answer area. Step 3: Set up a single-phase area and a control area respectively. At the same time, extract the logical connection words in the single-phase area, and then divide the single-phase area into several single-word groups according to the stop words in the stop word library. Then, according to the logical operation model corresponding to the logical connection words, perform vector operations on the single-word groups to obtain the feature vector of the single-phase area. Based on the feature vector of the single-phase area and the feature vector of the control area, calculate the logical similarity value of the single-phase area. Step 4: Set an order value for the functional area. At the same time, identify the single-word phrases in the single-phase area in the control area to obtain the total number of occurrences of the single-word phrases in the control area. Then, based on the position order of the single-word phrases and the positions of the corresponding text information in the control area, determine the number of overlaps in the position order. Calculate the text repetition value of the single-phase area based on the order value, the total number of occurrences of the single-word phrases in the control area, and the number of overlaps in the position order; Step 5: Calculate the area similarity value of the single-phase area based on the logical similarity value and the text repetition value. Then, obtain the area similarity value of each functional area of the target detection question. Next, integrate the area similarity values of all functional areas to determine the overall similarity value between the target detection question and the control question; Step 6: Based on the overall similarity value, determine the overlapping questions among the similar questions and transmit the overlapping questions to the display terminal.
[0006] As a further solution of the present invention, the specific method for partition storage in the question database includes: Obtain the application location of the collected historical questions, and determine the question nature of the historical questions according to the application location. Among them, the question nature includes professional tests and comprehensive tests; Obtain the question nature of the historical questions. If the question nature is a professional test, directly obtain the subject name of the professional test and store the historical questions in partitions according to the subject name; If the question nature is a comprehensive test, obtain all the questions in the comprehensive test. At the same time, mark each individual question as a single-phase exam question Pi, where i represents the question number corresponding to different questions; Then obtain the single-phase exam question Pi, extract the key vocabulary of the single-phase exam question Pi using the TF-IDF method, perform semantic analysis on the key vocabulary using the Word2Vec model, and according to the semantic analysis results, determine the information field of the single-phase exam question Pi. Then, store the single-phase exam question Pi in partitions according to the information field.
[0007] As a further solution of the present invention, the question application location is the corresponding exam subject. Then, use natural language processing algorithms to perform character recognition on the exam subject: if the exam subject is a specific subject name, mark the corresponding question nature as a professional test; if the exam subject is not a specific subject name, mark the corresponding question nature as a comprehensive test; Furthermore, the exam scope includes professional field tests and comprehensive knowledge tests. Among them, in the professional field test, all the questions appearing in a test paper are knowledge information related to the corresponding field. In the comprehensive knowledge test, a test paper will include knowledge information in multiple fields.
[0008] As a further solution of the present invention, the method for setting functional areas includes: Mark the test questions to be detected as target detection test questions, then use the TF-IDF method to extract the key vocabulary in the target detection test questions, and at the same time use the Word2Vec model to perform semantic analysis on the key vocabulary to obtain the information field of the target detection test questions; Then use the key vocabulary of the target detection test questions as retrieval keywords, input them into the storage space of the corresponding information field for information retrieval, and then mark the obtained retrieval results as similar test questions j, where j = 1, 2,..., J, indicating that there are a total of J similar test questions; Then use the language model in natural language processing technology to perform semantic analysis on the target detection test questions and similar test questions respectively, and set functional areas for the target detection test questions and similar test questions according to the semantic analysis results. Among them, the functional areas include a material area, a question stem area, an option area, and an answer area.
[0009] As a further solution of the present invention, in the process of extracting logical connection words in a single-phase area, the single-phase area refers to any one of the functional areas in the target detection test questions, and the comparison area refers to the area in the similar test questions that has the same function as the single-phase area; When extracting logical connection words in the single-phase area, logical connection words refer to words that play a connecting role in a sentence and represent various logical relationships, including parallel logical relationship words, progressive logical relationship words, transitional logical relationship words, conditional logical relationship words, and causal logical relationship words. At the same time, different logical operation models are established for different logical connection words, where one type of logical connection word corresponds to one logical operation model.
[0010] As a further solution of the present invention, the calculation method of the logical similarity value includes: Set a stop word library, where the stop word library contains various stop words. Specifically, stop words refer to words that have no actual lexical meaning; Based on the stop words stored in the stop word library, identify the stop words in the single-phase area. When a stop word is detected, use the corresponding stop word as a pause symbol to divide the single-phase area into several single-body phrases; Among them, when there are multiple consecutive stop words, merge the multiple consecutive stop words into an overall, and set a pause symbol according to the position of this overall; Use the bag-of-words model algorithm to convert the single-body phrases in the single-phase area into vectors respectively, then obtain the logical connection words in the single-phase area, and identify the logical operation model of the logical connection words; Take the vectors of the single-body phrases in the single-phase area as input values and input them into the corresponding logical operation model. The logical operation model performs active operation processing, and marks the operation result as the feature vector of the single-phase area. Obtain the control region Dj, and process the control region Dj according to the above method to obtain the feature vector Tj of the control region; After that, use the cosine similarity algorithm to calculate the cosine similarity between the feature vector Tj of the control region and the feature vector of the single-phase region respectively, and mark the calculation result as the logical similarity value LXm, where m represents the number of the functional region, and m = 1, 2,..., k, indicating that there are k functional regions in the target detection test question.
[0011] As a further solution of the present invention, the calculation method of the text repetition value includes: Set an order value for the functional region, where the order value refers to the influence of the order of keyword phrases in the functional region on the calculation of the coincidence rate of the functional region; Obtain all single-word phrases in the single-phase region again, identify the single-word phrases in the control region, and if there is text information in the control region that is the same as the single-word phrase, mark the corresponding text information as 1; After the single-word phrases are identified in the control region, first obtain the total number SL of the text information with the value of 1, and then extract the text information corresponding to 1 separately according to the position order in the control region, and at the same time identify the coincidence number WL of the position order between the single-word phrase and the text information 1; Use the formula to obtain the text repetition value NXm, where ZD is the total number of text information in the control region, is the order value of the single-phase region, and are weight factors respectively.
[0012] As a further solution of the present invention, the calculation method of the overall similarity value includes: Use the formula to calculate the regional similarity value Gm of the single-phase region, where, and are both proportionality coefficients; Take other functional regions as single-phase regions respectively, calculate the regional similarity value Gm of each functional region, and use the formula to calculate the overall similarity value XT between the target detection test question and the control test question, where, represents the similarity weight of the functional region m.
[0013] As a further solution of the present invention, the method for determining the overlapping test questions among the similar test questions includes: Compare the overall similarity value XT of the target detection test question with the similarity threshold Xy. If XT < Xy, mark the corresponding similar test question as a non-overlapping test question. On the contrary, if XT ≥ Xy, mark the corresponding similar test question as an overlapping test question; After all similar test questions are marked, identify the overlapping test questions among the similar test questions. If there are overlapping test questions among the similar test questions, obtain the corresponding overlapping test questions and transmit them to the display terminal. If there are no overlapping test questions among the similar test questions, generate test question independent information and transmit it to the display terminal.
[0014] Compared with the existing technology, the advantages of the present invention are as follows: The present invention sets multiple functional areas for the examination questions according to the structural characteristics of the examination questions, including a material area, a question stem area, an option area, and an answer area. Then, calculate the logical similarity value and content similarity value of each functional area to determine the area similarity value of a single functional area. After that, set the similarity weight according to the characteristics of each functional area, combine the area similarity value with the similarity weight, calculate the overall similarity value of the examination questions, and finally determine the overlapping test questions based on the overall similarity value, thereby enhancing the flexibility of duplicate checking and making the similarity calculation more scientific and reasonable. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic structural diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0017] Refer to Figure 1 , a method for duplicate checking of examination questions, which specifically includes the following steps: Step 1: Set up a test question database, and then collect historical test question information and store the historical test questions in the test question database; In this embodiment, when the test question database stores the collected historical test question information, the historical test questions are stored according to the corresponding disciplines. Among them, storing according to disciplines is beneficial to reducing the data retrieval volume; Further, the method for identifying the disciplines of historical test questions includes: Obtain the application location of the collected historical test questions, and determine the nature of the historical test questions according to the application location. Among them, the nature of the test questions includes professional tests and comprehensive tests; Among them, when collecting historical test questions, usually according to the overall test content, all the test questions in a test paper are collectively integrated and collected. At this time, the application position of the test question is the corresponding test subject. Then, using natural language processing algorithms, character recognition is performed on the test subject: if the test subject is a specific subject name, the corresponding test question property is marked as a professional test; if the test subject is not a specific subject name, the corresponding test question property is marked as a comprehensive test; It should be further noted that since the test fields applied in the present invention include industries such as education assessment tests, personnel recruitment tests, professional qualification tests, and talent assessment tests, the above test scope includes professional field tests and comprehensive knowledge tests. Among them, in the professional field test, all the test questions appearing in a test paper are knowledge information related to the corresponding field, while in the comprehensive knowledge test, a test paper will include knowledge information in multiple fields. For example: in a comprehensive knowledge test paper, knowledge information in various fields such as computer knowledge, legal knowledge, political knowledge, literary knowledge, historical knowledge, and common sense of life will exist simultaneously; Obtain the test question property of the historical test question. If the test question property is a professional test, directly obtain the subject name of the professional test and store the historical test questions in zones according to the subject name; If the test question property is a comprehensive test, obtain all the test questions in the comprehensive test, and at the same time mark each individual test question as a single-phase test question Pi, where i represents the test question number corresponding to different test questions; Then obtain the single-phase test question Pi, extract the keyword vocabulary of the single-phase test question Pi using the TF-IDF method, perform semantic analysis on the keyword vocabulary using the Word2Vec model, and determine the information field of the single-phase test question Pi according to the semantic analysis result. Then store the single-phase test question Pi in zones according to the information field; Among them, the specific processing processes of the TF-IDF method and the Word2Vec model are both existing technologies and will not be elaborated here; Step 2: Identify the test question to be detected, mark the test question to be detected as the target detection test question, extract the keyword vocabulary in the target detection test question using the TF-IDF method, and at the same time perform semantic analysis on the keyword vocabulary using the Word2Vec model to obtain the information field of the target detection test question; Then use the keyword vocabulary of the target detection test question as the retrieval keyword, input it into the storage space of the corresponding information field for information retrieval, and mark the obtained retrieval result as a similar test question j, where j = 1, 2,..., J, indicating that there are a total of J similar test questions; After that, semantic analysis is performed on the target detection questions and similar questions respectively, and functional areas are set for the target detection questions and similar questions according to the semantic analysis results. The functional areas include a material area, a question stem area, an option area, and an answer area. Further, the functional areas of the target detection questions and similar questions are divided by a language model in natural language processing technology, and the specific processing method of the language model is a prior art and will not be elaborated here; Step 3: Arbitrarily select a functional area in the target detection question and mark it as a single-phase area. At the same time, arbitrarily select a similar question and mark it as a control question, and then mark the corresponding functional area in the control question as a control area Dj; First, extract the logical connection words in the single-phase area. The logical connection words refer to the words that play a connecting role in a sentence and represent various logical relationships, including parallel logical relationship words, progressive logical relationship words, transitional logical relationship words, conditional logical relationship words, and causal logical relationship words, etc. For example, "both...and...", "while...while...", "not only...but also...", "not only...but...", "although...but...", etc.; After that, different logical operation models are established for different logical connection words. One type of logical connection word corresponds to one logical operation model. Specifically, the logical operation model corresponding to each logical connection word is set by those skilled in the art according to big data experience. For example, in the parallel logical relationship words, the corresponding logical operation model is and in the transitional logical relationship words, the corresponding logical operation model is ; Step 4: Set a stop word library. The stop word library contains various stop words. Specifically, stop words refer to the words that have no actual lexical meaning, such as prepositions, including "in", "at", "from", etc., conjunctions, including "and", "with", etc., auxiliary words, including "de", "di", "de", etc., and modal particles, including "ah", "ma", "ba", etc.; First, identify the stop words in the single-phase area based on the stop words stored in the stop word library. When a stop word is detected, the corresponding stop word is used as a pause symbol to divide the single-phase area into several single words; It should be further noted that when there are multiple consecutive stop words, the multiple consecutive stop words are combined into a whole, and a pause symbol is set according to the position of this whole; For example: The text content in the single-phase area is "What are the working systems of the community residents' committee?", among which, the detected stop words are "yǒu" and "nǎxiē", and the stop words are used as pause symbols. At this time, the single words existing in the text content of the single-phase area are "community residents' committee" and "working systems"; After that, based on the logical connectives and single-word phrases in the single-phase region, the logical similarity of the single-phase region is calculated. The specific calculation method includes: First, use the bag-of-words model algorithm to convert the single-word phrases in the single-phase region into vectors respectively, and then obtain the logical connectives in the single-phase region and identify the logical operation model of the logical connectives. Among them, the specific conversion process of the bag-of-words model algorithm is prior art and will not be elaborated here; Take the vectors of the single-word phrases in the single-phase region as input values and input them into the corresponding logical operation model. The logical operation model performs active operation processing and marks the operation result as the feature vector of the single-phase region; Then obtain the control region Dj, and process the control region Dj according to the above method to obtain the feature vector Tj of the control region; After that, use the cosine similarity algorithm to calculate the cosine similarity between the feature vector Tj of the control region and the feature vector of the single-phase region respectively, and mark the calculation result as the logical similarity value LXm, where m represents the number of the functional region, and m = 1, 2,..., k, indicating that there are k functional regions in the target detection test questions; Step Five: Set an order value for the functional region. Among them, the order value refers to the influence of the order of the keyword vocabulary in the functional region on the calculation of the coincidence rate of the functional region. The larger the order value, the greater the influence of the order of the keyword vocabulary on the calculation of the coincidence value, and the smaller the order value, the smaller the influence of the order of the keyword vocabulary on the calculation of the coincidence value; Among them, the specific order values of each functional region are obtained by those skilled in the art through big data operations; For example, in the option region, there is one keyword in each of options 1, 2, 3, and 4. In question type A, the order of the option region is 2, 1, 4, 3, and in question type B, the order of the option region is 3, 2, 1, 4. At this time, although the option regions of question type A and question type B are different, the keywords are the same. Therefore, the coincidence rate of the option regions of question type A and question type B is 1, so the order value of the option region is the smallest; Obtain all the single-word phrases in the single-phase region again, and identify them in the control region. If there is text information in the control region that is the same as the single-word phrase, mark the corresponding text information as 1; After the identification of the single-word phrases in the control region is completed, first obtain the total number SL of the text information with the value of 1, and then extract the text information corresponding to 1 separately according to the position order in the control region, and at the same time identify the coincidence number WL of the position order between the single-word phrase and the text information 1; After that, use the formula to obtain the text repetition value NXm, where ZD is the total number of text information in the control region, is the order value of the single-phase region, and are weight factors respectively, and and The specific values of are obtained by those skilled in the art through big data operations respectively; Step Six: Use the formula to calculate the regional similarity value Gm of the single-phase region. Among them, and are both proportionality coefficients, and and The specific values of are obtained by those skilled in the art through big data operations; After that, other functional regions are respectively used as single-phase regions, and the regional similarity value Gm of each functional region is calculated. Then, the formula is used again to calculate the overall similarity value XT between the target detection test question and the control test question. Among them, represents the similarity weight of the functional region m, and the similarity weight of each functional region m is obtained by those skilled in the art through big data operations; Step Seven: Compare the overall similarity value XT of the target detection test question with the similarity threshold Xy. If XT < Xy, the corresponding similar test question is marked as a non-coincident test question. On the contrary, if XT ≥ Xy, the corresponding similar test question is marked as a coincident test question. Among them, the specific value of the similarity threshold Xy is obtained by those skilled in the art through big data operations; After all similar test questions are marked, identify the coincident test questions among the similar test questions. If there are coincident test questions among the similar test questions, obtain the corresponding coincident test questions and transmit them to the display terminal. If there are no coincident test questions among the similar test questions, generate test question independent information and transmit it to the display terminal.
[0018] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A method for checking for duplicate examination questions, characterized in that: The method specifically comprises the following steps: Step 1: Set up a test question database, collect historical test question information, and partition and store the historical test question information in the test question database according to the corresponding information field; Step 2: Mark the test questions to be tested as target test questions, identify the information field of the target test questions, and perform information retrieval in the corresponding storage area in the test question database to obtain similar test questions, and then set functional areas. At the same time, the target test questions and similar test questions are divided according to the functional areas, where the functional areas include the material area, the question stem area, the option area, and the answer area; Step 3: Set up a single-phase area and a control area respectively, extract the logical associated words in the single-phase area at the same time, and then divide the single-phase area into several monomer phrases according to the stop words in the stop word library, and then perform vector operations on the monomer phrases according to the logical operation model corresponding to the logical associated words to obtain the feature vector of the single-phase area, and then calculate the logical similarity value of the single-phase area based on the feature vector of the single-phase area and the feature vector of the control area; Step 4: Set an order value for the functional area, and identify the monomer phrases in the single-phase area in the control area to obtain the total number of monomer phrases appearing in the control area, and then determine the number of overlaps in the position order based on the position order of the monomer phrases and the position of the text information corresponding to the control area, and calculate the text repetition value of the single-phase area based on the order value, the total number of monomer phrases appearing in the control area, and the number of overlaps in the position order; Step 5: Based on the logical similarity value and the text repetition value, the regional similarity value of the single-phase region is calculated, and then the regional similarity value of each functional region of the target test question is obtained, and then the regional similarity values of all functional regions are integrated to determine the overall similarity value between the target test question and the control question; Step 6: Based on the overall similarity value, the overlapping test questions are determined among the similar test questions, and the overlapping test questions are transmitted to the display terminal.
2. A method for checking duplicate examination questions according to claim 1, characterized in that: The specific methods for partitioning storage in the test database include: Obtaining application locations of the collected history test questions, and determining the test question nature of the history test questions according to the application locations, wherein the test question nature includes professional tests and comprehensive tests; Obtain the nature of the historical test questions. If the nature of the test questions is a professional test, directly obtain the subject name of the professional test, and partition and store the historical test questions according to the subject name; If the nature of the test question is a comprehensive test, all test questions in the comprehensive test are obtained, and each individual test question is marked as a single-phase test question Pi, where i represents the test question number corresponding to different test questions; After that, the single-phase test question Pi is obtained, and the key words of the single-phase test question Pi are extracted using the TF-IDF method. The Word2Vec model is then used to perform semantic analysis on the key words. Based on the results of the semantic analysis, the information field of the single-phase test question Pi is determined, and then the single-phase test question Pi is partitioned and stored according to the information field.
3. A method for checking duplicate examination questions according to claim 2, characterized in that: The application position of the test question is the corresponding test subject, and then the natural language processing algorithm is used to perform text recognition on the test subject: if the test subject is a specific subject name, the nature of the corresponding test question is marked as a professional test; if the test subject is not a specific subject name, the nature of the corresponding test question is marked as a comprehensive test; Furthermore, the scope of the examination includes professional field tests and comprehensive knowledge tests. In the professional field tests, the test questions appearing in a test paper are all knowledge information related to the corresponding field, and in the comprehensive knowledge test, one test paper will include knowledge information from multiple fields.
4. A method for checking duplicate examination questions according to claim 1, characterized in that: The setting methods of functional areas include: Mark the test questions to be tested as target test questions, then use the TF-IDF method to extract the key words in the target test questions, and use the Word2Vec model to perform semantic analysis on the key words to obtain the information field of the target test questions; Then, the key words of the target test questions are used as search keywords and input into the storage space of the corresponding information field for information retrieval. Then, the search results obtained are marked as similar test questions j, where j = 1, 2, ..., J, indicating that there are J similar test questions in total. Then, semantic analysis is performed on the target test questions and similar test questions using the language model in natural language processing technology, and functional areas are set for the target test questions and similar test questions according to the results of the semantic analysis, wherein the functional areas include the material area, the question stem area, the option area and the answer area.
5. A method for checking duplicate examination questions according to claim 1, characterized in that: In the process of extracting the logical associated words of the single-phase region, the single-phase region refers to any functional region in the target test question, and the control region refers to the region with the same function as the single-phase region in similar test questions; When extracting logical conjunctions in a single-phase region, logical conjunctions refer to words that play a connecting role in a sentence and express various logical relationships, including parallel logical relation words, progressive logical relation words, transitional logical relation words, conditional logical relation words, and causal logical relation words. At the same time, different logical operation models are established for different logical conjunctions, where one type of logical conjunction corresponds to one logical operation model.
6. A method for checking duplicate examination questions according to claim 5, characterized in that: The calculation method of logical similarity value includes: Set a stop word library, where the stop word library contains a variety of stop words. Specifically, stop words refer to words that have no actual lexical meaning; Based on the stop words stored in the stop word library, the stop words in the single-phase area are identified. When the stop words are detected, the corresponding stop words are used as pause symbols to divide the single-phase area into a number of monomer phrases. When there are multiple consecutive stop words, the multiple consecutive stop words are merged into a whole, and a pause symbol is set according to the position of the whole; The bag-of-words model algorithm is used to convert the monomer phrases in the single-phase area into vectors respectively, and then the logical associated words in the single-phase area are obtained to identify the logical operation model of the logical associated words; The vector of the monomer phrase in the single-phase region is used as an input value and input into the corresponding logic operation model, the logic operation model performs active operation processing, and the operation result is marked as the feature vector of the single-phase region; Then, the control region Dj is obtained, and the control region Dj is processed according to the above method to obtain a feature vector Tj of the control region; Then, the cosine similarity algorithm is used to calculate the cosine similarity between the feature vector Tj of the control area and the feature vector of the single-phase area, and the calculation result is marked as the logical similarity value LXm, where m represents the number of the functional area, and m=1, 2,..., k, indicating that there are k functional areas in the target test question.
7. A method for checking duplicate examination questions according to claim 6, characterized in that: The calculation method of text repetition value includes: Set an order value for the functional area, where the order value refers to the effect of the order of key words in the functional area on the calculation of the overlap rate of the functional area; All monomer phrases in the single-phase region are obtained again, and the monomer phrases are identified in the control region. If there is text information consistent with the monomer phrase in the control region, the corresponding text information is marked as 1; When the monomer phrase is identified in the control area, the total number SL of text information 1 is obtained first, and then the text information corresponding to 1 is extracted separately according to the position order in the control area, and the number of overlaps WL between the position order of the monomer phrase and the text information 1 is identified at the same time; Using the formula Get the text repetition value NXm, where ZD is the total number of text information in the control area, is the order value in the single-phase region, and are weight factors respectively.
8. A method for checking duplicate examination questions according to claim 7, characterized in that: The calculation method of the overall similarity value includes: Using the formula The regional similarity value Gm of the single-phase area is calculated, where: and All are proportionality coefficients; Treat other functional areas as single-phase areas, and calculate the regional similarity value Gm of each functional area, using the formula The overall similarity value XT between the target test questions and the control test questions is calculated, where represents the similarity weight of functional region m.
9. A method for checking duplicate examination questions according to claim 1, characterized in that: Methods for determining overlapping questions among similar questions include: Compare the overall similarity value XT of the target test question with the similarity threshold Xy. If XT < Xy, the corresponding similar test question is marked as a non-coincidence test question. Otherwise, if XT ≥ Xy, the corresponding similar test question is marked as a coincidence test question. When all similar test questions are marked, the overlapping test questions among the similar test questions are identified. If there are overlapping test questions among the similar test questions, the corresponding overlapping test questions are obtained and transmitted to the display terminal. If there are no overlapping test questions among the similar test questions, independent test question information is generated and transmitted to the display terminal.
Citation Information
Patent Citations
Judgment method for duplicate checking of similar test questions
CN114238737A