An adaptive question recognition method and system
By employing an adaptive question recognition method, utilizing a large language model and the MD5 algorithm, the problems of manual processing and complex rule matching in the question entry and recognition process are solved, achieving efficient and accurate question data processing and system resource optimization.
Patent Information
- Application Number
- CN202411744841.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-02
AI Technical Summary
In existing technologies, the process of entering and recognizing questions relies on manual processing and complex rule matching, resulting in poor accuracy. This is especially true in complex question banks, where the process is difficult, time-consuming, and has extremely poor fault tolerance.
An adaptive question recognition method is adopted, which uses a large language model to segment and analyze questions in documents. Question elements are extracted by inserting delimiters, classifiers and parameter extractors. Multiple parameters are processed by combining bag-of-words model and cosine similarity algorithm to realize structured data transmission, and data integrity is ensured by MD5 algorithm.
It significantly reduces data processing time, lowers system resource consumption, and improves the accuracy and flexibility of question recognition, making it suitable for various question management systems.
Smart Images

Figure CN119227672B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an adaptive question recognition method and system. BACKGROUND
[0002] In the field of education, doing exercises is a good way to understand knowledge and test learning effect. Whether it is a school or an educational institution, it needs to manage its own question bank to help students master knowledge. Question input is an important link of question bank management and the key to efficiency. Rule matching technology is a common analysis method. The current document analysis and template input technology on the market all need a lot of manual processing, and complex rule settings are also required in the system. The accuracy of manual data processing and the comprehensiveness of system rules jointly affect the correctness of question analysis. The more complex the question bank, the greater the difference between the elements of the questions, the more difficult the manual processing and system processing, and the more time-consuming.
[0003] The current rule matching technology has too strict requirements for the format of the questions, such as question identifiers, option identifiers, and analysis identifiers, which all need to be clearly specified. Data preprocessing needs to be strictly in accordance with the rules. If the content is not processed according to the rules, it will lead to input failure, and the fault tolerance is very poor. How to reduce the difficulty of manual processing, reduce the complexity of question recognition in the system, and improve the recognition accuracy is still a problem that needs to be solved urgently in the question bank management system. SUMMARY
[0004] The embodiments of the present application provide an adaptive question recognition method and system to solve the above technical problems in the prior art.
[0005] To provide a general summary, and an important part is not intended to be generic, nor is it intended to determine the key / important elements or delineate the scope of these embodiments. Its only purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0006] According to a first aspect of the embodiments of the present application, an adaptive question recognition method is provided.
[0007] In one embodiment, the adaptive question recognition method comprises:
[0008] segmenting the document into questions to obtain question data corresponding to the document;
[0009] analyzing the obtained question data using a large language model to obtain question data of different question types and question parameters corresponding to each question data;
[0010] According to a predetermined data structure, the obtained question data of different question types and corresponding question parameters are subjected to data structural processing to obtain structured and encapsulated question data.
[0011] In one embodiment, the format of the document includes: a word document format and an Excel document format.
[0012] In one embodiment, the document is subjected to question segmentation to obtain question data corresponding to the document, which includes:
[0013] A predetermined separator is inserted in the questions of the word document in advance, and the word document is identified based on the separator to obtain questions corresponding to the word document, and the questions are segmented and extracted into question data;
[0014] The row data of the Excel document is identified to determine the question position of each row data, and the questions are segmented and extracted into question data based on the determined question position.
[0015] In one embodiment, the obtained question data is analyzed by using a large language model to obtain question data of different question types and question parameters corresponding to each question data, which includes:
[0016] The question type features corresponding to different question types are obtained in advance, and the obtained question data is classified by using a question classifier of a large language model according to the question type features to obtain question data of different question types;
[0017] The content parameters of the question elements of each question data are extracted by using a parameter extractor of a large language model to obtain corresponding question parameters;
[0018] The question elements include: question type, stem, options, answer, question analysis, chapter classification, label and note.
[0019] In one embodiment, when the content parameters of the question elements of each question data are extracted by using a parameter extractor of a large language model, if multiple different content parameters of the same question element are extracted, the same question element with multiple different content parameters is bid by using a relative majority voting method, and the content parameter with the most number of bids is taken as the question parameter corresponding to the question element.
[0020] In one embodiment, when extracting the content parameters of the question elements of each question data by using the parameter extractor of the large language model, if multiple different content parameters of the same question element are extracted, the extracted content parameters are segmented by using the jieba library, and single characters in the segmented content parameters are deleted and kept as words. A vocabulary table is constructed based on the kept words. The words in the constructed vocabulary table are vectorized by using a bag-of-words model to obtain vectorized word data, and the cosine similarity of the vectorized word data is calculated by using a cosine similarity algorithm. The calculated cosine similarity is taken as a one-time content score, and the content parameter with the highest score is selected as the question parameter corresponding to the question element.
[0021] In one embodiment, the adaptive question identification method further comprises: performing data transmission on the structured and encapsulated question data by using an HTTP request node.
[0022] In one embodiment, the data transmission on the structured and encapsulated question data by using the HTTP request node comprises:
[0023] A check string is set in advance, and the check string is spliced with the structured and encapsulated question data to obtain spliced question data;
[0024] The hash value of the spliced question data is calculated by using an MD5 algorithm, and the hash value is taken as fingerprint data of the structured and encapsulated question data;
[0025] The structured and encapsulated question data and the fingerprint data are simultaneously transmitted, and the fingerprint data is used by a data receiving end to perform fingerprint verification on the structured and encapsulated question data to determine whether the structured and encapsulated question data is tampered with.
[0026] According to a second aspect of the embodiment of the present application, an adaptive question identification system is provided.
[0027] In one embodiment, the adaptive question identification system comprises:
[0028] A document processing module is configured to perform question segmentation on a document to obtain question data corresponding to the document;
[0029] A model analysis module is configured to analyze the obtained question data by using a large language model to obtain question data of different question types and question parameters corresponding to each question data;
[0030] A data processing module is configured to perform data structuring processing on the obtained question data of different question types and corresponding question parameters according to a predetermined data structure to obtain structured and encapsulated question data.
[0031] In one embodiment, the format of the document includes: a word document format and an Excel document format.
[0032] In one embodiment, when the document processing module performs title segmentation on the document to obtain the title data corresponding to the document, a predetermined separator is inserted in the title of the word document in advance, and the title of the word document is recognized based on the separator to obtain the title corresponding to the word document, and the title is segmented and extracted into title data; the row data of the Excel document is recognized to determine the title position of each row data, and the title is segmented and extracted into title data based on the determined title position.
[0033] In one embodiment, when the model analysis module analyzes the obtained title data using a large language model to obtain title data of different types of questions and question parameters corresponding to each title data, it obtains question type characteristics corresponding to different types of questions set in advance, and classifies the obtained title data using the question classifier of the large language model according to the question type characteristics to obtain title data of different types of questions; the parameter extractor of the large language model is used to extract the content parameters of the question elements of each title data to obtain the corresponding question parameters; wherein the question elements include: question type, stem, options, answer, question analysis, chapter classification, label and note.
[0034] In one embodiment, when the model analysis module uses the parameter extractor of the large language model to extract the content parameters of the question elements of each title data, if the same question element of multiple different content parameters is extracted, the same question element of multiple different content parameters is voted using the relative majority voting method, and the content parameter with the most votes is taken as the question parameter corresponding to the question element.
[0035] In one embodiment, when the model analysis module uses the parameter extractor of the large language model to extract the content parameters of the question elements of each title data, if the same question element of multiple different content parameters is extracted, the extracted content parameters are segmented using the jieba library, and single words in the segmented words are deleted and kept. The reserved words are used to build a vocabulary; the words in the constructed vocabulary are vectorized using the bag-of-words model to obtain vectorized word data, and the cosine similarity of the vectorized word data is calculated using the cosine similarity algorithm; the calculated cosine similarity is used as the content one-time score, and the content parameter with the highest score is selected as the question parameter corresponding to the question element.
[0036] In one embodiment, the adaptive question recognition method further comprises a data transmission module for transmitting the structured and encapsulated title data through an HTTP request node.
[0037] In one embodiment, the data transmission module, when performing data transmission on the structured encapsulated question data through the HTTP request node, sets a check string in advance, splices the check string with the structured encapsulated question data to obtain spliced question data, calculates a hash value of the spliced question data by using an MD5 algorithm, and takes the hash value as fingerprint data of the structured encapsulated question data; the structured encapsulated question data and the fingerprint data are simultaneously transmitted, and the data receiving end is prompted to perform fingerprint checking on the structured encapsulated question data by using the fingerprint data to determine whether the structured encapsulated question data is tampered.
[0038] According to a third aspect of the embodiments of the present application, a computer device is provided.
[0039] In one embodiment, the computer device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the above method.
[0040] According to a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided.
[0041] In one embodiment, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the above method.
[0042] The technical solutions provided by the embodiments of the present application can include the following beneficial effects:
[0043] The present application can adjust the complex rules in the data preprocessing stage to simple question segmentation identifiers, and analyze and process the questions by using a large language model, which can greatly reduce the time investment of the business in the data processing stage. At the same time, there is no need to set numerous question element matching rules in the system, which can effectively reduce the question data abnormal situation caused by missing rules and reduce system resource consumption. At the same time, the design of the module allows it to be flexibly integrated into common question management systems, which can effectively expand the capability boundary of the system, has a wide application prospect and practical value.
[0044] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0045] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.
[0046] Figure 1 is a flowchart of an adaptive question recognition method according to an exemplary embodiment;
[0047] Figure 2 is a structural block diagram of an adaptive question recognition system according to an exemplary embodiment;
[0048] Figure 3 is a structural diagram of a question classifier according to an exemplary embodiment;
[0049] Figure 4 is a diagram of a parameter extractor according to an exemplary embodiment;
[0050] Figure 5 is a structural diagram of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION
[0051] The following description and drawings are illustrative of the specific embodiments herein and are not intended to be limiting. Numerous specific details are described to provide a thorough understanding of the specific embodiments. However, in certain instances, well-known methods, procedures, components, and circuits have been omitted in order to avoid obscuring the concepts of the specific embodiments. Parts and features of some embodiments can be included or substituted for parts and features of other embodiments. The scope of the embodiments of this document includes the entire scope of the claims and all available equivalents thereof. In this document, the terms “first,” “second,” and the like, do not denote any order, quantity, or importance, but rather are used to distinguish one element from another. In this document, the terms “including,” “containing,” and the like, are intended to be open-ended terms that specifically permit the inclusion of more than one element, and that do not exclude additional elements. In this document, the term “associated with” is intended to mean that the associated elements are in some way connected, whether directly or indirectly, through other elements, or through the use of one or more wires, cables, and / or computer program elements, over a wired connection, and / or over a wireless connection.
[0052] The terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like as used herein to indicate orientation or positional relationships based on the orientations or positional relationships shown in the drawings, are for purposes of this description only, and are not intended to indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore should not be construed as limiting the application. In the description of the description, unless otherwise specified and limited, the terms "mount", "connect", "connection" should be interpreted broadly, for example, can be mechanical connection or electrical connection, can be internal communication of two elements, can be direct connection, or indirect connection through intermediate medium, the specific meaning of the above terms can be understood by the person skilled in the art according to the specific circumstances.
[0053] In this paper, unless otherwise specified, the term "a plurality of" means two or more.
[0054] In this paper, the character " / " represents that the front and rear objects are a kind of "or" relationship. For example, A / B represents: A or B.
[0055] In this paper, the term "and / or" is a description of the association between objects, which means that there can be three kinds of relationships. For example, A and / or B, which means: A or B, or, A and B, three kinds of relationships.
[0056] It should be understood that although each step in the flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified in this paper, the execution of these steps has no strict order limitation, and these steps can be executed in other order. Moreover, at least part of the steps in the figure can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed with at least part of other steps or other steps or stages of sub-steps or stages in turn or alternately.
[0057] Each module in the device or system of the present application can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above modules by the processor.
[0058] In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0059] Figure 1An embodiment of an adaptive question recognition method of the present application is shown.
[0060] In the optional embodiment, the adaptive question recognition method comprises:
[0061] In step S101, the document is subjected to question segmentation to obtain question data corresponding to the document.
[0062] In step S103, the obtained question data is analyzed by using a large language model to obtain question data of different question types and question parameters corresponding to each question data.
[0063] In step S105, the obtained question data of different question types and corresponding question parameters are subjected to data structural processing according to a predetermined data structure to obtain structured and encapsulated question data.
[0064] Figure 2 An embodiment of an adaptive question recognition system of the present application is shown.
[0065] In the optional embodiment, the adaptive question recognition system comprises:
[0066] The document processing module 201 is configured to subject the document to question segmentation to obtain question data corresponding to the document.
[0067] The model analysis module 203 is configured to analyze the obtained question data by using a large language model to obtain question data of different question types and question parameters corresponding to each question data.
[0068] The data processing module 205 is configured to subject the obtained question data of different question types and corresponding question parameters to data structural processing according to a predetermined data structure to obtain structured and encapsulated question data.
[0069] In the above optional embodiment, the format of the document comprises a word document format and an Excel document format.
[0070] In the above optional embodiment, when the document is subjected to question segmentation to obtain question data corresponding to the document, a predetermined separator is inserted in the question of the word document in advance, and the word document is recognized based on the separator to obtain the question corresponding to the word document, and the question is segmented and extracted into question data; the row data of the Excel document is recognized to determine the question position of each row of data, and the question is segmented and extracted into question data based on the determined question position.
[0071] In the optional embodiment described above, when the obtained question data is analyzed by using a large language model (LLM) to obtain question data of different question types and question parameters corresponding to each question data, the question type characteristics corresponding to different question types are obtained in advance, and the question data obtained is classified by using the question classifier of the large language model according to the question type characteristics to obtain question data of different question types; the content parameters of the question elements of each question data are extracted by using the parameter extractor of the large language model to obtain the corresponding question parameters; wherein the question elements include: question type, stem, options, answer, question analysis, chapter classification, label and note.
[0072] In addition, when the content parameters of the question elements of each question data are extracted by using the parameter extractor of the large language model, if the same question element of multiple different content parameters is extracted, the same question element of multiple different content parameters is bid by using the relative majority voting method, and the content parameter with the most bids is taken as the question parameter corresponding to the question element. Alternatively, when the content parameters of the question elements of each question data are extracted by using the parameter extractor of the large language model, if the same question element of multiple different content parameters is extracted, the extracted content parameters are segmented by using the jieba library, and single words in the segmented words are deleted and kept. The word table is constructed based on the kept words; the words in the constructed word table are vectorized by using the bag-of-words model to obtain vectorized word data, and the cosine similarity of the vectorized word data is calculated by using the cosine similarity algorithm; the calculated cosine similarity is taken as the content one-time score, and the content parameter with the highest score is selected as the question parameter corresponding to the question element.
[0073] Specifically, the question is divided and stored in the question list, the question is extracted from the question list and input into the question classifier. The question classifier is set as shown in the following table: Figure 3 In the question classification, the basic question types contained in the question bank are defined, which are commonly single-choice questions, multiple-choice questions, judgment questions, fill-in-the-blank questions, question-and-answer questions, and drag-and-drop questions. The principle of the question classifier is to classify each question type, and the characteristics of the question type need to be clearly defined in the classification. The large language model will combine its own knowledge with the characteristics defined in the classification prompt words to classify the question input into the question classifier.
[0074] Each classification of the question classifier is followed by a parameter extractor (such as Figure 4 ). The question classified by the question classifier is transmitted to the corresponding parameter extractor. In the parameter extractor, the large language model can extract the question elements of different question types according to the definition of the question elements in the prompt words.
[0075] When the parameter extractor extracts parameters, in addition to content extraction, it also needs to set strategies to ensure the correctness of content extraction. The strategies that can be used include the following two:
[0076] 1) Voting method. This strategy is suitable for the extraction of parameters with enumerated values, such as the enumeration values of the question type of multiple-choice questions. To prevent extraction errors, an absolute rule is combined with a relative majority voting method. First, set an absolute rule, and the question type obtained by hitting the absolute rule is the final question type. For example, the absolute rule for the question type of multiple-choice questions is "a. The beginning or end of the stem explicitly mentions one of single-choice or multiple-choice, then the question type is directly assigned, b. If there are multiple answers, it is a multiple-choice question." If no absolute rule is hit, use the relative majority voting method to vote for the multiple question types extracted by other rules, and the question type with the most votes is the final question type of the question.
[0077] 2) Consistency scoring. This strategy is suitable for consistency verification of text extraction.
[0078] ① Use a large language model to extract parameters multiple times for the same question, such as the extraction of stems, which can obtain multiple stem information for a question.
[0079] ② Use the jieba library to split the input parameter content, for example, input 3 stems of the same question as,
[0080] A: "Machine learning is a branch of artificial intelligence",
[0081] B: "Artificial intelligence has applications in various fields",
[0082] C: "Learning Python can help understand machine learning",
[0083] Through jieba splitting, the following content will be obtained: ['Machine learning is a branch of artificial intelligence', 'Artificial intelligence has applications in various fields', 'Learning Python can help understand machine learning'];
[0084] ③ Delete single words in the split and only keep words, then build a word list for the stem of the question. The above content will be retained as ['python' 'a' 'artificial intelligence' 'branch' 'can' 'various fields' 'learning' 'help' 'application''machine' 'understand'];
[0085] (4) Using the bag-of-words model to vectorize the extracted words, i.e., counting the number of occurrences of each word in the stem, and then representing each stem with a vector. A, B and C in step (2) will be represented as: [0 1 1 1 0 0 1 0 0 1 0], [0 0 1 0 0 1 0 0 1 0 0], [1 0 0 0 1 0 2 1 0 1 1];
[0086] (5) Using the cosine similarity algorithm (cosine similarity), the similarity of the text content is calculated. The cosine similarity of A and C is 0.447.
[0087] (6) The sum of the cosine similarity of each content is calculated as the consistency score of the text content, and the content with the highest score is the final parameter content. For example, the similarity of A and B in step (2) is 0.258, the similarity of A and C is 0.447, and the similarity of B and C is 0. Then the similarity score of content A is 0.705 (0.258+0.447), the similarity score of content B is 0.258, and the similarity score of content C is 0.447. Finally, content A is used as the stem of the question.
[0088] In the above optional embodiment, when performing data structuring processing, the input variable can be defined first, i.e., the parameters extracted by the parameter extractor, and then in the code module, the data processing function is defined, and finally the data is restructured, and the data structuring processing is completed. For example, the parameter extractor extracts question_type as a single-choice question and answer as A, and after structuring, it is {“question_type”: “single-choice question”, “answer”: “A”}. The purpose of structuring is to directly transmit the data extracted by the large language model to the background through the interface.
[0089] In the above optional embodiment, the adaptive question recognition method further comprises: performing data transmission on the structured and encapsulated question data through an HTTP request node. When performing data transmission on the structured and encapsulated question data through the HTTP request node, a verification string is set in advance, and the verification string is spliced with the structured and encapsulated question data to obtain spliced question data; the hash value of the spliced question data is calculated by using the MD5 algorithm, and the hash value is used as the fingerprint data of the structured and encapsulated question data; the structured and encapsulated question data and the fingerprint data are simultaneously transmitted, and the fingerprint data is used by the data receiving end to perform fingerprint verification on the structured and encapsulated question data to determine whether the structured and encapsulated question data has been tampered with.
[0090] Specifically, in the HTTP request node, set the API interface, access method, and verification string (the string agreed with the backend), and put the structured data processed in the previous step into the body. The specific steps are as follows:
[0091] ① Concatenate the body data (structured data) with the verification string. For example, if the verification string is "AFB65@#$$$$@", the concatenated data will be: {"question_type": "Single Choice", "answer": "A"} AFB65@#$$$$@;
[0092] ② Use the MD5 (Message Digest Algorithm 5) algorithm to calculate the hash value of the spliced data as the "fingerprint" of the data content. In this paper, the md5 method of the built-in hashlib module of Python is directly used to calculate the hash value. Since the hash value calculated by this algorithm is irreversible, the default method is sufficiently effective;
[0093] ③Send the structured data and "fingerprint" to the backend at the same time. Note that the verification string is not transmitted. Since the question information is not sensitive data, this process uses plain text transmission. To ensure that the transmitted content is not tampered with, "fingerprint" verification is required;
[0094] ④ After receiving the structured data, the background also concatenates the structured data with the verification string and calculates the "fingerprint" of the concatenated content.
[0095] ⑤Compare the "fingerprints" of ③ and ④. If the "fingerprints" are consistent, it means that the transmitted content has not been tampered with, and the incoming structured data can be used to create questions.
[0096] Based on the above technical solution, the present invention can adjust the complex rules in the data preprocessing stage into simple topic segmentation identifiers, and use a large language model to analyze and process the topics, which can greatly reduce the time investment of the business in the data processing stage. At the same time, there is no need to set up numerous topic element matching rules in the system, which can effectively reduce the abnormal situation of topic data caused by missing rules and reduce system resource consumption. At the same time, the design of this module allows it to be flexibly integrated into common topic management systems, which can effectively expand the capability boundary of the system and has broad application prospects and practical value.
[0097] Figure 5An embodiment of a computer device of the present application is shown, which can be a server, and the computer device comprises a processor, a memory and a network interface connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store static information and dynamic information data. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is configured to be executed by the processor to implement the steps in the above method embodiments.
[0098] Those skilled in the art can understand that, Figure 5 The structure shown in the above is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can comprise more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0099] In addition, the present application also provides the above computer device, which comprises a memory and a processor. The memory stores a computer program. The processor executes the computer program to implement the steps in the above method embodiments.
[0100] In addition, the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0101] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments of each method. In the embodiments of the present application, any reference to memory, storage, database or other medium can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0102] The present application is not limited to the structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is only limited by the appended claims.
Claims
1. An adaptive question recognition method, characterized in that: include: Segment the document by topic and obtain the topic data corresponding to the document; The document formats include: Word document format and Excel document format; segmenting the document title to obtain the title data corresponding to the document includes: inserting a predetermined delimiter in the title of the Word document in advance, performing Word document recognition based on the delimiter to obtain the title corresponding to the Word document, and segmenting the title into title data; identifying the row data of the Excel document, determining the title position of each row data, and segmenting the title into title data based on the determined title position; Use the large language model to analyze the acquired question data to obtain question data of different question types and the question parameters corresponding to each question data; According to a predetermined data structure, the obtained question data of different question types and corresponding question parameters are subjected to data structuring processing to obtain structured encapsulated question data; The method of analyzing the acquired question data using the large language model to obtain question data of different question types and question parameters corresponding to each question data includes: obtaining pre-set question type features corresponding to different question types, and classifying the acquired question data using the question classifier of the large language model according to the question type features to obtain question data of different question types; extracting content parameters of question elements of each question data using the parameter extractor of the large language model to obtain corresponding question parameters; wherein the question elements include: question type, question stem, options, answer, question analysis, chapter classification, tags, and notes; When extracting content parameters of the topic element of each topic data using the parameter extractor of the large language model, if multiple different content parameters of the same topic element are extracted, the multiple different content parameters of the same topic element are voted using the relative majority voting method, and the content parameter with the largest number of votes is used as the topic parameter corresponding to the topic element; or When extracting the content parameters of the topic element of each topic data using the parameter extractor of the large language model, if multiple different content parameters of the same topic element are extracted, the extracted content parameters are segmented using the Jieba vocabulary, and individual characters in the segmentation are deleted while retaining the words. A vocabulary table is constructed based on the retained words. The words in the constructed vocabulary table are vectorized using the bag-of-words model to obtain vectorized vocabulary data, and the cosine similarity of the vectorized vocabulary data is calculated using the cosine similarity algorithm. The calculated cosine similarity is used as a one-time content score, and the content parameter with the highest score is selected as the topic parameter corresponding to the topic element. The data transmission of the structured encapsulated question data is performed through the HTTP request node; the data transmission of the structured encapsulated question data through the HTTP request node includes: pre-setting a verification string, and splicing the verification string with the structured encapsulated question data to obtain the spliced question data; using the MD5 algorithm to calculate the hash value of the spliced question data, and using the hash value as the fingerprint data of the structured encapsulated question data; transmitting the structured encapsulated question data and the fingerprint data at the same time, and prompting the data receiving end to use the fingerprint data to perform fingerprint verification on the structured encapsulated question data to determine whether the structured encapsulated question data has been tampered with.
2. An adaptive question recognition system, characterized in that: include: The document processing module is used to segment documents into topics and obtain the topic data corresponding to the documents; The document formats include: Word document format and Excel document format; segmenting the document title to obtain the title data corresponding to the document includes: inserting a predetermined delimiter in the title of the Word document in advance, performing Word document recognition based on the delimiter to obtain the title corresponding to the Word document, and segmenting the title into title data; identifying the row data of the Excel document, determining the title position of each row data, and segmenting the title into title data based on the determined title position; The model analysis module is used to analyze the acquired question data using a large language model to obtain question data of different question types and question parameters corresponding to each question data; The data processing module is used to perform data structuring processing on the obtained question data of different question types and corresponding question parameters according to a predetermined data structure to obtain structured encapsulated question data; When the model analysis module uses the large language model to analyze the acquired question data to obtain question data of different question types and question parameters corresponding to each question data, it obtains pre-set question type features corresponding to different question types, and classifies the acquired question data based on the question type features using the question classifier of the large language model to obtain question data of different question types; and uses the parameter extractor of the large language model to extract content parameters of question elements of each question data to obtain corresponding question parameters; wherein the question elements include: question type, question stem, options, answer, question analysis, chapter classification, tags, and notes; When extracting content parameters of the topic element of each topic data using the parameter extractor of the large language model, if multiple different content parameters of the same topic element are extracted, the multiple different content parameters of the same topic element are voted using the relative majority voting method, and the content parameter with the largest number of votes is used as the topic parameter corresponding to the topic element; or When extracting the content parameters of the topic element of each topic data using the parameter extractor of the large language model, if multiple different content parameters of the same topic element are extracted, the extracted content parameters are segmented using the Jieba vocabulary, and individual characters in the segmentation are deleted while retaining the words. A vocabulary table is constructed based on the retained words. The words in the constructed vocabulary table are vectorized using the bag-of-words model to obtain vectorized vocabulary data, and the cosine similarity of the vectorized vocabulary data is calculated using the cosine similarity algorithm. The calculated cosine similarity is used as a one-time content score, and the content parameter with the highest score is selected as the topic parameter corresponding to the topic element. A data transmission module is used to transmit the structured encapsulated question data through an HTTP request node; when transmitting the structured encapsulated question data through an HTTP request node, the data transmission module pre-sets a verification string and concatenates the verification string with the structured encapsulated question data to obtain concatenated question data; calculates a hash value of the concatenated question data using the MD5 algorithm, and uses the hash value as fingerprint data of the structured encapsulated question data; transmits the structured encapsulated question data and the fingerprint data simultaneously, and prompts the data receiving end to perform fingerprint verification on the structured encapsulated question data using the fingerprint data to determine whether the structured encapsulated question data has been tampered with.
Citation Information
Patent Citations
Question input method, question input device, electronic equipment and computer readable storage medium
CN112861864A