High-quality data mining method, production method and system
By using the method of estimating high-quality test questions and the A/B-grade production strategy, the problem of efficient production of massive test question banks was solved, the accuracy and efficiency of the test question banks were improved, and user needs were met.
Patent Information
- Application Number
- CN202011530665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-12-22
AI Technical Summary
Existing technologies make it difficult to efficiently produce and maintain massive question banks, especially to quickly replenish high-quality questions and answers. Traditional crowdsourcing systems are inefficient and difficult to meet user needs.
By estimating high-quality test questions based on the same clusters or similar clusters of periodic characteristics and current hot spots, adopting A/B production methods and repeated delivery strategies, and using the crowdsourcing system to flexibly allocate tasks, priorities and pricing are automatically adjusted to ensure the production of high-quality test questions.
It improves the accuracy of photographed test questions and adapts to the production needs of test question banks of different sizes. The priority and pricing automatic adjustment strategies better concentrate user resources, ensure the return on investment of the production of high-quality test questions, and achieve efficient test question bank updates.
Smart Images

Figure CN112559821B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network education technology, in particular to the field of machine learning-assisted technology, and more specifically to a high-quality data mining method, production method and system. Background Art
[0002] With the development of the internet and artificial intelligence, online education is becoming increasingly popular, and the application of artificial intelligence and machine learning in online education is also becoming increasingly common. This is especially true in online exam-based education, where the traditional one-question-one-answer model is difficult to adapt to the evolving landscape due to the potentially massive number of questions involved. Consequently, machine learning technology has been rapidly developed, categorizing and answering various types of questions, simplifying the development of question banks and simplifying the difficulty of answering them.
[0003] In recent years, the development of image recognition technology has led to a new form of test-based online education, namely, taking photos of test questions and uploading them to search for answers. Due to the inherent deficiencies of the test question bank, it is impossible to include all questions in advance, so many questions will have no answers. Although the maintainers of the test question bank can manually answer them and include the answers in the test question bank, the speed of manual answering is limited after all, and it is difficult to significantly increase the number of questions and answers added every day.
[0004] In addition, due to the development of current technology, the question bank is particularly large, and crowdsourcing is usually used to maintain and distribute the production and answering tasks of the test questions. Therefore, the industry also needs to research and develop corresponding crowdsourcing systems to study how to more efficiently allocate and manage crowdsourcing tasks such as question selection, duplicate checking, answering, review, and typesetting. Summary of the Invention
[0005] In view of this, the main object of the present invention is to provide a high-quality data mining method, a production method and a mining system using the same, in order to at least partially solve at least one of the above-mentioned technical problems.
[0006] To achieve the above object, as a first aspect of the present invention, a method for mining high-quality data is provided, comprising the following steps:
[0007] Based on periodic features and / or based on the same cluster or similar cluster of the current hot spot, relevant data of high-quality test questions is estimated.
[0008] A second aspect of the present invention provides a method for producing high-quality test questions, comprising the following steps:
[0009] Use the mining method described above to estimate relevant information of high-quality test questions;
[0010] Evaluate the relevant information of the obtained high-quality test questions to determine whether further expansion of production is necessary;
[0011] If necessary, organize production.
[0012] A third aspect of the present invention provides a high-quality test question mining system, comprising:
[0013] The data integration and analysis module is used to estimate the relevant data of high-quality test questions using the mining method described above;
[0014] The data screening module is used to evaluate the relevant data of high-quality test questions identified by the data integration and analysis module to determine whether further expansion of production is needed;
[0015] The data production module organizes the production of data that the data screening module determines needs to be further expanded.
[0016] A fourth aspect of the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program. When the computer executable program is executed by the processor, the processor executes the high-quality data mining method as described above.
[0017] A fifth aspect of the present invention further proposes a computer-readable medium storing a computer-executable program, which, when executed, implements the high-quality data mining method described above.
[0018] Based on the above technical solutions, it can be seen that the high-quality data mining method, production method, and crowdsourcing system using the same of the present invention have at least one of the following beneficial effects compared to the prior art:
[0019] By estimating the same clusters and similar clusters of hot words, the present invention can produce high-quality test questions in advance and improve the recall rate of photo test questions;
[0020] The present invention can flexibly allocate tasks through a crowdsourcing system and adapt to the production tasks of test question banks of different sizes;
[0021] The present invention can expand the test question set as much as possible while ensuring the urgency of the task by setting different priorities and weights;
[0022] The A / B file production method of the present invention is more suitable for the production of high-quality test questions, can better concentrate crowdsourcing user resources, and better ensure the return on investment (ROI) of production;
[0023] The repeated production strategy of the present invention is an important guarantee for high-quality data to be put online as scheduled, and it is also a capability that ordinary crowdsourcing production does not have. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a block flow diagram of the high-quality data mining method of the present invention;
[0025] Figure 2 1 is a schematic diagram of a framework of a high-quality data mining system according to an embodiment of the present invention;
[0026] Figure 3 This is an overall flow chart of a high-quality data mining method according to an embodiment of the present invention;
[0027] Figure 4 is a schematic structural diagram of an electronic device according to an embodiment of the present invention;
[0028] Figure 5 is a schematic diagram of a computer-readable recording medium as an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In the introduction of specific embodiments, the detailed description of the structure, performance, effect or other features is intended to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from implementing the present invention with a technical solution that does not include the aforementioned structure, performance, effect or other features under specific circumstances.
[0030] The flowcharts in the accompanying drawings are merely illustrative of the process flow and do not necessarily include all of the content, operations, and steps in the flowcharts, nor do they necessarily imply that all of the steps in the flowcharts must be executed in the order shown. For example, some of the steps in the flowcharts may be separated, some may be combined or partially combined, and so on. The execution order shown in the flowcharts may be changed according to actual circumstances without departing from the spirit of the present invention.
[0031] Frames in the accompanying drawings Figure 1 The term "functional entity" generally refers to a functional entity and does not necessarily correspond to a physically independent entity. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processing unit devices and / or microcontroller devices.
[0032] The same reference numerals in the accompanying drawings represent the same or similar elements, components or parts, and thus repeated descriptions of the same or similar elements, components or parts may be omitted below. It should also be understood that although the first, second, third and other numbered adjectives may be used herein to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these adjectives. In other words, these adjectives are only used to distinguish one from another. For example, the first device may also be called the second device, but this does not deviate from the essential technical solution of the present invention. In addition, the terms "and / or" and "and / or" refer to all combinations including any one or more of the listed items.
[0033] The meanings of some technical terms in this manual are as follows:
[0034] Crowdsourcing refers to the practice of a company or organization outsourcing tasks previously performed by employees to a non-specific (and usually large) network of the general public. Therefore, a crowdsourcing system is an organizational model that implements this crowdsourcing approach. This model has had a disruptive impact on several industries in the United States. The test question production and answering tasks described in this invention can also be accomplished through a crowdsourcing system, thereby alleviating the recruitment challenges faced by companies and enabling flexibility for employees. It also allows for efficient test question production and answering, meeting growing demand.
[0035] High-quality test questions are a concept defined in the present invention to distinguish them from ordinary test questions. They refer to test questions that are highly requested by customers, have a large demand, and can address the knowledge points that customers are most concerned about. In other words, they are the test questions that customers most want to experience and that can most effectively improve their scores.
[0036] Compared with ordinary test data, high-quality test data has many obvious characteristics, such as:
[0037] (1) Periodic characteristics
[0038] Due to the cyclical nature of learning, students repeat the same knowledge every year. They have cyclical learning content throughout the academic year and quarter, and they also take exams at regular intervals throughout the year, such as weekly and monthly exams, the middle school entrance exam, and the college entrance exam. Therefore, high-quality exam questions are obviously something that every class of students will experience and attempt.
[0039] (2) Knowledge points and characteristics of the same and similar questions
[0040] Determining whether a problem is the same or similar is a rather unique problem. Take a math word problem, for example: "Xiaoming and Xiaohong bought 10 kilograms of apples. They each ate 2 kilograms. How many kilograms were left?" If Xiaoming and Xiaohong's purchase of apples is replaced by Xiaobai and Xiaohei's purchase of pears, the problem is essentially the same. If buying 10 kilograms of apples is replaced by buying 9 kilograms of apples, the problem is essentially different. However, in practice, making such a judgment based on semantics is extremely difficult, and the methods used vary across different academic levels.
[0041] Therefore, the question bank will treat exactly the same questions as the same questions, and the two situations with slight changes mentioned above will be recorded as similar questions.
[0042] In recent years, the emergence of the photo-taking question search method has greatly facilitated learners. However, due to the inherent deficiencies of the question bank, it is impossible to include all questions in advance, resulting in many questions without answers. Although the maintainer of the question bank can manually answer them and include the answers in the question bank, the speed of manual answering is limited after all, and it is difficult to significantly increase the number of questions and answers added every day. To address this problem, the present invention takes a different approach, circumventing the conventional question production logic, that is, discovering hot spots based on user needs and then organizing production based on hot spots. Instead, it estimates hot spots in advance, organizes production, and then goes online to wait for users to search, thereby improving the accuracy rate of photo-taking question search.
[0043] To make an advance estimate, it is necessary to evaluate the test questions and know which test questions may be in higher demand. Therefore, the present invention proposes the concept of high-quality test questions, that is, test questions that are frequently requested by customers, have a large demand, and can meet the knowledge points that customers are most concerned about. By estimating the possible high-quality test questions in advance, they are given priority in producing question stems and answers to meet user needs.
[0044] Therefore, the present invention introduces concepts such as knowledge point information, identical clusters and similar clusters, that is, knowledge points are divided into different clusters according to their relevance to the main information in the questions. Among them, knowledge point information is an attribute of the test question bank data, and the knowledge point data with a relatively high proportion are all formulated and supplemented by teaching and research teachers, rather than implemented by algorithms. The clustering technology for obtaining identical clusters and similar clusters is also one of the important capabilities and technical reserves of the applicant in the construction of the test question bank. The clustering service is a basic service that businesses such as question duplication detection, clustering, retrieval, and mounting rely on. Since the present invention directly uses the knowledge base and hot word library that have completed clustering, the specific clustering algorithm will not be discussed in detail here.
[0045] The inventors found that in order to estimate the data information of high-quality test questions in advance, the mining method needs to focus on the following two aspects: (1) making question production predictions based on periodic knowledge points and learning content; (2) finding hot similarity clusters and hot knowledge points based on current hot words, and then inferring the questions to be produced through the two.
[0046] Specifically, if Figure 1 As shown, the present invention discloses a high-quality data mining method, comprising the following steps:
[0047] Based on periodic features and / or based on the same cluster or similar cluster of the current hot spot, relevant data information of high-quality test questions is estimated.
[0048] The periodic feature is determined based on the semester and school year information of the knowledge point, the periodic correlation between the search questions and the student status information of the searcher, and the timing and frequency of the search questions.
[0049] The current hotspot is the relevant knowledge point information obtained by mapping the system to the matching content found in the question bank based on the relevant auxiliary information of the search topic (Query).
[0050] Among them, the same cluster or similar cluster of the current hot spot is also obtained based on the page views (pv) of the question, the auxiliary information of the question (such as subject section, knowledge point, quality score, source, etc.), the similar cluster of the question, etc., and its implementation method is also through the mapping relationship in the knowledge base.
[0051] Among them, the same cluster or similar cluster of the current hot spot is also based on the auxiliary information of the question (such as academic stage, knowledge point, quality score, source, etc.), matching the current time information: the academic year, academic quarter, key examinations, key subject progress and other information are searched in the question bank, and the corresponding hot knowledge points are inferred based on the search matching content. The corresponding key questions are retrieved based on the knowledge points, and then the hot cluster is found.
[0052] The present invention also discloses a method for producing high-quality test questions, comprising the following steps:
[0053] Use the mining method described above to estimate relevant data information of high-quality test questions;
[0054] Evaluate the relevant data and information of the obtained high-quality test questions to determine whether further expansion of production is necessary;
[0055] If necessary, organize production.
[0056] Among them, those that need to be further expanded need to be marked with timeliness according to the source of their data generation, and divided into high-timeliness questions and low-timeliness questions;
[0057] As a preference, during the data production process, if it is found that a high-timeliness test question exceeds the timeliness, it will be re-determined whether the question needs to be resubmitted to low-timeliness production or abandoned.
[0058] Among them, the questions in the hot question set are scored based on their same cluster, completeness and / or richness information;
[0059] Based on predetermined rules, some high-quality data do not need to be put into production, and some low-quality data do not need to be put into production. The remaining data will be put into production based on the scoring.
[0060] Among them, those that need to be organized for production are distributed to test question production personnel in a crowdsourcing mode for production.
[0061] Among them, for those that need to organize production, the data to be produced needs to be divided into A / B file data according to the scores of the hot question set, among which the A file data is the data that must be produced, and the B file data is the high-priority production data.
[0062] Among them, the A / B file data will determine the production priority and pricing based on the current production capacity of the question bank, the dynamic calculation of the current A / B file data volume, and the estimated task completion status: if the production capacity is insufficient, the priority will be increased; when the priority is high enough, the production price of this link will be increased; when the production capacity is still insufficient, the priority and pricing of file B will be lowered.
[0063] Among them, in order to ensure the production of A-level data, a repeated delivery mode is adopted for A-level data, so that more crowdsourcing system task acceptors can produce the topic at the same time, and after one crowdsourcing system task acceptor completes the production, others can manually choose to abandon the topic and obtain part of the benefits.
[0064] Among them, the questions after production are transferred to the question bank and sorted according to the page views of the questions in the question bank, and the top ranked ones are used as the basis for the next round of data mining.
[0065] like Figure 2 As shown, the present invention also discloses a high-quality test question mining system, comprising:
[0066] The data integration and analysis module is used to estimate the relevant data information of high-quality test questions using the mining method described above;
[0067] The data screening module is used to evaluate the relevant data information of high-quality test questions determined by the data integration and analysis module to determine whether further expansion of production is needed;
[0068] The data production module organizes the production of data that the data screening module determines needs to be further expanded.
[0069] The data screening module also assigns timeliness tags to test questions requiring further production (i.e., determines a production deadline based on a threshold) based on the data's source, classifying them into high-timeliness and low-timeliness questions. High-timeliness questions are prioritized for production; during data production, if a high-timeliness question is found to have exceeded its timeliness, the module will reassess whether to resubmit it to low-timeliness production or discard it.
[0070] Among them, the data screening module also scores the questions in the hot question set based on the same cluster and completeness and richness information of the questions, and then according to predetermined rules, some high-quality data do not need to be put into production, some low-quality data do not need to be put into production, and the remaining data is put into production according to the scores.
[0071] The data production module further includes a crowdsourcing platform, which distributes test questions that need to be further expanded to test question production personnel in a crowdsourcing mode.
[0072] The data production module divides the data to be produced into A / B file data according to the scores of the hot question set, wherein the A file data is the data that must be produced and the B file data is the high-priority production data.
[0073] A / B file data are all higher priority data, but the production pricing is not necessarily higher than that of ordinary production questions. A / B file data will also be dynamically calculated based on the current question bank's production capacity, the current amount of A / B file data, and the estimated task completion status. If the production capacity is insufficient, the priority will be increased; when the priority is high enough (higher than other production data, increasing the priority will not improve production efficiency), the production price of this link will be increased (there is an upper limit); when the production capacity is still insufficient, the priority and pricing of the B file will be lowered (there is a lower limit, higher than ordinary questions). This shows that each link in the production of the question bank has the corresponding ability to adjust priorities and prices in order to flexibly adapt to various production tasks.
[0074] To ensure the production of A-level data, the data production module adopts a repeated delivery mode for A-level data, allowing more crowdsourcing system task acceptors to produce the topic at the same time. After a crowdsourcing system task acceptor completes the production, others can manually choose to give up the topic and obtain part of the benefits.
[0075] The data production module will put the produced questions online and sort them according to the page views after going online. The top-ranked ones will be used as the basis for the next round of data mining.
[0076] The present invention also discloses an electronic device, which includes a processor and a memory, wherein the memory is used to store a computer executable program, wherein when the computer executable program is executed by the processor, the processor executes the high-quality data mining method as described above.
[0077] The present invention also discloses a computer-readable medium having a computer-executable program stored thereon, wherein when the computer-executable program is executed, the high-quality data mining method described above is implemented.
[0078] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0079] Figure 3 This is an overall flow chart of a high-quality data mining method according to an embodiment of the present invention. Figure 3 As shown, the process of this mining method is mainly divided into three parts: 1. Data integration and analysis; 2. Data production; 3. Question production.
[0080] 1. Data integration and analysis
[0081] Analyze the question data in the question bank and select the hot questions for production.
[0082] (1) Find hot knowledge points and hot clusters based on the question information: the question's page view volume (PV), the question's auxiliary information (subject section, knowledge point, quality score, source, etc.), and the question's similar clusters; among them, find high PV data based on the question's page views (PV), and find the corresponding hot clusters based on the question information; find the corresponding hot knowledge points based on the question's auxiliary information, and technically use a simple mapping relationship to achieve this.
[0083] (2) According to the auxiliary information of the questions (school stage, knowledge points, quality score, source, etc.), match the current time information: the school year, school quarter, key examinations, key subject progress, and then infer the hot knowledge points and hot clusters; specifically, by searching the above information in the question bank, the matching content can be obtained, and the corresponding knowledge points can be obtained. Then, according to the knowledge points, the corresponding key questions can be found, and then the hot clusters can be found.
[0084] (3) Based on the user query, relevant knowledge points are mined and hot knowledge points are found; for example, the corresponding knowledge point data is matched based on the relevant auxiliary information of the query.
[0085] (4) Reverse search to obtain hot question sets based on hot knowledge points and hot clusters; for example, reverse search is also achieved through mapping relationships.
[0086] (5) Based on the PV, quality, and richness of the questions, data is mined and added to the hot question set. This is essentially similar to the filtering strategy used in data production.
[0087] 2. Data production
[0088] In the hot topic set, the questions are scored based on their same clusters and completeness and richness information. Some high-quality data do not need to be put into production; some low-quality data do not need to be put into production; and the remaining data will be put into production based on the scores.
[0089] 3. Question production
[0090] According to the scores of the hot question set, the data is divided into A / B grades, among which A grade data is the data that must be produced, and B grade data is the high-priority production data.
[0091] (1) A / B file data flow
[0092] For hot topic data, the timeliness will be marked according to the source of the data generation (that is, the production deadline will be determined according to the threshold). During the data production process, if it is found that the timeliness has been exceeded, the question will be re-determined to resubmit it to a production with lower timeliness or abandoned.
[0093] (2) A / B file data production
[0094] ① Each link in the question bank production has the corresponding ability to adjust priority and price. A / B file data are both higher priority data, but the production pricing is not necessarily higher than that of ordinary production questions.
[0095] ② The A / B file data will be dynamically calculated based on the current question bank capacity, the current A / B file data volume, and the estimated task completion status. If capacity is insufficient, the priority of the item will be increased. When the priority is high enough (higher than other production data, increasing the priority will not improve production efficiency), the production price of that link will be increased (with an upper limit). If capacity is still insufficient, the priority and pricing of the B file will be lowered (with a lower limit, higher than the average question).
[0096] ③ To ensure the production of A-file data, a repeated release mode is adopted for A-file data, so that more teachers can produce the question at the same time. After one teacher completes the production, other teachers can manually choose to give up the question and obtain part of the profit.
[0097] (3) Data feedback
[0098] After the production, the questions will be put online, and the page views (PV) after the launch will serve as the basis for the next round of data mining.
[0099] The A / B-tier production method of the present invention is the core content of the present invention. Compared with the common production and production methods (submitting topics, setting priorities and pricing, and then producing), A / B-tier production is more suitable for the production of high-quality topics. The main reasons are as follows:
[0100] ① Automatic adjustment rules for priority and pricing can better concentrate crowdsourcing user resources and ensure the production of important data.
[0101] ② For data on estimated production, analyze its value curve, which differs from the typical gradual decline or cyclical rise and fall curves; the curve should be extremely high before the deadline and zero after the deadline. For this type of data, an A / B production flow strategy can better ensure production return on investment (ROI).
[0102] ③ The repeated production strategy is an important guarantee for high-quality data to be launched on schedule, and it is also a capability that ordinary crowdsourcing production does not have.
[0103] Figure 4 1 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes the high-quality data mining method described above.
[0104] like Figure 4 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.
[0105] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.
[0106] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).
[0107] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0108] It should be understood that Figure 4 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.
[0109] Figure 5 Schematic diagram of a computer-readable recording medium according to this embodiment of the present invention. Figure 5As shown, a computer-readable recording medium stores a computer-executable program, which, when executed, implements the high-quality data mining method described above in the present invention. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0110] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Python, Java, C++, and the like, as well as conventional procedural programming languages such as C or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0111] Through the above description of the embodiments, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by a vehicle containing at least a part of the above-mentioned system or components. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, electronic control unit, client, server, etc. on the locomotive side. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity. It can also be implemented in a distributed manner by unspecified hardware. For example, some method steps executed by the computer program can be executed on the locomotive side, and another part can be executed in a mobile terminal or smart helmet, etc. For computer software, the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), or it can be distributed and stored on a network, as long as it can enable electronic devices to execute the method according to the present invention.
[0112] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A high-quality data mining method, characterized in that: The steps include: High-quality data includes high-quality test question data, which are test questions that are frequently requested by customers, are in high demand, and are relevant to the knowledge points that customers are most concerned about; Based on periodic characteristics, and / or, based on the same cluster or similar cluster of the current hot spot, the existing high-quality test questions are estimated in advance and the relevant data information of the high-quality test questions is obtained to give priority to the production of corresponding questions and answers; wherein, the production prediction of questions is made according to the periodic knowledge points and learning content, and the hot similar clusters and hot knowledge points found according to the current hot words are reversely deduced to obtain the questions to be produced; the hot knowledge points and hot similar clusters are found through the question information; the same cluster or similar cluster of the current hot spot is obtained based on the auxiliary information of the question, matching the current time information, searching in the question bank and obtaining the search matching content, or based on the page browsing of the question The method is to obtain the hot clusters based on the page views, auxiliary information of the questions and / or similar clusters of the questions, specifically including: matching the current time information with the auxiliary information of the questions, retrieving the matching content in the question bank to obtain the corresponding knowledge points, and finding the corresponding key questions based on the knowledge points to find the corresponding hot clusters, or finding the high page views data based on the page views of the questions, finding the corresponding hot clusters based on the question information, and finding the corresponding hot knowledge points based on the auxiliary information of the questions; scoring the questions based on the same clusters and completeness and richness information of the questions, where some high-quality data do not need to be put into production; some low-quality data do not need to be put into production; and the remaining data are put into production based on the scores; For those that require production, data is divided into A and B categories based on the scores of the hot question sets. Category A data is required for production, and category B data is high-priority production data. Each link in the question bank production process has the corresponding ability to adjust priorities and prices. The A / B file data will be dynamically calculated based on the current question bank capacity and the current A / B file data volume to estimate the task completion status; if the capacity is insufficient, the priority will be increased; if the capacity is still insufficient, the priority and pricing of the B file will be lowered; to ensure the production of A file data, a repeated delivery model is adopted for A file data, allowing more crowdsourcing system task acceptors to produce the question at the same time, and after a crowdsourcing system task acceptor completes the production, others can manually choose to give up the question and obtain part of the profit; After the produced questions are transferred to the question bank, they are sorted according to the page views of the questions in the question bank and the top ranked ones are used as the basis for the next round of data mining.
2. The method according to claim 1, characterized in that The periodic feature is determined based on the semester and school year information of the knowledge point, the periodic correlation between the search questions and the student status information of the searcher, and the timing and frequency of the appearance of the search questions.
3. The method according to claim 1, characterized in that The current hotspot is the relevant knowledge point information mapped by the system after it retrieves matching content from the question bank based on the relevant auxiliary information of the search topic.
4. The method according to claim 1, wherein Auxiliary information of the question includes: subject section, knowledge point, quality score, and source.
5. The method according to claim 1, wherein Matching current time information includes: academic year, quarter, key exams, and key subject progress information.
6. A method for producing high-quality test questions, characterized in that: The following steps are involved: Using the high-quality data mining method according to any one of claims 1 to 5 to estimate relevant data information of high-quality test questions; Evaluate the relevant data and information of the obtained high-quality test questions to determine whether further expansion of production is necessary; If necessary, the questions in the hot question set will be scored based on the same cluster, completeness and / or richness information of the questions, and divided into A-level must-produce data and B-level high-priority production data according to the scores. Production will be organized by calculating and estimating the task completion status based on the current question bank production capacity, the current data volume of A-level must-produce data and B-level high-priority production data.
7. The method according to claim 6, characterized in that For those that need further expansion of production, the timeliness of the questions needs to be marked according to the source of their data generation, and they are divided into high-timeliness questions and low-timeliness questions.
8. The method according to claim 7, characterized in that During the data production process, if it is found that a high-timeliness test question exceeds the timeliness, it will be re-determined whether the question needs to be resubmitted to low-timeliness production or abandoned.
9. The method according to claim 6, characterized in that Also includes: For those that need to be organized for production, they are distributed to test question production personnel in a crowdsourcing manner for production.
Citation Information
Patent Citations
Evaluation method and device for question bank
CN106780204A
Test question intelligent pushing method and system
CN109002564A
Crowdsourcing topic recommendation method, device and apparatus and storage medium
CN111353015A