Data organization method, device and electronic equipment for question bank
By clustering questions in the question bank, the problem of low efficiency in processing questions with the same content is solved, resulting in more efficient question management and more accurate search results, thus improving the user experience.
Patent Information
- Application Number
- CN202110588181.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-05-27
AI Technical Summary
The existing question bank suffers from low efficiency in processing questions with identical content.
By clustering the questions in the question bank to form circular or star-shaped clusters, questions with the same content are grouped into a cluster, and the clusters and their correspondence with the questions are stored. Subsequent question bank management and data services are based on the clusters.
It improved the efficiency of handling identical questions in the question bank, reduced the cost of deduplication, improved the accuracy and recall rate of user question searches, and enhanced the user experience.
Smart Images

Figure CN113297381B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of the Internet, and is particularly suitable for online education technology of the Internet, and more particularly relates to a data organization method and device for a question bank, an electronic device, and a computer readable medium. BACKGROUND
[0002] At present, more and more products for searching questions by taking photos appear on the market. The core competitiveness of such products mainly lies in the accuracy of question searching. The higher the accuracy is, the more similar the searched questions are to the original questions, which is beneficial to improving user experience and increasing user stickiness.
[0003] The function of searching questions by taking photos is realized based on a question bank. At present, there are many repeated questions in the question bank. For data management and data services such as retrieval, there is potential for optimization. SUMMARY
[0004] (I) Technical problem to be solved
[0005] The present application aims to solve the technical problem of low processing efficiency of questions with the same content in the existing question bank.
[0006] (II) Technical solution
[0007] To solve the above technical problem, one aspect of the present application proposes a data organization method for a question bank, which comprises the following steps:
[0008] selecting questions with the same content from the question bank;
[0009] grouping each question with the same content into a cluster, storing each cluster and the corresponding relationship thereof with the question, and subsequently managing the question bank based on the cluster and providing data services externally.
[0010] According to a preferred embodiment of the present application, the questions with the same content are selected from the question bank, comprising:
[0011] performing clustering processing on the questions in the question bank to form a plurality of clusters, each cluster comprising one or more questions with the same content.
[0012] According to a preferred embodiment of the present application, the cluster structure is a ring-shaped cluster, the ring-shaped cluster comprising questions with the same content in pairs, and the questions with the same content in pairs having a corresponding relationship.
[0013] According to a preferred embodiment of the present invention, the cluster structure is a star cluster, which includes a virtual topic as a cluster head and various topics with the same content as cluster members. The virtual topic is selected from the various topics that are members of the cluster, or generated from the various topics that are members of the cluster. There is a correspondence between the cluster head and the various topics that are members of the cluster.
[0014] According to a preferred embodiment of the present invention, the virtual title serving as the cluster head is formed in one of the following ways:
[0015] The virtual question that serves as the cluster head is the highest quality among all questions with the same content located within the corresponding cluster.
[0016] The virtual question that serves as the cluster head includes multiple fields, each of which is the highest quality selected from the corresponding fields of all questions with the same content within the same cluster.
[0017] The virtual question that serves as the cluster head is the question with the highest confidence among all questions within the corresponding cluster. The question confidence is configured based on question attributes. Optionally, the question attributes include at least one of the following: the format of the question field, the source of the question, and the number of times the question has been searched.
[0018] According to a preferred embodiment of the present invention, the cluster head is obtained in the following manner:
[0019] Retrieve the specified fields that make up the virtual question;
[0020] Vote for the specified field of the virtual question from each specified field of all questions within the cluster;
[0021] A virtual question is constructed based on each specified field, and this virtual question is used as the cluster head; the cluster head has the same content as all questions in the corresponding cluster, and there is a corresponding relationship between them.
[0022] According to a preferred embodiment of the present invention, the step of selecting the designated field as the virtual question from each designated field of all questions within the cluster includes:
[0023] Configure the voting criteria for each specified field;
[0024] The specified field that constitutes the virtual question is selected based on the voting criteria corresponding to each specified field and the specified field itself;
[0025] Optionally, when the specified field is one or more of subject, grade level, and answer, the corresponding voting criteria are configured as follows: the subject, grade level, and answer corresponding to the question with the highest question confidence weighting value in the star cluster are used as the subject, grade level, and answer of the cluster head;
[0026] When the specified field is one or more of the answer, explanation, and key points, the corresponding voting criterion is configured as follows: the specified field with the highest confidence level among the specified fields of each question in the star cluster is used as the specified field of the cluster head;
[0027] When the specified field is one or more of the following: knowledge point, test paper information, book information and test question year, the corresponding voting criteria are configured as follows: configure the confidence threshold of each specified field, and concatenate the specified fields in each question in the star cluster whose confidence is higher than the corresponding confidence threshold, and use them as the specified field of the cluster head.
[0028] When the specified field is the question stem, the corresponding voting criterion is configured as follows: after deduplication of the specified field of the questions in the star cluster, they are merged and used as the question stem content of the cluster head.
[0029] According to a preferred embodiment of the present invention, when the questions in a cluster change or new questions are received, the method further includes: comparing the changed or newly added questions with the questions in each cluster to assign the changed or newly added questions to clusters.
[0030] Optionally, the method further includes: periodically detecting whether the titles in each cluster have changed;
[0031] Optionally, when the cluster is a star-shaped cluster, the process of determining duplicates between the changed or newly added questions and the questions within each cluster, and then assigning the changed or newly added questions to clusters, includes:
[0032] The questions that have changed or been newly added to the database will be checked for duplicates against the cluster heads of each star-shaped cluster.
[0033] If the deduplication result is that there are no overlaps, a new cluster is constructed for the changed or newly added questions, and the changed question is deleted from the star cluster in which the changed question was located before the change; if the deduplication result is that there are overlapping star clusters, the changed or newly added questions are added to the star cluster in which the overlapping cluster head is located, and the cluster head is updated.
[0034] Optionally, if the deduplication result indicates that there are two or more overlapping clusters, a cluster fusion step is initiated, wherein the cluster fusion step is used to merge the two or more overlapping clusters into one cluster;
[0035] Optionally, when the cluster is a star-shaped cluster, the cluster fusion step includes: selecting the target cluster head corresponding to the star-shaped cluster containing the most questions as the first target cluster head; adding the new questions to the star-shaped cluster where the first target cluster head is located; and adding all the questions from the star-shaped clusters where the other target cluster heads are located to the star-shaped cluster where the first target cluster head is located.
[0036] Optionally, a new cluster is constructed for questions that have changed or are newly added to the database, including: constructing a cluster head based on a specified field of the question that has changed or is newly added to the database, wherein the new question and the constructed cluster head form a new star cluster.
[0037] According to a preferred embodiment of the present invention, the method further includes:
[0038] If an unstable problem is detected in a star cluster where the correspondence with the cluster head has changed, the unstable problem is removed from the star cluster.
[0039] Optionally, after deleting the unstable question from the star cluster, the method further includes: constructing a new cluster head based on a specified field of the unstable question; comparing all other questions in the star cluster except the unstable question with the constructed new cluster head; if a question identical to the new cluster head is found, deleting the question from the star cluster and merging it with the new cluster head and the unstable question to form a new star cluster.
[0040] Optionally, if a cluster head in a star cluster is detected to not meet the cluster head requirements, the cluster head is deleted from the star cluster, and all questions in the star cluster corresponding to the cluster head are removed; the removed questions are compared with the cluster heads of other star clusters; and the removed questions are clustered according to the comparison results.
[0041] Optionally, if the number of questions in the star cluster changes, the method further includes: voting on a specified field of the cluster head for each specified field of all questions in the star cluster where the number of questions has changed, in order to update the cluster head.
[0042] According to a preferred embodiment of the present invention, the method further includes:
[0043] Receive the questions that the user wants to search for;
[0044] The search query is matched with the cluster heads of each star-shaped cluster;
[0045] Send at least one question from the star cluster containing at least one matched cluster head to the user;
[0046] Optionally, a virtual question for at least one matched cluster head is sent to the user.
[0047] A second aspect of the present invention provides a data organization apparatus for a question bank, the apparatus comprising:
[0048] The filtering module is used to filter out questions with the same content from the question bank;
[0049] The cluster module is used to group questions with the same content into a cluster, store each cluster and its correspondence with the questions, and then manage the question bank and provide data services to external parties based on the clusters.
[0050] A third aspect of the present invention provides a question bank, the question bank comprising:
[0051] A question table is used to store information for each question.
[0052] A cluster relationship table stores cluster information and the correspondence between each cluster and the questions.
[0053] A cluster status table is used to store the status of each cluster, which includes deleted, normal, and pending modification.
[0054] The management module manages the question bank and provides data services to external parties by querying the cluster table and the cluster status table based on the clusters.
[0055] A fourth aspect of the present invention provides an electronic device including a processor and a memory, the memory being used to store a computer-executable program, wherein when the computer program is executed by the processor, the processor performs the method described in any of the preceding claims.
[0056] The fifth aspect of the present invention also provides a computer-readable medium storing a computer-executable program, which, when executed, implements the method described in any of the preceding claims.
[0057] (III) Beneficial Effects
[0058] This invention filters out questions with identical content from a question bank; it defines each question with identical content as a cluster, stores the correspondence between each cluster and its corresponding questions, and performs subsequent question bank management and external data service provision based on these clusters. This improves the processing efficiency of questions with identical content in the question bank and facilitates question management. The cluster structure can be, for example, a ring cluster or a star cluster.
[0059] In this invention, the cluster head of the star-shaped cluster is determined to be identical to all the questions, indicating a correspondence. If a question in the star-shaped cluster changes, the correspondence between the cluster head and other questions remains unchanged; that is, the correspondence between the cluster head and other questions is still reliable. Only deduplication between the changed question and the cluster head of the star-shaped cluster is needed, and the star-shaped cluster can be modified based on the deduplication result. Modifying the star-shaped cluster can reduce the number of deduplication checks to 2^n, effectively reducing deduplication costs, supporting cluster changes, thereby improving the accuracy of question searches and enhancing the user experience.
[0060] This invention adjusts the cluster head of a star cluster based on specified fields of all questions in the changed star cluster and the voting criteria corresponding to those specified fields after the number of questions in the star cluster changes. This achieves dynamic cluster balance and ensures the reliability of the correspondence between the cluster head and each question in the star cluster after each change, thereby further improving the accuracy of deduplication.
[0061] This invention matches the user's search query with the cluster heads of various star-shaped clusters; it then sends the query and answer corresponding to at least one matched cluster head to the user, thereby improving the accuracy and recall rate of the user's search and enhancing the user experience. Attached Figure Description
[0062] Figure 1 This is a flowchart illustrating a data organization method for a question bank according to the present invention;
[0063] Figures 2a-2b This is a schematic diagram of the structure of a ring-shaped cluster;
[0064] Figure 3 This is a schematic diagram of a star-shaped cluster structure according to the present invention;
[0065] Figures 4a-4i This is a schematic diagram illustrating the modification of the star-shaped cluster in this invention;
[0066] Figure 5 This is a schematic diagram of the structure of a question bank data organization device according to the present invention;
[0067] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention;
[0068] Figure 7 This is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. Detailed Implementation
[0069] In the description of specific embodiments, detailed descriptions of structures, performance, effects, or other features are provided to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from implementing the present invention with technical solutions that do not contain the aforementioned structures, performance, effects, or other features under specific circumstances.
[0070] The flowchart in the accompanying drawings is merely an exemplary process demonstration and does not imply that the solution of this invention must include all the content, operations, and steps in the flowchart, nor does it imply that they must be executed in the order shown in the diagram. For example, some operations / steps in the flowchart can be decomposed, some operations / steps can be combined or partially combined, etc. Without departing from the inventive spirit of this invention, the execution order shown in the flowchart can be changed according to the actual situation.
[0071] The box in the attached diagram Figure 1 Generally, these refer to functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processing unit devices and / or microcontroller devices.
[0072] The same reference numerals in the accompanying drawings denote the same or similar elements, components, or parts, and therefore, repeated descriptions of the same or similar elements, components, or parts may be omitted below. It should also be understood that although terms such as first, second, third, etc., indicating numbers may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. That is, these terms are only used to distinguish one from another. For example, a first device may also be referred to as a second device, without departing from the essential technical solution of the invention. Furthermore, the terms "and / or" and "and / or" refer to all combinations including any one or more of the listed items.
[0073] To address the aforementioned technical problems, this invention proposes a data organization method for a question bank, which includes filtering out questions with identical content; grouping questions with identical content into a cluster; storing the correspondence between each cluster and its corresponding questions; and performing subsequent question bank management and providing data services to external parties based on the clusters, thereby improving the processing efficiency of questions with identical content in the question bank and facilitating question management.
[0074] Optionally, the questions in the question bank can be clustered to form multiple clusters, each cluster including one or more questions with the same content.
[0075] The cluster structure of the present invention can be a ring cluster, wherein the ring cluster includes pairs of identical topics, and the pairs of identical topics have a corresponding relationship.
[0076] The cluster structure of this invention can also be a star-shaped cluster. A star-shaped cluster includes a virtual question as the cluster head and various questions with identical content as cluster members. The virtual question is selected from or generated from the questions that are members within the cluster. The cluster head and the questions that are members within the cluster all have the same deduplication result, meaning there is a correspondence. If a question in the question bank changes, the correspondence between the cluster head and other questions remains unchanged; that is, the correspondence between the cluster head and other questions is still reliable. When changing a cluster, it is not necessary to re-deduplicate the cluster head with other questions. Only the changed question needs to be deduplicated with the cluster head, and the star-shaped cluster can be changed based on the deduplication result. Changing the star-shaped cluster can reduce the number of deduplication checks to 2^n, effectively reducing deduplication costs, supporting cluster changes, thereby improving the accuracy of user question searches and enhancing user experience.
[0077] During the cluster change process, the changed or newly added questions will be compared with the questions in each cluster to determine duplicates, so that the changed or newly added questions can be reassigned to clusters.
[0078] It should be noted that changes to a question include changes to the substantive content of the question (i.e., the question stem), as well as changes to other non-substantive information such as the source and format of the question.
[0079] In some embodiments, the deduplication step is triggered only when the question undergoes a substantial change in content, and the question with the substantial change is reassigned to a new cluster. If the question undergoes a non-substantial change in content, deduplication is not triggered, cluster reassignment is not performed, and only the cluster head is updated. Furthermore, for star clusters, the cluster head update process includes: re-voting the cluster head of the star cluster based on specified fields of all questions in the star cluster and the voting criteria corresponding to the specified fields, and updating the cluster head to achieve dynamic cluster balance. This ensures the reliability of the correspondence between the cluster head and the questions in the star cluster after each star cluster change, further improving the accuracy of deduplication.
[0080] The cluster referred to in this invention is a data organization structure for a problem.
[0081] The term "duplication judgment" in this invention refers to determining whether two questions are substantially the same.
[0082] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0083] Figure 1 This is a flowchart illustrating a data organization method for a question bank according to the present invention, as shown below. Figure 2a As shown, the method includes the following steps:
[0084] S1. Select questions with the same content from the question bank;
[0085] In this context, "same content" refers to the question stem being identical. As long as the substantive content of the questions is the same, regardless of whether other information about the questions (such as the source of the questions, the format of the questions, etc.) is the same, the questions are judged to be of the same content.
[0086] In this embodiment of the invention, the questions in the question bank can be clustered to form multiple clusters, each cluster including one or more questions with the same content. Various implementation methods exist, and this embodiment is not limited to any one. For example, existing machine learning clustering algorithms or existing question deduplication algorithms can be used to filter out questions with the same content.
[0087] S2. Group questions with the same content into a cluster, store each cluster and its correspondence with the questions, and then manage the question bank and provide data services to external parties based on the clusters.
[0088] In this invention, the cluster structure can be varied. In one cluster structure, to improve the accuracy of question search, questions in the question bank are typically pairwise checked for duplicates, and then questions with the same content are organized into ring clusters to ensure that all questions in the same ring cluster are identical. During the question search process, only the original question searched by the user needs to be checked for duplicates with one question in the ring cluster, and the identical questions are output based on the deduplication result. In this way, it is not necessary to check the original question searched by the user against all questions in the question bank individually; on the one hand, this reduces the computational load of deduplication; on the other hand, the questions output by deduplication through ring clusters come from different ring clusters, which can effectively avoid the situation where the output questions in the individual deduplication process may come from the same ring cluster, thereby improving the recall rate of question search.
[0089] The circular clusters consist of pairs of identical topics, indicating a correspondence between them. Each topic can be a member of the cluster, and connecting two duplicate results with a line segment represents the same topic, meaning there is a correspondence between the two topics.
[0090] Figure 2a The diagram shows the structure of a ring cluster. Problems B, C, and F are paired with Problem A and their results are identical. Problems E and F are also paired and their results are identical. Problem D and Problem E are paired and their results are identical. Since all problems are paired and their results are identical, it can be deduced that problems A and B are identical. Therefore, problems A and B form the following structure: Figure 2b The diagram shows a ring-shaped cluster. Each question can be a member of the cluster, and two questions with the same deduplication result are connected by line segments, indicating a correspondence between the two questions.
[0091] In practice, the relationships between questions may change, new questions may be added, or existing questions may be deleted. All of these changes can cause cluster imbalance, requiring specific strategies to restore balance. Clustering, splitting, and merging of questions are routine and frequent operations. For example... Figure 3 As shown, if the correspondence between question A and question F changes, the correspondence between all questions in the annular cluster becomes unreliable, requiring modification operations such as deduplication, re-clustering, splitting, or merging clusters. In this type of annular cluster, re-clustering requires n-squared, i.e., O(n^2)^2. 2 A reliable cluster is established by re-judging the number of questions, where n is the number of questions in the re-cluster.
[0092] To further reduce deduplication costs and support cluster changes, this embodiment of the invention also provides a star-shaped cluster structure. The star-shaped cluster includes a virtual title serving as the cluster head and individual titles with identical content serving as cluster members. The virtual title is selected from or generated from the individual titles of the cluster members, and there is a correspondence between the cluster head and the individual titles within the cluster.
[0093] For example, the virtual title serving as the cluster head can be formed in one of the following ways:
[0094] The virtual question that serves as the cluster head is the highest quality among all questions with the same content located within the corresponding cluster.
[0095] The virtual question that serves as the cluster head includes multiple fields, each of which is the highest quality selected from the corresponding fields of all questions with the same content within the same cluster.
[0096] The virtual question that serves as the cluster head is the question with the highest confidence among all questions within the corresponding cluster. The question confidence is configured based on question attributes. Optionally, the question attributes include at least one of the following: the format of the question field, the source of the question, and the number of times the question has been searched.
[0097] For example, the format of the question field includes: editable format and image format. The confidence level of questions in editable format can be configured to be higher than that of questions in image format. The question source includes: past exam questions and test questions. The confidence level of questions in past exam questions can be configured to be higher than that of test questions. Furthermore, the question confidence level can be configured based on the number of times a question is searched. If the question confidence level is configured based on multiple question attributes, the weight of each question attribute can be pre-set, and then the question confidence level can be determined based on the confidence level of the question in each question attribute and the weight of that question attribute.
[0098] Figure 3 Here is an example of a star-shaped cluster structure, such asFigure 4a As shown, the star-shaped cluster includes a cluster head X and various questions A to F with identical content. The cluster head X is determined to be identical to all questions A to F within the cluster, and there is a corresponding relationship between them. Each question can be a member of the cluster, and two questions with the same deduplication result are connected by a line segment, forming a correspondence between the two questions. The virtual question serving as the cluster head X includes multiple fields, each of which is the highest quality selected from the corresponding fields of the questions with identical content within the corresponding cluster.
[0099] In one example, the cluster head can be obtained in the following way:
[0100] S21. Obtain the specified fields that make up the virtual question;
[0101] A question consists of multiple fields. The fields that make up each question within a cluster are also different. The specified fields that make up a virtual question that represents a cluster can be preset. Those skilled in the art can set the specific fields that a virtual question includes according to actual requirements.
[0102] The specified fields divide the virtual question from multiple dimensions, meaning that the specified fields constitute the entire question. For example, the specified fields may include: question stem, answer, explanation, key points, subject, grade level, knowledge points, test paper information, book information, and test year, etc.
[0103] S22. Vote for each specified field of all questions within the cluster as the specified field of the virtual question;
[0104] For example, voting criteria corresponding to each specified field can be pre-configured; then, the specified field that makes up the virtual question can be selected based on the voting criteria corresponding to each specified field and the specified field itself.
[0105] In one example of voting for a specified field, different voting criteria can be set for different specified fields.
[0106] When the specified field is one or more of subject, grade level, and answer, the corresponding voting criteria are configured as follows: the subject, grade level, and answer corresponding to the question with the highest confidence weighting value in the star cluster are selected as the subject, grade level, and answer of the cluster head; for example, for subject, the subject content corresponding to the question with the highest confidence weighting value built into the star cluster can be selected as the subject content of the cluster head.
[0107] When the specified field is one or more of the answer, explanation, and key points, the corresponding voting criterion is configured as follows: the specified field with the highest confidence level among all questions in the star cluster is selected as the specified field of the cluster head; wherein, the confidence level of the specified field can be pre-configured. For example, for the answer, the answer content corresponding to the question with the highest answer confidence level in the star cluster can be selected as the answer content of the cluster head.
[0108] When the specified field is one or more of the following: knowledge point, exam paper information, book information, and exam year, the corresponding voting criteria are configured as follows: configure a confidence threshold for each specified field, and concatenate the specified fields in each question of the star cluster whose confidence is higher than the corresponding confidence threshold, and use this as the specified field of the cluster head; for example, for knowledge points, a confidence threshold for knowledge points can be preset, all questions with a confidence level higher than the knowledge point confidence threshold can be selected, and the content of all knowledge points in these questions can be concatenated as the knowledge point content of the cluster head. Preferably, before concatenating the specified fields, the specified fields of each question can be deduplicated to remove duplicate content in the specified fields.
[0109] When the specified field is the question stem, the corresponding voting criterion is configured as follows: after deduplication of the specified field of the questions in the star cluster, they are merged and used as the question stem content of the cluster head.
[0110] In some specific embodiments, the cluster head generation schemes include: 1) Voting scheme: selecting the highest vote based on the confidence level of each question, applicable to specified fields: subject, grade level, multiple choice answer; 2) Highest confidence level-based scheme: using the question with the highest confidence level in the field as the standard; applicable to specified fields: answer, explanation, key points, etc.; 3) Concatenation scheme: concatenating data with a confidence level higher than the threshold, applicable to specified fields: knowledge points, test paper information, book information, test question year; 4) Integration scheme: deduplicating and integrating the information of this field of all questions into multiple sets of data, applicable to specified field: question.
[0111] S23. Construct a virtual question based on each specified field, and use the virtual question as the cluster head;
[0112] The cluster head is identical to all the questions within the corresponding cluster, and there is a corresponding relationship between them.
[0113] In another example, the star-shaped cluster includes a virtual question as the cluster head and various questions with identical content. The cluster head is determined to be identical to all questions within the cluster, and there is a corresponding relationship between them. Each question can be an element or member of the cluster. The virtual question as the cluster head can be the question with the highest confidence among all questions within the corresponding cluster. The question confidence can be pre-configured, and the actual question with the highest confidence is selected as the cluster head through voting.
[0114] In the star-shaped clustering of this invention, the cluster head and all questions have the same deduplication result, meaning a correspondence exists. If a question in the question bank changes, the correspondence between the cluster head and other questions remains unchanged; that is, the correspondence between the cluster head and other questions is still reliable. Therefore, when changing a cluster, it is unnecessary to re-deduplicatively check the cluster head against other questions. Only the changed question needs to be deduplicated against the cluster head, and the star-shaped cluster can be modified based on the deduplication result. Changing the star-shaped cluster can reduce the number of deduplication checks to 2^n, effectively reducing deduplication costs, supporting cluster changes, thereby improving the accuracy of question searches and enhancing the user experience.
[0115] Furthermore, this invention also supports cluster changes. For example, adding a new question to the question bank, changes in questions within a cluster, or changes in the correspondence between questions within a cluster will all cause cluster changes. Specific examples may include adding a new question or removing a question.
[0116] When a question in a cluster changes or a new question is received, the method further includes: comparing the changed or newly added question with the questions in each cluster to determine if they are duplicates, so as to reallocate the changed or newly added question to a new cluster.
[0117] For example, if a changed or newly added question overlaps with a question in a certain cluster, the changed or newly added question is added to that cluster; if the deduplication result is that there are two or more overlapping clusters, the cluster fusion step is initiated, which merges the two or more overlapping clusters into one cluster.
[0118] If a question changes or is newly added to the question bank and does not overlap with any questions in any cluster in the question bank, then a new cluster is created for the question that has changed or been newly added to the question bank.
[0119] This invention can automatically trigger deduplication detection when questions change or new questions are added to the database. Optionally, it can also actively detect whether questions in each star-shaped cluster have changed to trigger deduplication detection. Therefore, the method further includes: periodically detecting whether questions in each cluster have changed.
[0120] The following uses a star-shaped cluster as an example to illustrate the cluster modification of the present invention. The method further includes the following steps:
[0121] S101. When the questions in a star cluster change, the changed questions are compared with the cluster heads of each star cluster for duplicates.
[0122] The change in the question refers to a change in the substantive content of the question. Other changes besides the substantive content, such as changes in the source or format of the question, will not trigger deduplication. During the deduplication process, the changed question is checked for deduplication with the cluster head of its star-shaped cluster and the cluster heads of other star-shaped clusters in the question bank. The deduplication results determine which cluster head the changed question shares content with.
[0123] S102. If the clusters do not overlap, the changed questions are reassigned to new clusters, and the changed questions are deleted from the star clusters. If the clusters overlap, the changed questions are added to the star clusters where the overlapping cluster heads are located.
[0124] Among them, the reassignment of clusters for changed questions refers to constructing a cluster head based on the specified fields of the changed questions, and generating a new star cluster by the cluster head and the changed questions.
[0125] If there is overlap, the modified problem will be added to the star cluster containing the overlapping cluster head. In a special case, if a modified problem overlaps with the cluster head of its star cluster, no changes will be made.
[0126] Optionally, when receiving newly added questions, the method further includes:
[0127] S201. Determine if the newly added questions are duplicated by comparing them with the cluster heads of each star cluster.
[0128] S202. If there is a target cluster head with the same content as the newly added question, add the newly added question to the star cluster where the target cluster head is located.
[0129] In one example, if the number of target cluster heads that are identical to the content of the newly added question is greater than or equal to 2, the method further includes:
[0130] Perform a star-shaped cluster fusion step that merges all target cluster heads that have the same content as the newly added questions.
[0131] The star cluster fusion step includes: selecting a first target cluster head from the at least two target cluster heads; adding a new topic to the first star cluster where the first target cluster head is located; and adding all topics from the star clusters where the other target cluster heads are located to the first star cluster.
[0132] For example, when a new question is received, it is compared with the cluster heads of each star-shaped cluster for deduplication; for example... Figure 4b The new problem Q is compared with the cluster heads X and Y of the two existing star clusters for deduplication, and the star clusters are modified according to the deduplication results.
[0133] likeFigures 4c-4d As shown, if a target cluster head X1 with the same deduplication result as the new question Q is found, the new question Q is added to the star cluster where the target cluster head X1 is located. If the number of target cluster heads with the same content as the newly added question is greater than or equal to 2, that is, at least two target cluster heads with the same content as the new question are found, the star cluster fusion step of merging the target cluster heads with the same content as the newly added question is executed.
[0134] In an exemplary star-shaped cluster fusion method, a first target cluster head is first selected from the at least two target cluster heads. The first target cluster head can be selected randomly, or the target cluster head corresponding to the star-shaped cluster containing the most questions can be used as the first target cluster head. The specific selection method is not limited in this invention, but the latter involves less computation. Then, the new questions are added to the first star-shaped cluster containing the first target cluster head. Finally, all questions from the star-shaped clusters containing the other target cluster heads are added to the first star-shaped cluster. See also... Figure 4c In target cluster heads X2 and X3, if the star cluster containing X2 contains more problems, then... Figure 4d As shown, X2 is taken as the first target cluster head, and the new question Q is added to the star cluster where X2 is located, i.e., the first star cluster. Then, as shown... Figure 4e As shown, all the questions in the star cluster containing X3 are added to the star cluster containing X2 to complete the cluster fusion.
[0135] S203. If there is no target cluster head with the same content as the newly added question, construct a cluster head according to the specified field of the new question, and the new question and the constructed cluster head form a new star cluster.
[0136] The method for constructing cluster heads and new star-shaped clusters is described in step S2 and its related detailed steps, which will not be repeated here.
[0137] Optionally, if an unstable question is detected in the star cluster where the correspondence between the question and the cluster head has changed, that is, the duplicate judgment result between the question and the cluster head changes from the same to different, the correspondence between the question and the cluster head can be detected by mining through preset strategies or based on user reports.
[0138] Furthermore, the method also includes:
[0139] Construct a new cluster head based on the specified fields of the unstable problem;
[0140] Each problem in the star cluster, except for the unstable problem, is compared with the new cluster head for deduplication.
[0141] If a topic identical to the new cluster head is found, that topic is removed from the star cluster and merged with the new cluster head and the unstable topic to form a new star cluster.
[0142] For example, in a star-shaped cluster, the correspondence between the title and the cluster head changes, such as... Figure 4f If an unstable problem S is detected in the star cluster where the correspondence with the cluster head X has changed, the unstable problem S is deleted from the star cluster. A new cluster head Y can be constructed based on specified fields of the unstable problem S; furthermore, such as Figure 4g All problems A through D in the star-shaped cluster, except for the unstable problem S, can be compared with the new cluster head Y for duplicates. If a problem identical to the new cluster head Y is found, for example... Figure 4h Problem B is removed from the star cluster and merged with the new cluster head Y and the unstable problem S to form a new star cluster.
[0143] Optionally, if a cluster head in a star cluster is detected to not meet the cluster head requirements, the cluster head will be removed from the star cluster. The cluster head requirements refer to the fact that the cluster head cannot be judged as an incorrect question, a question with political risk, or a question that violates public order and good morals.
[0144] Optionally, if a star cluster is detected as not meeting the cluster requirements, all questions in the star cluster are removed, and the cluster head is deleted. When all questions and cluster heads within a star cluster are determined to be incorrect, politically risky, or contrary to public order and good morals, the star cluster is deemed not to meet the cluster requirements.
[0145] Furthermore, the method also includes:
[0146] The removed questions will be considered duplicates of the cluster heads of other star-shaped clusters;
[0147] The removed questions will be processed based on the plagiarism detection results.
[0148] For example, such as Figure 4i If a star-shaped cluster is detected as not meeting cluster requirements, all questions A through D in the star-shaped cluster are removed. Specifically, this removal can involve deleting the correspondence between questions and cluster heads, and simultaneously deleting the cluster head X from the star-shaped cluster. Further, as... Figure 5 The removed questions A through D are compared with the cluster heads of other star-shaped clusters, such as cluster heads X2 and X3, for deduplication. Based on the deduplication results, the removed questions are processed. Specifically, the removed questions can be added to the star-shaped cluster containing the cluster head with the same deduplication result, and a new star-shaped cluster is constructed based on the specified field of the removed questions that are different from the deduplication results of all cluster heads.
[0149] To ensure the reliability of the correspondence between the cluster head and the various questions in the star cluster after each change in the star cluster structure, when the number of questions in the star cluster changes—that is, when new questions are added or deleted—a vote is held for each specified field in all questions within the star cluster whose number of questions has changed. A new cluster head is then constructed based on the virtual content, thereby adjusting the cluster head of the star cluster with the changed number of questions. This achieves dynamic cluster balance and further improves the accuracy of deduplication. The specific method for constructing the cluster head can be found in step S22, and will not be elaborated here.
[0150] Furthermore, if the cluster head is an actual question, the cluster head of the star cluster with the changing number of questions can be selected based on the confidence level of all questions (including the cluster head) in the star cluster with the changing number of questions, and the new cluster head can replace the original cluster head to achieve dynamic cluster balance.
[0151] The above methods, through constructing star-shaped clusters, dynamically modifying star-shaped clusters, and dynamically adjusting cluster heads, achieve the organization of question data in the question bank. The star-shaped clusters of this invention have the advantages of low deduplication cost, ease of modification, and high recall rate.
[0152] Furthermore, during the user's search process, the system can receive the questions the user wants to search for; match the questions with the cluster heads of each star-shaped cluster; and send at least one question from the star-shaped cluster containing at least one matched cluster head to the user.
[0153] Optionally, information related to the question, such as the question stem, answer, explanation, and knowledge points, corresponding to at least one matched cluster head virtual question can be sent to the user to improve the accuracy and recall rate of the user's question search and enhance the user experience.
[0154] Figure 5 This is a schematic diagram of the structural framework of a question bank data organization device provided by the present invention, as shown below. Column The device includes:
[0155] Filtering module 51 is used to filter out questions with the same content from the question bank;
[0156] Cluster module 52 is used to group questions with the same content into a cluster, store each cluster and its correspondence with the questions, and subsequently manage the question bank and provide data services to the outside world based on the clusters.
[0157] In one implementation, the filtering module 51 is used to cluster the questions in the question bank to form multiple clusters, each cluster including one or more questions with the same content.
[0158] In one example, the cluster module 52 groups topics with the same content into a circular cluster, wherein the circular cluster includes pairs of topics with the same content, and the topics with the same content have a corresponding relationship.
[0159] In one example, the cluster module 52 groups the questions with the same content into a star-shaped cluster. The star-shaped cluster includes a virtual question as the cluster head and the questions with the same content as the cluster members. The virtual question is selected from the questions in the cluster or generated by the questions in the cluster. There is a correspondence between the cluster head and the questions in the cluster.
[0160] In one example, the cluster module 52 forms a virtual title as the cluster head in one of the following ways:
[0161] The virtual question that serves as the cluster head is the highest quality among all questions with the same content located within the corresponding cluster.
[0162] The virtual question that serves as the cluster head includes multiple fields, each of which is the highest quality selected from the corresponding fields of all questions with the same content within the same cluster.
[0163] The virtual question that serves as the cluster head is the question with the highest confidence among all questions within the corresponding cluster. The question confidence is configured based on question attributes. Optionally, the question attributes include at least one of the following: the format of the question field, the source of the question, and the number of times the question has been searched.
[0164] In one example, the cluster module 52 obtains the cluster head in the following manner:
[0165] Retrieve the specified fields that make up the virtual question;
[0166] Voting is conducted from each specified field of all topics within the cluster to select the specified field for the virtual topic;
[0167] A virtual question is constructed based on each specified field, and this virtual question is used as the cluster head; the cluster head has the same content as all questions in the corresponding cluster, and there is a corresponding relationship between them.
[0168] For example, the cluster module 52 configures the voting criteria corresponding to each specified field; and selects the specified field that makes up the virtual question based on the voting criteria corresponding to each specified field and the specified field.
[0169] Optionally, when the specified field is one or more of subject, grade level, and answer, the corresponding voting criteria are configured as follows: the subject, grade level, and answer corresponding to the question with the highest question confidence weighting value in the star cluster are used as the subject, grade level, and answer of the cluster head;
[0170] When the specified field is one or more of the answer, explanation, and key points, the corresponding voting criterion is configured as follows: the specified field with the highest confidence level among the specified fields of each question in the star cluster is used as the specified field of the cluster head;
[0171] When the specified field is one or more of the following: knowledge point, test paper information, book information and test question year, the corresponding voting criteria are configured as follows: configure the confidence threshold of each specified field, and concatenate the specified fields in each question in the star cluster whose confidence is higher than the corresponding confidence threshold, and use them as the specified field of the cluster head.
[0172] When the specified field is the question stem, the corresponding voting criterion is configured as follows: after deduplication of the specified field of the questions in the star cluster, they are merged and used as the question stem content of the cluster head.
[0173] Furthermore, when the questions in the cluster change or new questions are received, the device also includes:
[0174] The modification module 53 is used to compare the changed or newly added questions with the questions in each cluster, so as to reassign the changed or newly added questions to clusters.
[0175] Optionally, the device further includes: a detection module for periodically detecting whether the questions in each cluster have changed;
[0176] Optionally, when the cluster is a star-shaped cluster, the change module 53 includes:
[0177] The deduplication module is used to check for duplicates between changed or newly added questions and the cluster heads of each star cluster.
[0178] The sub-modification module is used to construct a new cluster for the changed or newly added questions if the deduplication result is that they are all non-overlapping, and delete the changed question from the star cluster in which the changed question was located before the change; if the deduplication result is that there are overlapping star clusters, add the changed or newly added questions to the star cluster in which the overlapping cluster head is located, and update the cluster head.
[0179] Optionally, the sub-change module is further configured to initiate a cluster fusion step if the deduplication result indicates that there are two or more overlapping clusters, wherein the cluster fusion step merges the two or more overlapping clusters into one cluster;
[0180] Optionally, when the cluster is a star-shaped cluster, the cluster fusion step includes: selecting the target cluster head corresponding to the star-shaped cluster containing the most questions as the first target cluster head; adding the new questions to the star-shaped cluster where the first target cluster head is located; and adding all the questions from the star-shaped clusters where the other target cluster heads are located to the star-shaped cluster where the first target cluster head is located.
[0181] Optionally, the sub-change module is further configured to construct a new cluster for questions that have changed or are newly added to the database, including: constructing a cluster head based on a specified field of the question that has changed or is newly added to the database, wherein the new question and the constructed cluster head form a new star cluster.
[0182] Furthermore, the sub-change module is also used to delete the unstable problem from the star cluster if an unstable problem in the star cluster whose correspondence with the cluster head has changed is detected.
[0183] Optionally, after deleting the unstable question from the star cluster, the sub-change module is further configured to construct a new cluster head based on the specified field of the unstable question; to perform duplicate checks on all other questions in the star cluster except the unstable question with the new cluster head; if a question identical to the new cluster head is found, the question is deleted from the star cluster and merged with the new cluster head and the unstable question to form a new star cluster.
[0184] Optionally, the sub-modification module is further configured to: if it is detected that the cluster head in the star cluster does not meet the cluster head requirements, delete the cluster head from the star cluster and remove all the questions in the star cluster corresponding to the cluster head; determine the duplicates of the removed questions with the cluster heads of other star clusters; and perform cluster allocation processing on the removed questions according to the duplicate determination results.
[0185] Optionally, if the number of questions in a star cluster changes, the sub-change module is further configured to vote for each specified field of all questions in the star cluster where the number of questions has changed, and to construct a new cluster head based on the virtual content.
[0186] Furthermore, the device also includes:
[0187] The receiving module is used to receive the questions that the user wants to search for;
[0188] The matching module is used to match the search query with the cluster heads of each star cluster;
[0189] The sending module is used to send at least one topic from the star cluster containing at least one matched cluster head to the user;
[0190] Optionally, the sending module sends a virtual title to the user for at least one matched cluster head.
[0191] Those skilled in the art will understand that the modules in the above-described device embodiments can be distributed throughout the device as described, or they can be modified accordingly and distributed in one or more devices different from the above embodiments. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0192] The present invention also provides a question bank, the question bank comprising:
[0193] A question table is used to store information for each question.
[0194] The cluster relationship table stores cluster information and the correspondence between each cluster and the question; Table 1 is a schematic diagram of a cluster relationship table.
[0195] A cluster status table is used to store the status of each cluster, which includes deleted, normal, and pending modification.
[0196] The management module, when managing the question bank and providing data services to external parties, queries the cluster table and the cluster status table based on the cluster.
[0197] Table 1. Schematic diagram of a cluster relationship table
[0198] Type Comment id bigint(20) auto_increment Primary key tid bigint(20) unsinged [0] tid tid2 bigint(20) unsinged [0] Target tid cid bigint(20) unsinged [0] Judged cluster id ctime int(11) unsinged [0] Creation time mtime int(11) unsinged [0] Modification time status int(10) [0] Status: 0; normal: 1; deleted onlinestatus tinyint(2) [0] Whether to retrieve online source varchar(100) [] Judgment source Figure 6
[0199] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store a computer-executable program. When the computer program is executed by the processor, the processor executes a data organization method for a question bank.
[0200] like Figure 6 As shown, the electronic device is embodied in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not preclude distributed processing, meaning that processors can be distributed across different physical devices. The electronic device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0201] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some steps of the method.
[0202] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).
[0203] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0204] It should be understood that Figure 7 The electronic device shown is merely one example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as displays, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. Any electronic device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as an electronic device covered by the present invention.
[0205] Figure 7 This is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. As shown, a computer-readable recording medium stores a computer-executable program, which, when executed, implements the data organization method for the question bank described above. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0206] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0207] From the above description of the embodiments, those skilled in the art will readily understand that the present invention can be implemented by hardware capable of executing specific computer programs, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. included in the system. The present invention can also be implemented by a vehicle including at least a portion of the above-described system or components. The present invention can also be implemented by computer software that executes the methods of the present invention, for example, by control software executed by the microprocessor, electronic control unit, client, server, etc. of a live streaming device. However, it should be noted that the computer software that executes the methods of the present invention is not limited to execution in one or a specific set of hardware entities; it can also be implemented in a distributed manner by unspecified hardware. For computer software, the software product can be stored in a computer-readable storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or distributed across a network, as long as it enables electronic devices to execute the methods according to the present invention.
[0208] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data organization method of a question bank, characterized by, The method comprises the following steps: Clustering or deduplicating the same content of the questions to form multiple clusters, each cluster comprising one or more questions with the same content, the same content referring to the same stem of the question content; Grouping the questions with the same content into a cluster, and storing each cluster and the corresponding relationship between the cluster and the questions; The cluster structure comprises a star cluster, the star cluster comprising a virtual question as a cluster head and the questions with the same content as cluster members; The virtual question as the cluster head in the star cluster is formed in one of the following ways: the virtual question is the question with the best quality among the questions with the same content in the corresponding cluster; the virtual question comprises multiple fields, each of the multiple fields being the field selected from the corresponding fields of the questions with the same content in the corresponding cluster and having the best quality; the virtual question is the question with the highest confidence among the questions in the corresponding cluster, the confidence of the question being configured according to the question attributes; The cluster head is obtained in the following way: obtaining the specified fields of the virtual question; selecting the specified field of the virtual question from each specified field of all the questions in the cluster; constructing the virtual question according to the specified fields, and taking the virtual question as the cluster head; the cluster head being the same as the content of all the questions in the corresponding cluster and having the corresponding relationship; The method further comprises the following steps: periodically detecting whether the questions in each cluster change; when the questions in the cluster change or a newly entered question is received, deduplicating the changed or newly entered question from the questions in each cluster and performing cluster allocation; if there is no same star cluster head, constructing the cluster head according to the specified fields of the changed or newly entered question to form a new star cluster, if the deduplication result of the cluster head of each star cluster has two or more than two overlapping clusters, merging the overlapping clusters into one cluster, if an unstable question with a changed corresponding relationship with the cluster head is detected in the star cluster, deleting the unstable question, and if the number of questions in the star cluster changes, updating the cluster head according to the specified fields of the cluster head selected from each specified field of all the questions in the cluster with the changed number of questions.
2. The method of claim 1, wherein, The cluster structure further comprises a ring cluster: after deduplicating the questions two by two, organizing the questions with the same content in a ring cluster to ensure that all the questions in the same ring cluster are the same, and the questions with the same content two by two in the ring cluster having the corresponding relationship.
3. The method of claim 1, wherein, The question attributes comprise at least one of the format of the question field, the source of the question, and the number of times the question is searched.
4. The method of claim 1, wherein, The method of selecting the specified field of the virtual question from each specified field of all the questions in the cluster comprises the following steps: Configuring the selection standard corresponding to each specified field; Selecting the specified field of the virtual question according to the selection standard corresponding to each specified field and the specified field.
5. The method of claim 4, wherein, The method further comprises the following steps: When the specified field is one or more of the subject, the segment, and the answer, the selection standard corresponding to the specified field is configured as follows: the subject, the segment, and the answer of the question with the highest weighted value of the confidence in the star cluster are taken as the subject, the segment, and the answer of the cluster head. For the specified field being one or more of the answer, the resolution and the point, the corresponding ticket selection standard is configured as: the specified field with the highest confidence of the specified field of each question in the star cluster is taken as the specified field of the cluster head; For the specified field being one or more of the knowledge point, the test paper information, the book information and the test question year, the corresponding ticket selection standard is configured as: the confidence threshold of each specified field is configured, and the specified fields with a confidence higher than the corresponding confidence threshold in the specified fields of each question in the star cluster are spliced to be taken as the specified field of the cluster head; For the specified field being the stem, the corresponding ticket selection standard is configured as: the specified fields of the questions in the star cluster are removed and combined to be taken as the stem content of the cluster head.
6. The method of claim 1, wherein, The changed or newly stored question is compared with the questions in each cluster for duplication to cluster the changed or newly stored question, including: When the cluster is a star cluster, the changed or newly stored question is compared with the cluster heads of each star cluster for duplication; If the duplication result is that none of them coincides, a new cluster is constructed for the changed or newly stored question, and the changed question is deleted from the star cluster where the changed question is located before the change; If the duplication result is that there is a coincident star cluster, the changed or newly stored question is added to the star cluster where the coincident cluster head is located, and the cluster head is updated.
7. The method of claim 6, wherein, The cluster fusion includes: Selecting the target cluster head corresponding to the star cluster containing the most questions as the first target cluster head; adding the new question to the star cluster where the first target cluster head is located; Adding all the questions in the star clusters where the other target cluster heads are located to the star cluster where the first target cluster head is located.
8. The method of claim 6, wherein, Selecting the first target cluster includes: randomly selecting the first target cluster head, or selecting the target cluster head corresponding to the star cluster containing the most questions as the first target cluster head.
9. The method of claim 7, wherein, Further including: Constructing a new cluster for the changed or newly stored question, wherein the cluster head is constructed according to the specified field of the changed or newly stored question, and the new question and the constructed cluster head form a new star cluster.
10. The method of claim 1, wherein, After deleting the unstable question from the star cluster, further including: The new cluster head according to the specified field of the unstable question; Comparing each question in the star cluster except the unstable question with the constructed new cluster head for duplication; If the same question as the new cluster head is found, the question is deleted from the star cluster, and is combined with the new cluster head and the unstable question to form a new star cluster.
11. The method of claim 10, wherein, Further including: If it is detected that the cluster head in the star cluster does not meet the cluster head requirement, the cluster head is deleted from the star cluster, and all the questions in the star cluster corresponding to the cluster head are removed; Comparing the removed questions with the cluster heads of other star clusters for duplication; and performing cluster allocation processing on the removed questions according to the duplication result.
12. The method of claim 1, wherein, Further including: Receiving a question to be searched by a user; Matching the question to be searched with the cluster heads of each star cluster; Sending at least one question in the star cluster where the matched at least one cluster head is located to the user and / or sending a virtual question of the matched at least one cluster head to the user.
13. A data organization apparatus for a question bank, characterized by comprising: The device includes: The screening module screens the same content questions from the question bank cluster or duplicates to screen the same content questions to cluster to form a plurality of clusters, each cluster including one or more same content questions, same content referring to the same stem of the question content; The cluster module groups each same content question into a cluster, and stores each cluster and its corresponding relationship with the questions. The cluster structure includes a star cluster, which includes a virtual question as a cluster head and each same content question as a cluster member. The virtual question as the cluster head is formed in one of the following ways: the virtual question is the best quality one of each same content question in the corresponding cluster; the virtual question includes a plurality of fields, each of which is the best quality one selected from the corresponding fields of each same content question in the corresponding cluster; the virtual question is the question with the highest confidence in each question in the corresponding cluster, and the question confidence is configured according to the question attributes; the cluster head is obtained by: obtaining each specified field of the virtual question; selecting the specified field of the virtual question from each specified field of all questions in the cluster; constructing the virtual question according to each specified field, and taking the virtual question as the cluster head; the cluster head is the same content as all questions in the corresponding cluster and has a corresponding relationship; Periodically detecting whether the questions in each cluster change; when the questions in the cluster change or a newly entered question is received, the changed or newly entered question is compared with the questions in each cluster for duplicates and cluster allocation; if there is no same star cluster head, a new star cluster head is formed according to the specified fields of the changed or newly entered question; if there are two or more overlapping clusters in the duplicate comparison result of each star cluster head, the overlapping clusters are merged into one by cluster fusion; if an unstable question with a changed corresponding relationship with the cluster head is detected in the star cluster, the unstable question is deleted; if the number of questions in the star cluster changes, the specified fields of the cluster head are updated according to the specified fields of all questions in the cluster with the changed number of questions.
14. An electronic device comprising a processor and a memory, the memory being configured to store a computer executable program, characterized in that: When the computer executable program is executed by the processor, the processor executes the method of any one of claims 1-12.
15. A computer readable medium storing a computer executable program, characterized in that, The computer executable program is executed to implement the method of any one of claims 1-12.
Citation Information
Patent Citations
Cluster based service issuing and discovering method in self-organizing network facing to service
CN101163158A
Question meaning text-based same knowledge point test question grouping system and method
CN112256869A