Information acquisition method, device, storage medium, and electronic device
By using synonym lists and feature extraction technology in the scenario of deduplication of repeated educational questions, we can obtain question information with a similarity higher than a threshold, solving the problem of insufficient information acquisition caused by incomplete deduplication in existing technologies and achieving a more comprehensive information acquisition effect.
Patent Information
- Application Number
- CN202011212443.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-03
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-11-03
AI Technical Summary
Existing technologies are not thorough enough in removing duplicate questions in education, resulting in less comprehensive information acquisition.
By obtaining topic information, performing feature extraction, and using the configured synonym list to determine synonym relationships, we can obtain topic information with a similarity higher than a threshold, thereby achieving more comprehensive information acquisition.
It improves the comprehensiveness of information acquisition, reduces the exposure of duplicate questions, and solves the problem of insufficient information acquisition caused by incomplete deduplication.
Smart Images

Figure CN113392094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to an information acquisition method, device, storage medium, and electronic equipment. Background Art
[0002] Deduplication of duplicate questions in education can be considered a subfield of text deduplication. Deduplication is a fundamental task in the internet industry. For example, plagiarized news articles and identical advertising copy contribute to the homogeneity of online content and increase database storage burdens. Similarly, in the practical application of deduplication in education, obtaining a large amount of duplicate information is crucial. For example, in the list of similar questions recommended for consolidation exercises and test-taking, a large number of duplicate questions may appear, resulting in a poor user experience.
[0003] Existing techniques for deduplication are mostly unsupervised. While direct text matching relies less on data annotation and can be effective and iterative when data annotation is scarce, it does not adequately address the diverse expression of questions. For example, "AB is perpendicular to BC" and "the measure of angle ABC is 90 degrees," "the solution of the equation $\\frac{2}{x\\text{-}1}$=1 is ()," and "the solution of the equation $\\frac{2}{x-1}$=1 is ()." In these scenarios, direct keyword matching often yields low scores, making it impossible to identify similar or duplicate questions. In other words, existing techniques suffer from incomplete deduplication, resulting in less comprehensive information acquisition. Summary of the Invention
[0004] The embodiments of the present invention provide an information acquisition method, device, storage medium, and electronic device to at least solve the technical problem of low comprehensiveness of information acquisition caused by incomplete deduplication.
[0005] According to one aspect of an embodiment of the present invention, there is provided an information acquisition method, comprising: acquiring title information of a target title; performing feature extraction on the title information of the target title to obtain a first feature sequence; determining, from a synonym list corresponding to a configured target question bank, a second feature value having a synonym relationship with a first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values having the synonym relationship, and the feature values having the synonym relationship have the same semantics; acquiring a second feature sequence based on the second feature value and the first feature sequence; acquiring target title information corresponding to the target feature sequence from the target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0006] According to another aspect of an embodiment of the present invention, an information acquisition device is also provided, including: a first acquisition unit, used to acquire question information of a target question; a first extraction unit, used to perform feature extraction on the question information of the target question to obtain a first feature sequence; a first determination unit, used to determine, from a synonym list corresponding to a configured target question bank, a second feature value having a synonym relationship with the first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values having the synonym relationship, and the feature values having the synonym relationship have the same semantics; a second acquisition unit, used to acquire a second feature sequence based on the second feature value and the first feature sequence; and a third acquisition unit, used to acquire target question information corresponding to the target feature sequence from the target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0007] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned information acquisition method when running.
[0008] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the information acquisition method through the computer program.
[0009] In an embodiment of the present invention, the following steps are performed: obtaining the title information of a target title; performing feature extraction on the title information of the target title to obtain a first feature sequence; determining a second feature value having a synonymous relationship with the first feature value in the first feature sequence from a synonym list corresponding to a configured target title bank, wherein the synonym list records at least two groups of feature values having the synonymous relationship, and the feature values having the synonymous relationship have the same semantics; obtaining a second feature sequence based on the second feature value and the first feature sequence; obtaining target title information corresponding to the target feature sequence from the target title bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold; obtaining a second feature sequence having the same semantics as the first feature sequence obtained by feature extraction but richer in features based on the synonym list, and utilizing the more comprehensive second feature sequence to match more title information, and determining the title information having a similarity higher than the target threshold from the candidate title information as the target title information, thereby achieving the effect of improving the comprehensiveness of information acquisition, and thus solving the technical problem of low comprehensiveness of information acquisition due to incomplete deduplication. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0011] Figure 1 is a schematic diagram of an application environment of an optional information acquisition method according to an embodiment of the present invention;
[0012] Figure 2 is a schematic diagram of a flow chart of an optional information acquisition method according to an embodiment of the present invention;
[0013] Figure 3 is a schematic diagram of an optional information acquisition method according to an embodiment of the present invention;
[0014] Figure 4 is a schematic diagram of another optional information acquisition method according to an embodiment of the present invention;
[0015] Figure 5 is a schematic diagram of another optional information acquisition method according to an embodiment of the present invention;
[0016] Figure 6 is a schematic diagram of another optional information acquisition method according to an embodiment of the present invention;
[0017] Figure 7 is a schematic diagram of another optional information acquisition method according to an embodiment of the present invention;
[0018] Figure 8 is a schematic diagram of another optional information acquisition method according to an embodiment of the present invention;
[0019] Figure 9 is a schematic diagram of an optional information acquisition device according to an embodiment of the present invention;
[0020] Figure 10 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] In the embodiments of this application, the following technical terms may be used, but are not limited to:
[0024] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0025] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0026] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, and knowledge graphs.
[0027] According to one aspect of an embodiment of the present invention, a method for obtaining information is provided. Optionally, as an optional implementation, the information obtaining method can be applied to, but is not limited to, Figure 1 In the environment shown, it may include, but is not limited to, a user device 102, a network 110, and a server 112. The user device 102 may include, but is not limited to, a display 108, a processor 106, and a memory 104. The display 108 may be used, but is not limited to, presenting a human-computer interaction interface containing input and acquisition of topic information, and may also be used to provide a human-computer interaction interface for receiving human-computer interaction operations performed on the human-computer interaction interface to obtain messages to be interacted with. The memory 104 may be used, but is not limited to, storing input topic information, acquired topic information, or the aforementioned messages to be interacted with. The server 112 may include, but is not limited to, a database 114 and a processing engine 116. The database 114 may be used, but is not limited to, storing topic information sent by the user device 102 and pre-configured topic information. The processing engine 116 may be used, but is not limited to, storing and processing the acquired topic information sent by the user device 102 to obtain topic information to be output whose similarity to the topic information sent by the user device 102 meets preset conditions.
[0028] The specific process can be as follows:
[0029] Step S102 , the user device 102 receives a human-computer interaction operation performed on the human-computer interaction interface through the human-computer interaction interface on the display 108 to obtain the topic information of the target topic input by the user;
[0030] Steps S104-S106, the user device 102 sends the target topic information to the server 112 via the network 110;
[0031] In steps S108-S114, the server 112 performs feature extraction on the target question information through the processing engine 116 to obtain a first feature sequence, then uses the synonym list corresponding to the target question bank configured in the database 114 to search for a second feature value that has a synonym relationship with the first feature in the first feature sequence, and processes the second feature value and the first feature sequence through the processing engine 116 to obtain a second feature sequence, and calculates the similarity between the feature sequence in the database 114 and the second feature sequence to determine that the feature sequence whose similarity meets the preset conditions is the target feature sequence, and then searches the target question information corresponding to the above target feature sequence through the target question bank in the database 114;
[0032] Steps S116-S118, the server 112 sends the target topic information to the user device 102 via the network 110;
[0033] In step S120 , the processor 106 in the user device 102 displays the target topic information on the display 108 , and stores the target topic information and the target topic information in the memory 104 .
[0034] Optionally, in this embodiment, the terminal device may be a terminal device equipped with a conversation application client, and may include, but is not limited to, at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Device), a PAD, a desktop computer, a smart TV, etc. The conversation application client herein may be, but is not limited to, an application client equipped with a conversation function, such as an instant messaging application client, a game application client with a conversation function, a transaction application client with a conversation function, an online education application client with a conversation function, etc. The network may include, but is not limited to, a wired network and a wireless network, wherein the wired network includes a local area network, a metropolitan area network, and a wide area network, and the wireless network includes Bluetooth, Wi-Fi, and other networks that enable wireless communication. The server may be a single server, a server cluster consisting of multiple servers, or a cloud server. The above is merely an example and is not intended to be limiting in this embodiment.
[0035] Alternatively, as an optional implementation, Figure 2 As shown, the information acquisition method includes:
[0036] S202, obtaining the topic information of the target topic;
[0037] S204, performing feature extraction on the topic information of the target topic to obtain a first feature sequence;
[0038] S206, determining a second feature value that has a synonymous relationship with the first feature value in the first feature sequence from a synonym list corresponding to the configured target question bank, wherein the synonym list records at least two groups of feature values that have a synonymous relationship, and the feature values that have a synonymous relationship have the same semantics;
[0039] S208, obtaining a second feature sequence based on the second feature value and the first feature sequence;
[0040] S210 , acquiring target question information corresponding to a target feature sequence from a target question bank, wherein a similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0041] Optionally, in this embodiment, the above information acquisition method can be applied to, but not limited to, an educational client for obtaining repeated question information of an input question, such as Figure 1In the educational client running in the user device 102 shown, in particular, for educational scenarios such as the need to obtain multiple solutions to a question, the need to obtain the number of repetitions of a question, and the need to obtain similar question information for a question, for example, a large number of repeated questions will appear in the question list of similar question recommendations in consolidation exercises and test paper compilation scenarios, resulting in a poor user experience. The above information acquisition method can be used to remove duplicate questions and improve the user experience. For example, repeated question information can be understood as a kind of popularity, so the above information acquisition method can be used as a solution for generating popularity tags. For example, the above information acquisition method can be used to find different questions stored in the question bank, and then use the analysis method of the question itself to extract different solution methods, and then provide multiple solution tags for the question as the characteristics of the question for the user to answer and choose. In addition, the above information acquisition method also obtains and removes duplicate question information of the target question in the deduplication scenario. For example, when the question is selected when entering the database, the above information acquisition method is used to filter the duplicate questions so that the duplicate question information is not entered into the database, reducing the storage pressure of the question inventory. The above is only an example and is not limited in this embodiment.
[0042] Optionally, in this embodiment, the target topic may be, but is not limited to, a topic input by the user in the corresponding target client, and may also be, but is not limited to, any topic stored in the initial topic bank associated with the target client;
[0043] It should be noted that when the target question is a question entered by a user on a corresponding target client, the above information acquisition method can be applied, but is not limited to, in educational scenarios such as obtaining information about multiple solutions to a question, obtaining information about the number of repeated answers to a question, or obtaining information about similar questions to a question. Furthermore, when the target question is any question stored in an initial question bank associated with the target client, the above information acquisition method can be applied, but is not limited to, in scenarios where duplicates are removed from a large number of questions stored in the initial question bank. The above is merely an example and is not intended to be limiting in this embodiment.
[0044] Optionally, in this embodiment, the question information may be, but is not limited to, key information of the target question, or may be, but is not limited to, all information of the target question. For example, the question information may be, but is not limited to, the title information or stem information of the target question. For another example, the question information may be, but is not limited to, text information or formula information of the target question.
[0045] Optionally, in this embodiment, obtaining the target question information may include, but is not limited to, first obtaining the first question information of the target question, then normalizing the first question information to obtain the second question information, and using the second question information as the target question information. Optionally, the normalization process may include, but is not limited to, a method of simplifying calculations, that is, converting a dimensioned expression into a dimensionless expression, thus becoming a scalar. Optionally, in this embodiment, the normalization process may include, but is not limited to, converting the first question information of the target question into relatively unified question expression information.
[0046] To further illustrate, the optional normalization process can be based on but not limited to Figure 3 The normalization module shown is executed. Specifically, whether it is the expression of the question in the question bank or the question information input by the user, there are a variety of expressions, professional terms, Latex expressions, etc. Therefore, it is very necessary to introduce a normalization operation for the question. Furthermore, the main logic in the normalization of the question is to perform certain normalization operations on the question, such as normalization of quantity units and synonyms, and convert them into a single term expression, so that the expression of the question is more standardized and unified, providing an effective role for the subsequent formula and text feature extraction. For example, the normalization module obtains the original text 304 (i.e., the first question information) and performs a normalization operation on the original text, which can be but not limited to include Figure 3 The normalization module 302 performs full-width to half-width conversion, special character filtering, formula taf replacement, special error correction, and other normalization operations to obtain normalized text 306 (i.e., second topic information) after normalization. The above is only an example and is not limited in this embodiment.
[0047] Optionally, in this embodiment, the synonym list may be, but is not limited to, composed of multiple groups of synonyms obtained based on a synonym mining algorithm in feature mining, wherein the above-mentioned synonym mining algorithm is mainly mined from a configured target question bank using methods such as word2vec. Optionally, word2vec may be, but is not limited to, a group of related models for generating word vectors, which are shallow, two-layer neural networks used for training to reconstruct linguistic word texts;
[0048] In addition, while obtaining the synonym list, multiple stop words can also be obtained through the stop word fern algorithm in feature mining, and a stop word list consisting of multiple stop words can be obtained. Optionally, the source of the stop words can be, but is not limited to, some special symbols, meaningless words, and some open source stop words such as "Baidu Stop Word List", "Harbin Institute of Technology Stop Word List", "Sichuan University Machine Learning Laboratory Stop Word List", etc.
[0049] For further illustration, optionally, for example Figure 4 As shown, in the configured question bank 402, a large number of synonyms, stop words, Tags of Latex, etc. are mined. In addition to being used to form the synonym list 404 and the stop word list 406, it also facilitates the subsequent question normalization process and the logic of obtaining target question information with a similarity higher than the target threshold.
[0050] It should be noted that, from the synonym list corresponding to the configured target question bank, a second feature value having the same semantics as the first feature value in the first feature sequence is determined, and based on the second feature value and the first feature sequence, a second feature sequence for obtaining target question information is obtained. This can, but is not limited to, overcome the problem of the same semantic expression. For example, the similarity between similar words in the dictionary is quite large. For instance, an equilateral triangle and a regular triangle have the same semantic meaning in the mathematical concept, but their similarity is almost zero. As a result, two feature values with the same semantic meaning have a similarity of zero, leading to their unsuccessful matching. Therefore, by using the synonym list of the configured target question bank, for a specific concept such as a regular triangle, it is normalized to the semantic expression of an equilateral triangle, and based on the second feature value having the same semantics as the first feature value in the first feature sequence, the second feature sequence is obtained, thus solving the problem of the unmatched semantic expression of the same concept;
[0051] In addition, the prior art also has the problem of meaningless single - character expressions. For example, in the actually matched sequences, there are some special meaningless expressions. For example, in the expressions "AB = 6m, AB = _cm" and "AB = 6m, then AB = _cm", there are many meaningless words such as "then", and these words will also cause the matching score to be too low, resulting in the failure of similarity matching. Therefore, through the above - mentioned stop - word matching list, the feature sequence is filtered once, thereby increasing the matching score of the sequence and further improving the matching success rate.
[0052] It should be noted that, from the synonym list corresponding to the configured target question bank, a second feature value having a synonym relationship with the first feature value in the first feature sequence is determined, where the optional first feature value can be, but is not limited to, any feature value in the first feature sequence.
[0053] Optionally, in this embodiment, obtaining target question information corresponding to a target feature sequence from a target question bank may be, but is not limited to, performing similarity matching on the feature sequences corresponding to all question information in the target question bank and the second feature sequence. If the matching similarity between the feature sequence corresponding to the question information in the target question bank and the second feature sequence is higher than a target threshold, then it is determined that the two are successfully matched, and the successfully matched feature sequence is determined to be the target feature sequence, and the question information corresponding to the target feature sequence is determined to be the target question information. Optionally, the similarity matching may be, but is not limited to, using different similarity calculation methods, such as cosine distance, inner product, Jaccard distance, etc.
[0054] To further illustrate, the similarity matching can be achieved by referring to, but not limited to, the following formulas (1) and (2);
[0055] m i =Jaccard(Norm(Stopword(x i )),Norm(Stopword(x j ))) Formula (1);
[0056]
[0057] Wherein, Norm is a synonym normalization operation, Stopword is a synonym filtering operation, Intersect is an intersection operation, and Union is a set union operation. Optionally, the synonym filtering operation may be, but is not limited to, the operation in the above-mentioned information acquisition method of obtaining a second feature sequence based on a second feature value having a synonym relationship with the first feature value in the first feature sequence and the first feature sequence.
[0058] Optionally, in this embodiment, obtaining target question information corresponding to the target feature sequence from the target question bank may include, but is not limited to, first obtaining a third feature sequence having a similarity with the second feature sequence higher than a target threshold, and then determining the third feature sequence as the target feature sequence, thereby obtaining the target question information corresponding to the target feature sequence;
[0059] It should be noted that obtaining a third feature sequence whose similarity with the second feature sequence is higher than the target threshold may include, but is not limited to, first obtaining multiple third feature sequences corresponding to all question bank information in the target question bank, and then calculating the similarity between the multiple third feature sequences and the second feature sequence. Optionally, the similarity calculation may include, but is not limited to, matching multiple types of feature values in the feature sequence to obtain multiple scores, and then fusing the multiple scores into an overall score as the final matching score (similarity) of the two matching feature sequences. Specifically, for example, reference may be made to the following formula (3):
[0060]
[0061] in, is the score of the characteristic factor (eigenvalue) of the k-th match, λ k is the matching weight coefficient for the kth item, which can be, but is not limited to, a hyperparameter. Specifically, for example, first, through manual experience, in the matching sequence, the matching weight of the question stem is greater than the matching weight of the options, and the matching score of the text is greater than the matching score of the formula and analytical formula. Then, a grid search is used to determine the weight of each matching factor (the corresponding feature value of the two matching parties).
[0062] In addition, obtaining the target question information corresponding to the target feature sequence from the target question bank can be, but is not limited to, first obtaining a third feature sequence whose similarity with the second feature sequence is higher than the target threshold. A supervised matching scheme similar to the DSSM, Bi-LSTM, and BERT series can also be used to further improve the effect.
[0063] Optionally, in this embodiment, before obtaining the target question information corresponding to the target feature sequence from the target question bank, it can be, but is not limited to, also including obtaining the category information in the question information of the target question, wherein the category information is used to indicate the question category of the target question, for example, the target question is a multiple-choice question or a fill-in-the-blank question, and the target question is the calculation of the entity relationship (greater than or less than). In other words, considering that in the actual business logic of repeated questions, there are necessary business rules, such as the difference in formula numbers in repeated questions (such as x=y+1 and x=y+2) cannot be considered as repeated questions, the difference in question types (multiple-choice questions and fill-in-the-blank questions) is not considered as repeated questions, and the change in the direction of the entity relationship (greater than and less than) is not considered as a repeated question. For this reason, for all the question information in the target question bank, it is also necessary to filter the formula judgment, question type comparison, entity relationship and other modules, and perform ReRanking re-sorting according to certain business logic. Finally, the system outputs candidate question information that meets the business logic, and then performs matching operations with the second sequence features. Optionally, after filtering the question information in the target question bank, it is also possible but not limited to filtering the feature sequence corresponding to the question information in the target question bank;
[0064] To further illustrate, the above filtering can be based on, but not limited to, Figure 5 The filtering model 502 shown is executed to obtain question information 504 in the target question bank, and undergoes probability filtering, question type filtering, other strategy filtering and other operations to obtain candidate question information 506 that meets the business logic.
[0065] Through the embodiments provided by the present application, the title information of the target title is obtained; feature extraction is performed on the title information of the target title to obtain a first feature sequence; from the synonym list corresponding to the configured target question bank, a second feature value having a synonym relationship with the first feature value in the first feature sequence is determined, wherein the synonym list records at least two groups of feature values having a synonym relationship, and the feature values having a synonym relationship have the same semantics; based on the second feature value and the first feature sequence, a second feature sequence is obtained; target question information corresponding to the target feature sequence is obtained from the target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold, based on the synonym list, a second feature sequence having the same semantics as the first feature sequence obtained by feature extraction but richer in internal features is obtained, and the more comprehensive second feature sequence is used to match more question information, and among the candidate question information, question information having a similarity higher than the target threshold is determined as the target question information, thereby achieving the effect of improving the comprehensiveness of information acquisition.
[0066] As an optional solution, obtaining a second feature sequence based on the second feature value and the first feature sequence includes:
[0067] S1, sequentially merging a first eigenvalue and a second eigenvalue having a synonymous relationship with the first eigenvalue to obtain a target eigenvalue corresponding to the first eigenvalue, wherein the target eigenvalue includes one of the following: the first eigenvalue or the second eigenvalue, or a eigenvalue obtained by fusion of the first eigenvalue and the second eigenvalue;
[0068] S2, combining the target eigenvalues corresponding to the respective first eigenvalues into a second eigenvalue sequence.
[0069] It should be noted that the first eigenvalue and the second eigenvalue having a synonymous relationship with the first eigenvalue are merged in sequence to obtain the target eigenvalue corresponding to the first eigenvalue, and the target eigenvalues corresponding to each first eigenvalue are combined into a second eigensequence.
[0070] Optionally, in this embodiment, for example Figure 6 As shown, the target question information "the solution to the right angle is" is obtained on the client, and then the target question information is subjected to feature extraction to obtain a first feature sequence 602, wherein the first feature sequence 602 is used to represent the target question information "the solution to the right angle is";
[0071] To further illustrate, the first feature sequence 602 can be optionally processed as a synonym. Specifically, the first feature value in the first feature sequence 602, such as the first feature value corresponding to a right angle, and the second feature value that has a synonym relationship with the first feature value are merged, wherein the second feature value is used to represent the same semantic expression of the right angle, for example, the second feature value is used to represent 90°, which is consistent with the meaning of the right angle. Then, the second feature sequence 604 obtained after the merging process takes into account the semantic expression of both the right angle and 90°. Therefore, the target topic information obtained based on the second feature sequence 604 has advantages such as being richer than the topic information obtained based on the first feature sequence 602. For example, the topic information obtained based on the first feature sequence 602 may only display the relevant content of the right angle △ABC on the client, while the target topic information obtained based on the second feature sequence 604 not only displays the relevant content of the right angle △ABC, but also displays the relevant content of △ABC=90°.
[0072] Through the embodiments provided in the present application, the first feature value and the second feature value having a synonymous relationship with the first feature value are merged in sequence to obtain a target feature value corresponding to the first feature value, wherein the target feature value includes one of the following: the first feature value or the second feature value, or a feature value obtained by feature fusion of the first feature value and the second feature value; the target feature values corresponding to each first feature value are combined into a second feature sequence, thereby achieving the purpose of improving the comprehensiveness of the feature value and realizing the effect of improving the comprehensiveness of the question information matched based on the feature value.
[0073] As an optional solution, in the process of obtaining the second feature sequence based on the second feature value and the first feature sequence, the following is further included:
[0074] S1, determining a first feature value associated with a stop feature value in a filter word list from a first feature sequence, wherein the filter word list contains at least two stop feature values, and the stop feature values are feature values unrelated to the topic information;
[0075] S2: Delete the first feature value associated with the disabled feature value in the first feature sequence.
[0076] It should be noted that the first feature value associated with the disabled feature value in the filter word list is determined from the first feature sequence, and the first feature value associated with the disabled feature value in the first feature sequence is deleted to obtain an updated first feature sequence.
[0077] Optionally, in this embodiment, for example Figure 7As shown, the question information of the target question "Then, the solution of the equation of X is" is obtained on the client, and then the question information of the target question is feature extracted to obtain a first feature sequence 702, wherein the first feature sequence 702 is used to represent the question information of the target question "Then, the solution of the equation of X is".
[0078] To further illustrate, the first feature value associated with the disabled feature value in the filter word list can be optionally deleted from the first feature sequence, for example, the first feature value corresponding to "then," "yes" is deleted, and then the updated first feature sequence 704 is obtained, and the target title information corresponding to the content "Solution of the equation of X" is obtained based on the updated first feature sequence 704.
[0079] Through the embodiment provided by the present application, a first feature value having an association relationship with a disabled feature value in a filter word list is determined from a first feature sequence, wherein at least two disabled feature values are recorded in the filter word list, and the disabled feature values are feature values irrelevant to the topic information; the first feature value having an association relationship with the disabled feature value in the first feature sequence is deleted, thereby achieving the purpose of reducing the inability to obtain information due to invalid feature values and realizing the effect of improving the accuracy of obtaining topic information.
[0080] As an optional solution, feature extraction is performed on the topic information of the target topic, including:
[0081] S1, when the title information of the target topic includes text information, performing text feature extraction on the title information of the target topic;
[0082] S2: When the topic information of the target topic includes formula information, extract formula features from the topic information of the target topic.
[0083] It should be noted that when the target topic information includes text information, text feature extraction is performed on the target topic information; when the target topic information includes formula information, formula feature extraction is performed on the target topic information.
[0084] To further illustrate, the optional title information, such as the target title, may carry a variety of information types. If only one feature extraction method is used, it may be impossible to quickly extract features with high accuracy due to the singleness of the method.
[0085] Through the embodiments provided in the present application, when the title information of the target topic includes text information, text feature extraction is performed on the title information of the target topic; when the title information of the target topic includes formula information, formula feature extraction is performed on the title information of the target topic. By utilizing two different feature extraction methods, the purpose of extracting features based on different extraction methods is achieved, and a comprehensive feature extraction effect is realized.
[0086] As an optional solution, text feature extraction is performed on the title information of the target title, including at least one of the following:
[0087] S1, extracting text features from the text characters contained in the title information of the target title to obtain text character sequence features;
[0088] S2, extracting text features from the segmentation results obtained after segmenting the text information in the target question to obtain segmentation sequence features;
[0089] In addition, formula feature extraction is performed on the topic information of the target topic, including at least one of the following:
[0090] S1, extract formula features from the formula character information in the target question information to obtain the formula original string sequence features;
[0091] S2, extracting formula features from the parsed results obtained after parsing the formula information in the question information of the target question, and obtaining formula parsing result sequence features.
[0092] Optionally, in this embodiment, feature extraction can be, but is not limited to, extracting features for matching modeling from the title information of the target title. Specifically, for the acquired title information of the target title, extract the sequence of word segmentation, the sequence after formula parsing, the sequence of original formula strings, and the sequence of text character n-grams. Among them, the text word segmentation sequence can be, but is not limited to, referring to a list of word segmentation results obtained after word segmentation of the title information of the target title, the original string sequence of the formula can be, but is not limited to, referring to the sequence of original character strings obtained by the formula extractor, and the sequence after formula parsing can be, but is not limited to, referring to the sequence after formula parsing obtained by the formula extractor, wherein the feature extractor can be, but is not limited to, used to convert the original features into a set of features with obvious physical meaning or statistical meaning or kernel. Optionally, n-gram can be, but is not limited to, an algorithm based on a statistical language model, the basic idea of which is to perform a sliding window operation of size N on the content of the text according to bytes, forming a sequence of byte fragments of length N;
[0093] To further illustrate, the original text string of the optional title information of the target title is "In the following formula $\\frac{2}{3}$a+b, S=$\\frac{1}{2}$ab, 5, m, 8+y, m+3=2, $\\frac{2}{3}$≥$\\frac{5}{7}$, the algebraic expression is ()", then the 1-gram sequence of the text word segmentation obtained after word segmentation is: [(,), 1, 2, 3, 5, 7, 8+y, =, a+b, ab, frac, m, m+3, s, ≥, the following, in, algebraic expression, the formula, is].
[0094] Furthermore, optionally, the formula original string sequence obtained by the formula extractor may be, but is not limited to: [$\frac{2}{3}$a+b, $\frac{2}{3}$≥$\frac{5}{7}$, 5, 8+y, S=$\frac{1}{2}$ab, m, m+3=2];
[0095] The sequence of formulas obtained after formula extraction can be, but is not limited to, [2 / 3*a+b, 2 / 3>=5 / 7, 5, 8+y, S=1 / 2*(a*b), m, m+3=2].
[0096] In addition, feature extraction can also include, but is not limited to, the expression of feature sequences of semantic embeddings of text characters and words, so that in addition to the different character expressions in the text, there are also some deep semantic information expressions.
[0097] Through the embodiments provided by the present application, text feature extraction is performed on the title information of the target title, including at least one of the following: text feature extraction is performed on the text characters contained in the title information of the target title to obtain text character sequence features; text feature extraction is performed on the word segmentation results obtained after word segmentation of the text information in the title information of the target title to obtain word segmentation sequence features; formula feature extraction is performed on the title information of the target title, including at least one of the following: formula feature extraction is performed on the formula character information in the title information of the target title to obtain formula original string sequence features; formula feature extraction is performed on the parsed results obtained after parsing the formula information in the title information of the target title to obtain formula parsing result sequence features, thereby achieving the purpose of refining the feature extraction granularity and realizing the effect of improving the accuracy of feature extraction.
[0098] As an optional solution, before obtaining the topic information of the target topic, include:
[0099] S1, extracting features from the candidate question information recorded in the target question bank to obtain the index feature sequence corresponding to each candidate question information;
[0100] S2, establishing an index feature sequence library based on the index feature sequence.
[0101] It should be noted that due to the large amount of data in a question bank of millions, the amount of calculation required for direct matching will be very large and time-consuming. For example, a single match takes 2ms, and after matching 1 million questions, it takes about 2*1 million = 2 million / (1000*3600) = 0.55 hours. Although matching only takes 2ms, it takes more than half an hour to fully match the entire question bank, and the user experience will be very poor. To avoid saving time, it is possible but not limited to introducing an efficient recall module in a manner similar to a recommendation system to recall more candidate duplicate questions, thereby saving matching time. Here, it is possible but not limited to establishing an index feature sequence library through a search server (such as a Lucene-based search server elastic search). In the established index feature sequence library, index feature sequences corresponding to the question information of the target question bank are stored, and then candidate questions are retrieved through the questions.
[0102] Optionally, to further illustrate, the questions and options in each question information of the target question bank are spliced together, and then word segmentation and formula extraction are performed to obtain a feature sequence corresponding to each of the above question information, and then the word segmentation and formula feature sequences are spliced with spaces and placed into the index feature sequence library.
[0103] Through the embodiments provided in the present application, feature extraction is performed on the candidate question information recorded in the target question bank to obtain the index feature sequence corresponding to each candidate question information; an index feature sequence library is established based on the index feature sequence, thereby achieving the purpose of quickly searching for the corresponding candidate question information through the pre-established index feature sequence library, and realizing the effect of improving the overall efficiency of information acquisition.
[0104] As an optional solution, target question information corresponding to the target feature sequence is obtained from the target question bank, including:
[0105] S1, traverse the index feature sequences of the index feature sequence library, and obtain the similarity between each index feature sequence and the second feature sequence in turn;
[0106] S2, determining N index feature sequences with similarities higher than a target threshold as target feature sequences, where N is an integer greater than or equal to 0;
[0107] S3, obtaining target question information corresponding to the target feature sequence from all candidate question information in the target question bank.
[0108] It should be noted that the index feature sequences of the index feature sequence library are traversed, and the similarity between each index feature sequence and the second feature sequence is obtained in turn. The N index feature sequences with similarity higher than the target threshold are determined as target feature sequences, and the target question information corresponding to the target feature sequence is obtained from all the candidate question information in the target question library.
[0109] To further illustrate, optionally, when the second feature sequence is obtained, the second feature sequence can be directly matched with all feature sequences in the index feature sequence library for similarity to obtain the top-K target feature sequences with a similarity greater than the target threshold. Furthermore, since a question index has been established in the index feature sequence library, the corresponding target question information can be directly found in the target question library through the target feature sequence of the index feature sequence library.
[0110] Through the embodiments provided by the present application, the index feature sequences of the index feature sequence library are traversed, and the similarity between each index feature sequence and the second feature sequence is obtained in turn; N index feature sequences with similarities higher than the target threshold are determined as target feature sequences, where N is an integer greater than or equal to 0; and the target question information corresponding to the target feature sequence is obtained from all the candidate question information in the target question library, thereby achieving the purpose of efficiently obtaining the target question information in the feature sequence dimension and realizing the effect of improving the efficiency of obtaining question information.
[0111] As an optional solution, N index feature sequences with similarities higher than a target threshold are determined as target feature sequences, including:
[0112] S1, determining, from the synonym list, a fourth feature value that has a synonym relationship with the third feature value in the N index feature sequences;
[0113] S2, obtaining N target index feature sequences based on the fourth eigenvalue and the N index feature sequences;
[0114] S3: Determine the N target index feature sequences as target feature sequences.
[0115] Optionally, in this embodiment, synonym filtering based on the synonym list is performed not only on the acquired target title information, but also on the index feature sequence of the index feature sequence library, and the N target index feature sequences are accurately determined as target feature sequences by using a two-way synonym filtering method.
[0116] Through the embodiments provided in the present application, a fourth eigenvalue having a synonymous relationship with the third eigenvalue in N index feature sequences is determined from a synonym list; based on the fourth eigenvalue and the N index feature sequences, N target index feature sequences are obtained; the N target index feature sequences are determined as target feature sequences, thereby achieving the purpose of accurately obtaining the N target index feature sequences and realizing the effect of improving the accuracy of obtaining the target feature sequences determined based on the N target index feature sequences.
[0117] As an optional solution, for ease of understanding, a target system of the above information acquisition method is also provided, for example Figure 8 As shown, the input of the target system 802 is a new question, and the output is a list of repeated exercises related to the question. The overall framework process is as follows: First, relevant normalization operations are performed on the query question 804 to convert it into a relatively unified question expression; second, the processed question information is input into the matching module, and a specific query feature sequence is extracted based on the unified expression of the question; third, through this feature sequence, some candidate repeated questions are recalled from the entire question bank based on the recall module as a candidate question bank, wherein the recall module can also be used but not limited to establish a question index to form a full question bank; third, based on the candidate feature sequences extracted from the candidate repeated questions in the candidate question bank and the extracted query feature sequences, competitive matching is performed to calculate the sequence matching score, third, a comprehensive weighted calculation is performed on the specific matching score to obtain the total score, third, based on the ranking module, the repeated scores are ranked according to the obtained total score, and finally, the scores of the questions are ranked according to the level of the repetition score through a re-ranking model, and the general threshold size determines the number of repeated question lists output, and then the repeated question list 806 is output, wherein the candidate question bank is used to represent the question information set at different stages, which is not limited here.
[0118] Through the above-mentioned target system 802, a set of algorithms for identifying repeated questions is provided, and the techniques of question normalization, recall, sorting, and re-sorting are used to identify repeated questions in the question bank. Specifically, certain normalization operations are performed on synonym expressions, formula expressions, measurement units, common errors, etc. of the questions, and the general expressions of the questions are transformed to adapt to different expressions to improve the effect of feature extraction. Secondly, feature representations of different levels of granularity such as text, formulas, and n-grams are extracted for the general expressions of the questions, and then the feature representations are used for recall and matching. A comprehensive scoring is performed on the matching scores of different feature representations, and then they are sorted in order of high and low scores. Finally, a re-sorting module based on business logic is used to find a specific duplicate list for a specific question.
[0119] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0120] According to another aspect of the embodiments of the present invention, there is also provided an information acquisition device for implementing the above information acquisition method. Figure 9 As shown, the device includes:
[0121] A first acquiring unit 902 is configured to acquire the topic information of the target topic;
[0122] A first extraction unit 904 is configured to perform feature extraction on the topic information of the target topic to obtain a first feature sequence;
[0123] A first determining unit 906 is configured to determine, from a synonym list corresponding to a configured target question bank, a second feature value that is synonymous with the first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values that are synonymous, and the feature values that are synonymous have the same semantics;
[0124] A second acquiring unit 908 is configured to acquire a second feature sequence based on the second feature value and the first feature sequence;
[0125] The third acquiring unit 910 is configured to acquire target question information corresponding to a target feature sequence from a target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0126] Optionally, in this embodiment, the information acquisition device can be applied to, but not limited to, an educational client for obtaining repeated question information of an input question, such as Figure 1In the educational client running in the user device 102 shown, in particular, for educational scenarios such as the need to obtain multiple solutions to a question, the need to obtain the number of repetitions of a question, and the need to obtain similar question information for a question, for example, a large number of repeated questions will appear in the question list of similar question recommendations in the consolidation practice and test paper setting scenarios, resulting in a poor user experience. The above-mentioned information acquisition device can be used to remove duplicate questions to improve the user experience. For example, the information of repeated questions can be understood as a kind of popularity, so the above-mentioned information acquisition device can be used as a solution for generating popularity tags. For example, the above-mentioned information acquisition device can be used to find different questions stored in the question bank, and then use the parsing device of the question itself to extract different solution methods, and then provide multiple solution tags for the question as the characteristics of the question for the user to answer and choose. In addition, the above-mentioned information acquisition device also obtains and removes duplicate question information of the target question in the deduplication scenario. For example, when the question is stored in the database, the question is selected and the above-mentioned information acquisition device is used to filter the duplicate questions so that the duplicate question information is not stored in the database, reducing the storage pressure of the question inventory. The above is only an example and is not limited in this embodiment.
[0127] Optionally, in this embodiment, the target topic may be, but is not limited to, a topic input by the user in the corresponding target client, and may also be, but is not limited to, any topic stored in the initial topic bank associated with the target client;
[0128] It should be noted that when the target question is a question entered by a user on a corresponding target client, the information acquisition device can be used, but is not limited to, in educational scenarios such as obtaining information about multiple solutions to a question, obtaining information about the number of repeated answers to a question, or obtaining information about similar questions to a question. When the target question is any question stored in an initial question bank associated with the target client, the information acquisition device can be used, but is not limited to, in scenarios where duplicates are removed from a large number of questions stored in the initial question bank. The above is merely an example and is not intended to be limiting in this embodiment.
[0129] Optionally, in this embodiment, the question information may be, but is not limited to, key information of the target question, or may be, but is not limited to, all information of the target question. For example, the question information may be, but is not limited to, the title information or stem information of the target question. For another example, the question information may be, but is not limited to, text information or formula information of the target question.
[0130] Optionally, in this embodiment, obtaining the target question information may include, but is not limited to, first obtaining the first question information of the target question, then normalizing the first question information to obtain the second question information, and using the second question information as the target question information. Optionally, normalization may include, but is not limited to, a method of simplifying calculations, i.e., converting a dimensioned expression into a dimensionless expression, thus becoming a scalar. Optionally, in this embodiment, the normalization may include, but is not limited to, converting the first question information of the target question into relatively unified question expression information.
[0131] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0132] As an optional solution, the second obtaining unit 908 includes:
[0133] a first processing module, configured to sequentially merge the first eigenvalue and a second eigenvalue having a synonymous relationship with the first eigenvalue to obtain a target eigenvalue corresponding to the first eigenvalue, wherein the target eigenvalue includes one of the following: the first eigenvalue or the second eigenvalue, or a eigenvalue obtained by fusion of the first eigenvalue and the second eigenvalue;
[0134] The second processing module is configured to combine target feature values corresponding to the first feature values into a second feature sequence.
[0135] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0136] As an optional solution, it also includes:
[0137] a second determining unit configured to determine, from the first feature sequence, a first feature value associated with a stop feature value in a filter word list during a process of obtaining the second feature sequence based on the second feature value and the first feature sequence, wherein the filter word list includes at least two stop feature values, and the stop feature values are feature values unrelated to the topic information;
[0138] The deleting unit is configured to delete the first feature value associated with the disabled feature value in the first feature sequence during the process of obtaining the second feature sequence based on the second feature value and the first feature sequence.
[0139] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0140] As an optional solution, the first extraction unit 904 includes:
[0141] A first extraction module is used to perform text feature extraction on the title information of the target topic when the title information of the target topic includes text information;
[0142] The second extraction module is configured to perform formula feature extraction on the target topic information when the target topic information includes formula information.
[0143] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0144] As an optional solution, the first extraction module includes at least one of the following: a first extraction submodule for performing text feature extraction on text characters contained in the title information of the target question to obtain text character sequence features; a second extraction submodule for performing text feature extraction on a segmentation result obtained after segmenting the text information in the title information of the target question to obtain segmentation sequence features;
[0145] Furthermore, the second extraction module includes at least one of the following: a third extraction sub-module, which is used to extract formula features from the formula character information in the title information of the target question to obtain the formula original string sequence features; a fourth extraction sub-module, which is used to extract formula features from the parsing results obtained after parsing the formula information in the title information of the target question to obtain the formula parsing result sequence features.
[0146] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0147] As an optional solution, it includes:
[0148] The second extraction unit is used to extract features from the candidate question information recorded in the target question bank before obtaining the question information of the target question, so as to obtain an index feature sequence corresponding to each candidate question information;
[0149] The establishing unit is used to establish an index feature sequence library based on the index feature sequence before obtaining the topic information of the target topic.
[0150] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0151] As an optional solution, the third obtaining unit 910 includes:
[0152] A first acquisition module is used to traverse the index feature sequences of the index feature sequence library and sequentially obtain the similarity between each index feature sequence and the second feature sequence;
[0153] a determination module, configured to determine N index feature sequences having a similarity greater than a target threshold as target feature sequences, where N is an integer greater than or equal to 0;
[0154] The second acquisition module is used to acquire target question information corresponding to the target feature sequence from all candidate question information in the target question bank.
[0155] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0156] As an optional solution, it includes:
[0157] A first determining submodule is configured to determine, from the synonym list, a fourth feature value that has a synonym relationship with the third feature value in the N index feature sequences;
[0158] An acquisition submodule, configured to acquire N target index feature sequences based on the fourth eigenvalue and the N index feature sequences;
[0159] The second determining submodule is configured to determine N target index feature sequences as target feature sequences.
[0160] For specific embodiments, reference may be made to the examples shown in the above information acquisition method, which will not be repeated here.
[0161] According to another aspect of the embodiments of the present invention, an electronic device for implementing the above information acquisition method is also provided. Figure 10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0162] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0163] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0164] S1, obtain the topic information of the target topic;
[0165] S2, performing feature extraction on the topic information of the target topic to obtain a first feature sequence;
[0166] S3, determining, from a synonym list corresponding to the configured target question bank, a second feature value that has a synonym relationship with the first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values that have a synonym relationship, and the feature values that have a synonym relationship have the same semantics;
[0167] S4, obtaining a second feature sequence based on the second feature value and the first feature sequence;
[0168] S5. Obtain target question information corresponding to the target feature sequence from the target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0169] Alternatively, those skilled in the art will appreciate that Figure 10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 10 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 10 Different configurations shown.
[0170] Among them, the memory 1002 can be used to store software programs and modules, such as program instructions / modules corresponding to the information acquisition method and device in the embodiment of the present invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, realizes the above-mentioned information acquisition method. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely located relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1002 can be used to store, but is not limited to, information such as the topic information of the target topic, the first feature sequence, the second feature sequence, and the target topic information. As an example, if Figure 10 As shown, the memory 1002 may include, but is not limited to, the first acquisition unit 902, the first extraction unit 904, the first determination unit 906, the second acquisition unit 908, and the third acquisition unit 910 in the information acquisition device. In addition, the memory 1002 may also include, but is not limited to, other module units in the information acquisition device, which will not be described in detail in this example.
[0171] Optionally, the transmission device 1006 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1006 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1006 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0172] In addition, the electronic device further includes: a display 1008 for displaying information such as the target topic information, the first feature sequence, the second feature sequence, and the target topic information; and a connection bus 1010 for connecting the various module components in the electronic device.
[0173] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0174] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned information acquisition method, wherein the computer program is configured to perform the steps of any of the aforementioned method embodiments when executed.
[0175] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0176] S1, obtain the topic information of the target topic;
[0177] S2, performing feature extraction on the topic information of the target topic to obtain a first feature sequence;
[0178] S3, determining, from a synonym list corresponding to the configured target question bank, a second feature value that has a synonym relationship with the first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values that have a synonym relationship, and the feature values that have a synonym relationship have the same semantics;
[0179] S4, obtaining a second feature sequence based on the second feature value and the first feature sequence;
[0180] S5. Obtain target question information corresponding to the target feature sequence from the target question bank, wherein the similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
[0181] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0182] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0183] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods of various embodiments of the present invention.
[0184] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, there may be other division methods, such as combining or integrating multiple units or components into another system, or ignoring or not implementing some features. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of units or modules, and may be electrical or other forms.
[0186] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0187] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0188] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. An information acquisition method, characterized in that: include: Get the first question information of the target question; Normalizing the first topic information to convert the first topic information into unified topic expression information to obtain second topic information; performing feature extraction on the second topic information to obtain a first feature sequence; Determining, from a synonym list corresponding to a configured target question bank, a second feature value having a synonym relationship with a first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values having the synonym relationship, and the feature values having the synonym relationship have the same semantics; sequentially merging the first feature value and the second feature value having the synonymous relationship with the first feature value to obtain a target feature value corresponding to the first feature value, wherein the target feature value includes: the first feature value, the second feature value, and a feature value obtained by fusion of the first feature value and the second feature value; combining the target feature values corresponding to the respective first feature values into a second feature sequence; Target question information corresponding to a target feature sequence is acquired from the target question bank, wherein a similarity between the target feature sequence and the second feature sequence is higher than a target threshold.
2. The method according to claim 1, characterized in that In the step of sequentially merging the first feature value and the second feature value having the synonymous relationship with the first feature value to obtain a target feature value corresponding to the first feature value, wherein the target feature value includes: the first feature value, the second feature value, and a feature value obtained by fusion of the first feature value and the second feature value; and combining the target feature values corresponding to the respective first feature values into a second feature sequence, the step further includes: Determining, from the first feature sequence, the first feature value associated with a disabled feature value in a filter word list, wherein the filter word list contains at least two disabled feature values, each of which is irrelevant to the topic information; The first feature value having the association relationship with the disabled feature value in the first feature sequence is deleted.
3. The method according to claim 1, characterized in that The performing feature extraction on the second topic information includes: In a case where the second topic information includes text information, performing text feature extraction on the second topic information; In a case where the second topic information includes formula information, formula feature extraction is performed on the second topic information.
4. The method according to claim 3, characterized in that The performing text feature extraction on the second topic information includes at least one of the following: performing text feature extraction on text characters included in the second topic information to obtain text character sequence features; performing text feature extraction on a segmentation result obtained after segmenting the text information in the second topic information to obtain a segmentation sequence feature; The performing of formula feature extraction on the second topic information includes at least one of the following: performing formula feature extraction on formula character information in the second topic information to obtain formula original string sequence features; The formula feature extraction is performed on the parsed result obtained after parsing the formula information in the second question information to obtain a formula parsing result sequence feature.
5. The method according to claim 1, wherein Performing feature extraction on the candidate question information recorded in the target question bank to obtain an index feature sequence corresponding to each candidate question information; An index feature sequence library is established based on the index feature sequence.
6. The method according to claim 5, characterized in that The step of obtaining target question information corresponding to a target feature sequence from the target question bank includes: Traversing the index feature sequences in the index feature sequence library, and sequentially obtaining the similarity between each index feature sequence and the second feature sequence; Determining N index feature sequences whose similarity is higher than the target threshold as the target feature sequences, where N is an integer greater than or equal to 0; The target question information is obtained from all the candidate question information in the target question bank.
7. The method according to claim 6, characterized in that Determining the N index feature sequences having similarities higher than the target threshold as the target feature sequences includes: Determining, from the synonym list, a fourth feature value having the synonym relationship with the third feature value in the N index feature sequences; Based on the fourth eigenvalue and the N index feature sequences, obtaining N target index feature sequences; The N target index feature sequences are determined as the target feature sequences.
8. An information acquisition device, characterized in that: include: A first acquiring unit, configured to acquire first topic information of a target topic; Normalizing the first topic information to convert the first topic information into unified topic expression information to obtain second topic information; a first extraction unit, configured to perform feature extraction on the second topic information to obtain a first feature sequence; a first determining unit configured to determine, from a synonym list corresponding to a configured target question bank, a second feature value having a synonym relationship with the first feature value in the first feature sequence, wherein the synonym list records at least two groups of feature values having the synonym relationship, and the feature values having the synonym relationship have the same semantics; A second acquiring unit, configured to acquire a second feature sequence based on the second feature value and the first feature sequence; a third acquiring unit, configured to acquire target question information corresponding to a target feature sequence from the target question bank, wherein a similarity between the target feature sequence and the second feature sequence is higher than a target threshold; The second acquiring unit includes: a first processing module, configured to sequentially merge the first feature value and the second feature value having the synonym relationship with the first feature value to obtain a target feature value corresponding to the first feature value, wherein the target feature value includes: the first feature value, the second feature value, and a feature value obtained by fusion of the first feature value and the second feature value; The second processing module is configured to combine the target feature values corresponding to the first feature values into a second feature sequence.
9. The device according to claim 8, characterized in that Also includes: a second determining unit configured to determine, from the first feature sequence, during the process of obtaining the second feature sequence based on the second feature value and the first feature sequence, the first feature value associated with a stop feature value in a filter word list, wherein the filter word list contains at least two stop feature values, each of which is irrelevant to the topic information; The deleting unit is configured to delete the first feature value associated with the disabled feature value in the first feature sequence during the process of acquiring the second feature sequence based on the second feature value and the first feature sequence.
10. The device according to claim 8, characterized in that The first extraction unit includes: a first extraction module, configured to perform text feature extraction on the second topic information when the second topic information includes text information; The second extraction module is configured to perform formula feature extraction on the second topic information when the second topic information includes formula information.
11. The device according to claim 10, characterized in that The first extraction module includes at least one of the following: a first extraction submodule, configured to perform text feature extraction on the text characters included in the second topic information to obtain text character sequence features; A second extraction submodule is configured to perform text feature extraction on a segmentation result obtained after segmenting the text information in the second topic information to obtain a segmentation sequence feature; The second extraction module includes at least one of the following: a third extraction submodule, configured to perform the formula feature extraction on the formula character information in the second title information to obtain the formula original string sequence feature; The fourth extraction submodule is configured to perform the formula feature extraction on the parsing result obtained after parsing the formula information in the second question information, and obtain a formula parsing result sequence feature.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.
13. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Approximate question push method and system
CN106651696A
Interactive question and answer control method and device
CN109670013A