A knowledge system construction method based on text-based minimal information units

By utilizing full-stack full-text search technology and a term node network, the inefficiency of linear reading methods is solved, enabling personalized knowledge management and rapid searching, thus meeting the high-density reading needs of modern society.

CN113886561BActive Publication Date: 2026-02-06SHANGSHUTAI TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111159569.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2026-02-06
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing technologies lack efficient knowledge management methods, making it difficult to transform the massive amounts of information and book texts in the Internet age into digital knowledge that can be learned quickly. Linear reading methods do not conform to the brain's habit of jumping around in search.

Method used

By using full-stack full-text search technology to sort text by keywords, establish a network of term nodes, and generate a network knowledge system, personalized knowledge indexes are built using keyword entries and user editing needs.

Benefits of technology

It enables non-linear and personalized knowledge management, improves reading speed and information retrieval efficiency, conforms to brain cognitive habits, and supports multi-dimensional search and personalized knowledge construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886561B_ABST
    Figure CN113886561B_ABST
Patent Text Reader

Abstract

The application provides a knowledge system construction method based on text minimum information units. The text is subjected to keyword sorting based on full-stack full-text search technology, and the sorted keywords are subjected to auditing through text source auditing rules to determine a text term library; the relevance between terms is judged according to the text term library, and a term node network is established based on the relevance; a multi-text network is established according to the term node network, and a network knowledge system is generated. The application has the beneficial effect that the application understands the definition of key concepts in the text through keyword interconnection, so as to summarize the core content of a book as quickly as possible, and realizes super-fast reading; the term cloud of the application can stimulate creativity, produce relevance between two adjacent but seemingly irrelevant terms, and inspire new knowledge of innovation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital books, and particularly relates to a knowledge system construction method based on minimized information units of texts. BACKGROUND

[0002] At present, in the face of massive information contents and book texts in the Internet era, how to convert them into digital knowledge that can be quickly learned, the industry has not proposed an efficient knowledge management method with the help of modern technology.

[0003] At the same time, the contradiction between the increasing digital publications and the increasingly occupied reading time is becoming increasingly acute, and an innovative method and tool for converting reading into digital knowledge is urgently needed.

[0004] In the scenario of learning through reading, users read texts, and often sequentially browse according to the text logic and layout, that is, the so-called linear reading mode. The notes made by users are also recorded and summarized in the original order of the text. However, this does not conform to the habit of jumping eyes to search for things that can cause brain excitement. Therefore, a learning method of discrete point recording aiming to improve learning efficiency is more suitable for the needs of the Internet. SUMMARY

[0005] The present application provides a knowledge system construction method based on minimized information units of texts, to solve the problem that the existing method lacks efficient knowledge management.

[0006] A knowledge system construction method based on minimized information units of texts, comprising:

[0007] Based on the full-stack full-text search technology, the keywords of the text are sorted, and the sorted keywords are audited through the text source auditing rules to determine the text term library;

[0008] According to the text term library, the relevance between terms is judged, and based on the relevance, a term node network is established;

[0009] According to the term node network, a multi-text network is established, and a network knowledge system is generated.

[0010] In an embodiment of the present application: the full-stack full-text search technology is used to sort the keywords of the text, comprising:

[0011] The target text is obtained, and the keywords of the target text are determined through the full-stack full-text search technology;

[0012] According to the keywords, the corresponding keyword paragraphs are determined;

[0013] According to the keyword paragraph, determine the keyword in the paragraph;

[0014] According to the keyword, determine the keyword ranking.

[0015] In an embodiment of the present application: the sorted keywords are audited by the text source audit rule to determine the text term library, comprising:

[0016] Get the sorted keywords and determine the corresponding keyword;

[0017] According to the keyword, determine the key definition of each keyword, and determine the key definition;

[0018] Based on the preset audit editing rule, the key definition is audited to determine the corresponding original text of each key definition;

[0019] Get the publication time of the original text and perform publication sorting;

[0020] According to the publication sorting, establish a text term library.

[0021] In an embodiment of the present application: the sorted keywords are audited by the text source audit rule to determine the text term library, further comprising:

[0022] Get the request information of the user searching for the target text, and feedback the request result;

[0023] According to the request result, get the text content of the target text;

[0024] According to the text content, receive the user's editing demand;

[0025] According to the editing demand, authorize the user's editing right;

[0026] According to the editing right, receive the user's added user term to the target text, and add the user term to the text term library.

[0027] In an embodiment of the present application: according to the text term library, determine the relevance between the terms, comprising:

[0028] According to the text term library, determine the corresponding term sub-set of each text;

[0029] According to the term sub-set, get the key definition of each term in the term sub-set;

[0030] According to the key definition, based on the correlation calculation, judge the relevance of different terms in the term sub-set;

[0031] According to the relevance of the different entries, relevance of different texts is determined;

[0032] According to the relevance of the different texts, relevance between different entries in the text entry library is determined.

[0033] In an embodiment of the present application, the entry node network is established based on the relevance, comprising:

[0034] According to the relevance, a defined distance between different entries in the text entry library is determined.

[0035] According to the defined distance, the different entries in the text entry library are connected respectively to determine a connection relationship.

[0036] According to the connection relationship, an entry node network is constructed; wherein,

[0037] The entry node network comprises a single text entry node network and a multi-text entry node network.

[0038] In an embodiment of the present application, the entry node network is established based on the relevance, further comprising:

[0039] According to the text entry library, an entry parent group constructed by entries of multiple texts is determined.

[0040] It is judged whether there are same entries associated with different texts in the entry parent group; wherein,

[0041] When there are same entries in the different texts, the same entries are deleted.

[0042] When there are no same entries in the different texts, an entry cloud is generated according to the entry parent group.

[0043] In an embodiment of the present application, the entry node network is established based on the relevance, further comprising:

[0044] According to the relevance, a correlation between different entries is determined; wherein,

[0045] The correlation comprises bidirectional strong correlation, unidirectional strong correlation, bidirectional weak correlation and unidirectional weak correlation.

[0046] In an embodiment of the present application, the entry node network is established based on the relevance, further comprising:

[0047] According to the text entry library, entries are audited based on preset auditing rules respectively; wherein,

[0048] The auditing rules comprise objective definition auditing and academic value auditing.

[0049] Obtaining an audit result after the audit, and judging the legitimacy between the terms according to the audit result and the relevance.

[0050] In an embodiment of the present application: the multi-text network is established according to the term node network, and a network knowledge system is generated, comprising:

[0051] According to the term node network, the calling relationship between the terms and the texts is established;

[0052] According to the calling relationship, the relationship graph and the directory graph between the terms and the texts are established, and the individualized unconventional connection relationship between the terms is set;

[0053] According to the relationship graph, the directory graph and the individualized unconventional connection relationship, the multi-text network is determined;

[0054] Through the multi-text network, the network knowledge system of different texts is established.

[0055] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims.

[0056] The technical solutions of the present application will be further described in detail below with the help of the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0057] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application serve to explain the present application, and do not constitute a limitation on the present application. In the drawings:

[0058] Figure 1 A method flow chart of a text-based minimum information unit knowledge system construction method in an embodiment of the present application;

[0059] Figure 2 A network knowledge correlation graph in an embodiment of the present application;

[0060] Figure 3 A star cloud graph of different term correlations in an embodiment of the present application;

[0061] Figure 4 A step graph of keyword sorting on text by the full-stack full-text search technology of the present application;

[0062] Figure 5 A step graph of establishing a term text library of the present application

[0063] Figure 6 An interface graph of book search of the present application. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present application will be described herein below with reference to the accompanying drawings, in which it should be understood that the preferred embodiments described herein are intended for the purpose of explanation only and are not intended to limit the present application.

[0065] The present application is a solution developed for the current information overload situation, because of the homogenization of the multi-fusion of text language, the efficient extraction of the core content of the text, by simplifying the text into the smallest unit of knowledge, and connecting into a set, by integrating the same similar terms between different texts, all the terms are connected into a unified network. Knowledge interconnection, thus establishing a knowledge system.

[0066] Noun definition:

[0067] Term: Term is also called term, entry, dictionary, is a dictionary of language, refers to the collection of dictionary editing, words and their explanations. Term can be a word, or can be composed of words.

[0068] Term mother group: The term set composed of terms of different books is called term mother group.

[0069] Term cloud: A term set composed of terms of different books, and no repeated terms.

[0070] Key words: Specifically refers to the words used by the media in the production and use of index. It is a library science vocabulary. Key word search is one of the main methods of network search index, that is, the specific name of the product, service and company that the visitor wants to understand.

[0071] Text term library: A database formed by combining terms of multiple texts.

[0072] Full-text retrieval technology: It is a retrieval technology that takes data such as text, sound, image as the main content, and retrieves the content of the document rather than the external characteristics. The main full-text retrieval system includes TRS system, Tianyu system, etc.

[0073] Minimum information unit: The minimum information unit is a term, and each term represents a minimum information unit.

[0074] As shown in the accompanying Figure 1 The present application is a knowledge system construction method based on the minimum information unit of text, which comprises:

[0075] Based on the full-stack full-text search technology, the keywords of the text are sorted, and the sorted keywords are audited by the text source auditing rules to determine the text term library;

[0076] According to the text term library, the relevance between terms is determined, and a term node network is established based on the relevance;

[0077] According to the term node network, a multi-text network is established, and a mesh knowledge system is generated.

[0078] The principle of the above technical solution is that the present application is connected by terms based on keywords through text (including books, journals, documents and other text reading materials, of course, the present application also relates to the technical field of video, sound and picture, through the key elements in the picture or the words on the picture, the key words in the sound and the subtitle words in the video), to form an auxiliary knowledge system which can facilitate fast query of books and reading through terms, and can improve the reading speed of users. The full-stack full-text search technology of the present application is based on full-text retrieval technology, that is, a retrieval technology which takes data such as text, sound, image, etc. as the main content, and retrieves the content of the document material rather than the external characteristics. The present application finds the paragraph with the keyword in this way, then realizes the sequential ordering of the keyword through the paragraph, that is, the front-back order of reading, and finally the term needs to be audited by the text source, that is, the corresponding original text is audited as the term source, and only after the audit is passed can it be used as a term to form a term database of text. The relevance between terms is determined because there are many terms in the text, and through the relevance of these terms, fast reading can be assisted. This reading is based on reading guidance between different terms, and at the same time, because of the relevant relationship, fast text searching can also be realized. Finally, the present application constructs a term node network through the relevance of these terms, which is achieved by taking the same text as a term set, taking multiple texts as a term mother group, realizing the comprehensive use of all texts, and then realizing the construction of a knowledge system based on terms and mesh association architecture.

[0079] The beneficial effects of the above technical solution are that the user can use the term as a search keyword, use general text full-stack search technology to obtain target text, and also use the search of the target term to find related terms and thus find the associated books. The application simulates the brain cognitive mechanism, allows the user to establish the connection between terms according to his own cognition, and since the connection between the terms and the text (books) is rational, the user also obtains a personalized knowledge index. The relevance of the search statement recommends the search intention to the user and supports multi-dimensional, fuzzy / accurate matching of search statements and other functions. This cross-text and cross-field association between terms provides the possibility of creating new knowledge for individual users. This knowledge management method is heuristic and creative for individual knowledge construction. Thinking is nonlinear and networked. Restating in the manner consistent with the brain habit can more efficiently reorganize information. The difficulty of memory is formally reflected in the process of actively reorganizing information. The linear presentation connection mode of the existing text follows the syntax structure (tree structure) and arranges the character string (linear order). This data presentation mode with words as the carrier does not adapt to the fast pace and high density of reading requirements in modern society. A nonlinear note visualization tool can embody this theory. It helps users effectively call notes, strengthens memory, and quickly builds a basic knowledge framework.

[0080] In one embodiment of the application: the keyword sorting of the text based on the full-stack full-text search technology comprises:

[0081] Obtaining the target text and determining the keywords of the target text through the full-stack full-text search technology;

[0082] Determining the corresponding keyword paragraph according to the keywords;

[0083] Determining the keyword term in the paragraph according to the keyword paragraph;

[0084] Determining the keyword sorting according to the keyword term.

[0085] The principle and beneficial effects of the above technical solution are:

[0086] In an actual scenario: as shown in the accompanying drawings Figure 4As shown, in actual implementation of the present application, a text is first determined based on the full-stack full-text search technology to determine the keywords of the target text, that is, the keywords are marked by a marking method, and then the paragraph where the keywords are located is determined, and the keywords may exist in different paragraphs but the same keywords exist in the paragraphs. Then in the paragraphs, each keyword has a corresponding keyword, for example, "just before he returned to the international stage in the second stage of his career", the internal keywords are career and international stage. After the keywords are determined, the paragraph can be located. Then the keyword is collected, for example, the second stage of the career. Because the keyword contains the keywords, the keyword can be directly read through the keyword, and therefore, the order of the keywords can be determined through the keyword. Figure 6 The interface diagram of the present application when searching for books.

[0087] First, the full-stack full-text search technology of the present application is suitable for finding the corresponding paragraph of the keywords of the text, and then the keywords can be sorted through the keyword in the paragraph, which is also a sort of reading order of the text. The full-stack full-text search technology has two functions, that is, the keywords paragraph and the keyword can be determined through the text, and the book retrieval can be realized through the keyword, and one technology has two applications.

[0088] In an embodiment of the present application, the sorted keywords are audited by the text source auditing rule to determine the text keyword library, including:

[0089] The sorted keywords are obtained, and the corresponding keyword is determined.

[0090] According to the keyword, the key definition of each keyword is determined, and the key definition is determined.

[0091] Based on the preset auditing and editing rule, the key definition is audited to determine the original text corresponding to each key definition.

[0092] The publication time of the original text is obtained, and the publication order is determined.

[0093] According to the publication order, the text keyword library is established.

[0094] The principle and beneficial effects of the above technical solution are that:

[0095] In an actual scenario, for example, Figure 5As shown, the present application is not only to identify the key words of the target text, but also to expand the word library according to the text. The present application is provided with a word database, and after the text source is determined, the corresponding word database can be determined according to the content of the text source, for example: art has art data database, computer has computer database. Before determining the text database, each key word has its necessary key definition, for example: computer has the definition of computer, which belongs to the computer database. The pre-audit editing rule is a kind of artificial and computer auditing rule. The computer is responsible for outputting the word form and determining which book the word belongs to, and then determining the book where the word is located. The original text of the book is when the book is sold, that is, the publication time. After the publication time of multiple books is determined, the publication order is determined.

[0096] In the process of text word library of the present application, that is, the process of word acquisition and word audit, the key words and the text are corresponded, which is a kind of text source based audit mechanism, mainly to ensure the correctness of the word; secondly, the text word library is also the basis of the knowledge system, which contains all the text words, and is convenient for the correlation calculation and judgment between the words.

[0097] In an embodiment of the present application, the sorted key words are audited by the text source audit rule to determine the text word library, further comprising:

[0098] Obtain the request information of the user searching the target text, and feed back the request result;

[0099] According to the request result, obtain the text content of the target text;

[0100] According to the text content, receive the editing demand of the user;

[0101] According to the editing demand, authorize the editing right to the user;

[0102] According to the editing right, receive the user word added to the target text by the user, and add the user word to the text word library.

[0103] The principle and beneficial effects of the above technical scheme are:

[0104] In an actual scenario: in the micro letter application or system of the application, a user's request in the form of voice or text is collected, for example, "find Outlaws of the Marsh"; then the system will display the corresponding book on the feedback interface. After the book is determined, the specific content of the book will appear in the form of an electronic book. If the text needs to be edited, that is, some terms are added, the user will have the right to edit, and after having the right, the user can directly write or select some words in the book as terms.

[0105] The term of the application is not only determined based on the auditing mechanism of the system itself, but also can set personalized terms of the user through the editing demand of the user. If the personalized term is better than the system, it will replace the term in the system. However, the editing right of the term needs to be opened to the user when adding the term. The key term is an important minimum unit for describing the content of the text. Usually, the term in scientific papers, monographs and popular science books is in the form of concept. The key term of the system also follows this academic habit, which is objectively summarized by the user according to the actual content of the target text (book). In order to ensure objectivity and seriousness, the term must come from the target text (book), and the original text is cited as the definition.

[0106] In an embodiment of the application: the correlation between the terms is determined according to the text term library, including:

[0107] According to the text term library, the term sub-set corresponding to each text is determined;

[0108] According to the term sub-set, the key definition of each term in the term sub-set is obtained;

[0109] According to the key definition, the correlation between different terms in the term sub-set is determined based on the correlation calculation;

[0110] According to the correlation of different terms, the correlation of different texts is determined;

[0111] According to the correlation of different texts, the correlation between different terms in the text term library is determined.

[0112] The principle and beneficial effects of the above technical solution are:

[0113] In an actual scenario: a text term library will exist multiple text terms, each text has its corresponding term sub-set, for example: in the novel book database, there is a corresponding term sub-set in "The Three Kingdoms", and there is a corresponding term sub-set in "Outlaws of the Marsh". The correlation of the terms is mainly that one term can lead to the next term, or both belong to the same text, which means that there is a correlation. The closer the two terms are, the stronger the correlation is. For example:

[0114] "Noise" is a major new work by Nobel laureate in economics and the "father of behavioral economics," Daniel Kahneman, in collaboration with decision-making experts Olivier Syboni and Cass Sunstein. It is a globally acclaimed milestone, the culmination of Kahneman's decade-long reflection following his bestseller, "Thinking, Fast and Slow," and represents another significant discovery in the field of behavioral science. For decades, it has been believed that bias is the key factor leading to human errors in judgment. However, today, Kahneman systematically points out that noise is the black hole affecting human judgment. Through systematic research, "Noise" reveals the essential quantity of "judgment errors" using two formulas.

[0115] In the above case, the entries for the names Olivier Syboni and Cass Sunstein are quite related.

[0116] In determining the relevance of terms, this invention considers that a text may contain multiple keyword terms, and the more terms a text contains, the more accurate it is. This invention uses relevance calculation to determine the relevance between terms. During the calculation process, pre-inertial parameters such as reading order parameters, semantic parameters, and importance parameters are introduced. Of course, different texts may also contain the same terms, and this invention will also use higher-level relevance calculation to determine the relevance between different texts.

[0117] In one specific embodiment, the importance parameter is determined using the following scheme:

[0118] Step 1: Determine the semantic parameters for each text element based on the text itself.

[0119] ;

[0120] in, Indicates the first The semantic features of each term; Indicates type coefficient; Indicates the first The order parameters of each term; Indicates the first The number of characters in each entry; Indicates the first The weight of each term in the text; Indicates the first The first entry The semantic features of each character; Indicates the first One text feature parameter; , Indicates the total number of entries; , Indicates the total number of texts;

[0121] In the present invention is used to determine the semantic parameters of the term according to the term; is used to determine the specific content total of the current text, and is divided by is used to extract the ratio of the number of characters of the current term and the total number of characters of the text; and by accumulating the sum of the ratio of the number of characters of each term and the total number of characters of the text in this text, the capacity parameter of the overall subject text is determined; and is used to calculate the product of each character of each term and the text feature, and by the content parameter of each term relative to the text can be determined; the sum of the character number parameter and the content parameter of the text is set as the subject parameter.

[0122] Step 2: Determine the term content feature according to the feature of each term in the text:

[0123] ;

[0124] wherein, represents the type coefficient difference of the term; represents the type coefficient of the term; represents the range parameter of the knowledge involved in the text of the th term; represents the average range of the knowledge involved in the same type of term; represents the total range coefficient of the knowledge involved in the th term; represents the content total range of the th term belonging to the

[0125] th text; The calculation of the text content feature introduces the type coefficient difference of the term in the present invention, and by the range coefficient of the content involved in the current type of term is calculated; determines the total range coefficient of the knowledge involved in the current type of term; according to the product of the ratio of the range coefficient and the term content, the content feature of the current term of the term is determined, and by accumulating the calculation, the content feature of the term is determined;

[0126] Step 3: Determine the importance of the term in all terms according to the type parameter of the term and the content feature of the term:

[0127] ;

[0128] wherein, represents the importance of the term.

[0129] In calculating the importance of the text, after the content feature and the type parameter are determined, by the ratio of the specific content coefficient of the main content of each term and the total content of the text is calculated, the importance is determined.

[0130] The working principle of the technical solution is that the calculation of the content parameters is to determine the type attribute of each term, and the invention can only find any related text in the question bank through the term when searching for the text.

[0131] In an embodiment of the present application, the term node network is established based on the relevance, comprising:

[0132] According to the relevance, the defined distance between different terms in the text term library is determined;

[0133] According to the defined distance, the different terms in the text term library are connected respectively to determine the connection relationship;

[0134] According to the connection relationship, a term node network is constructed; wherein,

[0135] The term node network comprises a single text term node network and a multi-text term node network.

[0136] The principle and beneficial effects of the technical solution are:

[0137] In an actual scenario, the present application determines the defined distance between different terms according to the relevance, which adopts Mahalanobis distance algorithm, Manhattan distance or Chebyshev distance, mainly to determine the correlation degree between different terms, and constructs a node network through the distance between them, such as Figure 3 .

[0138] In the process of establishing the term node network, the defined distance between different terms, that is, the distance in meaning, is mainly used to connect the terms and finally form the term node network, realize the construction of the term node network, and finally determine the relationship between each term and all other terms, thereby realizing the construction of the network knowledge system. In this process, the distribution of the term node is as shown in Figure 3 , and the system will preset a (term) dot distance classification table. The distance between the dots is inversely proportional to the relevance between the corresponding terms. For example, the distance between the dots "France - French" is the smallest, the distance between "France - poetry" is the second, and the distance between "France - rugby" is the farthest. However, the distance mentioned here is irrelevant to the connection line. The purpose is only to realize the density of the term "star cloud".

[0139] In an embodiment of the present application, the term node network is established based on the relevance, further comprising:

[0140] According to the text term library, a term parent group constructed by the terms of multiple texts is determined;

[0141] judging whether the same word exists in the word group of different texts; wherein,

[0142] deleting the same word when the different texts have the same word;

[0143] generating a word cloud according to the word group when the different texts do not have the same word.

[0144] The principle and beneficial effects of the above technical solution are that:

[0145] In an actual scenario, the word group is composed of word sub-sets of multiple texts, that is, word sub-sets of multiple books, and in the word group, there may be repeated words that occupy space memory, but they are in multiple texts, so the invention deletes these words. For example, both the word group of electrical engineering foundation and the word group of power cable design have the word chopper circuit, so we can only leave one word chopper circuit.

[0146] The invention calls words in the cloud in the form of a word cloud, and in this process, the word cloud is expanded by new texts, so the word group is different texts, and the expansion of the word group is realized.

[0147] In an embodiment of the invention, the word node network is established based on the association, and further comprises:

[0148] judging the correlation between different words according to the association; wherein,

[0149] The correlation includes bidirectional strong correlation, unidirectional strong correlation, bidirectional weak correlation and unidirectional weak correlation.

[0150] The principle and beneficial effects of the above technical solution are that: in the process of establishing the node network, the judgment of word association includes various correlations.a. Bidirectional strong correlation (referring to the association that conforms to the association habit of most people, for example, France-Paris); b. Unidirectional strong correlation (referring to the single-directional relationship derived from A to B, for example, seed-sprout); c. Bidirectional weak correlation (referring to indirect and non-universal correlation relationship, which makes it possible to generate creativity and new knowledge); d. Unidirectional weak correlation (referring to indirect and unconventional single-directional derivation relationship, which makes it possible to generate creativity and new knowledge). Through the above-mentioned manner, the final judgment of word correlation can be realized, and the division of correlation relationship can also be accurately realized to determine how to connect the words connected in a network form.

[0151] In an embodiment of the invention, the word node network is established based on the association, and further comprises:

[0152] According to the text term library, the terms are audited based on preset auditing rules respectively, wherein,

[0153] The auditing rules include objective definition auditing and academic value auditing.

[0154] The auditing results after auditing are obtained, and the legality between terms is judged according to the auditing results and relevance.

[0155] The principle and beneficial effects of the above technical solution are that:

[0156] In an actual scenario: objective definition auditing is what the meaning of the term expressed in the objective aspect is? Academic value auditing is how high the value of the term in the corresponding text book is.

[0157] For example: alternating current, the objective definition is current, the current direction changes periodically with time, and the average current in a period is zero.

[0158] Academic value is how many parts of the alternating current are mentioned in the book, how many key knowledge parts in the application will use alternating current, and then calculate the weight in the text as the academic value.

[0159] The present application also audits the legality of the term, which is based on the objective definition and academic value of the term from two aspects, so that not only the legality of the term can be judged, but also the similarity of the term can be judged, and the similar terms can be merged.

[0160] In an embodiment of the present application: according to the term node network, a multi-text network is established, and a network knowledge system is generated, comprising:

[0161] According to the term node network, the calling relationship between terms and texts is established;

[0162] According to the calling relationship, the relationship graph and directory graph between terms and texts are established, and the individualized unconventional connection relationship between terms is set;

[0163] According to the relationship graph, directory graph and individualized unconventional connection relationship, the multi-text network is determined;

[0164] The principle and beneficial effects of the above technical solution are that: in an actual scenario: the knowledge finally realized by the present application exists in the relationship graph and directory graph, and in actual implementation, as follows: Figure 2The shown: Graphical Zone: Represented by the round dot form of the entry unit. Its size is directly related to the number of books in the index. The association of the entry is represented by: a. solid double-headed arrow represents bidirectional strong correlation; solid single-headed arrow represents unidirectional strong correlation; c. dotted double-headed arrow represents bidirectional weak correlation; d. dotted single-headed arrow represents unidirectional weak correlation. The round dot and the line constitute the "nebula" form of the entire entry group. The text (book) is classified into 20 categories, and is represented by 20 Chinese traditional colors (the specific classification and color list will be set in the appendix). The color of the dot is the color of the type of the associated book. The presentation of the search entry: the search result will be presented in the form of the magnification of the dot corresponding to the target entry. It will automatically move to the center of the screen, and the other entry dots associated with it will also move. The system will present the above behavior in the form of animation; the system requires them to automatically change the coordinate position to achieve the best visual effect, that is, the self-adaptive function of the figure.

[0165] In the face of the massive information content and book texts in the Internet era, how to convert them into knowledge, the industry has not yet proposed an efficient knowledge management method with the help of modern technology. At the same time, the contradiction between the increasing number of digital publications and the increasingly squeezed reading time is becoming increasingly acute, and an innovative method and tool for reading and converting knowledge is urgently needed. In the scenario of learning through reading, users often read texts in the order of the text's writing logic and layout, that is, the so-called linear reading method. The notes made by users are also recorded and summarized in the original order of the text. However, this does not conform to the habit of jumping eyes to search for things that can stimulate the brain. Therefore, the appearance of the present invention, in the face of the current information glut, the homogenization of the multi-fusion of text language, the efficient extraction of the core content of the text, by simplifying the text into the smallest unit of knowledge and connecting it into a set, by integrating the same similar entries between different texts, all entries are connected into a unified network. Knowledge interconnection, thus establishing a knowledge system.

[0166] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. A knowledge system construction method based on text-minimized information units, characterized by, The method comprises the following steps: Based on the full-stack full-text search technology, the keywords of the text are sorted, and the sorted keywords are audited by the text source auditing rules to determine the text term library; Wherein, the step of determining the text term library by auditing the sorted keywords through the text source auditing rules comprises the following steps: Obtain the sorted keywords and determine the corresponding keyword terms; According to the keyword terms, determine the key definition of each keyword term, and determine the key definition; Based on the preset auditing and editing rules, the key definition is audited to determine the original text corresponding to each key definition; Obtain the publication time of the original text and perform publication sorting; According to the publication sorting, a text term library is established; The step of determining the text term library by auditing the sorted keywords through the text source auditing rules further comprises the following steps: Obtain the request information of the user searching for the target text, and feed back the request result; According to the request result, the text content of the target text is obtained; According to the text content, the user's editing demand is received; According to the editing demand, the user is authorized with editing permission; According to the editing permission, the user term added by the user to the target text is received, and the user term is added to the text term library; According to the text term library, the relevance between the terms is judged, and based on the relevance, a term node network is established; According to the term node network, a multi-text network is established, and a network knowledge system is generated; Wherein, the step of establishing a multi-text network according to the term node network and generating a network knowledge system comprises the following steps: According to the term node network, the calling relationship between the terms and the text is established; According to the calling relationship, the relationship graph and the directory graph between the terms and the text are established, and the individualized unconventional connection relationship between the terms is set; According to the relationship graph, the directory graph and the individualized unconventional connection relationship, a multi-text network is determined; Through the multi-text network, a network knowledge system of different texts is established.

2. The knowledge system construction method of claim 1, wherein, The step of sorting the keywords of the text based on the full-stack full-text search technology comprises the following steps: Obtain the target text, and determine the keywords of the target text through the full-stack full-text search technology; According to the keywords, the corresponding keyword paragraphs are determined; According to the keyword paragraphs, the keyword terms within the paragraphs are determined; According to the keyword terms, the keyword sorting is determined.

3. The knowledge system construction method based on text-based minimized information units of claim 1, wherein, The step of judging the relevance between the terms according to the text term library comprises the following steps: According to the text term library, the term sub-set corresponding to each text is determined; According to the term sub-set, the key definition of each term in the term sub-set is obtained; According to the key definition, the relevance of different terms in the term sub-set is judged based on the correlation calculation; According to the relevance of different texts, the relevance between different terms in the text term library is determined. The step of establishing a term node network based on the relevance comprises the following steps:

4. The knowledge system building method based on text-based minimized information unit according to claim 1, characterized in that, According to the relevance, the definition distance between different terms in the text term library is judged; According to the definition distance, the connection relationship of different terms in the text term library is determined. ​ According to the connection relationship, a word entry node network is constructed; wherein The word entry node network comprises a single text word entry node network and a multi-text word entry node network.

5. The knowledge system building method based on text-based minimized information unit of claim 1, wherein, The word entry node network is established based on the relevance, and further comprises: According to the text word entry library, a word entry mother group constructed by word entries of multiple texts is determined; It is judged whether there is a same word entry associated with different texts in the word entry mother group; wherein When the same word entry exists in the different texts, the same word entry is deleted; When the same word entry does not exist in the different texts, a word entry cloud is generated according to the word entry mother group.

6. The knowledge system building method based on text-based minimized information unit of claim 1, wherein, The word entry node network is established based on the relevance, and further comprises: According to the relevance, a correlation between different word entries is judged; wherein The correlation comprises bidirectional strong correlation, unidirectional strong correlation, bidirectional weak correlation and unidirectional weak correlation.

7. The knowledge system building method based on text-based minimized information unit as claimed in claim 1, wherein, The word entry node network is established based on the relevance, and further comprises: According to the text word entry library, word entries are audited based on preset auditing rules respectively; Wherein, The auditing rules comprise objective definition auditing and academic value auditing; An auditing result after auditing is obtained, and the legality between word entries is judged according to the auditing result and the relevance.

Citation Information

Patent Citations

  • Entry recommending method and device

    CN102831185A

  • Domain knowledge base establishing method applied to discrete manufacturing industry production process

    CN111737498A