Book retrieval optimization method based on big data
By building a book feature data tree and establishing a book index pool, the problem that the book search system in the prior art is difficult to understand user intentions, and high-relevance and efficient book search results are achieved.
Patent Information
- Application Number
- CN202510107496.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing book search system is difficult to accurately understand the user's true intentions, resulting in low relevance and accuracy of the search results.
By setting keyword extraction pointers, extracting book feature words and feature sentences, building a book feature data tree, and generating feature sequences based on the number of occurrences of feature words, comparing and setting feature marks, establishing a book index pool and generating a book index tree, and matching to select index book data.
It effectively reflects the content and structure of the book, improves the relevance of search results, and reduces the search time through rapid matching, and improves the user experience.
Smart Images

Figure CN120030175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of book indexing, and in particular to a book retrieval optimization method based on big data. Background Art
[0002] In the digital age, libraries and bookstores are no longer just places to store physical books. They have evolved into a diversified source of book data, including e-books, literature databases, multimedia resources, etc. With the development of information technology, massive amounts of data information are generated and stored. How to efficiently and accurately retrieve the specific content required by users from this information has become an important issue in library management and services.
[0003] Existing book retrieval systems mostly use simple keyword matching algorithms. When the query terms entered by users are vague or ambiguous, it is difficult for the system to accurately understand the user's true intention, resulting in low relevance and accuracy of the retrieval results. Therefore, a book retrieval optimization method based on big data is provided. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a book retrieval optimization method based on big data.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] 1. A book retrieval optimization method based on big data, comprising the following steps:
[0007] Step S1, setting a number of keyword extraction pointers, extracting book feature words and book feature sentences from the book data in each book data source through the keyword extraction pointers, and then generating a book feature data tree based on the order of distribution of the book feature words and book feature sentences in each book data;
[0008] Step S2, generating a feature sequence of each book data according to the number of occurrences of book feature words and book feature sentences in the book data source, comparing the feature sequences of any two book data in the same book data source, and setting multiple feature annotations for the corresponding book feature data tree according to the comparison results;
[0009] Step S3: according to the book feature words of each book data source, set the book source feature words and the conventional book feature words for each book data source, and then establish a book index pool according to the book feature word types of all book data sources;
[0010] Step S4, obtain a book search request, generate a book index tree according to the book search request, and then match the book index tree with the book index pool, and send the book index tree to the corresponding book data source according to the matching result, and then the book data source matches the book index tree with the feature annotations in each book feature data tree, and obtains the total index degree between the book index tree and each book data according to the matching result, and then selects the indexed book data according to the total index degree value.
[0011] Furthermore, the process of extracting the book feature words and book feature sentences includes:
[0012] Setting a number for each book data source and setting a number of keyword extraction pointers, wherein the keyword extraction pointers are used to traverse all book data in the book data source, and then extract a number of book feature words and book feature sentences from each book data;
[0013] Then, each book data source sets a non-repeating book number for each book data stored therein. It should be noted that the book numbers of the book data in different book data sources are also non-repeating.
[0014] Through each keyword extraction pointer, each book data is traversed in turn, and after the traversal is completed, a number of book feature words and book feature sentences are generated, and the book feature words uploaded by each book data source are compared with each other. According to the comparison results, the repeated book feature words are eliminated, and a specific representation number is set for each book feature word, and a corresponding feature word-representation digital dictionary is generated, and then the feature word-representation digital dictionary is sent to each book data source;
[0015] According to the position distribution of each book feature word in each book feature sentence and the feature word-representation digital dictionary, each book data source generates a feature matrix for each book feature sentence.
[0016] Furthermore, the process of generating the book feature data tree of the book data includes:
[0017] The order of distribution of book feature words and book feature sentences in each book data generates a corresponding book feature data tree, wherein the book feature data tree is provided with a name relationship cluster, a first-level relationship cluster and a second-level relationship cluster;
[0018] The book feature data tree is composed of a plurality of feature nodes, each of which contains a characterizing number or a feature matrix, and the connection order of each feature node is consistent with the order in which the book feature words or book feature sentences appear in the book data;
[0019] At the same time, according to the distribution of names, chapters and paragraphs in the book data, the feature nodes corresponding to the book feature words or book feature sentences in the same name, chapter or paragraph are in the same first-level relationship cluster or second-level relationship cluster.
[0020] Furthermore, the process of generating the feature sequence of the book data includes:
[0021] Count the number of occurrences of various book feature words generated by the book data source, set a common feature threshold according to the total number of books in the book data source, compare the number of occurrences of the book feature words with the common feature threshold, record the book feature words whose number of occurrences is greater than or equal to the common feature threshold as preliminary public book feature words, and send the preliminary public book feature words to the index data terminal, otherwise do nothing;
[0022] Compare the book feature word types of each book data with each other. If there is a book feature word type that is only occupied by one book data, then record the book feature word type as a local book feature word of the corresponding book data, otherwise do not perform any operation;
[0023] At the same time, according to the representation numbers or feature matrices in the feature nodes of each name relationship cluster, first-level relationship cluster and second-level relationship cluster in the book feature data tree, the corresponding name feature sequence, first-level feature sequence and second-level feature sequence are generated.
[0024] Furthermore, the process of setting multiple feature annotations for the book feature data tree includes:
[0025] Starting from the first primary feature sequence of any book data, the primary feature sequence is compared with each primary feature sequence of another book data until the comparison operation of the first primary feature sequence and the primary feature sequence of another book data is completed;
[0026] The proportion of identical and continuous feature column segments in each comparison result is obtained, recorded as the comparison matching degree, and a comparison matching degree threshold is set;
[0027] If the comparison match degree is less than or equal to the comparison match degree threshold, the comparison result between the corresponding primary feature sequences is ignored;
[0028] If the comparison match degree is less than or equal to the comparison match degree threshold, a first-level similarity mark is set in another book data corresponding to the corresponding sequence segment in the first-level feature sequence;
[0029] Then select the first primary feature sequence, compare the primary feature sequence with each primary feature sequence of another book data, and set a primary similarity mark in each primary feature sequence according to the comparison result, and so on, until all the primary feature sequences are compared with each primary feature sequence of another book data;
[0030] Then select each primary feature sequence in another book data and repeat the above comparison operation;
[0031] When all the book data in the book data source have completed the comparison operation of the first-level feature sequence, for any first-level feature sequence of the book data, the first-level similar identifications obtained in each comparison result are overlapped, and a first-level overlap threshold is set, and then the sequence fragments corresponding to the characterization numbers or feature matrices with the number of overlapping first-level similar identifications less than or equal to the first-level overlap threshold in each first-level feature sequence are recorded as first-level feature sequence segments, otherwise no operation is performed;
[0032] Then, according to the process of setting the primary feature sequence segment for the primary feature sequence of each book data, the name feature sequence segment and the secondary feature sequence segment are set for the name feature sequence and the secondary feature sequence of each book data;
[0033] According to the position distribution of the name feature sequence segment, the primary feature sequence segment and the secondary feature sequence segment, the name feature annotation, the primary feature annotation and the secondary feature annotation are set at the corresponding feature nodes in the book feature data tree.
[0034] Furthermore, the process of establishing the book index pool includes:
[0035] Compare the types of prepared public book feature words of each book data source with each other. If there is a type of book feature word that is only occupied by one book data, then record the type of book feature word as the book source feature word of the corresponding book data source; otherwise, record the corresponding prepared public book feature word as the public book feature word;
[0036] A book index pool is established according to the book source feature words and conventional book feature words of all book data sources, and the book index pool stores a plurality of book source feature words, conventional book feature words and book clusters.
[0037] Furthermore, the process of matching the book index tree with the book index pool includes:
[0038] Retrieving a feature word-representation digital dictionary to traverse the book search request, and generating a book index tree according to the traversal result, wherein the book index tree is composed of a plurality of index nodes, each of which contains an index representation number or an index feature matrix;
[0039] Input each index node in the book index tree into the book index pool, obtain the book feature words corresponding to the index node according to the feature word-representation digital dictionary, and then first match each index node with the book source feature words in each book cluster. If the match is successful, send the book index tree to the corresponding book data source according to the number of the book cluster;
[0040] If the index node does not match any of the book source feature words, each index node is matched with a regular book feature word in each book cluster, and the index establishment threshold is set according to the number of book feature word types associated with the book index tree;
[0041] If the number of matches between the index node and the regular book feature words in the book cluster is greater than or equal to the index establishment threshold, the book index tree is sent to the corresponding book data source according to the number of the corresponding book cluster.
[0042] Furthermore, the process of obtaining the total index value between the book index tree and each book data includes:
[0043] When any book data source receives a book index tree, it first generates an index sequence from the representation numbers and feature matrices in the index nodes in the book index tree;
[0044] Starting from the name relationship cluster, the name feature sequence segments with name feature annotations in the name relationship clusters in each book feature data tree are compared with the index sequence. Whenever there is a part in the index sequence that is identical to the name feature sequence segment, the index degree is recorded once;
[0045] Then, the first-level characteristic sequence segments and the second-level characteristic sequence segments in the first-level relationship cluster and the second-level relationship cluster are compared in turn, and the index degree is recorded according to the comparison result;
[0046] The total index values between each book feature data tree and the book index tree are counted respectively, and the index threshold is set according to the total number of index nodes in the book index tree. Then, the book data source marks the total index value of the book data whose total index value is greater than or equal to the index threshold, and then sends the book data to the index data terminal, otherwise, the corresponding book data is ignored;
[0047] The index data terminal integrates the book data sent by each book data source, and for the same book data, the book data with the highest total index value is retained as the index book data, and the others are automatically eliminated;
[0048] Then, each indexed book data source is sorted according to the total index value, and the sorting result is sent to the user, with the books with the highest total index value being ranked at the front.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. The present invention extracts feature words and feature sentences of books and constructs a book feature data tree. It generates feature sequences of each book data according to the number of occurrences of book feature words and book feature sentences in the book data source, compares the feature sequences of any two book data in the same book data source, and sets multiple feature annotations for the corresponding book feature data tree according to the comparison results, thereby effectively reflecting the content and structure of the book and improving the relevance of the search results.
[0051] 2. The present invention establishes a book index pool and generates a book index tree, matches the book index tree with the book index pool, and selects indexed book data from the corresponding book data source according to the matching result, thereby achieving rapid matching with multiple book data sources, reducing retrieval time, and improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention.
[0053] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION
[0054] To make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described in detail below. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other implementation methods obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.
[0055] Example 1
[0056] like Figure 1 As shown, a book retrieval optimization method based on big data includes the following steps:
[0057] Step S1, setting a number of keyword extraction pointers, extracting book feature words and book feature sentences from the book data in each book data source through the keyword extraction pointers, and then generating a book feature data tree based on the order of distribution of the book feature words and book feature sentences in each book data;
[0058] Step S2, generating a feature sequence of each book data according to the number of occurrences of book feature words and book feature sentences in the book data source, comparing the feature sequences of any two book data in the same book data source, and setting multiple feature annotations for the corresponding book feature data tree according to the comparison results;
[0059] Step S3: according to the book feature words of each book data source, set the book source feature words and the conventional book feature words for each book data source, and then establish a book index pool according to the book feature word types of all book data sources;
[0060] Step S4, obtain a book search request, generate a book index tree according to the book search request, and then match the book index tree with the book index pool, and send the book index tree to the corresponding book data source according to the matching result, and then the book data source matches the book index tree with the feature annotations in each book feature data tree, and obtains the total index degree between the book index tree and each book data according to the matching result, and then selects the indexed book data according to the total index degree value.
[0061] Example 2
[0062] This embodiment is a further limitation of Embodiment 1, and Step S1 is implemented by the following process:
[0063] An index data terminal is set up, and the index data terminal is connected to each book data source for communication, and then the index data terminal sets a number a for each book data source. 1 、a 2 、a 3 ,……,a n , where n is a positive integer greater than 50;
[0064] For any book data source, a number of keyword extraction pointers are set, and the keyword extraction pointers are used to traverse all book data in the book data source, and then extract a number of book feature words and book feature sentences from each book data;
[0065] Then, each book data source sets a non-repeating book number for each book data stored therein. It should be noted that the book numbers of the book data in different book data sources are also non-repeating.
[0066] Through each keyword extraction pointer, each book data is traversed in turn, and after the traversal is completed, a number of book feature words and book feature sentences are generated, and all book feature words are sent to the index data terminal;
[0067] The index data terminal compares the book feature words uploaded by each book data source, removes the repeated book feature words according to the comparison result, sets a specific representation number for each book feature word, and generates a corresponding feature word-representation number dictionary, and then sends the feature word-representation number dictionary to each book data source;
[0068] According to the position distribution of each book feature word in each book feature sentence and the feature word-representation digital dictionary, each book data source generates a feature matrix of each book feature sentence;
[0069] The order of distribution of book feature words and book feature sentences in each book data generates a corresponding book feature data tree, wherein the book feature data tree is provided with a name relationship cluster, a first-level relationship cluster and a second-level relationship cluster;
[0070] The book feature data tree is composed of a plurality of feature nodes, each of which contains a characterizing number or a feature matrix, and the connection order of each feature node is consistent with the order in which the book feature words or book feature sentences appear in the book data;
[0071] At the same time, according to the name, chapter and paragraph distribution in the book data, the feature nodes corresponding to the book feature words or book feature sentences in the same name, chapter or paragraph are in the same first-level relationship cluster or second-level relationship cluster, where the name relationship cluster corresponds to the book name, the first-level relationship cluster corresponds to the chapter distribution, and the second-level relationship cluster corresponds to the paragraph distribution.
[0072] Example 3
[0073] This embodiment is a further limitation of Embodiment 1, and Step S2 is implemented by the following process:
[0074] Count the number of occurrences of various book feature words generated by the book data source, set a common feature threshold according to the total number of books in the book data source, compare the number of occurrences of the book feature words with the common feature threshold, record the book feature words whose number of occurrences is greater than or equal to the common feature threshold as preliminary public book feature words, and send the preliminary public book feature words to the index data terminal, otherwise do nothing;
[0075] Compare the book feature word types of each book data with each other. If there is a book feature word type that is only occupied by one book data, then record the book feature word type as a local book feature word of the corresponding book data, otherwise do not perform any operation;
[0076] At the same time, according to the representation numbers or feature matrices in the feature nodes of each name relationship cluster, first-level relationship cluster and second-level relationship cluster in the book feature data tree, the corresponding name feature sequence, first-level feature sequence and second-level feature sequence are generated.
[0077] For any two book data, feature sequence segments are extracted from the name feature sequences, primary feature sequences and secondary feature sequences of the two book data respectively. The feature sequence segment extraction process includes:
[0078] Starting from the first primary feature sequence of any book data, the primary feature sequence is compared with each primary feature sequence of another book data until the comparison operation of the first primary feature sequence and the primary feature sequence of another book data is completed;
[0079] The proportion of identical and continuous feature column segments in each comparison result is obtained, recorded as the comparison matching degree, and a comparison matching degree threshold is set;
[0080] If the comparison match degree is less than or equal to the comparison match degree threshold, the comparison result between the corresponding primary feature sequences is ignored;
[0081] If the comparison match degree is less than or equal to the comparison match degree threshold, a first-level similarity mark is set in another book data corresponding to the corresponding sequence segment in the first-level feature sequence;
[0082] Then select the first primary feature sequence, compare the primary feature sequence with each primary feature sequence of another book data, and set a primary similarity mark in each primary feature sequence according to the comparison result, and so on, until all the primary feature sequences are compared with each primary feature sequence of another book data;
[0083] Then select each primary feature sequence in another book data and repeat the above comparison operation;
[0084] When all the book data in the book data source have completed the comparison operation of the first-level feature sequence, for any first-level feature sequence of the book data, the first-level similar identifications obtained in each comparison result are overlapped, and a first-level overlap threshold is set, and then the sequence fragments corresponding to the characterization numbers or feature matrices with the number of overlapping first-level similar identifications less than or equal to the first-level overlap threshold in each first-level feature sequence are recorded as first-level feature sequence segments, otherwise no operation is performed;
[0085] Then, according to the process of setting the primary feature sequence segment for the primary feature sequence of each book data, the name feature sequence segment and the secondary feature sequence segment are set for the name feature sequence and the secondary feature sequence of each book data;
[0086] According to the position distribution of the name feature sequence segment, the primary feature sequence segment and the secondary feature sequence segment, the name feature annotation, the primary feature annotation and the secondary feature annotation are set at the corresponding feature nodes in the book feature data tree.
[0087] Example 4
[0088] This embodiment is a further limitation of Embodiment 1, and step S3 is implemented by the following process:
[0089] The index data terminal integrates the prepared public book feature words sent by each book data source, compares the types of the prepared public book feature words of each book data source, and if there is a book feature word type that is only occupied by one book data, then the book feature word type is recorded as the book source feature word of the corresponding book data source, otherwise the corresponding prepared public book feature word is recorded as a public book feature word;
[0090] A book index pool is established according to the book source feature words and conventional book feature words of all book data sources, wherein the book index pool stores a plurality of book source feature words, conventional book feature words and book clusters;
[0091] The book clusters are marked with book data source numbers, and the associated book source feature words and conventional book feature words are framed in each book cluster, wherein conventional book feature words can be framed by multiple book clusters at the same time, but book source feature words can only be framed by one book cluster.
[0092] Example 5
[0093] This embodiment is a further limitation of Embodiment 1, and step S4 is implemented by the following process:
[0094] The user uploads a book search request to the index data terminal, wherein the book search request includes, for example, the book title, chapter content or fragment content;
[0095] Retrieving a feature word-representation digital dictionary to traverse the book search request, and generating a book index tree according to the traversal result, wherein the book index tree is composed of a plurality of index nodes, each of which contains an index representation number or an index feature matrix;
[0096] Input each index node in the book index tree into the book index pool, obtain the book feature words corresponding to the index node according to the feature word-representation digital dictionary, and then first match each index node with the book source feature words in each book cluster. If the match is successful, send the book index tree to the corresponding book data source according to the number of the book cluster;
[0097] If the index node does not match any of the book source feature words, each index node is matched with a regular book feature word in each book cluster, and the index establishment threshold is set according to the number of book feature word types associated with the book index tree;
[0098] If the number of matches between the index node and the regular book feature words in the book cluster is greater than or equal to the index establishment threshold, the book index tree is sent to the corresponding book data source according to the number of the corresponding book cluster.
[0099] When any book data source receives a book index tree, it first generates an index sequence from the representation numbers and feature matrices in the index nodes in the book index tree;
[0100] Starting from the name relationship cluster, the name feature sequence segments with name feature annotations in the name relationship clusters in each book feature data tree are compared with the index sequence. Whenever there is a part in the index sequence that is identical to the name feature sequence segment, the index degree is recorded once;
[0101] Then, the first-level characteristic sequence segments and the second-level characteristic sequence segments in the first-level relationship cluster and the second-level relationship cluster are compared in turn, and the index degree is recorded according to the comparison result;
[0102] The total index values between each book feature data tree and the book index tree are counted respectively, and the index threshold is set according to the total number of index nodes in the book index tree. Then, the book data source marks the total index value of the book data whose total index value is greater than or equal to the index threshold, and then sends the book data to the index data terminal, otherwise, the corresponding book data is ignored;
[0103] The index data terminal integrates the book data sent by each book data source, and for the same book data, the book data with the highest total index value is retained as the index book data, and the others are automatically eliminated;
[0104] Then, each indexed book data source is sorted according to the total index value, and the sorting result is sent to the user, with the books with the highest total index value being ranked at the front.
[0105] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A book retrieval optimization method based on big data, characterized in that: The following steps are involved: Step S1, setting a number of keyword extraction pointers, extracting book feature words and book feature sentences from the book data in each book data source through the keyword extraction pointers, and then generating a book feature data tree based on the order of distribution of the book feature words and book feature sentences in each book data; Step S2, generating a feature sequence of each book data according to the number of occurrences of book feature words and book feature sentences in the book data source, comparing the feature sequences of any two book data in the same book data source, and setting multiple feature annotations for the corresponding book feature data tree according to the comparison results; Step S3: according to the book feature words of each book data source, set the book source feature words and the conventional book feature words for each book data source, and then establish a book index pool according to the book feature word types of all book data sources; Step S4, obtain a book search request, generate a book index tree according to the book search request, and then match the book index tree with the book index pool, and send the book index tree to the corresponding book data source according to the matching result, and then the book data source matches the book index tree with the feature annotations in each book feature data tree, and obtains the total index degree between the book index tree and each book data according to the matching result, and then selects the indexed book data according to the total index degree value.
2. A book retrieval optimization method based on big data according to claim 1, characterized in that: The process of extracting the book feature words and book feature sentences includes: Set a number for each book data source, and set a number of keyword extraction pointers, traverse each book data in turn through each keyword extraction pointer, and generate a number of book feature words and book feature sentences after the traversal is completed, compare the book feature words uploaded by each book data source with each other, remove the repeated book feature words according to the comparison results, set a specific representation number for each book feature word, and generate a corresponding feature word-representation number dictionary; According to the position distribution of each book feature word in each book feature sentence and the feature word-representation digital dictionary, each book data source generates a feature matrix for each book feature sentence.
3. A book retrieval optimization method based on big data according to claim 2, characterized in that: The process of generating the book feature data tree of the book data includes: Generate a corresponding book feature data tree according to the order of distribution of book feature words and book feature sentences in each book data, wherein the book feature data tree is provided with a name relationship cluster, a first-level relationship cluster and a second-level relationship cluster; The book feature data tree is composed of a number of feature nodes, each of which contains a representation number or a feature matrix. According to the distribution of names, chapters and paragraphs in the book data, the feature nodes corresponding to the book feature words or book feature sentences in the same name, chapter or paragraph are in the same first-level relationship cluster or second-level relationship cluster.
4. A book retrieval optimization method based on big data according to claim 3, characterized in that: The process of generating the feature sequence of the book data includes: Count the number of occurrences of various book feature words generated by the book data source, set a common feature threshold according to the total number of books in the book data source, compare the number of occurrences of book feature words with the common feature threshold, and record the book feature words with a number of occurrences greater than or equal to the common feature threshold as preliminary common book feature words, otherwise do nothing; Compare the book feature word types of each book data with each other. If there is a book feature word type that is only occupied by one book data, then record the book feature word type as a local book feature word of the corresponding book data, otherwise do not perform any operation; At the same time, according to the representation numbers or feature matrices in the feature nodes of each name relationship cluster, first-level relationship cluster and second-level relationship cluster in the book feature data tree, the corresponding name feature sequence, first-level feature sequence and second-level feature sequence are generated.
5. The method for optimizing book retrieval based on big data according to claim 4, characterized in that: The process of setting multiple feature annotations for the book feature data tree includes: Starting from the first primary feature sequence of any book data, the primary feature sequence is compared with each primary feature sequence of another book data until the comparison operation of the first primary feature sequence and the primary feature sequence of another book data is completed; Obtain the proportion of identical and continuous feature column segments in each comparison result, record it as the comparison matching degree, and set the comparison matching degree threshold. If the comparison matching degree is less than or equal to the comparison matching degree threshold, ignore the comparison results between the corresponding primary feature sequences; If the comparison match degree is less than or equal to the comparison match degree threshold, a first-level similarity mark is set for the corresponding sequence fragment in the first-level feature sequence in the other book data, and so on, until all the first-level feature sequences are compared with each first-level feature sequence of the other book data; When all the book data in the book data source have completed the comparison operation of the first-level feature sequence, for any first-level feature sequence of the book data, the first-level similar identifications obtained in each comparison result are overlapped, and a first-level overlap threshold is set, and then the sequence fragments corresponding to the characterization numbers or feature matrices with the number of overlapping first-level similar identifications less than or equal to the first-level overlap threshold in each first-level feature sequence are recorded as first-level feature sequence segments, otherwise no operation is performed; Then, according to the process of setting the first-level feature sequence segment for the first-level feature sequence of each book data, the name feature annotation, the first-level feature annotation and the second-level feature annotation are set at the corresponding feature node in the book feature data tree.
6. A book retrieval optimization method based on big data according to claim 5, characterized in that: The process of establishing the book index pool includes: The types of prepared public book feature words of various book data sources are compared with each other. If there is a type of book feature word that is only occupied by one book data, the book feature word type is recorded as the book source feature word of the corresponding book data source, otherwise the corresponding prepared public book feature word is recorded as the public book feature word, and a book index pool is established based on the book source feature words of all book data sources and conventional book feature words, and the book index pool stores a number of book source feature words, conventional book feature words and book clusters.
7. The method for optimizing book retrieval based on big data according to claim 6, characterized in that: The process of matching the book index tree with the book index pool includes: Retrieving a feature word-representation digital dictionary to traverse the book search request, and generating a book index tree according to the traversal result, wherein the book index tree is composed of a plurality of index nodes, each of which contains an index representation number or an index feature matrix; Input each index node in the book index tree into the book index pool, obtain the book feature words corresponding to the index node according to the feature word-representation digital dictionary, and then first match each index node with the book source feature words in each book cluster. If the match is successful, send the book index tree to the corresponding book data source according to the number of the book cluster; If the index node does not match any of the book source feature words, each index node is matched with a regular book feature word in each book cluster, and the index establishment threshold is set according to the number of book feature word types associated with the book index tree; If the number of matches between the index node and the regular book feature words in the book cluster is greater than or equal to the index establishment threshold, the book index tree is sent to the corresponding book data source according to the number of the corresponding book cluster.
8. The method for optimizing book retrieval based on big data according to claim 7, characterized in that: The process of obtaining the total index value between the book index tree and each book data includes: When any book data source receives a book index tree, an index sequence is generated from the representation numbers and feature matrix in the index nodes in the book index tree; starting from the name relationship cluster, the name feature sequence segments with name feature annotations in the name relationship clusters in each book feature data tree are compared with the index sequence; Whenever there is a part in the index sequence that is identical to the name feature sequence segment, the index degree is recorded once, and then it is compared with the first-level feature sequence segments and the second-level feature sequence segments in the first-level relationship cluster and the second-level relationship cluster in turn, and the index degree is recorded according to the comparison results; The total index degree between each book feature data tree and the book index tree is counted respectively, and the index degree threshold is set according to the total number of index nodes in the book index tree. Then the book data source marks the total index degree value of the book data whose total index degree value is greater than or equal to the index degree threshold, otherwise the corresponding book data is ignored.