Artificial Intelligence-Based E-book Retrieval Methods and Systems

By constructing and debugging an e-book retrieval model, and utilizing a global text semantic relationship network and text semantic invertible moments, the problem of inaccurate identification of key content in e-book retrieval was solved, achieving efficient and accurate key content capture and retrieval.

CN116701567BActive Publication Date: 2025-10-28赵安祺
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310508159.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-10-28
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing e-book retrieval methods struggle to efficiently and accurately identify and extract key content, hampered by inaccuracies in semantic mining due to differences in word frequency data.

Method used

By constructing and debugging an e-book retrieval model, utilizing a global text semantic relationship network and text semantic invertible moments, the capture window of the target summary text in the candidate summary text set is determined. Combined with content-related weights, the model configuration parameters are optimized to achieve accurate capture of key content.

Benefits of technology

It improves the recognition accuracy of key content in e-book text, reduces interference from non-key content, and achieves fast and accurate retrieval processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116701567B_ABST
    Figure CN116701567B_ABST
Patent Text Reader

Abstract

This invention relates to an AI-based electronic book retrieval method and system. Before retrieving electronic book text, a debugged electronic book retrieval model can be used to determine at least one target summary text capture window. Since the electronic book retrieval model can reduce the impact of variations in the content of the target summary text caused by differences in word frequency data on semantic mining, it improves the accuracy of identifying key content (target summary text) and non-key content (non-target summary text) of electronic book text samples. Therefore, the determined capture window can cover the corresponding target summary text as completely and accurately as possible. This allows for the rapid and accurate determination of the retrieval summary semantic features corresponding to the electronic book text, thereby achieving precise and efficient retrieval processing based on the retrieval summary semantic features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to an electronic book retrieval method and system based on artificial intelligence. Background Technology

[0002] E-books share many characteristics with traditional books: they contain a certain amount of information, such as text and color pages; their layout follows the format of traditional books to suit readers' reading habits; and they convey information through being read. However, as a new form of book, e-books also possess many characteristics that differ from or are absent from traditional books: they are read through computer devices and displayed on a screen; they offer the advantages of combining text, images, and audio; they are searchable; they are copyable; they offer higher cost-effectiveness; they contain more information; and they have more diverse distribution channels. With the increasing popularity of e-books, efficient and accurate retrieval of e-books has become a key requirement. Summary of the Invention

[0003] In a first aspect, embodiments of the present invention provide an artificial intelligence-based electronic book retrieval method, applied to an electronic book retrieval system, the method comprising:

[0004] Obtain the e-book text to be searched and the corresponding word frequency distribution of the e-book text;

[0005] The text frequency distribution corresponding to the e-book text and the e-book text are input into the debugged e-book retrieval model. The e-book retrieval model processes the e-book text to obtain the global text semantic relationship network corresponding to the e-book text, and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network.

[0006] Based on the certainty index of the target summary text in the e-book text to be retrieved corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network, and the content involvement weight between several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network, at least one target summary text set to be processed is determined from the several candidate summary text sets, wherein the content involvement weight between the target summary text sets to be processed of different target summary texts is less than a preset threshold.

[0007] The capture window for the minimum target summary text is determined by the set of summary texts to be processed for the minimum target summary text and the global text semantic relationship network corresponding to the e-book text.

[0008] The capture window is used to determine the semantic features of the e-book text for retrieval summary; the e-book text is then processed for retrieval based on the semantic features of the retrieval summary.

[0009] This design allows for the determination of at least one target summary text capture window using a debugged e-book retrieval model before processing the e-book text. Given that the e-book retrieval model can reduce the impact of variations in the content of the target summary text caused by differences in word frequency data on semantic mining, it improves the accuracy of distinguishing between key content (target summary text) and non-key content (non-target summary text) in e-book text samples. Therefore, the determined capture window can cover the corresponding target summary text as completely and accurately as possible. This enables the rapid and accurate determination of the retrieval summary semantic features corresponding to the e-book text, thereby achieving precise and efficient retrieval processing based on these semantic features.

[0010] Under some possible design approaches, the debugging steps of the electronic book retrieval model include:

[0011] Semantic mining is performed on the word frequency distribution of the text corresponding to the e-book text sample using the e-book retrieval model to be debugged, to obtain at least one global word frequency semantic relationship network. The e-book text sample includes prior annotations of the target summary text.

[0012] By using each global word frequency semantic relationship network, determine the text semantic invertible moments corresponding to that global word frequency semantic relationship network;

[0013] Using the aforementioned e-book retrieval model, semantic mining is performed on the e-book text samples based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network to obtain the global text semantic relationship network samples corresponding to the e-book text samples.

[0014] Based on the global text semantic relationship network example corresponding to the e-book text example, determine the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example;

[0015] Based on the prior annotations of the target summary text of the e-book text example, the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, the model configuration parameters of the e-book retrieval model are optimized.

[0016] This design, using the above-mentioned approach to debug the e-book retrieval model, can determine several text semantic invertible moments based on several global word frequency semantic relationship networks obtained from the text word frequency distribution. Semantic mining is then performed on e-book text samples based on these individual text semantic invertible moments. Since different text semantic invertible moments contain word frequency data of the target summary text in the e-book text samples, semantic mining of e-book text samples based on these individual text semantic invertible moments will be combined with word frequency data of the target summary text in the e-book text samples. This reduces the impact of variations in the content of the target summary text caused by differences in word frequency data on semantic mining, thereby improving the accuracy of identifying key content (target summary text) and non-key content (non-target summary text) in e-book text samples. This enhances the accuracy and reliability of capturing and extracting the target summary text, thus providing an accurate and reliable analytical basis for subsequent retrieval processing.

[0017] Under some possible design approaches, determining the text semantic invertible moments corresponding to each global word frequency semantic relationship network includes:

[0018] For each of the given semantic adjustment vector sets, the word frequency semantic vector of the global word frequency semantic relationship network is multiplied by the semantic adjustment vectors in that set to obtain the corresponding word frequency semantic adjustment vectors for the global word frequency semantic relationship network. The semantic adjustment vectors in the set are used to adjust the global word frequency semantic relationship network. Different semantic adjustment vectors within the same set have different adjustment strategies, and different adjustment weights are applied to the adjustments made to different sets.

[0019] By using the semantic adjustment vectors corresponding to each of the several semantic adjustment vector sets, the text semantic invertible moments corresponding to the global word frequency semantic relationship network are determined.

[0020] Under some possible design approaches, the step of determining the text semantic invertible moments corresponding to the global word frequency semantic relationship network by using the word frequency semantic adjustment vectors corresponding to each of the several semantic adjustment vector sets includes:

[0021] For each feature label in one of the several word frequency semantic adjustment vectors corresponding to each semantic adjustment vector set, determine the semantic description mean variable of the semantic description variable at the feature label in the several word frequency semantic adjustment vectors, and use the obtained semantic description mean variable as the semantic description variable of the transition text semantic invertible moment corresponding to the semantic adjustment vector set at the feature label;

[0022] The semantic invertible moments of the transition text corresponding to several sets of semantic adjustment vectors are summed to obtain the semantic invertible moments of the text.

[0023] This design updates the word frequency semantic vector adjustment strategy and adjustment weights of the global text semantic relationship network through the semantic adjustment vector in the semantic adjustment vector set. After strengthening and summing the transition text semantic invertible moments to obtain the text semantic invertible moments, it can be understood as first updating the position of the invertible block, and then multiplying it with the e-book text example, thereby realizing semantic mining of the e-book text example. This can identify key content (target summary text) and non-key content (non-target summary text) in the e-book text example. In addition, by updating the adjustment weights of the word frequency semantic vector of the global text semantic relationship network through the semantic adjustment vector, the size of the invertible block can be updated. The e-book retrieval model obtained by this debugging can match different types of e-book text and adaptively adjust the size of the invertible block.

[0024] Under some possible design approaches, the semantic mining of the word frequency distribution corresponding to the e-book text sample is performed through the e-book retrieval model to be debugged, to obtain at least one global word frequency semantic relationship network, including:

[0025] The text word frequency distribution is processed by X-order invertible processing using the electronic book retrieval model to obtain the Xth global word frequency semantic relationship network of the text word frequency distribution; X is a positive integer not less than 1.

[0026] The step of determining the text semantic invertible moments corresponding to each global word frequency semantic relationship network includes:

[0027] Based on the Xth global word frequency semantic relationship network, determine the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network;

[0028] The step of using the e-book retrieval model to perform semantic mining on the e-book text examples based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network, to obtain the global text semantic relationship network examples corresponding to the e-book text examples, includes:

[0029] The electronic book retrieval model is used to reversibly process the electronic book text sample to obtain the original global text semantic relationship network corresponding to the electronic book text sample.

[0030] Based on the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network, semantic mining is performed on the Xth global text semantic relationship network sample corresponding to the e-book text sample to obtain the (X+1)th global text semantic relationship network sample corresponding to the e-book text sample; wherein, the Xth global text semantic relationship network sample is obtained by semantic mining on the (X-1)th global text semantic relationship network sample through the text semantic invertible moments corresponding to the (X-1)th global word frequency semantic relationship network, and the first global text semantic relationship network sample is the original global text semantic relationship network.

[0031] Under some possible design approaches, the step of performing semantic mining on the Xth global text semantic relationship network example corresponding to the Xth global text semantic relationship network based on the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text example includes:

[0032] The description dimension variables of the Xth global text semantic relationship network example under different description dimensions are optimized at least once to obtain the global text semantic optimized relationship network example after each optimization.

[0033] Summing the description dimension variables of each optimized global text semantic optimization network example and the Xth global text semantic optimization network example under the same description dimension yields the target global text semantic optimization network example.

[0034] The product of the word frequency semantic vector corresponding to the target global text semantic optimization relationship network example and the value of the corresponding feature label in the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network is obtained to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text example.

[0035] This design, by optimizing the description dimension variables of global text semantic relationship network examples under different description dimensions, can achieve the aggregation of detailed features of the same global text semantic relationship network examples under different description dimensions, thereby making the semantic details contained in the global text semantic relationship network examples as rich and complete as possible.

[0036] Under some possible design approaches, the process of optimizing the description dimension variables of the Xth global text semantic relationship network example under different description dimensions at least once to obtain the optimized global text semantic relationship network example for each optimization includes:

[0037] In response to YZ being not less than 1, the optimized description dimension variable of the Xth global text semantic relationship network example under the YZ description dimension is determined to be the description dimension variable before optimization under the Yth description dimension; Y is any positive integer less than Q, Q is the total number of description dimensions of the Xth global text semantic relationship network example, and Z is the set description dimension adjustment value.

[0038] In response to YZ being less than 1, the description dimension variable under the Q-Z+1th description dimension is determined to be the description dimension variable before optimization under the Y-Z+1th description dimension.

[0039] Under some possible design approaches, each candidate summary text set in the ebook text set corresponding to each text semantic unit in the global text semantic relationship network example corresponds to one or more deterministic indices of target summary texts; when the same candidate summary text set corresponds to several deterministic indices of target summary texts, the several deterministic indices of target summary texts are used to indicate the possibility that the candidate summary text set has several different target summary texts.

[0040] Under some possible design approaches, the optimization of the model configuration parameters of the e-book retrieval model based on the prior annotation of the target summary text of the e-book text example, the certainty index of the target summary text in the e-book text example corresponding to several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, includes:

[0041] Based on the prior annotation of the target summary text of the e-book text example and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example corresponding to each text semantic unit in the global text semantic relationship network example, a first model cost index is determined for the recognition result of the target summary text corresponding to each candidate summary text set in the e-book text example corresponding to the text semantic unit; the recognition result of the target summary text is used to characterize whether the target summary text exists.

[0042] Based on the semantic description variables corresponding to the global text semantic relationship network examples, a second model cost index is determined for the capture labels of the target summary text corresponding to several candidate summary text sets in the ebook text examples corresponding to each text semantic unit in the global text semantic relationship network examples; the capture labels of the target summary text are used to characterize the regional adjustment weight of each candidate summary text set relative to the target summary text region;

[0043] Based on the first model cost index and the second model cost index, the model configuration parameters of the electronic book retrieval model are optimized.

[0044] When optimizing the model configuration parameters of the e-book retrieval model, several model cost indicators can be determined based on the prior annotation of the target summary text of the e-book text example, the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example. This can improve the timeliness of e-book retrieval model debugging.

[0045] Under some possible design approaches, the first model cost index for determining the recognition result of the target summary text corresponding to each candidate summary text set in the ebook text example based on the prior annotation of the target summary text of the ebook text example and the certainty index of the target summary text corresponding to each candidate summary text set in the ebook text example corresponding to each text semantic unit in the global text semantic relationship network example includes:

[0046] Based on several candidate summary text sets in the e-book text sample corresponding to each text semantic unit, and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the e-book text sample and the topic of the target summary text, keywords are added for each candidate summary text set. The keywords are used to represent the topic of the target summary text contained in the candidate summary text set.

[0047] For each candidate summary text set corresponding to each text semantic unit in the global text semantic relationship network example, the topic of the target summary text corresponding to the target summary text with the largest certainty index among the certainty indices of the least estimated target summary text corresponding to the candidate summary text set is determined as the topic of the target predicted summary text corresponding to the candidate summary text set.

[0048] The first model cost metric is determined by the generated keywords and the topic of the target prediction summary text.

[0049] Under some possible design approaches, the second model cost metric, which determines the capture label of the target summary text corresponding to several candidate summary text sets in the ebook text sample for each text semantic unit in the global text semantic relationship network sample based on the semantic description variables corresponding to the global text semantic relationship network sample, includes:

[0050] Determine the content involvement weight between each candidate summary text set in the ebook text sample corresponding to each text semantic unit and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the ebook text sample.

[0051] The candidate summary text set whose corresponding content-related weights meet the set requirements is determined as the current summary text set, and the global annotation content set corresponding to the current summary text set is determined.

[0052] The second model cost index is determined based on the current summary text set, the global annotation content set corresponding to the current summary text set, and the word frequency semantic vector corresponding to the global text semantic relationship network sample.

[0053] Under some independent design approaches, the prior annotation of the target summary text of the e-book text example also includes several global annotation content sets containing the target summary text; considering that the global annotation content sets in the prior annotation of the target summary text in the e-book text example need to include the text content set corresponding to the target summary text in the e-book text example, based on this, the size of the global annotation content set in the prior annotation of the target summary text in the e-book text example can be proportional to the size of the target summary text.

[0054] The global annotation content set corresponding to the current summary text set is selected from the global annotation content set corresponding to each candidate summary text set that has been determined in advance.

[0055] The global annotation content set corresponding to each candidate summary text set is determined based on the following approach:

[0056] For each global annotation content set of the ebook text example, determine a mapping annotation content set that includes the mapping of the global annotation content set on the ebook text example;

[0057] Based on the e-book retrieval model, a global text semantic relationship network example of the e-book text example is determined, and several candidate summary text sets in the e-book text example corresponding to each text semantic unit in the global text semantic relationship network example are determined.

[0058] For each candidate summary text set, the global annotation content set corresponding to the candidate summary text set is determined by the content involvement weights between the candidate summary text set and several mapping annotation content sets.

[0059] Under some independent design approaches, the global annotation content set corresponding to the candidate summary text set is determined by the content involvement weights between several mapping annotation content sets and the candidate summary text set, including:

[0060] By using the content involvement weights between several mapping annotation content sets and the candidate summary text set, at least one target mapping annotation content set whose corresponding content involvement weights meet the set requirements is selected from several mapping annotation content sets.

[0061] Based on the global annotation content sets corresponding to the minimum target mapping annotation content sets, the global annotation content set corresponding to the candidate summary text set is determined.

[0062] This design allows us to determine the global annotation content set corresponding to each candidate summary text set. After filtering out the target summary text set from the candidate summary text set of the ebook text to be retrieved, we can annotate the ebook text to be retrieved according to its corresponding global annotation content set, thereby obtaining the capture window of the target summary text.

[0063] Under some independent design approaches, several candidate summary text sets in the ebook text samples corresponding to each text semantic unit in the global text semantic relationship network sample are determined in the following way:

[0064] Using the text unit in the ebook text example corresponding to each text semantic unit in the global text semantic relationship network example as the core unit of the candidate summary text set, and determining several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network example according to at least one set window size and at least one information carrying capacity corresponding to the candidate summary text set.

[0065] This design allows for the flexible and accurate determination of several candidate summary text sets in the ebook text sample corresponding to each text semantic unit.

[0066] Secondly, embodiments of the present invention also provide an electronic book retrieval system, including a processing engine, a network module, and a memory. The processing engine and the memory communicate through the network module. The processing engine is used to read and run a computer program from the memory to implement the above-described method.

[0067] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-described method when it is executed.

[0068] Other features will be described in part in the following description. These features will be partially discovered by those skilled in the art upon examination of the following content and figures, or may be learned through production or application. The features of the present invention can be implemented and obtained by practice or use of various aspects of the methods, tools, and combinations listed in the detailed examples described below. Attached Figure Description

[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0070] The methods, systems, and / or procedures shown in the accompanying drawings will be further described with reference to exemplary embodiments. These exemplary embodiments will be described in detail with reference to the drawings. These exemplary embodiments are non-limiting exemplary embodiments, wherein reference numerals in the various views of the drawings represent similar mechanisms.

[0071] Figure 1 This is a schematic diagram of the hardware and software components of an exemplary electronic book retrieval system according to some embodiments of the present invention.

[0072] Figure 2 This is a flowchart illustrating an exemplary AI-based electronic book retrieval method and / or process according to some embodiments of the present invention. Detailed Implementation

[0073] To better understand the above technical solutions, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0074] In the following detailed description, numerous specific details are illustrated by example to provide a comprehensive understanding of the relevant guidance. However, it will be apparent to those skilled in the art that the invention can be practiced without these details. In other instances, well-known methods, procedures, systems, components, and / or circuits have been described at a relatively high level without detail to avoid unnecessarily obscuring aspects of the invention.

[0075] These and other characteristics, the functions disclosed in the present invention, the methods of execution, the functions of related elements in the structure, the combination of components, and the economics of production, will become more apparent in consideration of the following description with reference to the accompanying drawings, all of which form part of this invention. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of the invention. It should be understood that these drawings are not drawn to scale.

[0076] This invention uses flowcharts to illustrate the execution process performed by a system according to embodiments of the invention. It should be clearly understood that the execution processes in the flowchart may not be executed sequentially. Instead, these execution processes may be executed in reverse order or simultaneously. Additionally, at least one other execution process may be added to the flowchart. One or more execution processes may be deleted from the flowchart.

[0077] Figure 1 This is a structural block diagram of an electronic book retrieval system 100 according to some embodiments of the present invention. The electronic book retrieval system 100 may include a processing engine 110, a network module 120 and a memory 130. The processing engine 110 and the memory 130 communicate through the network module 120.

[0078] Processing engine 110 can process relevant information and / or data to perform one or more functions described in this invention. For example, in some embodiments, processing engine 110 may include at least one processing engine (e.g., a single-core processing engine or a multi-core processor). By way of example only, processing engine 110 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction-set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction-set computer (RISC), a microprocessor, or any combination thereof.

[0079] Network module 120 facilitates the exchange of information and / or data. In some embodiments, network module 120 can be any type of wired or wireless network or a combination thereof. By way of example only, network module 120 may include cable networks, wired networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public telephone switched networks (PSTNs), Bluetooth networks, wireless personal area networks, near field communication (NFC) networks, etc., or any combination of the foregoing examples. In some embodiments, network module 120 may include at least one network access point. For example, network module 120 may include wired or wireless network access points, such as base stations and / or network access points.

[0080] The memory 130 may be, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 130 is used to store programs, which the processing engine 110 executes upon receiving an execution instruction.

[0081] I understand. Figure 1 The structure shown is for illustrative purposes only; the electronic book retrieval system 100 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0082] Figure 2 This is a flowchart illustrating an exemplary AI-based electronic book retrieval method and / or process according to some embodiments of the present invention. The AI-based electronic book retrieval method is applied to... Figure 1 The electronic book retrieval system 100 may further include S1-S5.

[0083] S1. Obtain the e-book text to be searched and the word frequency distribution of the corresponding e-book text.

[0084] In this embodiment of the invention, the e-book text to be searched can be a complete e-book text or an incomplete e-book text. The search requirement for this e-book text can be to search for similar e-book texts or to search for derivative texts corresponding to this e-book text (similar to the search approach for derivative TV series of a movie).

[0085] Furthermore, text frequency distribution can be a set of text frequency data for e-book text, used to reflect the frequency and heat value of words and phrases in e-book text.

[0086] S2. Input the text word frequency distribution corresponding to the e-book text and the e-book text into the debugged e-book retrieval model. After processing by the e-book retrieval model, obtain the global text semantic relationship network corresponding to the e-book text, and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network.

[0087] In this embodiment of the invention, the e-book retrieval model can be a deep convolutional neural network, a residual network, or a deep structured semantic network, and is not limited thereto. Based on this, the e-book retrieval model can extract textual semantic features from e-book text to obtain a global textual semantic relationship network (a set of textual semantic features of e-book text). Furthermore, each textual semantic unit in the global textual semantic relationship network can be understood as a corresponding textual semantic feature member (which can correspond to a combination of characters, words, and sentences). On this basis, a candidate summary text set corresponding to each textual semantic unit can be determined, and then a deterministic index can be output for the target summary text of the candidate summary text set.

[0088] In other words, the candidate summary text set can be understood as the extracted text region or the extracted text set of the target summary text, and the certainty index is used to reflect the probability that the target summary text exists in the candidate summary text set.

[0089] S3. Based on the certainty index of the target summary text in the e-book text to be retrieved corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network, and the content involvement weight between several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network, determine at least one target summary text set to be processed from the several candidate summary text sets, wherein the content involvement weight between the target summary text sets to be processed is less than a preset threshold.

[0090] In this embodiment of the invention, the content involvement weight is used to reflect the content redundancy among several candidate summary text sets corresponding to each text semantic unit. The content involvement weight ranges from 0 to 1. The higher the content involvement weight, the higher the content redundancy among several candidate summary text sets corresponding to each text semantic unit. Based on this, the content involvement weight can be combined to determine at least one target summary text set to be processed, thereby achieving accurate determination and differentiation of the summary text sets to be processed for different target summary texts.

[0091] S4. Determine the capture window for the minimum target summary text by using the set of summary texts to be processed for the minimum target summary text and the global text semantic relationship network corresponding to the e-book text.

[0092] In this embodiment of the invention, the target summary text corresponding to the e-book text can be multiple (i.e., at least one). Based on this, the capture window of the minimum target summary text can be determined by combining the summary text set to be processed of the minimum target summary text and the global text semantic relationship network corresponding to the e-book text. The capture window is used to mark the text region or text paragraph where the target summary text exists, thereby providing guidance for subsequent summary feature extraction.

[0093] S5. Use the capture window to determine the retrieval summary semantic features of the e-book text; perform retrieval processing on the e-book text using the retrieval summary semantic features.

[0094] Based on the above, the semantic features of the search summary can be mined based on the text area or text paragraph marked by the capture window. These semantic features can reflect the central idea of ​​the e-book text at the overall level. By using the semantic features of the search summary, the interference of other non-summary content on the search process can be effectively reduced.

[0095] For example, text corresponding to target semantic features that have a similarity greater than a set similarity to the retrieved summary semantic features can be retrieved from the feature library as the retrieved text for the ebook.

[0096] As can be seen, by applying S1-S5, before retrieving ebook text, the debugged ebook retrieval model can be used to determine at least one target summary text capture window. Given that the ebook retrieval model can reduce the impact of variations in the content of the target summary text caused by differences in word frequency data on semantic mining, it can improve the accuracy of identifying the key content (target summary text) and non-key content (non-target summary text) of ebook text samples. Therefore, the determined capture window can cover the corresponding target summary text as completely and accurately as possible. This allows for the rapid and accurate determination of the retrieval summary semantic features corresponding to the ebook text, thereby achieving accurate and efficient retrieval processing based on the retrieval summary semantic features.

[0097] In an optional embodiment, the debugging method of the above-mentioned electronic book retrieval model may include steps 100-500, and the examples in the embodiments of the present invention can be understood as training samples or debugging samples, and the prior annotations can be understood as annotation information or annotation results.

[0098] Step 100: Using the electronic book retrieval model to be debugged, perform semantic mining on the word frequency distribution of the text corresponding to the electronic book text sample to obtain at least one global word frequency semantic relationship network. The electronic book text sample includes prior annotations of the target summary text.

[0099] Among them, the global word frequency semantic relationship network can be understood as the word frequency semantic feature map or word frequency semantic feature set corresponding to the word frequency distribution of the text.

[0100] Step 200: Determine the text semantic invertible moments corresponding to each global word frequency semantic relationship network.

[0101] The invertible moment of text semantics can be understood as the convolutional feature matrix of text semantics.

[0102] Step 300: Using the electronic book retrieval model, semantic mining is performed on the electronic book text samples based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network to obtain the global text semantic relationship network samples corresponding to the electronic book text samples.

[0103] Step 400: Based on the global text semantic relationship network sample corresponding to the e-book text sample, determine the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text sample for each text semantic unit in the global text semantic relationship network sample.

[0104] Step 500: Based on the prior annotation of the target summary text of the e-book text example, the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, optimize the model configuration parameters of the e-book retrieval model.

[0105] Among them, the semantic description variable can be understood as the feature vector corresponding to the global text semantic relationship network sample. Based on this, by optimizing and adjusting the model configuration parameters of the electronic book retrieval model, the accuracy of the electronic book retrieval model in extracting and capturing target summary text can be improved.

[0106] The e-book retrieval model obtained through steps 100-500 can determine several text semantic invertible moments based on several global word frequency semantic relationship networks obtained from text word frequency distribution. Semantic mining is then performed on e-book text samples based on these text semantic invertible moments. Since different text semantic invertible moments contain word frequency data of the target summary text in the e-book text samples, semantic mining of e-book text samples based on these text semantic invertible moments is combined with word frequency data of the target summary text in the e-book text samples. This reduces the impact of variations in the content of the target summary text caused by differences in word frequency data on semantic mining, thereby improving the accuracy of identifying key content (target summary text) and non-key content (non-target summary text) in e-book text samples. This enhances the accuracy and reliability of capturing and extracting the target summary text, thus providing an accurate and reliable analytical basis for subsequent retrieval processing.

[0107] Under some possible design approaches, step 200 involves determining the text semantic invertible moments corresponding to each global word frequency semantic relationship network, including steps 210 and 220.

[0108] Step 210: For each of the set semantic adjustment vector sets, multiply the word frequency semantic vector of the global word frequency semantic relationship network with the semantic adjustment vectors in the set to obtain the word frequency semantic adjustment vectors corresponding to the global word frequency semantic relationship network.

[0109] The semantic adjustment vectors in the semantic adjustment vector set are used to adjust the global word frequency semantic relationship network. Different semantic adjustment vectors in the same semantic adjustment vector set have different adjustment strategies, and different adjustment weights are corresponding to different semantic adjustment vector sets.

[0110] In this embodiment of the invention, the semantic adjustment vector can be understood as a semantic offset update vector, which is used to adjust and update the semantic frequency of words. The adjustment processing strategy can be understood as the update method, and the adjustment weight can be in the range of 0 to 1.

[0111] Step 220: Determine the text semantic invertible moment corresponding to the global word frequency semantic relationship network by using the several semantic adjustment vectors corresponding to each semantic adjustment vector set in the several semantic adjustment vector sets.

[0112] Under some possible design approaches, step 220 involves determining the text semantic invertible moments corresponding to the global word frequency semantic relationship network by using the word frequency semantic adjustment vectors corresponding to each of the several semantic adjustment vector sets, including steps 221 and 222.

[0113] Step 221: For each feature label in one of the several word frequency semantic adjustment vectors corresponding to each semantic adjustment vector set, determine the semantic description mean variable of the semantic description variable at the feature label in the several word frequency semantic adjustment vectors, and use the obtained semantic description mean variable as the semantic description variable of the transition text semantic invertible moment corresponding to the semantic adjustment vector set at the feature label.

[0114] Step 222: Strengthen and sum the semantic invertible moments of the transition text corresponding to several semantic adjustment vector sets to obtain the semantic invertible moments of the text.

[0115] Applying steps 221 and 222, the strategy and adjustment weights of adjusting the word frequency semantic vector of the global text semantic relationship network are updated through the semantic adjustment vector in the semantic adjustment vector set. After strengthening and summing the semantic invertible moments of the transitional text to obtain the text semantic invertible moments, it can be understood that the position of the invertible block is updated first, and then multiplied with the e-book text example, thereby realizing the semantic mining of the e-book text example. This can identify the key content (target summary text) and non-key content (non-target summary text) in the e-book text example. In addition, by updating the adjustment weights of the word frequency semantic vector of the global text semantic relationship network through the semantic adjustment vector, the size of the invertible block can be updated. The e-book retrieval model obtained by this debugging can match different types of e-book text and adaptively adjust the size of the invertible block.

[0116] Under some possible design approaches, step 100 involves semantic mining of the text frequency distribution corresponding to the e-book text sample using the e-book retrieval model to be debugged, obtaining at least one global word frequency semantic relationship network. This includes: performing X-order invertible processing (convolution operation) on the text word frequency distribution using the e-book retrieval model to obtain the Xth global word frequency semantic relationship network of the text word frequency distribution; where X is a positive integer not less than 1. Based on this, step 200 involves determining the text semantic invertible moment corresponding to each global word frequency semantic relationship network, including: determining the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network based on the Xth global word frequency semantic relationship network. Further, step 300 involves using the e-book retrieval model to perform semantic mining on the e-book text sample based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network, obtaining a global text semantic relationship network sample corresponding to the e-book text sample, including steps 310 and 320.

[0117] Step 310: Use the electronic book retrieval model to perform reversible processing on the electronic book text sample to obtain the original global text semantic relationship network corresponding to the electronic book text sample.

[0118] Step 320: Based on the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network, perform semantic mining on the Xth global text semantic relationship network sample corresponding to the e-book text sample to obtain the (X+1)th global text semantic relationship network sample corresponding to the e-book text sample; wherein, the Xth global text semantic relationship network sample is obtained by performing semantic mining on the (X-1)th global text semantic relationship network sample through the text semantic invertible moments corresponding to the (X-1)th global word frequency semantic relationship network, and the first global text semantic relationship network sample is the original global text semantic relationship network.

[0119] Under some possible design approaches, step 320 involves semantic mining of the Xth global text semantic relationship network sample corresponding to the Xth global text semantic relationship network based on the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network, to obtain the (X+1)th global text semantic relationship network sample corresponding to the e-book text sample, including steps 321-323.

[0120] Step 321: Optimize the description dimension variables of the Xth global text semantic relationship network example under different description dimensions at least once to obtain the optimized global text semantic relationship network example after each optimization.

[0121] Step 322: Sum the description dimension variables of each optimized global text semantic optimization relationship network example and the Xth global text semantic relationship network example under the same description dimension to obtain the target global text semantic optimization relationship network example.

[0122] Step 323: Multiply the word frequency semantic vector corresponding to the target global text semantic optimization relationship network example with the value of the corresponding feature label in the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text example.

[0123] Applying to steps 321-323, by optimizing the description dimension variables of global text semantic relationship network examples under different description dimensions (description level), it is possible to aggregate the detailed features of the same global text semantic relationship network examples under different description dimensions, thereby making the semantic details contained in the global text semantic relationship network examples as rich and complete as possible.

[0124] Under some possible design approaches, step 321, which optimizes the description dimension variables of the Xth global text semantic relationship network example under different description dimensions at least once to obtain the optimized global text semantic relationship network example for each optimization, includes: in response to YZ not being less than 1, determining the optimized description dimension variable of the Xth global text semantic relationship network example under the YZth description dimension, which is the description dimension variable before optimization under the Yth description dimension; Y is any positive integer less than Q, Q is the total number of description dimensions of the Xth global text semantic relationship network example, and Z is the set description dimension adjustment value; in response to YZ being less than 1, determining the description dimension variable under the Q-Z+1th description dimension, which is the description dimension variable before optimization under the Y-Z+1th description dimension.

[0125] Under some possible design approaches, each candidate summary text set in the ebook text set corresponding to each text semantic unit in the global text semantic relationship network example corresponds to one or more deterministic indices of target summary texts; when the same candidate summary text set corresponds to several deterministic indices of target summary texts, the several deterministic indices of target summary texts are used to indicate the possibility that the candidate summary text set has several different target summary texts.

[0126] Under some possible design approaches, step 500 optimizes the model configuration parameters of the e-book retrieval model based on the prior annotation of the target summary text of the e-book text example, the certainty index of the target summary text in the e-book text example corresponding to several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, including steps 510-530.

[0127] Step 510: Based on the prior annotation of the target summary text of the e-book text example and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example corresponding to each text semantic unit in the global text semantic relationship network example, determine the first model cost index of the recognition result of the target summary text corresponding to each candidate summary text set in the e-book text example corresponding to the text semantic unit; the recognition result of the target summary text is used to characterize whether the target summary text exists.

[0128] The model cost metric can be understood as the loss function value during the model training and debugging process.

[0129] Step 520: Based on the semantic description variables corresponding to the global text semantic relationship network example, determine the second model cost index of the capture label of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example; the capture label of the target summary text is used to characterize the regional adjustment weight of each candidate summary text set relative to the target summary text region.

[0130] Among them, the region adjustment weight can be understood as the text set adjustment coefficient of the target summary text region (target summary text set), which is used for position offset adjustment and update processing.

[0131] Step 530: Optimize the model configuration parameters of the electronic book retrieval model based on the first model cost index and the second model cost index.

[0132] Applying to steps 510-530, when optimizing the model configuration parameters of the e-book retrieval model, several model cost indicators can be determined based on the prior annotation of the target summary text of the e-book text example, the deterministic index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example. This can improve the timeliness of debugging the e-book retrieval model.

[0133] Under some possible design approaches, step 510 determines the first model cost index of the recognition result of the target summary text corresponding to each candidate summary text set in the e-book text sample based on the prior annotation of the target summary text of the e-book text sample and the deterministic index of the target summary text corresponding to each candidate summary text set in the e-book text sample corresponding to each text semantic unit in the global text semantic relationship network sample, including steps 511-513.

[0134] Step 511: Based on several candidate summary text sets in the e-book text sample corresponding to each text semantic unit, and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the e-book text sample and the topic of the target summary text, add keywords for each candidate summary text set. The keywords are used to represent the topic of the target summary text contained in the candidate summary text set.

[0135] Step 512: For each candidate summary text set corresponding to each text semantic unit in the global text semantic relationship network example, determine the topic of the target summary text corresponding to the target predicted summary text corresponding to the candidate summary text set as the topic of the target summary text corresponding to the minimum estimated target summary text corresponding to the candidate summary text set.

[0136] Step 513: Determine the first model cost index based on the generated keywords and the topic of the target prediction summary text.

[0137] The aforementioned topics can be understood as the categories of the corresponding summary text. Based on this, the first model cost index can be determined by the generated keywords (tags) and the topics of the target prediction summary text, thereby ensuring the credibility of the first model cost index.

[0138] Under some possible design approaches, step 520 includes steps 521-523, which are based on the semantic description variables corresponding to the global text semantic relationship network examples and determine the second model cost index of the capture label of the target summary text corresponding to several candidate summary text sets in the e-book text examples corresponding to each text semantic unit in the global text semantic relationship network examples.

[0139] Step 521: Determine the content involvement weight between each candidate summary text set in the e-book text sample corresponding to each text semantic unit and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the e-book text sample.

[0140] Step 522: Determine the candidate summary text set whose corresponding content-related weights meet the set requirements as the current summary text set, and determine the global annotation content set corresponding to the current summary text set.

[0141] Step 523: Determine the second model cost index based on the current summary text set, the global annotation content set corresponding to the current summary text set, and the word frequency semantic vector corresponding to the global text semantic relationship network sample.

[0142] In some possible design approaches, the prior annotation of the target summary text of the ebook text example also includes several global annotation content sets containing the target summary text. Considering that the global annotation content sets in the prior annotation of the target summary text in the ebook text example need to include the text content set corresponding to the target summary text in the ebook text example, the size of the global annotation content set in the prior annotation of the target summary text in the ebook text example can be proportional to the size of the target summary text. The global annotation content set corresponding to the current summary text set is selected from the global annotation content set corresponding to each pre-determined candidate summary text set. Specifically, the global annotation content set (text annotation region) corresponding to each candidate summary text set is determined based on the following approach: For each global annotation content set of the e-book text example, a mapping annotation content set containing the mapping of the global annotation content set on the e-book text example is determined; a global text semantic relationship network example of the e-book text example is determined according to the e-book retrieval model, and several candidate summary text sets in the e-book text example corresponding to each text semantic unit in the global text semantic relationship network example are determined; for each candidate summary text set, the global annotation content set corresponding to the candidate summary text set is determined by the content involvement weight between the several mapping annotation content sets and the candidate summary text set.

[0143] Under some possible design approaches, the global annotation content set corresponding to the candidate summary text set is determined by the content involvement weights between several mapping annotation content sets and the candidate summary text set, including: selecting at least one target mapping annotation content set from several mapping annotation content sets whose content involvement weights satisfy a set requirement; and determining the global annotation content set corresponding to the candidate summary text set based on the global annotation content sets corresponding to the at least one target mapping annotation content set.

[0144] This design allows for the determination of the global annotation content set corresponding to each candidate summary text set. After filtering out the target summary text set from the candidate summary text set of the ebook text to be retrieved, the ebook text to be retrieved can be annotated according to its corresponding global annotation content set, thereby accurately obtaining the capture window of the target summary text.

[0145] Under some possible design approaches, several candidate summary text sets in the e-book text sample corresponding to each text semantic unit in the global text semantic relationship network example are determined in the following way: taking the text unit in the e-book text sample corresponding to each text semantic unit in the global text semantic relationship network example as the core unit of the candidate summary text set, and determining several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network example according to at least one set window size and at least one information carrying capacity corresponding to the candidate summary text set.

[0146] This design allows for the flexible and accurate determination of several candidate summary text sets in the ebook text sample corresponding to each text semantic unit.

[0147] The basic concepts have been described above. It is obvious that the detailed disclosure above is merely illustrative and does not constitute a limitation of the present invention. Although not explicitly stated herein, various modifications, improvements, and alterations can be made to the present invention by those skilled in the art. Such modifications, improvements, and alterations are suggested in this invention and therefore remain within the spirit and scope of the exemplary embodiments of the present invention.

[0148] Furthermore, this invention uses specific terminology to describe embodiments of the invention. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the invention. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different parts of this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in at least one embodiment of the invention can be appropriately combined.

[0149] Furthermore, it will be understood by those skilled in the art that various aspects of the present invention can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, various aspects of the present invention can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software can be referred to as a “unit,” “component,” or “system.” Moreover, various aspects of the present invention can be embodied as a computer product located on at least one computer-readable medium, said product comprising computer-readable program code.

[0150] A computer-readable signal medium may contain a propagated data signal containing computer program encoding, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and so on, or suitable combinations thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program encoding located on the computer-readable signal medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.

[0151] The computer program code required for the execution of various aspects of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., or similar conventional programming languages ​​such as C, Visual Basic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages ​​such as Python, Ruby, and Groovy, or other programming languages. The program code can be executed entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).

[0152] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numerals, or other names described in this invention are not intended to limit the order of the processes and methods of this invention. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the additional claims are not limited to the disclosed embodiments. Rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments of this invention. For example, while the system components described above can be implemented by hardware devices, they can also be implemented solely by software solutions, such as installing the described system on an existing server or mobile device.

[0153] It should also be understood that, in order to simplify the description of the invention and thus aid in the understanding of at least one embodiment, multiple features may sometimes be grouped into a single embodiment, drawing, or description thereof in the foregoing description of the embodiments of the invention. However, this method of disclosure does not imply that the subject matter of the invention requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiment disclosed above.

Claims

1. An electronic book retrieval method based on artificial intelligence, characterized in that, The method, applied to an electronic book retrieval system, includes: Obtain the e-book text to be searched and the corresponding word frequency distribution of the e-book text; The text frequency distribution corresponding to the e-book text and the e-book text are input into the debugged e-book retrieval model. The e-book retrieval model processes the e-book text to obtain the global text semantic relationship network corresponding to the e-book text, and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network. Based on the certainty index of the target summary text in the e-book text to be retrieved corresponding to several candidate summary text sets in the e-book text to be retrieved for each text semantic unit in the global text semantic relationship network, and the content involvement weight between several candidate summary text sets corresponding to each text semantic unit in the global text semantic relationship network, at least one target summary text set to be processed is determined from the several candidate summary text sets, wherein the content involvement weight between the target summary text sets to be processed of different target summary texts is less than a preset threshold. The capture window for the minimum target summary text is determined by the set of summary texts to be processed for the minimum target summary text and the global text semantic relationship network corresponding to the e-book text. The capture window is used to determine the semantic features of the e-book text for retrieval; the e-book text is then processed for retrieval based on the semantic features of the retrieval summary. The debugging steps for the electronic book retrieval model include: Semantic mining is performed on the word frequency distribution of the text corresponding to the e-book text sample using the e-book retrieval model to be debugged, to obtain at least one global word frequency semantic relationship network. The e-book text sample includes prior annotations of the target summary text. By using each global word frequency semantic relationship network, determine the text semantic invertible moments corresponding to that global word frequency semantic relationship network; Using the aforementioned e-book retrieval model, semantic mining is performed on the e-book text samples based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network to obtain the global text semantic relationship network samples corresponding to the e-book text samples. Based on the global text semantic relationship network example corresponding to the e-book text example, determine the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example; Based on the prior annotations of the target summary text of the e-book text example, the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, the model configuration parameters of the e-book retrieval model are optimized.

2. The method according to claim 1, characterized in that, The step of determining the text semantic invertible moments corresponding to each global word frequency semantic relationship network includes: For each of the given semantic adjustment vector sets, the word frequency semantic vector of the global word frequency semantic relationship network is multiplied by the semantic adjustment vectors in that set to obtain the corresponding word frequency semantic adjustment vectors for the global word frequency semantic relationship network. The semantic adjustment vectors in the set are used to adjust the global word frequency semantic relationship network. Different semantic adjustment vectors within the same set have different adjustment strategies, and different adjustment weights are applied to the adjustments made to different sets. By using the semantic adjustment vectors corresponding to each of the several semantic adjustment vector sets, the text semantic invertible moments corresponding to the global word frequency semantic relationship network are determined.

3. The method according to claim 2, characterized in that, The step of determining the text semantic invertible moments corresponding to the global word frequency semantic relationship network by using the word frequency semantic adjustment vectors corresponding to each of the several semantic adjustment vector sets includes: For each feature label in one of the several word frequency semantic adjustment vectors corresponding to each semantic adjustment vector set, determine the semantic description mean variable of the semantic description variable at the feature label in the several word frequency semantic adjustment vectors, and use the obtained semantic description mean variable as the semantic description variable of the transition text semantic invertible moment corresponding to the semantic adjustment vector set at the feature label; The semantic invertible moments of the transition text corresponding to several sets of semantic adjustment vectors are summed to obtain the semantic invertible moments of the text.

4. The method according to claim 1, characterized in that, The step of performing semantic mining on the text frequency distribution corresponding to the e-book text sample using the e-book retrieval model to be debugged, to obtain at least one global word frequency semantic relationship network, includes: using the e-book retrieval model to perform X-order invertible processing on the text word frequency distribution to obtain the Xth global word frequency semantic relationship network of the text word frequency distribution; X is a positive integer not less than 1; The step of determining the text semantic invertible moment corresponding to each global word frequency semantic relationship network includes: determining the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network based on the Xth global word frequency semantic relationship network; The step of using the e-book retrieval model to perform semantic mining on the e-book text examples based on the text semantic invertible moments corresponding to each global word frequency semantic relationship network to obtain global text semantic relationship network examples corresponding to the e-book text examples includes: using the e-book retrieval model to perform reversible processing on the e-book text examples to obtain the original global text semantic relationship network corresponding to the e-book text examples; performing semantic mining on the Xth global text semantic relationship network example corresponding to the e-book text examples based on the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text examples; wherein, the Xth global text semantic relationship network example is obtained by performing semantic mining on the (X-1)th global text semantic relationship network example through the text semantic invertible moments corresponding to the (X-1)th global word frequency semantic relationship network, and the first global text semantic relationship network example is the original global text semantic relationship network.

5. The method according to claim 4, characterized in that, The step of performing semantic mining on the Xth global text semantic relationship network example corresponding to the Xth global text semantic relationship network based on the text semantic invertible moments corresponding to the Xth global word frequency semantic relationship network to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text example includes: The description dimension variables of the Xth global text semantic relationship network example under different description dimensions are optimized at least once to obtain the global text semantic optimized relationship network example after each optimization. Summing the description dimension variables of each optimized global text semantic optimization network example with the Xth global text semantic optimization network example under the same description dimension, we obtain the target global text semantic optimization network example. The product of the word frequency semantic vector corresponding to the target global text semantic optimization relationship network example and the value of the corresponding feature label in the text semantic invertible moment corresponding to the Xth global word frequency semantic relationship network is used to obtain the (X+1)th global text semantic relationship network example corresponding to the e-book text example. The step of optimizing the description dimension variables of the Xth global text semantic relationship network example under different description dimensions at least once to obtain the optimized global text semantic relationship network example after each optimization includes: in response to YZ not being less than 1, determining the optimized description dimension variables of the Xth global text semantic relationship network example under the YZ description dimension, which are the description dimension variables before optimization under the Y description dimension; Y is any positive integer less than Q, Q is the total number of description dimensions of the Xth global text semantic relationship network example, and Z is the set description dimension adjustment value; in response to YZ being less than 1, determining the description dimension variables under the Q-Z+1 description dimension, which are the description dimension variables before optimization under the Y-Z+1 description dimension.

6. The method according to claim 1, characterized in that, The optimization of the model configuration parameters of the e-book retrieval model based on the prior annotation of the target summary text of the e-book text example, the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example for each text semantic unit in the global text semantic relationship network example, and the semantic description variables corresponding to the global text semantic relationship network example, includes: Based on the prior annotation of the target summary text of the e-book text example and the certainty index of the target summary text corresponding to several candidate summary text sets in the e-book text example corresponding to each text semantic unit in the global text semantic relationship network example, a first model cost index is determined for the recognition result of the target summary text corresponding to each candidate summary text set in the e-book text example corresponding to the text semantic unit; the recognition result of the target summary text is used to characterize whether the target summary text exists. Based on the semantic description variables corresponding to the global text semantic relationship network examples, a second model cost index is determined for the capture labels of the target summary text corresponding to several candidate summary text sets in the ebook text examples corresponding to each text semantic unit in the global text semantic relationship network examples; the capture labels of the target summary text are used to characterize the regional adjustment weight of each candidate summary text set relative to the target summary text region; Based on the first model cost index and the second model cost index, the model configuration parameters of the electronic book retrieval model are optimized.

7. The method according to claim 6, characterized in that, The first model cost index for determining the recognition result of the target summary text corresponding to each candidate summary text set in the ebook text example based on the prior annotation of the target summary text of the ebook text example and the certainty index of the target summary text corresponding to each candidate summary text set in the ebook text example corresponding to each text semantic unit in the global text semantic relationship network example includes: Based on several candidate summary text sets in the e-book text sample corresponding to each text semantic unit, and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the e-book text sample and the topic of the target summary text, keywords are added for each candidate summary text set. The keywords are used to represent the topic of the target summary text contained in the candidate summary text set. For each candidate summary text set corresponding to each text semantic unit in the global text semantic relationship network example, the topic of the target summary text corresponding to the target summary text with the largest certainty index among the certainty indices of the least estimated target summary text corresponding to the candidate summary text set is determined as the topic of the target predicted summary text corresponding to the candidate summary text set. The first model cost metric is determined by the generated keywords and the topic of the target prediction summary text.

8. The method according to claim 6, characterized in that, The second model cost metric, which determines the capture labels of the target summary text corresponding to several candidate summary text sets in the ebook text examples corresponding to each text semantic unit in the global text semantic relationship network examples based on the semantic description variables corresponding to the global text semantic relationship network examples, includes: Determine the content involvement weight between each candidate summary text set in the ebook text sample corresponding to each text semantic unit and the annotation content set of the target summary text corresponding to the prior annotation of the target summary text of the ebook text sample. The candidate summary text set whose corresponding content-related weights meet the set requirements is determined as the current summary text set, and the global annotation content set corresponding to the current summary text set is determined. The second model cost index is determined based on the current summary text set, the global annotation content set corresponding to the current summary text set, and the word frequency semantic vector corresponding to the global text semantic relationship network sample.

9. An electronic book retrieval system, characterized in that, The method includes a processing engine, a network module, and a memory, wherein the processing engine and the memory communicate through the network module, and the processing engine is used to read and run a computer program from the memory to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-subject extraction method based on concept vector model

    CN104008090A

  • Multisource semantic analysis based information retrieval method

    CN106156272A