Electronic book management method and system
By collecting and analyzing the reading data and text characteristics of e-books, and using machine learning models for classification and management, the problem of classification chaos on e-book platforms is solved, and efficient e-book storage and management is achieved.
Patent Information
- Application Number
- CN202510741700.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing e-book platform lacks an effective classification system, cannot distinguish between different types of e-books, and fails to classify and manage according to users' reading data, resulting in storage confusion and cannot meet users' reading needs.
By collecting reading data and text data of e-books, using natural language processing and random forest models to determine the initial evaluation value, generating reading marks and adjusting evaluation value, performing cluster analysis and storage management, and setting access permissions and backup modes.
It realizes orderly classification of e-books, accurately reflects users' reading habits and interests, ensures data security and efficient utilization of storage resources, and improves management reliability and intelligence.
Smart Images

Figure CN120256629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of e - book management, and in particular, to a management method and system for e - books. Background Art
[0002] An e - book is a digital book, generally storing text content in electronic text or picture format for users to read on an e - reading device. With the development of the information industry and the increasing popularity of electronic devices and the Internet, the demand for reading e - books on the Internet and mobile terminals is growing. However, there are certain drawbacks in the storage management of existing e - book platforms. On the one hand, the types of various e - books are rich and diverse and cover multiple different fields. The e - book platform lacks an effective classification system and cannot distinguish different types of e - books. On the other hand, the reading behaviors of users on the platform contain users' reading habits and interest tendencies. The e - book platform has not established a matching analysis and classification management mechanism and cannot classify and manage e - books based on reading data. Eventually, the storage of e - books on the e - book platform is chaotic and cannot meet the reading needs of users.
[0003] Therefore, it is necessary to design a management method and system for e - books to solve the problems existing in the current technology. Summary of the Invention
[0004] In view of this, the present invention proposes a management method and system for e - books, aiming to solve the above problems.
[0005] On the one hand, the present invention proposes a management system for e - books, including: An acquisition unit, configured to obtain the reading data and text data of all e - books, collect the text features of each text data, and determine an initial evaluation value for each e - book based on the text features and a feature model; A judgment unit, configured to generate a reading mark according to each reading data, determine a reading identifier based on the number of marks of the reading mark, and judge whether to adjust the initial evaluation value according to the reading identifier; A first management unit, configured to, when it is determined to adjust the initial evaluation value, compare the number of marks with a historical data set, determine an evaluation adjustment factor according to the comparison result, determine a target evaluation value according to the evaluation adjustment factor and the initial evaluation value, perform clustering analysis on all e - books according to the text features, divide a first book set according to the clustering result, and compare each target evaluation value in the first book set with a target evaluation threshold, and construct a second book set according to the comparison result; The second management unit is configured to store the first book set and the second book set, determine access permissions, and determine the backup modes of the first book set and the second book set according to the book storage capacity.
[0006] Further, when collecting the text features of each text data, it includes: The collection unit analyzes the text data of all e-books based on natural language processing, performs word segmentation, stop word removal, and stemming on each text data, and uses a dependency syntax analysis model to parse the relationships between words to determine book phrases, and takes the book phrases that conform to the syntactic structure as keywords; Take all the keywords of each text data and the text data size as the text features.
[0007] Further, when determining the initial evaluation value of each e-book based on the text features and the feature model, it includes: Obtain a feature data set, sample it according to the sampling ratio to obtain a training set and a test set, use grid search to find the hyperparameters of the model, and establish a random forest model; Use the training set to train the random forest model, substitute the test set into the trained random forest model and determine the accuracy rate of the model prediction. When the accuracy rate is greater than or equal to the accuracy rate threshold, determine the trained random forest model as the feature model. Otherwise, continue to train the random forest model until it is greater than or equal to the accuracy rate threshold; Substitute the text features into the feature model to determine the initial evaluation value of each e-book.
[0008] Further, when generating a reading mark according to each reading data and determining a reading identifier based on the number of marks of the reading mark, it includes: Each text data corresponds to a reading data; The judgment unit obtains the interval time between each reading and the previous reading in the reading data, generates the reading mark for the interval time greater than the interval time threshold, and counts the number of marks generated for the reading mark; Preset a first preset number of marks and a second preset number of marks, and the first preset number of marks is greater than the second preset number of marks; When the number of marks is greater than or equal to the first preset number of marks, determine the reading identifier as a high-frequency reading identifier; When the number of marks is less than the first preset number of marks and greater than the second preset number of marks, determine the reading identifier as a medium-frequency reading identifier; When the number of marks is less than or equal to the second preset number of marks, determine the reading identifier as a low-frequency reading identifier.
[0009] Further, when determining whether to adjust the initial evaluation value according to the reading identifier, it includes: When the high-frequency reading identifier is recognized, the judgment unit determines to adjust the initial evaluation value; otherwise, it determines not to adjust the initial evaluation value and determines the initial evaluation value as the target evaluation value.
[0010] Further, when comparing the marked quantity with the historical data set, determining the evaluation adjustment factor according to the comparison result, and determining the target evaluation value according to the evaluation adjustment factor and the initial evaluation value, it includes: The historical data set includes a number of historical marked quantities and a number of historical evaluation adjustment factors, and each historical marked quantity corresponds to a historical evaluation adjustment factor; When there is a historical marked quantity in the historical data set whose similarity to the high-frequency reading identifier is greater than the similarity threshold, the first management unit uses the historical evaluation adjustment factor corresponding to the historical marked quantity with the maximum similarity as the evaluation adjustment factor; otherwise, the first management unit determines the evaluation adjustment factor according to a pre-trained evaluation model; The target evaluation value is the product value of the evaluation adjustment factor and the initial evaluation value.
[0011] Further, when performing clustering analysis on all e-books according to the text features and dividing the first book set according to the clustering result, it includes: The first management unit normalizes each parameter in the text features to form a text feature vector, combines the text feature vectors of all e-books into a feature matrix, randomly selects K clustering centers, and determines the clustering distance from each text feature vector to each clustering center; S1: Assign the e-books to the set corresponding to the clustering center with the closest clustering distance; S2: Re-determine the position of each clustering center; S3: Repeat S1 and S2 until the positions of the clustering centers no longer change or reach the preset number of iterations, and divide all e-books into K first book sets.
[0012] Further, when comparing each target evaluation value in the first book set with the target evaluation threshold and constructing the second book set according to the comparison result, it includes: The first management unit extracts the e-books in the first book set whose target evaluation value is greater than or equal to the target evaluation threshold, constructs the second book set, and constructs the remaining e-books in the first book set as the third book set.
[0013] Further, when storing the first book set and the second book set, determining access permissions, and determining the backup modes of the first book set and the second book set according to the book storage capacity, the following steps are included: The second management unit establishes a first repository and a second repository; The second management unit stores all of the second book set in the first repository, sets a first access permission for the first repository, stores all of the third book set in the second repository, and sets a second access permission for the second repository. The level of the first access permission is higher than that of the second access permission; Obtain the first book storage capacity of the first repository. When the first book storage capacity is greater than or equal to half of the storage capacity of the first repository, determine the backup mode of the first repository as incremental backup; When the first book storage capacity is less than half of the storage capacity of the first repository, determine the backup mode of the first repository as full backup; Determine the backup mode of the second repository as periodic backup.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: Optimize the classification system: Obtain e-book text data, extract text features, and combine with a feature model to determine an initial evaluation value, ensuring the reliability of quantifying each e-book. Perform clustering analysis on all e-books, and be able to classify e-books covering different fields and rich in variety in an orderly manner according to the text features of the e-books, avoiding the phenomenon of chaotic storage of e-books due to the lack of an effective classification system. The first management unit determines an evaluation adjustment factor through comparison and integrates the user's reading data into the evaluation system of e-books, so that the classification management of e-books can accurately reflect the user's reading habits and interest tendencies. The second management unit stores the first book set and the second book set, determines access permissions and dynamically determines backup modes according to the book storage capacity, not only ensuring the security of e-book storage and avoiding data loss, but also being able to dynamically plan backup strategies according to actual storage needs, improving the utilization efficiency of storage resources, and ensuring the reliability and intelligence of e-book management.
[0015] On the other hand, the present application also provides a management method for e-books, which is applied to the above-mentioned e-book management system and includes: Obtain the reading data and text data of all e-books, collect the text features of each text data, and determine the initial evaluation value of each e-book based on the text features and the feature model; Generate a reading mark according to each reading data, determine a reading identifier based on the number of marks of the reading mark, and determine whether to adjust the initial evaluation value according to the reading identifier; When it is determined to adjust the initial evaluation value, the number of tags is compared with the historical data set, an evaluation adjustment factor is determined based on the comparison result, a target evaluation value is determined according to the evaluation adjustment factor and the initial evaluation value, all e-books are clustered and analyzed according to the text features, a first book set is divided according to the clustering result, and each target evaluation value in the first book set is compared with a target evaluation threshold, and a second book set is constructed according to the comparison result; The first book set and the second book set are stored and access permissions are determined, and backup modes of the first book set and the second book set are determined according to the book storage capacity.
[0016] It can be understood that the above-mentioned management method and system for e-books have the same beneficial effects, which will not be elaborated here. Description of the Drawings
[0017] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a functional block diagram of a management system for e-books provided by an embodiment of the present invention; Figure 2 It is a flowchart of a management method for e-books provided by an embodiment of the present invention. Detailed Embodiments
[0018] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in conjunction with the embodiments.
[0019] In some embodiments of the present application, referring to Figure 1 as shown, a management system for e-books includes: An acquisition unit, configured to acquire reading data and text data of all e-books, collect text features of each text data, and determine an initial evaluation value of each e-book based on the text features and a feature model.
[0020] A judgment unit, configured to generate a reading mark according to each piece of reading data, determine a reading identifier based on the number of marks of the reading mark, and judge whether to adjust an initial evaluation value according to the reading identifier.
[0021] A first management unit, configured to, when it is determined to adjust the initial evaluation value, compare the number of marks with a historical data set, determine an evaluation adjustment factor according to the comparison result, determine a target evaluation value according to the evaluation adjustment factor and the initial evaluation value, perform cluster analysis on all e-books according to text features, divide a first book set according to the clustering result, and compare each target evaluation value in the first book set with a target evaluation threshold, and construct a second book set according to the comparison result.
[0022] A second management unit, configured to store the first book set and the second book set and determine access permissions, and determine backup modes for the first book set and the second book set according to the book storage capacity.
[0023] Specifically, the acquisition unit uses web crawlers and sniffing tools to call on the platform's server to obtain the reading data and text data of all e-books. Each e-book corresponds to a set of reading data and a set of text data. The reading data includes parameters such as starting to read and pausing reading of a certain e-book. The text data includes parameters such as the word count of a certain e-book and the related fields it represents. The text features reflect the keywords related to the text data. For example, words such as crimes, mysterious events, and clues in a mystery novel. By combining the text features with the established feature model, the initial evaluation value of each e-book is determined. The initial evaluation value reflects the content quality and characteristics of an e-book. The higher the initial evaluation value, the wider the coverage of its related keywords. The judgment unit generates reading marks based on the reading data. For example, if a user frequently reads an e-book, a reading mark can be established for the user's reading behavior. Based on the number of these reading marks, a reading identifier is determined. A larger number of marks indicates that the user has a strong interest in it, and the reading identifier tends to be "popular" or "highly concerned". According to the reading identifier, it is judged whether the initial evaluation value needs to be adjusted. When the reading identifier shows that a certain book has a high degree of attention, its initial evaluation value needs to be adjusted. When it is determined that the initial evaluation value needs to be adjusted, the first management unit compares the number of marks with the historical data set to determine the evaluation adjustment factor, avoiding blind adjustment of the initial evaluation value, and thus obtaining the target evaluation value. At the same time, clustering analysis is performed on all e-books according to the text features. E-books with similar text features are classified into the first book set. For example, all martial arts e-books are grouped into one category. By clustering, all e-books are divided into multiple categories or sets, so that similar or related e-books are grouped into the same set, forming a stable and reliable classification management basis. The target evaluation value of each e-book in the first book set is compared with the target evaluation threshold to construct the second book set. The second book set represents a category of e-books that are similar or related and are "popular" or "highly concerned" in terms of reading. The second management unit is responsible for storing the first book set and the second book set and determining the relevant access rights. At the same time, the backup mode is determined according to the book storage capacity. For popular and large-capacity e-books, a more complex backup mode is required to prevent the leakage of some exclusive e-books, thus ensuring the data security of e-books.
[0024] It can be understood that through the clustering analysis of text features and the adjustment of the initial evaluation value based on reading data, the system can classify e-books of different types and different reading levels into corresponding sets respectively, avoiding a chaotic storage state, improving the search efficiency of users when searching for specific books, determining the backup mode according to the book storage capacity, ensuring the security of e-book data, avoiding data loss caused by storage failures, and at the same time, setting relevant access permissions to protect the book copyright, thus realizing the orderly management of e-books.
[0025] In some embodiments of the present application, when collecting the text features of each text data, it includes: the collection unit analyzes the text data of all e-books based on natural language processing, performs word segmentation, stop word removal, and stemming on each text data, and uses a dependency syntax analysis model to parse the relationships between words to determine book phrases, and takes the book phrases that conform to the syntactic structure as keywords, and takes all the keywords of each text data and the text data size as text features.
[0026] Specifically, when collecting the text features of text data, first analyze the text data of all e-books through natural language processing (NLP) technology. The analysis includes the following three main steps: Word segmentation: Cut the continuous text in the text data into independent words or phrases for convenient subsequent processing. Stop word removal: Remove the common words that do not affect the actual semantics in the text data (such as "of", "is", etc.) to improve the effectiveness of the analysis. Stemming: Restore the words to their root forms (such as unifying "suspense" and "suspenseful" into "suspense") to reduce semantic redundancy. Use a dependency syntax analysis model to parse the grammatical and semantic relationships between words and extract noun phrases from the document. Dependency syntax analysis can accurately identify noun phrases (such as "ancient legend", "space exploration"), which reflect the themes of e-books and are the key basis for keyword extraction. The dependency syntax analysis model is mature and lengthy, so it will not be described in detail here. By combining NLP technology and dependency syntax analysis, the extraction and effective identification of text features are realized, thus improving the intelligent level of e-book management.
[0027] In some embodiments of the present application, when determining the initial evaluation value of each e-book based on text features and a feature model, it includes: obtaining a feature data set and sampling it according to a sampling ratio to obtain a training set and a test set, using grid search to find the hyperparameters of the model, establishing a random forest model, training the random forest model with the training set, substituting the test set into the trained random forest model and determining the accuracy rate predicted by the model. When the accuracy rate is greater than or equal to the accuracy rate threshold, determine the trained random forest model as the feature model. Otherwise, continue to train the random forest model until it is greater than or equal to the accuracy rate threshold, and substitute the text features into the feature model to determine the initial evaluation value of each e-book.
[0028] Specifically, the feature dataset contains all keyword data and the data evaluation values matched by the data samples. The feature dataset is divided into a training set and a test set. The division ratio is usually 4:1 to ensure that both the training set and the test set contain various data, so as to improve the generalization ability of the model. Grid search exhaustively searches for hyperparameters in the parameter space, such as the number of trees, the maximum depth of the trees, etc. The random forest model is trained using the training set. The random forest improves the accuracy and stability of the model by integrating multiple decision trees and averaging the prediction results of multiple trees. During the training process, the model attempts to learn the patterns in the data and the relationships between the data to improve the ability of prediction or classification. The test set is input into the trained random forest model to calculate the accuracy of the model. The accuracy reflects the performance of the model on unknown data and is an important indicator for evaluating the model performance. After the model is greater than or equal to the accuracy threshold, it is considered that the model has been able to stably approach the global optimal solution, and then the trained random forest model is determined as the feature model. The feature model is used to output the initial evaluation value. When the model cannot reach the accuracy threshold, the random forest model is continuously trained until it is greater than or equal to the accuracy threshold. The accuracy threshold in this embodiment is preferably 80%. The text features are substituted into the feature model to determine the initial evaluation value of each e-book, and machine learning models are used for prediction, ensuring the reliability and stability of the e-book classification management.
[0029] In some embodiments of the present application, when generating a reading mark according to each reading data and determining a reading identifier based on the number of marks of the reading mark, it includes: each text data corresponds to a reading data. The judgment unit obtains the interval time between each reading and the previous reading in the reading data, and generates a reading mark for the interval time greater than the interval time threshold, and counts the number of marks generated for the reading mark. A first preset number of marks and a second preset number of marks are preset in advance. The first preset number of marks is greater than the second preset number of marks. When the number of marks is greater than or equal to the first preset number of marks, the reading identifier is determined as a high-frequency reading identifier. When the number of marks is less than the first preset number of marks and greater than the second preset number of marks, the reading identifier is determined as a medium-frequency reading identifier. When the number of marks is less than or equal to the second preset number of marks, the reading identifier is determined as a low-frequency reading identifier.
[0030] In some embodiments of the present application, when judging whether to adjust the initial evaluation value according to the reading identifier, it includes: when a high-frequency reading identifier is recognized, the judgment unit determines to adjust the initial evaluation value; otherwise, it determines not to adjust the initial evaluation value and determines the initial evaluation value as the target evaluation value.
[0031] Specifically, each e-book corresponds to a reading data and a text data. The judgment unit focuses on the interval time between each reading and the previous reading. This indicator can effectively reflect the user's reading habits and interest intensity in the book. The interval time threshold is preferably 60s to avoid the influence of misjudgment on the statistical results caused by the user's behaviors such as screen switching or clicking on advertisements. When the interval time exceeds the preset interval time threshold, a reading mark is generated for this interval time, and then the abstract reading behavior is converted into quantifiable data that can be statistically analyzed. By counting the number of reading marks and comparing it with the first preset mark number and the second preset mark number set in advance, the reading behavior is divided into three levels: high-frequency reading identifier, medium-frequency reading identifier, and low-frequency reading identifier. The first preset mark number is preferably 10, and the second preset mark number is preferably 5. The high-frequency reading identifier means that the user is highly interested in the e-book, and multiple readings are an intuitive manifestation of the attractiveness of the e-book. The medium-frequency reading identifier represents that the user maintains a certain degree of attention to the book, but the interest intensity is slightly weaker. The low-frequency reading identifier indicates that the user has a low reading willingness. When a high-frequency reading identifier is recognized, it is determined that the initial evaluation value needs to be adjusted because the user's reading behavior verifies the actual value of the e-book and the initial evaluation value needs to be changed. In other cases, the initial evaluation value is maintained, making the scattered reading data structured and ensuring the intelligent level of classifying and managing e-books.
[0032] In some embodiments of the present application, when comparing the number of marks with the historical data set and determining the evaluation adjustment factor according to the comparison result, and determining the target evaluation value according to the evaluation adjustment factor and the initial evaluation value, it includes: the historical data set includes a number of historical mark numbers and a number of historical evaluation adjustment factors, and each historical mark number corresponds to a historical evaluation adjustment factor. When there is a historical mark number in the historical data set whose similarity to the high-frequency reading identifier is greater than the similarity threshold, the first management unit uses the historical evaluation adjustment factor corresponding to the historical mark number with the maximum similarity as the evaluation adjustment factor. Otherwise, the first management unit determines the evaluation adjustment factor according to the pre-trained evaluation model, and the target evaluation value is the product value of the evaluation adjustment factor and the initial evaluation value.
[0033] Specifically, the matching degree between the current number of tags and historical conditions is judged through a similarity threshold. When historical data with a high similarity is found, these data can be directly used to determine the evaluation adjustment factor, thereby ensuring the reliability and consistency of the adjustment operation. For the situation where the current number of tags does not fully match the historical data, the evaluation adjustment factor is determined through a pre-trained evaluation model to cope with the change in the number of tags. Through data-driven automatic adjustment, the automation level of the system and the accuracy of e-book classification are improved. By comprehensively using a large amount of historical data, rich reference information is provided for the determination of the evaluation adjustment factor, and it can learn and optimize from historical experience, thereby continuously improving the accuracy of the evaluation adjustment factor. The evaluation model is obtained by training based on different data samples and the evaluation adjustment factors matched by the data samples. The specific training process is the same as that of the feature model and will not be repeated here. The larger the number of tags, the larger the corresponding evaluation adjustment factor.
[0034] It can be understood that the initial evaluation value is adjusted according to the evaluation adjustment factor. When the number of tags is larger, the corresponding evaluation adjustment factor will increase accordingly. By establishing the product relationship between the evaluation adjustment factor and the initial evaluation value, the control of the initial evaluation value is realized, and the intelligent level of e-book classification management is improved.
[0035] In some embodiments of the present application, when all e-books are subjected to clustering analysis according to text features and the first book set is divided according to the clustering result, it includes: the first management unit normalizes each parameter in the text features to form a text feature vector, combines the text feature vectors of all e-books into a feature matrix, randomly selects K clustering centers, and determines the clustering distance from each text feature vector to each clustering center.
[0036] S1: Assign the e-books to the set corresponding to the clustering center with the closest clustering distance.
[0037] S2: Re-determine the position of each clustering center.
[0038] S3: Repeat S1 and S2 until the position of the clustering center no longer changes or reaches the preset number of iterations, and divide all e-books into K first book sets.
[0039] Specifically, normalizing each parameter in the text features (such as size and all keywords) ensures that all feature values are within the same magnitude range (such as [0, 1]), avoiding distortion in the clustering results due to scale differences in the feature vectors. According to the k-means clustering algorithm, the reliability and stability of e-book classification management are achieved. Based on the multi-dimensional feature vectors and normalization processing, the similarity between text features is effectively captured, ensuring the accuracy of the division results. The clustering center K and the number of iterations are adjusted according to the specific application scenario of the e-books. The set of text feature divisions provides data support for subsequent storage. The method for determining the clustering distance is lengthy and mature, and will not be described in detail here.
[0040] In some embodiments of the present application, when comparing each target evaluation value in the first book set with the target evaluation threshold and constructing a second book set according to the comparison results, it includes: The first management unit extracts the e-books in the first book set whose target evaluation values are greater than or equal to the target evaluation threshold, and constructs a second book set, and constructs the remaining e-books in the first book set into a third book set.
[0041] In some embodiments of the present application, when storing the first book set and the second book set, determining the access rights, and determining the backup modes of the first book set and the second book set according to the book storage capacity, it includes: The second management unit establishes a first storage repository and a second storage repository. The second management unit stores all the second book sets in the first storage repository, sets a first access right for the first storage repository, stores all the third book sets in the second storage repository, and sets a second access right for the second storage repository. The level of the first access right is greater than the level of the second access right. Obtain the first book storage capacity of the first storage repository. When the first book storage capacity is greater than or equal to half of the storage capacity of the first storage repository, determine the backup mode of the first storage repository as incremental backup. When the first book storage capacity is less than half of the storage capacity of the first storage repository, determine the backup mode of the first storage repository as full backup, and determine the backup mode of the second storage repository as periodic backup.
[0042] Specifically, the first management unit compares the target evaluation value with the target evaluation threshold, achieving a secondary screening and grading of e-books. The target evaluation value comprehensively considers the text features of e-books and the reading data of users, and is a quantitative manifestation of the book quality and popularity. E-books with a target evaluation value greater than or equal to the target evaluation threshold are extracted to construct a second book set. The keywords of the e-books in the second book set cover a wide range of fields, the content is relatively comprehensive, and there are many repeated readings by users. The remaining e-books in the first book set are grouped into a third book set, whose overall content and popularity are relatively lower compared to the second book set. By dividing according to the target evaluation value, a refined classification management of e-books is achieved. Since all e-books are divided into K first book sets during clustering, when extracting e-books, multiple second book sets and multiple third book sets will also be formed. The second management unit establishes a first storage repository and a second storage repository. When establishing the first storage repository and the second storage repository, it is based on MongoDB or Firebase, and users can traverse and search for the required e-books in the first storage repository and the second storage repository. The second book set is stored in the first storage repository and a higher first access permission is set. These e-books have higher value and popularity and are suitable for opening to specific users to avoid the risk of leakage of exclusive e-books. The third book set is stored in the second storage repository and a lower second access permission is set to balance resource allocation and user needs.
[0043] It can be understood that the first book storage capacity represents the actual capacity size of the second book set. In the selection of the backup mode, a dynamic decision is made based on the book storage capacity of the first storage repository. When the first book storage capacity is greater than or equal to half of the storage capacity of the first storage repository, incremental backup is adopted, and only the newly added or modified e-books are backed up, thereby improving the backup efficiency and saving storage resources. When the capacity is less than half, full backup is adopted to ensure the integrity and reliability of the e-book backup. For the second storage repository, periodic backup is adopted to back up the e-books at a fixed cycle, reducing the backup cost and resource burden while ensuring the safe management of e-books, and improving the level of intelligent management of e-books.
[0044] In summary, the beneficial effects of the present invention are as follows: Optimize the classification system: Obtain the e-book text data, extract the text features, and combine with the feature model to determine the initial evaluation value, ensuring the reliability of quantifying each e-book. Conduct clustering analysis on all e-books, and be able to classify the e-books covering different fields and rich in variety in an orderly manner according to the text features of the e-books, avoiding the phenomenon of chaotic storage of e-books due to the lack of an effective classification system. The first management unit determines the evaluation adjustment factor through comparison, integrates the user's reading data into the evaluation system of the e-books, so that the classification management of the e-books can accurately reflect the user's reading habits and interest tendencies. The second management unit stores the first book set and the second book set, determines the access rights, and dynamically determines the backup mode according to the book storage capacity, not only ensuring the security of e-book storage and avoiding data loss, but also dynamically planning the backup strategy according to the actual storage requirements, improving the utilization efficiency of storage resources, and ensuring the reliability and intelligence of e-book management.
[0045] In another preferred embodiment based on the above embodiments, refer to Figure 2 As shown, this embodiment provides a management method for e-books, which is used to apply the above management system for e-books, and includes: S100: Obtain the reading data and text data of all e-books, collect the text features of each text data, and determine the initial evaluation value of each e-book based on the text features and the feature model.
[0046] S200: Generate a reading mark according to each reading data, determine the reading identifier based on the number of marks of the reading mark, and determine whether to adjust the initial evaluation value according to the reading identifier.
[0047] S300: When it is determined to adjust the initial evaluation value, compare the number of marks with the historical data set, determine the evaluation adjustment factor based on the comparison result, determine the target evaluation value according to the evaluation adjustment factor and the initial evaluation value, conduct clustering analysis on all e-books according to the text features, divide the first book set according to the clustering result, and compare each target evaluation value in the first book set with the target evaluation threshold, and construct the second book set according to the comparison result.
[0048] S400: Store the first book set and the second book set and determine the access rights, and determine the backup mode of the first book set and the second book set according to the book storage capacity.
[0049] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0050] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0051] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0052] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. An e-book management system, characterized in that, Including: A collection unit configured to obtain the reading data and text data of all e-books, collect the text features of each text data, and determine the initial evaluation value of each e-book based on the text features and the feature model; A judgment unit configured to generate a reading mark according to each reading data, determine a reading identifier based on the number of marks of the reading mark, and judge whether to adjust the initial evaluation value according to the reading identifier; A first management unit configured to, when it is determined to adjust the initial evaluation value, compare the number of marks with a historical data set, determine an evaluation adjustment factor according to the comparison result, determine a target evaluation value according to the evaluation adjustment factor and the initial evaluation value, perform clustering analysis on all e-books according to the text features, divide a first book set according to the clustering result, and compare the target evaluation value of each e-book in the first book set with a target evaluation threshold, and construct a second book set according to the comparison result; A second management unit configured to store the first book set and the second book set and determine access permissions, and determine the backup modes of the first book set and the second book set according to the book storage capacity.
2. The management system for an e-book according to claim 1, wherein, When collecting the text features of each text data, it includes: The collection unit analyzes the text data of all e-books based on natural language processing, performs word segmentation, stop word removal, and stemming extraction on each text data, and uses a dependency syntax analysis model to analyze the relationships between words to determine book phrases, and uses the book phrases that conform to the syntactic structure as keywords; Taking all the keywords of each text data and the text data size as the text features.
3. The management system for an e-book according to claim 2, characterized in that, When determining the initial evaluation value of each e-book based on the text features and the feature model, it includes: Obtaining a feature data set and sampling it according to a sampling ratio to obtain a training set and a test set, using grid search to find the hyperparameters of the model, and establishing a random forest model; Training the random forest model with the training set, substituting the test set into the trained random forest model and determining the accuracy rate predicted by the model. When the accuracy rate is greater than or equal to the accuracy rate threshold, determining the trained random forest model as the feature model, otherwise, continuing to train the random forest model until it is greater than or equal to the accuracy rate threshold; Substituting the text features into the feature model to determine the initial evaluation value of each e-book.
4. The management system for an e-book according to claim 3, wherein When generating a reading mark according to each reading data and determining a reading identifier based on the number of marks of the reading mark, it includes: Each text data corresponds to a reading data; The judgment unit obtains the interval time between each reading and the previous reading in the reading data, generates the reading mark for the interval time greater than the interval time threshold, and counts the number of marks generated for the reading mark; Presetting a first preset number of marks and a second preset number of marks, the first preset number of marks being greater than the second preset number of marks; When the number of marks is greater than or equal to the first preset number of marks, determining the reading identifier as a high-frequency reading identifier; When the number of the marks is less than the first preset number of marks and greater than the second preset number of marks, determine that the reading identifier is a medium-frequency reading identifier; When the number of the marks is less than or equal to the second preset number of marks, determine that the reading identifier is a low-frequency reading identifier.
5. The management system for e-books according to claim 4, characterized in that, When determining whether to adjust the initial evaluation value according to the reading identifier, it includes: When the high-frequency reading identifier is recognized, the judgment unit determines to adjust the initial evaluation value; otherwise, it determines not to adjust the initial evaluation value and determines the initial evaluation value as the target evaluation value.
6. The management system for e-books according to claim 5, characterized in that, When comparing the number of the marks with the historical data set, determining an evaluation adjustment factor according to the comparison result, and determining a target evaluation value according to the evaluation adjustment factor and the initial evaluation value, it includes: The historical data set includes a plurality of historical numbers of marks and a plurality of historical evaluation adjustment factors, and each historical number of marks corresponds to a historical evaluation adjustment factor; When there is a historical number of marks in the historical data set whose similarity to the high-frequency reading identifier is greater than the similarity threshold, the first management unit uses the historical evaluation adjustment factor corresponding to the historical number of marks with the maximum similarity as the evaluation adjustment factor; otherwise, the first management unit determines the evaluation adjustment factor according to a pre-trained evaluation model; The target evaluation value is the product value of the evaluation adjustment factor and the initial evaluation value.
7. The management system for e-books according to claim 6, characterized in that, When performing clustering analysis on all e-books according to the text features and dividing a first book set according to the clustering result, it includes: The first management unit performs normalization processing on each parameter in the text features to form a text feature vector, combines the text feature vectors of all e-books into a feature matrix, randomly selects K clustering centers, and determines the clustering distance from each text feature vector to each clustering center; S1: Assign the e-books to the set corresponding to the clustering center with the closest clustering distance; S2: Re-determine the positions of each clustering center; S3: Repeat S1 and S2 until the positions of the clustering centers no longer change or reach a preset number of iterations, and divide all e-books into K first book sets.
8. The management system for an e-book according to claim 7, wherein, When comparing each target evaluation value in the first book set with a target evaluation threshold and constructing a second book set according to the comparison result, it includes: The first management unit extracts the e-books in the first book set whose target evaluation values are greater than or equal to the target evaluation threshold, constructs the second book set, and constructs the remaining e-books in the first book set into a third book set.
9. The management system for an e-book according to claim 8, characterized in that, When storing the first book set and the second book set and determining access permissions, and determining the backup modes of the first book set and the second book set according to the book storage capacity, it includes: The second management unit establishes a first storage repository and a second storage repository; The second management unit stores all the second book sets in the first storage repository, sets a first access permission for the first storage repository, stores all the third book sets in the second storage repository, and sets a second access permission for the second storage repository. The level of the first access permission is higher than that of the second access permission; Obtain the first book storage capacity of the first repository. When the first book storage capacity is greater than or equal to half of the storage capacity of the first repository, determine the backup mode of the first repository as incremental backup; When the first book storage capacity is less than half of the storage capacity of the first repository, determine the backup mode of the first repository as full backup; Determine the backup mode of the second repository as periodic backup.
10. A management method for an e-book, which is used for applying the management system for an e-book as described in any one of claims 1-9, characterized in that, Include: Obtain the reading data and text data of all e-books, collect the text features of each text data, and determine the initial evaluation value of each e-book based on the text features and the feature model; Generate a reading mark according to each reading data, determine a reading identifier based on the number of marks of the reading mark, and determine whether to adjust the initial evaluation value according to the reading identifier; When it is determined to adjust the initial evaluation value, compare the number of marks with the historical data set, determine an evaluation adjustment factor based on the comparison result, determine a target evaluation value according to the evaluation adjustment factor and the initial evaluation value, perform clustering analysis on all e-books according to the text features, divide the first book set according to the clustering result, and compare each target evaluation value in the first book set with a target evaluation threshold, and construct a second book set according to the comparison result; Store the first book set and the second book set and determine the access rights, and determine the backup modes of the first book set and the second book set according to the book storage capacity.
Citation Information
Patent Citations
Electronic book recommendation method and device, and server
CN106611050A
E-book recommendation method and device
CN110096644A
Library electronic book intelligent borrowing and returning system and method
CN118350907A
Reading analysis processing method of electronic book
CN118626804A
Book recommendation method and system based on artificial intelligence
CN119128282A
Cited By
Page turning control method and system for electronic reader
CN120447748A
Page turning control method and system for electronic reader
CN120447748B