Chinese language and literature classification management system in data management
Through data integration and real-time update of the data management system, the automated classification and personalized recommendation of Chinese language and literature are realized, and the problem of inefficient classification management of Chinese language and literature in the existing technology is solved, and the efficiency of literature retrieval and user experience are improved.
Patent Information
- Application Number
- CN202510529699.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology cannot achieve complete classification management of Chinese language and literature, resulting in low literature retrieval efficiency, difficulty in integrating cross-media resources, insufficient mining of implicit knowledge, and inability to meet the needs of efficient utilization in the information age.
It provides a Chinese language and literature classification management system in data management, including a data service platform, data update module, literature classification module, database, search and analysis module and user interaction module. Through data integration, real-time update, automated classification and personalized recommendation, multi-dimensional literature management and accurate retrieval can be realized.
The management efficiency and response speed of Chinese language and literature documents have been improved, and users can quickly retrieve the required documents, meet personalized needs, and avoid unauthorized users from disrupting the system through permission management.
Smart Images

Figure CN120448540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Chinese language management, and in particular to a Chinese language and literature classification management system in data management. Background Art
[0002] Chinese language and literature are the core carriers of Chinese cultural heritage, carrying the genetic code of national thought, historical memory, aesthetic values, and language evolution. Their role is to preserve cultural heritage, disseminate humanistic spirit, promote language research, and serve educational practice. Classification management is an inevitable requirement for efficiently utilizing the vast literary resources of the information age. Through systematic classification indexing and dynamic management, accurate retrieval and multi-dimensional analysis (such as by theme, genre, era, or author) of classics, works, and research materials can be achieved, addressing the inefficiency of traditional manual organization, the difficulty of integrating cross-media resources, and the inadequate mining of implicit knowledge. At the same time, classification management can support paradigm shifts in academic research (such as big data text analysis), optimized sharing of educational resources (such as hierarchical recommendations), and the digital preservation of cultural heritage (such as standardized storage of heterogeneous resources), thereby promoting the transition of literary resources from static storage to intelligent services, serving the diverse needs of cultural inheritance, academic innovation, and social education.
[0003] Chinese invention patent CN108830754A discloses a Chinese language and literature classification management system for library management, comprising a Chinese language book management platform, a management terminal, and a client. The management terminal and the client are respectively connected to the Chinese language book management platform. The Chinese language book management platform includes a book adding module, a book deleting module, a book query module, and a permission management module. The client includes a recommended book viewing module and a book information query module. The management terminal includes a book category management module, a book information management module, a user information management module, and a book borrowing management module. It integrates the client, the Chinese language book management platform, and the management terminal into a library management system via a network, ultimately enabling users to query and borrow books via a wireless network, thereby improving the efficiency of library management. However, it does not achieve complete classification management of Chinese language and literature. Summary of the Invention
[0004] The present invention provides a Chinese language and literature classification management system for data management to solve the technical problems of low flight stability and monitoring effect, lagging data processing capability and inability to view status in real time in existing drones used for forest fire prevention.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] The present invention provides a Chinese language and literature classification management system for data management, comprising:
[0007] Data service platform: responsible for providing data interfaces and integrating multi-source data, realizing multi-modal resource linkage through data analysis and providing API services;
[0008] Data update module: responsible for real-time capture and synchronous update of resources, and dynamic maintenance of data by combining version control and user feedback mechanism;
[0009] Document classification module: used to automatically classify and index documents and build a dynamic tag system to ensure the standardization of literary resources;
[0010] Database: As the foundation of the Chinese language and literature classification management system, it is used to carry, organize and manage all data;
[0011] Retrieval and analysis module: supports semantic retrieval and combined queries, assisting in in-depth research and management of documents and decision support;
[0012] User interaction module: responsible for providing personalized recommendations, virtual bookshelves and community collaboration platforms, supporting online reading annotations and special discussions;
[0013] User management module: It realizes classification management of users by granting corresponding permissions through user identity information authentication.
[0014] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0015] The present invention can improve the response speed of the management system through the integration and diversified analysis of data, facilitate users to quickly search for corresponding documents, and improve the efficiency of Chinese language and literature classification management.
[0016] The present invention enables users to obtain the latest Chinese language literature in a timely manner by updating data in real time, and at the same time can improve the accuracy of Chinese language and literature data updates through user feedback.
[0017] The present invention can classify and manage Chinese language and literature documents by classifying and labeling Chinese language and literature, and at the same time, by optimizing the search method, enable users to find similar documents according to their own needs.
[0018] The present invention builds a user profile based on the user's information, which can meet the user's needs when pushing documents and facilitate the user's search.
[0019] The present invention can grant users different operation permissions by classifying and managing users, thus preventing unauthorized users from disrupting the management system and further improving the efficiency of the classification management of Chinese language and literature. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a system diagram of the Chinese language and literature classification management system in data management provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0023] This embodiment provides a Chinese language and literature classification management system for data management.
[0024] 1. Data Service Platform
[0025] The data service platform includes a data integration unit, a data analysis unit, and a data interface unit; the data integration unit is responsible for collecting, cleaning, and standardizing Chinese language book data; the data analysis unit is responsible for identifying the genre of Chinese language books; and the data interface unit is responsible for providing data access and supporting a variety of data service requests.
[0026] The data collection includes text extraction, traditional Chinese character recognition, and variant character differentiation. The data collection is compared with a preset character library through OCR technology to identify the corresponding text information. The data cleaning includes text error correction and verification of rare characters. The data cleaning introduces a rule engine and connects with the library system and academic data system to check the rationality of the text through comparison of book data. The data standardization is used for text format conversion, and the data standardization is used to convert the PDF format into a searchable XML / TXT format to facilitate user retrieval.
[0027] The text style recognition includes text feature extraction and multimodal analysis. The text feature extraction includes topic identification, key word extraction, and sentiment style analysis. The text features select high-frequency words through the TextRank algorithm, and the high-frequency keywords are combined with the data in the database for comparison as the judgment basis. The multimodal analysis associates text and images through computer vision technology;
[0028] It should be noted that the association between text and images is used to match text and images, for example, automatically recommending related images when a user searches for a poem, or recommending descriptions of related scenes when a user searches for a poem.
[0029] The access portal includes a collection data system and an academic data system. The collection data system includes libraries and cultural institutions, and the academic data system includes network databases and academic forums.
[0030] It should be noted that cultural institutions include museums and memorial halls, online databases include academic platforms such as CNKI and Wanfang, and academic forums refer to online sites such as academic journals and academic websites.
[0031] 2. Data Update Module
[0032] The data update module includes a data capture unit, a data update unit, and a user feedback unit. The data capture unit is used to search and mine literary data and implement on-demand updates through scheduled tasks and special events. The data update unit is responsible for updating and synchronizing multi-source data. The user feedback unit allows users to submit data errors.
[0033] The scheduled task is used to automatically obtain literary data within a fixed time. The period of the scheduled task is set by the administrator and can be divided into daily, weekly, and monthly periods. The special events include API push and user upload. The API push is implemented in the form of Webhook;
[0034] It should be noted that user identity verification is required for user uploading, and only users who meet the identity verification requirements can upload documents.
[0035] The data update unit is used for incremental data update and cross-platform data synchronization. The incremental data update includes synchronizing new data and modified data. The cross-platform data synchronization realizes real-time synchronization between the online database and local data through the Kafka message queue;
[0036] The user feedback unit needs to verify the user's identity and grant the user corresponding data feedback rights based on the user's identity verification information. The user feedback rights include feedback data errors and supplementing missing data.
[0037] It should be noted that after the user reports a data error, the user feedback unit will send an error message to the other 20 users. When 10 users agree to the data feedback information, the data update unit will compare the data information in the database, make modifications, and send the modification information to other users. When more than 10 users agree to the modification, the data will be updated. If the user objects, the database will be compared again and modified.
[0038] When users supplement missing data, they also need verification from other users. If more than 10 out of 20 users agree to supplement the missing data, it will be adopted. If less than 10 users agree, it will not be adopted.
[0039] 3. Document Classification Module
[0040] The document classification module includes an automatic classification unit, an audit and correction unit, and a dynamic optimization unit. The automatic classification unit is used to classify and label documents. The audit and correction unit is used to correct the classification and labeling of documents and manually review them. The dynamic optimization unit is responsible for optimizing the label classification method.
[0041] The document classification includes CLC rules and custom extension rules. The document classification is fine-tuned using the BERT pre-trained model. The classification tasks include novels, poems, and essays. The document tags are classified using a multi-dimensional tag system. The multi-dimensional tag system includes topic tags and personalized tags. The topic tags are automatically generated using the LDA topic model. The personalized tags are generated based on the user's classification characteristics, which include age group, gender, and authority.
[0042] It should be noted that topic tags refer to the core themes in the document, such as local narratives, event discussions, and factual analysis. A document can have multiple topic tags to facilitate user retrieval.
[0043] Among the classification features, age group labels are used to classify document labels such as poetry, classical Chinese, and novels; gender labels are used to classify document labels such as war, feminism, and unknown exploration; and permission labels are used to hierarchically manage document permissions and to protect documents such as master's and doctoral theses and core journals.
[0044] The review and correction unit is managed by an administrator, who reviews the classification label errors reported by users. If the user feedback is correct, the automated classification unit is optimized using machine learning. If the user feedback is incorrect, an explanation is sent to the user.
[0045] The dynamic optimization unit uses the Context-Gloss Pair training model and constructs positive and negative context labels for the core sentences in the document in combination with the database. The positive label pair represents the label with the correct meaning of the target word, and the negative label pair represents the negative candidate label. By combining the context labels with the same context and target word into a training instance, the corresponding relevance score is calculated, and the relevance scores of the same group are normalized to obtain the relevance between the label and the document, as shown in the following formula:
[0046]
[0047] Where L represents the correlation score; N represents the batch size; m i represents the number of candidate words for the i-th training instance; l(s i , j) represents the index j and the forward context label s iA binary indicator when the index of ij The probability of the jth candidate word meaning of the i-th training instance; Score (context i , Gloss ij ) represents the relevance score of the context label; n i represents the total number of samples; k represents a constant.
[0048] It should be noted that the Gloss model calculates the correlation probability between the label and the context based on the context information of the core sentence, and then sorts each label in order according to the score. Each document can store up to 5 labels.
[0049] 4. Database
[0050] The database adopts a hybrid storage architecture to manage structured metadata and achieves precise matching through multimodal association and index optimization. The hybrid storage architecture includes structured storage and multimodal storage. The structured storage includes a document database, a relational database, and a graph database. The document database stores full-text content and classification labels in JSON format. The relational database establishes text relationship data according to the classification method. The graph database is used to construct author-work-genre association data. The multimodal storage is used to uniformly manage heterogeneous data of text and image types.
[0051] It should be noted that text relational data includes author information, bibliographic classification, and foreign key associations, which are used to associate document-related data to facilitate users to quickly review the general content of the document.
[0052] 5. Retrieval and Analysis Module
[0053] The retrieval analysis module includes a query parsing unit, a retrieval execution unit, and a retrieval optimization unit; the query parsing unit is responsible for converting the query statement input by the user into a standardized format that can be processed by the system; the retrieval execution unit supports multi-modal retrieval and multi-modal joint retrieval; the retrieval optimization unit is used to optimize the retrieval method;
[0054] It should be noted that the query statement includes tags, keywords, and authors, and the standardized format includes Boolean expressions and vectorized queries.
[0055] The query parsing unit uses the NLP model to analyze and understand the user's implicit needs, and accurately understands the user's needs by automatically expanding the search terms and distinguishing polysemous words;
[0056] The multimodal retrieval of the retrieval execution unit includes precise retrieval, fuzzy retrieval, and semantic retrieval. The multimodal joint retrieval supports retrieval from local databases and API service interfaces.
[0057] The retrieval optimization unit constructs a CNN convolutional neural network model and uses filtering and pooling operations to better learn the information representation of each sequence context through sequence recommendation, and uses the learned information representation for subsequent recommendations, as shown in the following formula:
[0058]
[0059] e c =max{max(α 1 ), max(α 2 ),…,max(α z )}
[0060] Where, Represents the convolution value of feature x at position m; φ α represents the activation function of the convolutional layer; E m:m+h-1 represents the embedding vector starting from position m and with a length of h; E represents the feature vector; F x represents the feature vector represented by feature x; α x represents the convolution result; |c| represents the absolute value of the number of sequences; e c is the representation of the context information sequence c, which is used for subsequent recommendations.
[0061] 6. User Interaction Module
[0062] The user interaction module includes a user portrait unit, a personalized recommendation unit, and an online forum unit. The user portrait unit can construct a user portrait based on the user's search information. The personalized recommendation unit is used to push document information based on the user portrait. The online forum unit supports multiple users to log in and opens user comments for discussing document information.
[0063] The construction of the user portrait includes defining user behavior and behavior characteristics. The user behavior includes borrowing history, annotation feedback, and forum posting content. The behavior characteristics include high-frequency search information, daily active information, and click preferences.
[0064] The personalized recommendation unit uses a collaborative filtering algorithm to recommend resources that coincide with the preferences of similar user groups, and optimizes the recommended information by analyzing the matching degree between resource features and user tags;
[0065] It should be noted that users can manually select recommended tags or modify the system recommended tags to facilitate users to find corresponding resources according to their own needs.
[0066] The online forum unit supports an academic exchange unit, a resource crowdsourcing unit, and a data feedback unit. The academic exchange unit is used to provide a topic discussion area to support scholars, students, and enthusiasts to share views, ask questions, and collaborate. The resource crowdsourcing unit can initiate crowdsourcing projects and rank them based on user contribution data. The data feedback unit is used to analyze forum discussion hotspots and synchronize them to the user portrait unit, triggering the system to add new categories or recommend related resources.
[0067] It should be noted that users can obtain corresponding titles based on their rankings, such as the title of Best Contributor, Literary Correction Master, and Field Analysis Specialist, which are used to motivate users, and specific permissions can be granted to users based on their corresponding titles. For example, if a user obtains the title of Literary Correction Master, he or she will be given the permission to modify documents.
[0068] 7. User Management Module
[0069] The user management module includes an information management unit and a user authentication and authorization unit. The information management unit is responsible for managing the import and export of user information. The user authentication and authorization unit is responsible for classifying and managing users according to their role identity information and controlling user access rights.
[0070] The information management includes user registration, information entry, and password modification. The information management uses a real-name authentication mechanism to ensure the authenticity of user information and a hash algorithm to protect user passwords to prevent data leakage;
[0071] The user authentication and authorization unit uses Token tokens to classify the user's identity information. The classification information includes administrators, ordinary users, advanced users, and expert users. The administrator is responsible for the overall management and maintenance of the management system, has the highest authority, and can directly manage other modules. The ordinary user can search and find documents, and has the authority to post comments and feedback erroneous data. The advanced user can access privacy platforms, such as unpublished papers and journals. The expert user is usually responsible for the review and classification of professional knowledge, and can organize, classify, and approve documents.
[0072] It should be noted that administrators are generally operation and maintenance personnel, who obtain permissions through Token authentication. Ordinary users can upgrade to advanced users based on their contribution. Advanced users can modify, delete, and supplement literature data in corresponding fields. Expert users are qualified and certified industry experts.
[0073] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Thus, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.
[0074] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0075] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0076] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0077] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be noted that although the preferred embodiment of the present invention has been described, it is clear that those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered as within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.
Claims
1. A Chinese language and literature classification management system for data management, characterized by: include: Data service platform: responsible for providing data interfaces and integrating multi-source data, realizing multi-modal resource linkage through data analysis and providing API services; Data update module: responsible for real-time capture and synchronous update of resources, and dynamic maintenance of data by combining version control and user feedback mechanism; Document classification module: used to automatically classify and index documents and build a dynamic tag system to ensure the standardization of literary resources; Database: As the foundation of the Chinese language and literature classification management system, it is used to carry, organize and manage all data; Retrieval and analysis module: supports semantic retrieval and combined queries, assisting in in-depth research and management of documents and decision support; User interaction module: responsible for providing personalized recommendations, virtual bookshelves and community collaboration platforms, supporting online reading annotations and special discussions; User management module: It realizes classification management of users by granting corresponding permissions through user identity information authentication.
2. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The data service platform is responsible for providing data interfaces and integrating multi-source data, and realizing multi-modal resource linkage through data analysis and API services, including: The data service platform includes a data integration unit, a data analysis unit, and a data interface unit; the data integration unit is responsible for collecting, cleaning, and standardizing Chinese language book data; the data analysis unit is responsible for identifying the genre of Chinese language books; and the data interface unit is responsible for providing data access and supporting a variety of data service requests. The data collection includes text extraction, traditional Chinese character recognition, and variant character differentiation. The data collection is compared with a preset character library through OCR technology to identify the corresponding text information. The data cleaning includes text error correction and verification of rare characters. The data cleaning introduces a rule engine and connects with the library system and academic data system to check the rationality of the text through comparison of book data. The data standardization is used for text format conversion, and the data standardization is used to convert the PDF format into a searchable XML / TXT format to facilitate user retrieval. The text style recognition includes text feature extraction and multimodal analysis. The text feature extraction includes topic identification, key word extraction, and sentiment style analysis. The text features select high-frequency words through the TextRank algorithm, and the high-frequency keywords are combined with the data in the database for comparison as the judgment basis. The multimodal analysis associates text and images through computer vision technology; The access portal includes a collection data system and an academic data system. The collection data system includes libraries and cultural institutions, and the academic data system includes network databases and academic forums.
3. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The data update module is responsible for real-time capture and synchronous update of resources, and dynamically maintains data by combining version control and user feedback mechanisms, including: The data update module includes a data capture unit, a data update unit, and a user feedback unit. The data capture unit is used to search and mine literary data and implement on-demand updates through scheduled tasks and special events. The data update unit is responsible for updating and synchronizing multi-source data. The user feedback unit allows users to submit data errors. The scheduled task is used to automatically obtain literary data within a fixed time. The period of the scheduled task is set by the administrator and can be divided into daily, weekly, and monthly periods. The special events include API push and user upload. The API push is implemented in the form of Webhook; The data update unit is used for incremental data update and cross-platform data synchronization. The incremental data update includes synchronizing new data and modified data. The cross-platform data synchronization realizes real-time synchronization between the online database and local data through the Kafka message queue; The user feedback unit needs to verify the user's identity and grant the user corresponding data feedback rights based on the user's identity verification information. The user feedback rights include feedback data errors and supplementing missing data.
4. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The document classification module is used to automatically classify and index documents and build a dynamic tag system to ensure the standardization of literary resources, including: The document classification module includes an automatic classification unit, an audit and correction unit, and a dynamic optimization unit. The automatic classification unit is used to classify and label documents. The audit and correction unit is used to correct the classification and labeling of documents and manually review them. The dynamic optimization unit is responsible for optimizing the label classification method. The document classification includes CLC rules and custom extension rules. The document classification is fine-tuned with the BERT pre-trained model. The classification tasks include novels, poems, and essays. The document tags are classified using a multi-dimensional tag system. The multi-dimensional tag system includes topic tags and personalized tags. The topic tags are automatically generated using the LDA topic model. The personalized tags are generated based on the user's classification characteristics, which include age group, gender, and authority. The review and correction unit is managed by an administrator, who reviews the classification label errors reported by users. If the user feedback is correct, the automated classification unit is optimized using machine learning. If the user feedback is incorrect, an explanation is sent to the user. The dynamic optimization unit uses the Context-Gloss Pair training model and constructs positive and negative context labels for the core sentences in the document in combination with the database. The positive label pair represents the label with the correct meaning of the target word, and the negative label pair represents the negative candidate label. By combining the context labels with the same context and target word into a training instance, the corresponding relevance score is calculated, and the relevance scores of the same group are normalized to obtain the relevance between the label and the document, as shown in the following formula: Where L represents the correlation score; N represents the batch size; m i represents the number of candidate words for the i-th training instance; l(s i , j) represents the index j and the forward context label s i A binary indicator when the index of ij The probability of the jth candidate word meaning of the i-th training instance; Score (context i , Gloss ij ) represents the relevance score of the context label; n i represents the total number of samples; k represents a constant.
5. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The database serves as the foundation of the Chinese language and literature classification management system and is used to carry, organize, and manage all data, including: The database adopts a hybrid storage architecture to manage structured metadata and achieves precise matching through multimodal association and index optimization. The hybrid storage architecture includes structured storage and multimodal storage. The structured storage includes a document database, a relational database, and a graph database. The document database stores full-text content and classification labels in JSON format. The relational database establishes text relationship data according to the classification method. The graph database is used to construct author-work-genre association data. The multimodal storage is used to uniformly manage heterogeneous data of text and image types.
6. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The retrieval analysis module supports semantic retrieval and combined queries, assisting in in-depth research and management of documents, including: The retrieval analysis module includes a query parsing unit, a retrieval execution unit, and a retrieval optimization unit; the query parsing unit is responsible for converting the query statement input by the user into a standardized format that can be processed by the system; the retrieval execution unit supports multi-modal retrieval and multi-modal joint retrieval; the retrieval optimization unit is used to optimize the retrieval method; The query parsing unit uses the NLP model to analyze and understand the user's implicit needs, and accurately understands the user's needs by automatically expanding the search terms and distinguishing polysemous words; The multimodal retrieval of the retrieval execution unit includes precise retrieval, fuzzy retrieval, and semantic retrieval. The multimodal joint retrieval supports retrieval from local databases and API service interfaces. The retrieval optimization unit constructs a CNN convolutional neural network model and uses filtering and pooling operations to better learn the information representation of each sequence context through sequence recommendation, and uses the learned information representation for subsequent recommendations, as shown in the following formula: e c =max{max(α 1 ),max(α 2 ),…,max(α z )} Where, Represents the convolution value of feature x at position m; φ α represents the activation function of the convolutional layer; E m:m+h-1 represents the embedding vector starting from position m and with a length of h; E represents the feature vector; F x represents the feature vector represented by feature x; α x represents the convolution result; |c| represents the absolute value of the number of sequences; e c is the representation of the context information sequence c, which is used for subsequent recommendations.
7. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The user interaction module is responsible for providing personalized recommendations, virtual bookshelves, and community collaboration platforms, supporting online reading annotations and topic discussions, including: The user interaction module includes a user portrait unit, a personalized recommendation unit, and an online forum unit. The user portrait unit can construct a user portrait based on the user's search information. The personalized recommendation unit is used to push document information based on the user portrait. The online forum unit supports multiple users to log in and opens user comments for discussing document information. The construction of the user portrait includes defining user behavior and behavior characteristics. The user behavior includes borrowing history, annotation feedback, and forum posting content. The behavior characteristics include high-frequency search information, daily active information, and click preferences. The personalized recommendation unit uses a collaborative filtering algorithm to recommend resources that coincide with the preferences of similar user groups, and optimizes the recommended information by analyzing the matching degree between resource features and user tags: The online forum unit supports an academic exchange unit, a resource crowdsourcing unit, and a data feedback unit. The academic exchange unit is used to provide a topic discussion area to support scholars, students, and enthusiasts to share views, ask questions, and collaborate. The resource crowdsourcing unit can initiate crowdsourcing projects and rank them based on user contribution data. The data feedback unit is used to analyze forum discussion hotspots and synchronize them to the user portrait unit, triggering the system to add new categories or recommend related resources.
8. The Chinese language and literature classification management system for data management according to claim 1, characterized in that: The user management module grants corresponding permissions to users through identity authentication to achieve user classification management, where: The user management module includes an information management unit and a user authentication and authorization unit. The information management unit is responsible for managing the import and export of user information. The user authentication and authorization unit is responsible for classifying and managing users according to their role identity information and controlling user access rights. The information management includes user registration, information entry, and password modification. The information management uses a real-name authentication mechanism to ensure the authenticity of user information and a hash algorithm to protect user passwords to prevent data leakage; The user authentication and authorization unit uses Token tokens to classify the user's identity information. The classification information includes administrators, ordinary users, advanced users, and expert users. The administrator is responsible for the overall management and maintenance of the management system, has the highest authority, and can directly manage other modules. The ordinary user can search and find documents, and has the authority to post comments and feedback erroneous data. The advanced user can access privacy platforms, such as unpublished papers and journals. The expert user is usually responsible for the review and classification of professional knowledge, and can organize, classify, and approve documents.
Citation Information
Patent Citations
Chinese language and literature classification management system in library management
CN108830754A