Ancient book information intelligent processing method and device based on artificial intelligence
Through artificial intelligence-based methods, scanning, mapping, classification and multi-information network processing of ancient books, the difficult problems of ancient book information integration and utilization have been solved, and efficient digitization and in-depth understanding of ancient book information have been achieved.
Patent Information
- Application Number
- CN202411394210.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-08
AI Technical Summary
The existing technology for processing ancient book information is limited by digitization and fragmentation, which makes it impossible to effectively integrate and utilize ancient book information. Moreover, due to the diversity and complexity of ancient book documents, it is difficult to fully cover and efficiently utilize them.
An artificial intelligence-based method is used to obtain editable text by scanning ancient books, establish a knowledge graph of ancient books, classify them into domain databases, introduce multi-form historical carriers, build a diversified information network, and conduct deep learning and optimization adjustments through machine learning models. An ancient book information processing dialog box is established to achieve multi-round interactive feedback.
A hierarchically structured knowledge system has been constructed, which has achieved efficient integration and utilization of ancient book information, provided a means of rapid and in-depth exploration, and improved the efficiency and accuracy of ancient book digitization.
Smart Images

Figure CN119669478B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to digital data processing, and in particular to an ancient book information intelligent processing method and device based on artificial intelligence. Background Art
[0002] The rapid development of science and technology, particularly the widespread application of artificial intelligence and big data, has ushered in unprecedented opportunities for the preservation, compilation, and utilization of ancient texts, significantly advancing the digitization of ancient books. Ancient books, as the product of a nation's civilizational heritage, carry a wealth of historical and cultural information. However, conventional compilation methods are not only inefficient but also often limited by physical resources, making it difficult to fully capture this vast collection of ancient texts. Furthermore, the text and symbols in ancient books are often complex and diverse, involving multiple languages, fonts, and writing styles, posing significant challenges to the compilation and interpretation of this information.
[0003] In summary, the existing technology faces limitations in digitization and fragmentation in the processing of ancient book information, and due to the diversity and complexity of ancient book documents, there are technical problems that lead to the inability to effectively integrate and utilize ancient book information. Summary of the Invention
[0004] This application provides an artificial intelligence-based intelligent processing device for ancient book information, aiming to solve the limitations of digitization and fragmentation faced in the processing of ancient book information in the existing technology, as well as the technical problems that ancient book information cannot be effectively integrated and utilized due to the diversity and complexity of ancient book documents.
[0005] In view of the above problems, the technical solution to implement this application is:
[0006] On the one hand, the present application provides an intelligent processing method for ancient book information based on artificial intelligence, wherein the method comprises: scanning ancient books to obtain scanned text, wherein the scanned text is in an editable text format; establishing an ancient book knowledge graph based on the scanned text, wherein the ancient book knowledge graph sets triples with text punctuation as constraints, and the triples are limited to word segmentation and part of speech for text editing; through the ancient book knowledge graph, M domain databases are classified, and a machine learning model is used to determine M vertical pre-trained ancient book information processing models, wherein the classification indicators corresponding to the M domain databases include age, core issues, and representative original ideas; introducing multi-form historical carriers, wherein the multi-form historical carriers include cultural relics, relics, and intangible cultural heritage; based on the multi-form historical carriers and the age and core issues in the classification indicators, a first multimodal perception set is determined; based on the The multi-form historical carriers, and the original thought representatives and core topics in the classification indicators, determine the second multimodal perception set; based on the multi-form historical carriers, and the era and original thought representatives in the classification indicators, determine the third multimodal perception set; based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set, adopt a streaming generation mode to configure a multi-layer branch guidance structure, and embed the M vertical pre-trained ancient book information processing models into the multi-layer branch guidance structure to establish a multi-dimensional information network; establish an ancient book information processing dialog box, receive the first round of input information, perform interactive connection in the multi-dimensional information network, output the first round of response information, and respond to the second round of input information, the third round of input information, ..., the Nth round of input information with a context joint guidance mechanism, and optimize and adjust in combination with the minimum difference threshold and the maximum difference threshold.
[0007] On the other hand, the present application provides an intelligent processing device for ancient book information based on artificial intelligence, wherein the device includes: an ancient book scanning module for scanning ancient books and obtaining scanned text, wherein the scanned text is in an editable text format; a knowledge graph establishment module for establishing an ancient book knowledge graph based on the scanned text, wherein the ancient book knowledge graph sets triples with text punctuation as constraints, and the triples are limited to word segmentation and part of speech for text editing; a database classification module for classifying M field databases through the ancient book knowledge graph, and using a machine learning model to determine M vertical pre-trained ancient book information processing models, wherein the classification indicators corresponding to the M field databases include age, core issues, and representative original ideas; a historical carrier introduction module for introducing multi-form historical carriers, wherein the multi-form historical carriers include cultural relics, relics, and intangible cultural heritage; a multimodal perception module for combining the multi-form historical carriers with the age and core issues in the classification indicators, Determine a first multimodal perception set; determine a second multimodal perception set based on the multi-form historical carriers and the original thought representatives and core topics in the classification indicators; determine a third multimodal perception set based on the multi-form historical carriers and the era and original thought representatives in the classification indicators; a multi-information network establishment module is used to configure a multi-layer branch guidance structure based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set using a streaming generation mode, and embed the M vertical pre-trained ancient book information processing models into the multi-layer branch guidance structure to establish a multi-information network; an optimization and adjustment module is used to establish an ancient book information processing dialog box, receive a first round of input information, perform interactive connections in the multi-information network, output a first round of response information, and respond to the second round of input information, the third round of input information,..., the Nth round of input information with a context-joint guidance mechanism, and perform optimization and adjustment in combination with the minimum difference threshold and the maximum difference threshold.
[0008] In summary, the one or more technical solutions provided in this application solve the limitations of digitization and fragmentation faced in the processing of ancient book information, and the technical problem that ancient book information cannot be effectively integrated and utilized due to the diversity and complexity of ancient books and documents. It realizes the construction of knowledge graphs and the division of domain databases on the basis of digitization of ancient books and documents, forming a knowledge system with hierarchical structure and logical association, and through training vertical processing models, conducts deep learning and optimization for the specific needs of the ancient book field, and establishes dialog boxes to assist users in quickly and deeply exploring the technical effects of ancient books through multiple rounds of interactive feedback with users. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of an intelligent processing method for ancient book information based on artificial intelligence is provided for this application;
[0010] Figure 2 A structural diagram of an ancient book information intelligent processing device based on artificial intelligence is provided for this application.
[0011] Explanation of the accompanying drawings: ancient book scanning module M100, knowledge graph establishment module M200, database classification module M300, historical carrier introduction module M400, multimodal perception module M500, multi-information network establishment module M600, optimization and adjustment module M700. DETAILED DESCRIPTION
[0012] Example 1
[0013] The present application is described in detail below with reference to the accompanying drawings. Figure 1 As shown, the present application provides an ancient book information intelligent processing method based on artificial intelligence, wherein the method includes:
[0014] S1: Scan ancient books to obtain scanned text, which is in an editable text format; S2: Build an ancient book knowledge graph based on the scanned text, wherein the ancient book knowledge graph sets triples based on text punctuation as constraints, and the triples are limited to segmentation and part of speech for text editing.
[0015] Choose equipment suitable for scanning ancient books to ensure that every detail of the ancient books can be captured with high quality; place the ancient book on the scanner and scan it page by page or in batches to obtain the scanned image of the ancient book; use optical character recognition (OCR) technology to convert the scanned image into an editable text format (such as TXT, DOCX, etc.). It is necessary to select or train an OCR model suitable for ancient book texts to improve the recognition accuracy; since OCR recognition has a certain probability of error, the recognized text needs to be proofread to correct typos, format errors and other problems to ensure the accuracy of the text.
[0016] Remove irrelevant information from the text, such as page numbers, notes, illustrations, etc., and only retain the main text content of the ancient book; use the word segmentation tool to segment the cleaned text, and divide the continuous text into independent words or phrases. This is especially important for ancient Chinese books because Chinese text does not have natural spaces to separate words; perform part-of-speech tagging on the results after word segmentation, and assign a part-of-speech label (such as noun, verb, adjective, etc.) to each word or phrase, which will help understand the role and relationship of words in sentences when constructing knowledge graphs later.
[0017] Define the triple structure. In the ancient book knowledge graph, triples are the basic knowledge representation units, usually composed of "entity-relationship-entity" or "entity-attribute-value" forms. Triples are set with text punctuation as constraints, which means that punctuation marks (such as periods, commas, semicolons, etc.) can be used to identify independent information units in sentences or clauses, and then construct triples; based on the results of word segmentation and part-of-speech tagging, combined with text punctuation, triples are extracted from the text. It is necessary to develop special algorithms or use natural language processing (NLP) technology to identify entities, relationships and attributes.
[0018] The extracted triples are stored and displayed in the form of a graph structure to form an ancient book knowledge graph. In the process of establishing the ancient book knowledge graph, it is necessary to assign a unique identifier to each entity and establish connections for the relationships between entities; the constructed knowledge graph is optimized, such as removing redundant information, merging duplicate entities, etc. At the same time, it is verified through automated methods to ensure the accuracy and completeness of the knowledge graph.
[0019] S3: Through the ancient book knowledge graph, M domain databases are classified, and a machine learning model is used to determine M vertical pre-trained ancient book information processing models. The classification indicators corresponding to the M domain databases include age, core issues, and representative ideas of the original works.
[0020] Clarifying the specific categories of the M fields can be based on subject classification (such as history, literature, philosophy, medicine, etc.), time span (such as dynasties, eras, etc.), subject content (such as politics, economy, culture, etc.) or any other standards that help to subdivide the content of ancient books. The classification indicators corresponding to the M field databases include era (dividing according to the dynasty of creation of ancient books, which helps to understand the background and value of ancient books in different historical periods), core issues (analyzing the main discussion content of ancient books and classifying them into corresponding issues, such as political systems, economic thoughts, literature and art, etc.), and original thought representatives (evaluating the representativeness of ancient books in a certain ideological system or school, such as Confucian classics, Taoist works, etc.); according to the defined fields and classification indicators, the information in the ancient book knowledge map is classified into the corresponding field database.
[0021] Clean the data in each domain database to remove duplicate, erroneous or irrelevant information; extract features useful for machine learning models from the cleaned data, such as keywords, subject terms, sentiment tendencies, etc.; if a supervised learning model is required, some data needs to be labeled to provide the labels required for training; according to the specific tasks of ancient book information processing (such as classification, clustering, summary generation, sentiment analysis, etc.), select appropriate machine learning models, such as neural networks, decision trees, support vector machines, etc.; use data from the domain database to pre-train the model. The purpose of pre-training is to enable the model to learn the knowledge and patterns unique to each domain, thereby improving the accuracy and efficiency of subsequent task processing; evaluate the performance of the model through methods such as cross-validation, and adjust parameters or optimize the model as needed.
[0022] Deploy the trained vertical pre-trained ancient book information processing model to corresponding application scenarios, such as ancient book digitization platforms, academic research tools, cultural heritage projects, etc.; use the model to process and analyze new ancient book information in real time to provide fast and accurate results; with the continuous addition of new data and the continuous development of technology, regularly update and retrain the model to maintain its adaptability and accuracy; collect user feedback on the model processing results, and regularly evaluate the model's performance, including indicators such as accuracy, recall rate, and F1 score; based on user feedback and performance evaluation results, iteratively optimize the model to continuously improve its processing capabilities and user experience.
[0023] S4: Introduce multi-form historical carriers, which include cultural relics, relics, and intangible cultural heritage; S5: Determine a first multi-modal perception set based on the multi-form historical carriers and the era and core issues in the classification indicators; determine a second multi-modal perception set based on the multi-form historical carriers and the original ideological representatives and core issues in the classification indicators; determine a third multi-modal perception set based on the multi-form historical carriers and the era and original ideological representatives in the classification indicators.
[0024] In the process of constructing the knowledge graph of ancient books, introducing multiple forms of historical carriers (such as cultural relics, relics, and intangible cultural heritage) is an effective way to enrich and deepen knowledge representation. Introducing multiple forms of historical carriers not only provides the visual, physical, and cultural background of ancient book texts, but also enhances the understanding and interpretation of the content of ancient books.
[0025] The first multimodal perception set focuses on the specific era in which the ancient books were written and its core issues. By combining physical evidence such as cultural relics and relics with the era information and core issue descriptions in the ancient book texts, a comprehensive perception of the historical background and thematic content of the ancient books is formed. Furthermore, archaeological discoveries, historical documents and other materials are used to determine the approximate era range of the ancient books, and cultural relics and relics related to this era are selected as auxiliary materials; core issues in ancient books such as political systems, economic development, and cultural changes are extracted through text analysis, and cultural relics, relics or intangible cultural heritage related to these issues are sought as examples; information from multiple modalities such as text descriptions, image displays (photos or restorations of cultural relics and relics), and sound records (such as related commentaries and historical audio materials) are integrated to form a comprehensive perception of the era and core issues.
[0026] The second multimodal perception set focuses on the original ideas represented by ancient books and their embodiment in core issues. By deeply exploring the ideological connotations of ancient books and their contributions to specific issues, combined with relevant cultural relics, intangible cultural heritage, etc., a deep understanding of the original ideas and their influence is formed. Furthermore, through methods such as text analysis and intellectual history research, the original ideas in ancient books are refined, such as philosophical viewpoints, political propositions, cultural concepts, etc.; the extracted original ideas are correlated with the core issues in ancient books to clarify on which issues the original ideas have played an important role; physical evidence such as cultural relics, intangible cultural heritage, and related audio and video materials are used to demonstrate the inheritance and influence of the original ideas in history, and enhance the understanding of the original ideas and their core issues.
[0027] The third multimodal perception set combines the era of ancient books with the original ideas they represent. Through the display of multiple forms of historical carriers, it reveals the ideological value and cultural significance of ancient books in a specific historical context. Furthermore, it deeply analyzes the historical context of the ancient books, including changes and developments in politics, economy, culture, etc.; combined with the era background, it explores the original ideas in ancient books and understands their unique status and influence in the society at that time; uses various forms of historical carriers such as cultural relics, relics, intangible cultural heritage, as well as related text, images, audio and other multimedia materials to present the original ideas and cultural significance of ancient books in a specific era. At the same time, it can provide users with a more immersive perceptual experience through modern technologies such as virtual reality (VR) and augmented reality (AR).
[0028] After introducing multi-form historical carriers, the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set are determined to more comprehensively and deeply understand and display the historical value and cultural connotation of ancient books, providing a richer and more diverse source of information for the construction of the ancient book knowledge graph.
[0029] S6: Based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set, a streaming generation mode is adopted to configure a multi-layer branch guidance structure, and the M vertical pre-trained ancient book information processing models are embedded in the multi-layer branch guidance structure to establish a multi-dimensional information network.
[0030] The streaming generation mode is a method of sequentially processing data, allowing data to be continuously transmitted and processed in the system in the form of a stream. In this scenario, the data in the multimodal perception set is regarded as an input stream, and the required output information is gradually generated through a series of processing steps (i.e., each layer in the multi-layer branch-guided structure).
[0031] A multi-layer branch-guided structure is configured, including an input layer that receives data from the first, second, and third multimodal perception sets, including information in various forms such as text, images, and audio; a preprocessing layer that preprocesses the input data, such as text cleaning, image denoising, and audio conversion, to ensure data quality and consistency; a feature extraction layer that uses specialized algorithms or models to extract features from the preprocessed data, which can be keywords in the text, visual features in the image, spectral features in the audio, etc.; a branching layer that distributes feature data to different branches based on the characteristics of the data and processing requirements. Each branch corresponds to one or more vertical pre-trained ancient book information processing models, which specialize in processing specific types of information or tasks; a model processing layer in which each vertical pre-trained model receives corresponding feature data and performs its specific processing tasks, such as classification, clustering, and sentiment analysis; a fusion layer that fuses the processing results of each branch model to form a comprehensive understanding and representation of ancient book information. Special fusion algorithms or strategies need to be developed to ensure effective integration between different modalities and different types of information; and an output layer that outputs the fused results in an appropriate form, such as generating reports, visual displays, and providing decision support.
[0032] Based on the multi-layer branch guidance structure, in each branch layer, the corresponding vertical pre-trained ancient book information processing model is selected and embedded according to the data type and task requirements processed by the branch. The vertical pre-trained ancient book information processing model has been pre-trained on data in related fields, so it can efficiently process specific types of ancient book information. By embedding M vertical pre-trained ancient book information processing models into the multi-layer branch guidance structure, parallel processing and cross-modal fusion of multimodal information are realized.
[0033] With the establishment of a multi-layer branch-guided structure and the embedding of models, a multi-dimensional information network has gradually taken shape. The multi-dimensional information network can receive historical carrier data from different channels and forms, and generate a comprehensive and in-depth understanding and representation of ancient book information through a multi-level and multi-branch processing flow. At the same time, due to the embedding of multiple vertical pre-training models, the multi-dimensional information network also has a high degree of professionalism and flexibility, and can be customized and optimized according to different task requirements.
[0034] S7: Establish an ancient book information processing dialog box, receive the first round of input information, perform interactive connection in the multi-information network, output the first round of response information, and respond to the second round of input information, the third round of input information, ..., the Nth round of input information with a context-joint guidance mechanism, and perform optimization and adjustment in combination with the minimum difference threshold and the maximum difference threshold.
[0035] Load the established multivariate information network, including a multi-layer branch guidance structure and an embedded vertical pre-trained ancient book information processing model; set the minimum difference threshold and the maximum difference threshold to control the similarity and difference of the response information; provide an intuitive user interface to allow users to input queries, questions or instructions about ancient books; parse the information input by the user to identify its type (such as text, image, etc.) and main content; map the parsed input information to the corresponding nodes or branches in the multivariate information network; use the embedded vertical pre-trained model to process the input information and generate preliminary response information; establish or update the context information library to store the context information of the current conversation, including historical input, response and user interaction behavior.
[0036] Based on the results of model processing and contextual information, the first round of output information is generated; the output information is fed back to the user in an appropriate form (such as text, image, voice, etc.); when the user inputs subsequent information such as the second round and the third round, the association between this information and the current context is identified; using the context joint guidance mechanism, historical information and current input are combined to perform joint reasoning to generate a more accurate response.
[0037] When the current cumulative difference index is less than the minimum difference threshold, the current cumulative difference index is ignored, which also indicates that continuing pre-training can enhance the model's ability to process domain data; when the current cumulative difference index is greater than the maximum difference threshold, roll back, which also indicates that when there is a large difference in language features from general data, only domain data cannot be used for training, otherwise catastrophic forgetting may occur.
[0038] For subsequent input information such as the second and third rounds, repeat the above steps to generate and optimize response information; maintain continuous interaction with the user until the user ends the conversation or reaches the preset interaction round limit; collect user feedback on system responses, including satisfaction, accuracy, etc.; based on user feedback and system performance, regularly adjust and optimize the model parameters and structure in the multivariate information network; according to actual application conditions, adjust the minimum difference threshold and maximum difference threshold in a timely manner to further improve system performance; establish an efficient and intelligent ancient book information processing dialog box to provide users with convenient and accurate ancient book information query and interactive experience.
[0039] Furthermore, a machine learning model is used to determine M vertical pre-trained ancient book information processing models. The method of this application includes:
[0040] Perform fine-grained semantic analysis on the scanned text and add fine-grained annotation information, wherein the fine-grained annotation information includes the core vocabulary and the first internal logical relationship between the core vocabulary and the core topic, and the core phrase and the second internal logical relationship between the core topic; use a machine learning model to perform topic modeling, and continue to use M domain databases for pre-training to ensure that the pre-training data has fine-grained annotation information to support the training of the semantic analysis layer.
[0041] The scanned text is segmented according to language rules to form a sequence of words or phrases; each word is assigned a part-of-speech tag, such as noun, verb, adjective, etc.; core vocabulary in the text is extracted using algorithms such as TF-IDF and TextRank; and core phrases are identified through methods such as N-gram models and dependency syntax analysis.
[0042] Furthermore, the first internal logical relationship refers to analyzing the direct relationship between each core word and the core topic, such as causal relationships, parallel relationships, and modifying relationships, and annotating them. The second internal logical relationship refers to analyzing the logical relationship between each core phrase and the core topic, and annotating them in a more detailed manner. This may involve complex relationships between the internal components of the phrase, such as subject-predicate relationships and verb-object relationships. A annotation system, such as XML or JSON, is designed to store text, core words, core phrases, and their internal logical relationships in a structured manner.
[0043] Collect a large amount of text data related to ancient books from M domain databases; ensure that these pre-training data have also undergone fine-grained semantic analysis and annotation, including core vocabulary, core phrases and the inherent logical relationship between them and core topics; select topic modeling models suitable for processing text data, such as LDA (latent Dirichlet allocation) and NMF (non-negative matrix factorization); based on the need for fine-grained annotation information, select appropriate deep learning models, such as BERT and GPT, for training the semantic analysis layer.
[0044] The model is pre-trained using pre-training data with fine-grained annotations, so that the model can learn the fine-grained semantic features in the text; based on the pre-training, the model is fine-tuned using task-specific datasets to improve the model's performance on specific tasks; the results of topic modeling (such as topic distribution) are combined with the results of fine-grained semantic parsing to form a more comprehensive and in-depth understanding of the text; users are allowed to conduct interactive queries based on the results of topic modeling, and the system uses the semantic parsing layer to return more detailed and relevant text fragments or explanations.
[0045] We evaluated the model's performance in fine-grained semantic parsing and topic modeling through methods like cross-validation and A / B testing. Based on these evaluation results and user feedback, we continuously optimized the model structure and parameter settings to improve overall performance and user experience. Through these steps, we built an ancient book information processing system capable of handling fine-grained semantic parsing and topic modeling, providing users with more accurate and in-depth information query and analysis services.
[0046] Furthermore, the scanned text is subjected to fine-grained semantic analysis and fine-grained annotation information is added. The method of the present application further includes:
[0047] Using NLP technology, a set of ancient Chinese linguistics rules and a set of modern Chinese linguistics rules are defined; in the ancient book information processing dialog box, a first task instruction and a first text to be processed are generated using the ancient Chinese linguistics rule set and the modern Chinese linguistics rule set, wherein the first task instruction activates the NLP dynamic response mechanism; utilizing the NLP dynamic response mechanism, in accordance with the text processing requirements corresponding to the first task instruction, any one of the core vocabulary extraction mode and the core phrase extraction mode is selectively triggered.
[0048] Using NLP technology, we define a set of classical Chinese linguistic rules and a set of modern Chinese linguistic rules. The classical Chinese linguistic rule set includes: vocabulary rules, which refer to the unique vocabulary, homophones, variant characters, and word usage in classical Chinese; syntactic rules, which refer to the unique sentence structures in classical Chinese, such as inverted sentences, omitted sentences, and judgment sentences; semantic rules, which refer to metaphors, allusions, and cultural background knowledge in classical Chinese, which are crucial for understanding the deep meaning of the text.
[0049] A collection of modern Chinese linguistic rules, including vocabulary rules, which refer to commonly used words, new words, foreign words, etc. in modern Chinese; syntactic rules, which refer to the grammatical structure of modern Chinese, including basic sentence patterns such as subject, predicate, object, attributive, adverbial, and complement; semantic rules refer to contextual understanding, polysemous word analysis, and sentiment analysis of modern Chinese.
[0050] In the ancient book information processing dialog box, based on user input or a pre-set scenario, the system first generates a first task instruction. This first task instruction clearly specifies the text processing task to be performed (such as translation, summarization, information extraction, etc.). Simultaneously, the system also generates or receives a first text to be processed: the ancient or modern text to be processed. This first task instruction is designed to be the key to activating the NLP dynamic response mechanism, dynamically selecting and configuring the most appropriate NLP processing flow based on the specific content of the task instruction.
[0051] Within the NLP dynamic response mechanism, the system will analyze and determine the core text processing mode to be adopted based on the text processing requirements corresponding to the first task instruction, including but not limited to the core vocabulary extraction mode: when the task requirements focus on identifying key words in the text, such as names of people, places, proper nouns, etc., the system will trigger this mode and use predefined vocabulary rules and context analysis technology to extract core words from the text; core phrase extraction mode: when the task needs to identify and extract key phrases or sentences in the text, such as parts that express core ideas and important information, the system will select this mode. Through syntactic analysis and semantic understanding technology, the system can identify and extract core phrases that meet the task requirements.
[0052] After selecting the appropriate text processing mode, the system processes the first text to be processed and generates processing results, including the translated text, summary, keyword list, and core phrase list, depending on the task instructions. This processing result is then fed back to the user, completing the entire processing process. Through these steps, the Ancient Books Information Processing Dialog System leverages NLP technology to efficiently process both ancient and modern texts, meeting users' diverse information needs.
[0053] Furthermore, the present application method also includes:
[0054] With a multi-level attention mechanism and a long short-term memory network, a gated recurrent unit is added to enhance the recognition of multiple text sequences of different lengths and types; in the ancient book information processing dialog box, a second task instruction and a second text to be processed are generated through the gated recurrent unit, and the second task instruction activates a grammatical analysis mechanism; using the grammatical analysis mechanism, in accordance with the text processing requirements corresponding to the second task instruction, any one of the part-of-speech tag marking mode and the punctuation tag marking mode is selectively triggered.
[0055] A multi-level attention mechanism is used to enhance the ability of LSTM to better capture key information in the text. The multi-level attention mechanism allows the model to focus on different parts of the text at different levels, thereby improving the ability to understand complex text structures. LSTM is responsible for processing long-term dependencies in text sequences, retaining important information and forgetting irrelevant information through its internal "gate" structure (forget gate, input gate, output gate).
[0056] In the Gated Recurrent Unit (GRU), by simplifying the LSTM structure (for example, by combining the forget gate and input gate into a single update gate), the model's computational efficiency and training speed are improved. In ancient text information processing, the GRU's flexibility enables it to effectively process text sequences of various lengths, from short verses to long ancient texts. Through the GRU's update and reset gates, the model can dynamically adjust its internal state to adapt to the characteristics of different text sequences.
[0057] In the ancient book information processing dialog box, when the user raises a new query or instruction, the system processes the input text through GRU to generate the second task instruction and the second text to be processed. At the same time, the GRU's attention mechanism will help the system identify the key parts of the user's intention, thereby generating more accurate task instructions. The second task instruction is designed to be a key signal that can activate the grammatical analysis mechanism; the grammatical analysis mechanism is a core module in the system, responsible for parsing the grammatical structure of the text sequence. When the second task instruction is received, the grammatical analysis mechanism is activated and prepares to further process the second text to be processed.
[0058] According to the text processing requirements corresponding to the second task instruction, the grammatical analysis mechanism will selectively trigger either the part-of-speech tag marking mode or the punctuation tag marking mode, wherein the part-of-speech tag marking mode: when the task needs to identify the part-of-speech information in the text (such as nouns, verbs, adjectives, etc.), the system will start the part-of-speech tag marking mode. Using the pre-trained part-of-speech tagging model or rule set, the system can assign the correct part-of-speech tag to each word in the text. Punctuation tag marking mode: when the task focuses on punctuation marks or sentence boundaries in the text, the system will trigger the punctuation tag marking mode, including identifying sentence terminators in the text (such as periods, question marks, exclamation marks, etc.), and processing the impact of other punctuation marks (such as commas, semicolons, quotation marks, etc.) on the text structure.
[0059] After selecting the appropriate tagging mode, the system processes the second text to be processed and generates the processing results, including part-of-speech tagged text and text structure with punctuation tags, depending on the task instructions. The processing results are then fed back to the user, completing the entire processing flow. Through these steps, combined with a multi-level attention mechanism, LSTM, GRU, and grammatical analysis mechanisms, the Ancient Book Information Processing Dialog System can efficiently process text sequences of various lengths and types, meeting the diverse needs of users.
[0060] Furthermore, the context-based joint guidance mechanism is used to respond to the second round of input information, the third round of input information, ..., and the Nth round of input information, and optimize and adjust in combination with the minimum difference threshold and the maximum difference threshold. The method of this application includes:
[0061] A minimum difference threshold and a maximum difference threshold are set, and the first-round output information is determined through the first-round input information; the second-round output information is determined through the first-round input information, the first-round output information, and the second-round input information using the context joint guidance mechanism; the third-round output information is determined through the first-round input information, the first-round output information, the second-round input information, the second-round output information, and the third-round input information using the context joint guidance mechanism; the context joint guidance mechanism is repeated N-1 times, and a difference evaluation is performed at the same time. When the current cumulative difference index meets the minimum difference threshold but does not meet the maximum difference threshold, the M vertical pre-trained ancient book information processing models are synchronously enhanced with the highest similarity with the pre-training stage as the learning direction.
[0062] Setting the minimum difference threshold and the maximum difference threshold is an important step to ensure that the model can be optimized gradually and smoothly during the learning process while avoiding overfitting. At the same time, the context joint guidance mechanism is used to continuously generate multiple rounds of output information, and guide the synchronous enhancement of the model based on the differences in this information. Specifically, setting the difference threshold includes the following: the minimum difference threshold: defines a minimum difference standard. When the output change of the model is less than this threshold, the model is considered to be close to convergence; the maximum difference threshold: defines a maximum difference upper limit to prevent the model from fluctuating too much during the learning process, which helps to avoid the model from falling into an unstable state or overfitting.
[0063] The first round of input information is the initial query or text provided by the user; the first round of output information is the preliminary result generated based on the first round of input information using the current state of M vertical pre-trained ancient book information processing models; the context joint guidance mechanism uses the previous input and output information, as well as the current input information, to jointly guide the generation of the next round of output information, including: combining the first round of input information, the first round of output information, and the second round of input information, and generating the second round of output information through the context joint guidance mechanism; repeating the above process, and using more context information in turn to generate the third round, fourth round... until the Nth round of output information.
[0064] After each round of output information is generated, the difference between the output information of the current round and the previous round is calculated and accumulated to form the current cumulative difference index; when the current cumulative difference index meets the minimum difference threshold (indicating that the model is learning stably) but does not reach the maximum difference threshold (to avoid over-adjustment), the model enhancement process is triggered.
[0065] The learning direction with the highest similarity to the pre-training stage is selected, and M vertical pre-trained ancient book information processing models are synchronously enhanced. The learning direction with the highest similarity evaluation is used to ensure that the model can be gradually and smoothly optimized during the learning process; based on the performance of the model on the validation set or a specific test set, or based on some measurement of the model's internal parameters or output; the above process (from the second round to the Nth round) will be repeated N-1 times to gradually optimize the output of the model; as each round of output information is generated and the difference evaluation is carried out, the model will continue to adjust in a better direction until the stopping condition is met (such as reaching the preset number of iterations, the performance improvement is no longer significant, etc.), which can gradually improve the processing ability of ancient book information while ensuring stability, and provide users with more accurate and useful output information.
[0066] Furthermore, while performing a difference assessment, the present application method also includes:
[0067] Configure the single-round difference measurement formula: Where ΔT round The difference index used to characterize a round of dialogue, t i,input Used to represent the weight of the i-th topic in the input information, p i,output Used to represent the weight of the i-th topic in the output information; configure the multi-round difference measurement formula: ΔT total It is used to represent the current cumulative difference index corresponding to multiple rounds of dialogue, where R is the number of rounds.
[0068] Through the single-round difference measurement formula: Calculate the sum of the absolute differences in the weights of all topics in the input and output information in a round of dialogue, reflecting the degree of information change in that round of dialogue; configure the multi-round difference measurement formula: Obtaining the degree of information change accumulated throughout the conversation helps to evaluate the coherence of the entire conversation process and the effectiveness of information processing.
[0069] In the dialog box for intelligent processing of ancient book information, when the user asks a question or enters information, the difference evaluation formula calculates the difference index of each round of dialogue and the cumulative difference index of the entire dialogue; if the current cumulative difference index meets the minimum difference threshold but does not reach the maximum difference threshold, it is considered that the information processing process is in a good learning state. Through difference evaluation, the intelligent processing process of ancient book information can be monitored and adjusted in real time to ensure that the system can respond to user needs accurately and coherently, and improve the quality and efficiency of information processing.
[0070] Furthermore, the present application method includes:
[0071] Connect to the user terminal device, obtain user location information, and match it with the location tag. After the match is successful, establish an interactive interface using Vue.js technology, and the interactive interface includes an output reading plug-in; request access to the audio output permission of the user terminal device. After the user accepts, the output reading plug-in provides audio guide services.
[0072] Front-end (user terminal device): Use the HTML5 Geolocation API to obtain the user's current location information (latitude and longitude); considering user privacy, when requesting location information, the user should be clearly informed and their consent should be obtained.
[0073] Backend part: receives the location information sent by the frontend; matches the location information with the preset location tag (which can be the location ID and the corresponding latitude and longitude range in the database); and returns the matching result to the frontend.
[0074] Front-end part (Vue.js): Use Vue.js framework to build user interface; receive the location matching results returned by the back-end and display relevant content accordingly; introduce output reading plug-in (such as using HTML5 <audio>tag or third-party libraries like howler.js).
[0075] Before trying to play audio, make sure you have obtained the user's audio output permission; before playing audio, request access to the audio output permission of the user's terminal device (corresponding to audio modules such as speakers), which needs to be requested when the application starts or before trying to play; considering user experience and privacy, be sure to clearly inform the user how the location information will be used, and test compatibility on various devices and browsers.
[0076] In summary, the beneficial effects of the embodiments of the present application are:
[0077] 1. Accelerate the digitization of ancient books and improve information processing speed through OCR recognition and NLP technology. Use semantic parsing and topic modeling to enhance the ability to understand and interpret the content of ancient books.
[0078] 2. Through fine-grained annotation and topic modeling, the ancient book knowledge graph is automatically constructed to facilitate knowledge query and association analysis. It is also combined with historical carriers such as cultural relics and relics to form a multimodal perception set, enriching the dimensions of ancient book information.
[0079] 3. Establish an ancient book information processing dialog box to realize natural language interaction with users, provide more personalized information services, and continuously optimize the information processing model through context-based joint guidance mechanism and difference evaluation algorithm to improve response quality and accuracy.
[0080] 4. By setting minimum and maximum difference thresholds, determining the first-round output information from the first-round input information; determining the second-round output information from the first-round input information, the first-round output information, and the second-round input information using a context-based joint guidance mechanism; determining the third-round output information from the first-round input information, the first-round output information, the second-round input information, the second-round output information, and the third-round input information using a context-based joint guidance mechanism; repeating this context-based joint guidance mechanism N-1 times while performing difference assessments, and when the current cumulative difference index meets the minimum difference threshold but does not meet the maximum difference threshold, the M vertical pre-trained ancient book information processing models are synchronously enhanced based on the learning direction with the highest similarity to the pre-training stage. This ensures that the processing capabilities of ancient book information can be gradually improved while ensuring stability, providing users with more accurate and useful output information.
[0081] Example 2
[0082] Based on the same inventive concept as the ancient book information intelligent processing method based on artificial intelligence in the aforementioned embodiment, Figure 2 As shown, the embodiment of the present application provides an intelligent processing device for ancient book information based on artificial intelligence, wherein the device includes:
[0083] The ancient book scanning module M100 is used to scan ancient books and obtain scanned text, which is in an editable text format;
[0084] The knowledge graph establishment module M200 is used to establish an ancient book knowledge graph based on the scanned text, wherein the ancient book knowledge graph sets triples based on text punctuation as constraints, and the triples are used for text editing based on segmentation and part of speech;
[0085] The database classification module M300 is used to classify M domain databases through the ancient book knowledge graph and use a machine learning model to determine M vertical pre-trained ancient book information processing models. The classification indicators corresponding to the M domain databases include age, core issues, and representative ideas of the original works;
[0086] The historical carrier introduction module M400 is used to introduce various historical carriers, including cultural relics, relics, and intangible cultural heritage;
[0087] Multimodal perception module M500 is configured to determine a first multimodal perception set based on the multi-format historical carriers and the era and core topics in the classification indicators; determine a second multimodal perception set based on the multi-format historical carriers and the representative ideas and core topics of the original works in the classification indicators; and determine a third multimodal perception set based on the multi-format historical carriers and the era and representative ideas of the original works in the classification indicators.
[0088] A multivariate information network establishing module M600 is configured to configure a multi-layer branch guidance structure based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set using a streaming generation mode, and embed the M vertical pre-trained ancient book information processing models into the multi-layer branch guidance structure to establish a multivariate information network;
[0089] The optimization and adjustment module M700 is used to establish an ancient book information processing dialog box, receive the first round of input information, perform interactive connections in the multi-information network, output the first round of response information, and respond to the second round of input information, the third round of input information,..., the Nth round of input information with a context-joint guidance mechanism, and perform optimization and adjustment in combination with the minimum difference threshold and the maximum difference threshold.
[0090] Furthermore, the database classification module M300 is used to perform the following method:
[0091] Performing fine-grained semantic analysis on the scanned text and adding fine-grained annotation information, wherein the fine-grained annotation information includes a core word and a first internal logical relationship between the word and the core topic, and a core phrase and a second internal logical relationship between the word and the core topic;
[0092] Use machine learning models for topic modeling, and continue to use M domain databases for pre-training to ensure that the pre-training data has fine-grained annotation information to support the training of the semantic parsing layer.
[0093] Furthermore, the database classification module M300 is further configured to execute the following method:
[0094] Use NLP technology to define the classical Chinese linguistics rule set and the modern Chinese linguistics rule set;
[0095] In the ancient book information processing dialog box, a first task instruction and a first to-be-processed text are generated by using the ancient Chinese linguistics rule set and the modern Chinese linguistics rule set, wherein the first task instruction activates an NLP dynamic response mechanism;
[0096] By utilizing the NLP dynamic response mechanism and comparing the text processing requirements corresponding to the first task instruction, any one of the core vocabulary extraction mode and the core phrase extraction mode is selectively triggered.
[0097] Furthermore, the database classification module M300 is further configured to execute the following method:
[0098] A multi-level attention mechanism is used, along with a long short-term memory network and a gated recurrent unit (GRU) to enhance the recognition of text sequences of varying lengths and types.
[0099] In the ancient book information processing dialog box, a second task instruction and a second to-be-processed text are generated through a gated loop unit, wherein the second task instruction activates a grammar analysis mechanism;
[0100] By utilizing the grammatical analysis mechanism and comparing with the text processing requirements corresponding to the second task instruction, any one of the part-of-speech tag marking mode and the punctuation tag marking mode is selectively triggered.
[0101] Furthermore, the optimization and adjustment module M700 is used to perform the following method:
[0102] Set the minimum difference threshold and the maximum difference threshold, and determine the first round of output information through the first round of input information;
[0103] Determine the second round output information by using the first round input information, the first round output information, and the second round input information using the context joint guidance mechanism;
[0104] Determine the third round output information by using the context-joint guidance mechanism based on the first round input information and the first round output information, the second round input information and the second round output information, and the third round input information;
[0105] The context-joint guidance mechanism is repeated N-1 times, and difference evaluation is performed simultaneously. When the current cumulative difference index meets the minimum difference threshold but does not meet the maximum difference threshold, the M vertical pre-trained ancient book information processing models are synchronously enhanced with the highest similarity with the pre-training stage as the learning direction.
[0106] Furthermore, the optimization and adjustment module M700 is further configured to execute the following method:
[0107] Configure the single-round difference measurement formula:
[0108] Where ΔT round The difference index used to characterize a round of dialogue, t i,input Used to represent the weight of the i-th topic in the input information, p i,output Used to represent the weight of the i-th topic in the output information;
[0109] Configure the multi-round difference measurement formula: ΔT total It is used to represent the current cumulative difference index corresponding to multiple rounds of dialogue, where R is the number of rounds.
[0110] Furthermore, the optimization and adjustment module M700 is further configured to execute the following method:
[0111] Connect to the user terminal device, obtain the user's location information, and match it with the location tag. After the match is successful, use Vue.js technology to build an interactive interface, which includes an output reading plug-in;
[0112] Request access to the audio output permission of the user terminal device. After the user accepts, the output reading plug-in provides audio guide service.
[0113] In summary, any step can be stored as a computer instruction or program in an unlimited computer memory and can be called and recognized by an unlimited computer processor, without any unnecessary restrictions.
[0114] Furthermore, the above technical solution only reflects the preferred technical solution of the technical solution of the embodiment of the present application. Some changes that may be made to certain parts thereof by technical personnel in this technical field all reflect the novel principles of the embodiment of the present application. Obviously, technical personnel in this field can make various changes and modifications to the present application without departing from the scope of the present application.< / audio>
Claims
1. An intelligent processing method for ancient book information based on artificial intelligence, characterized in that: The method comprises: Scanning ancient books to obtain scanned text, wherein the scanned text is in an editable text format; Based on the scanned text, an ancient book knowledge graph is established, wherein the ancient book knowledge graph sets triples based on text punctuation as constraints, and the triples are used for text editing based on segmentation and part of speech; Through the ancient book knowledge graph, M domain databases are classified, and a machine learning model is used to determine M vertical pre-trained ancient book information processing models. The classification indicators corresponding to the M domain databases include age, core topics, and representative ideas of the original works; Introducing multiple forms of historical carriers, including cultural relics, relics, and intangible cultural heritage; Based on the multi-form historical carriers, the era and core topics in the classification indicators, a first multi-modal perception set is determined; based on the multi-form historical carriers, the representative ideas of the original works and the core topics in the classification indicators, a second multi-modal perception set is determined; based on the multi-form historical carriers, the era and representative ideas of the original works in the classification indicators, a third multi-modal perception set is determined; Based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set, a streaming generation mode is adopted to configure a multi-layer branch guidance structure, and the M vertical pre-trained ancient book information processing models are embedded in the multi-layer branch guidance structure to establish a multivariate information network; Establish an ancient book information processing dialog box, receive the first round of input information, perform interactive connection in the multi-information network, output the first round of response information, and respond to the second round of input information, the third round of input information, ..., the Nth round of input information with a context-joint guidance mechanism, and perform optimization and adjustment in combination with the minimum difference threshold and the maximum difference threshold.
2. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 1, characterized in that: Using a machine learning model, M vertical pre-trained ancient book information processing models are determined, and the method includes: Performing fine-grained semantic analysis on the scanned text and adding fine-grained annotation information, wherein the fine-grained annotation information includes a core word and a first internal logical relationship between the word and the core topic, and a core phrase and a second internal logical relationship between the word and the core topic; Use machine learning models for topic modeling, and continue to use M domain databases for pre-training to ensure that the pre-training data has fine-grained annotation information to support the training of the semantic parsing layer.
3. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 2, characterized in that: Performing fine-grained semantic analysis on the scanned text and adding fine-grained annotation information, the method further includes: Use NLP technology to define the classical Chinese linguistics rule set and the modern Chinese linguistics rule set; In the ancient book information processing dialog box, a first task instruction and a first to-be-processed text are generated by using the ancient Chinese linguistics rule set and the modern Chinese linguistics rule set, wherein the first task instruction activates an NLP dynamic response mechanism; By utilizing the NLP dynamic response mechanism and comparing the text processing requirements corresponding to the first task instruction, any one of the core vocabulary extraction mode and the core phrase extraction mode is selectively triggered.
4. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 3, characterized in that: The method further comprises: A multi-level attention mechanism is used, along with a long short-term memory network and a gated recurrent unit (GRU) to enhance the recognition of text sequences of varying lengths and types. In the ancient book information processing dialog box, a second task instruction and a second to-be-processed text are generated through a gated loop unit, wherein the second task instruction activates a grammar analysis mechanism; By utilizing the grammatical analysis mechanism and comparing with the text processing requirements corresponding to the second task instruction, any one of the part-of-speech tag marking mode and the punctuation tag marking mode is selectively triggered.
5. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 4, characterized in that: Using a context-based joint guidance mechanism, in response to second-round input information, third-round input information, ..., and N-round input information, optimizing and adjusting in combination with a minimum difference threshold and a maximum difference threshold, the method includes: Set the minimum difference threshold and the maximum difference threshold, and determine the first round of output information through the first round of input information; Determine the second round output information by using the first round input information, the first round output information, and the second round input information using the context joint guidance mechanism; Determine the third round output information by using the context-joint guidance mechanism based on the first round input information and the first round output information, the second round input information and the second round output information, and the third round input information; The context-joint guidance mechanism is repeated N-1 times, and difference evaluation is performed simultaneously. When the current cumulative difference index meets the minimum difference threshold but does not meet the maximum difference threshold, the M vertical pre-trained ancient book information processing models are synchronously enhanced with the highest similarity with the pre-training stage as the learning direction.
6. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 5, characterized in that: While performing a difference assessment, the method further comprises: Configure the single-round difference measurement formula: Where, ΔT round The difference index used to characterize a round of dialogue, t i,input Used to represent the weight of the i-th topic in the input information, p i,output Used to represent the weight of the i-th topic in the output information; Configure the multi-round difference measurement formula: ΔT total It is used to represent the current cumulative difference index corresponding to multiple rounds of dialogue, where R is the number of rounds.
7. The method for intelligent processing of ancient book information based on artificial intelligence according to claim 6, characterized in that: The method comprises: Connect to the user terminal device, obtain the user's location information, and match it with the location tag. After the match is successful, establish an interactive interface using Vue.js technology, and the interactive interface includes an output reading plug-in; Request access to the audio output permission of the user terminal device. After the user accepts, the output reading plug-in provides audio guide service.
8. An intelligent processing device for ancient book information based on artificial intelligence, characterized in that: The device for implementing the ancient book information intelligent processing method based on artificial intelligence according to any one of claims 1 to 7 comprises: An ancient book scanning module is used to scan ancient books and obtain scanned text, which is in an editable text format; A knowledge graph building module is used to build an ancient book knowledge graph based on the scanned text, wherein the ancient book knowledge graph sets triples based on text punctuation as constraints, and the triples are used for text editing based on segmentation and part of speech; A database classification module is used to classify M domain databases through the ancient book knowledge graph and use a machine learning model to determine M vertical pre-trained ancient book information processing models. The classification indicators corresponding to the M domain databases include age, core topics, and representative ideas of the original works; A historical carrier introduction module is used to introduce various historical carriers, including cultural relics, relics, and intangible cultural heritage; The multimodal perception module is configured to determine a first multimodal perception set based on the multi-format historical carriers and the era and core topics in the classification indicators; determine a second multimodal perception set based on the multi-format historical carriers and the representative ideas and core topics of the original works in the classification indicators; and determine a third multimodal perception set based on the multi-format historical carriers and the era and representative ideas of the original works in the classification indicators; a multivariate information network establishment module, configured to configure a multi-layer branch guidance structure based on the first multimodal perception set, the second multimodal perception set, and the third multimodal perception set using a streaming generation mode, and embed the M vertical pre-trained ancient book information processing models into the multi-layer branch guidance structure to establish a multivariate information network; The optimization and adjustment module is used to establish an ancient book information processing dialog box, receive the first round of input information, conduct interactive connections in the multi-information network, output the first round of response information, and respond to the second round of input information, the third round of input information,..., the Nth round of input information with a context-based joint guidance mechanism, and perform optimization and adjustment in combination with the minimum difference threshold and the maximum difference threshold.
Citation Information
Patent Citations
Knowledge graph construction method of ancient book mathematical theory
CN116450845A
Method and device for retrieving infectious diseases based on infectious disease knowledge graph
CN117009408A