Intelligent questioning and answering method and system in material field based on vector database

Optimizing the material science literature search process through vector database and multi-process parallel processing technology, solving the problem of low search efficiency and accuracy in the existing system, achieving fast and accurate literature search and reply generation, which is suitable for large-scale materials science literature processing.

CN120256567APending Publication Date: 2025-07-04TIANMUSHAN LABORATORY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510319606.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing RAG system has low literature search efficiency and accuracy in the field of materials science, making it difficult to quickly locate relevant literature resources and generate text responses that meet user habits.

Method used

The intelligent question-and-answer method based on vector database is adopted to identify problem keywords through language big models, replace multi-language synonyms, sort literature lists using cross-language Reranker model, and generate problem responses through language big models, combining OCR processing and multi-process parallel processing technology to optimize literature management and retrieval processes.

Benefits of technology

It realizes efficient and accurate literature search and reply generation, meets the special needs of the field of materials science, improves the system's response speed and stability, can generate text that meets human reading habits, and supports large-scale literature processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256567A_ABST
    Figure CN120256567A_ABST
Patent Text Reader

Abstract

The invention discloses a material field intelligent question and answer method and system based on a vector database, and the method comprises the steps: recognizing and extracting a question input by a user through a language large model to obtain a question keyword, determining Chinese vocabularies in the question keyword, determining whether the Chinese vocabularies have multi-language synonyms, and if yes, determining that the Chinese vocabularies have multi-language synonyms; replacing the Chinese vocabulary with the multi-language synonym to obtain a multi-language question; retrieving in a vector database based on the multi-language question to obtain a first literature list; performing cross matching on the multi-language question and the first literature list through a cross-language Reranker model, and sorting the first literature list to obtain a second literature list; and forming a question reply based on the literature of the second literature list through the language large model, aiming at the special requirements in the field of material science, the method optimizes the question and literature processing flow, can quickly and accurately retrieve the literature required by the user, and generates the question reply conforming to the habit of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information technology, and particularly relates to an intelligent question-answering method and system in the field of materials based on a vector database. Background Art

[0002] In today's rapidly developing technological era, materials science plays an increasingly important role. The discovery and application of new materials are driving innovation in multiple fields such as industry, healthcare, and energy. However, with the in-depth research of materials, the number of related papers and literatures has increased exponentially, bringing huge challenges in information retrieval and integration to college students and researchers. Traditional literature search methods are often time-consuming and laborious, and it is difficult to effectively handle the vast amount of information.

[0003] Traditional literature search methods have low retrieval efficiency and high requirements for the comprehension ability of practitioners. The vast amount of information is prone to information overload, making it difficult to locate the most relevant literature resources. Even if relevant resources are found, it takes a large amount of time to analyze the resources and form summaries and insights. Against this background, the progress of artificial intelligence and natural language processing technologies provides new possibilities for solving this problem. In particular, the emergence of the Retrieval-Augmented Generation (RAG) technology based on large language models has brought revolutionary changes to the literature research in the field of materials science.

[0004] However, for the vast amount of literature resources in the field of materials, the current RAG systems still have problems of low retrieval efficiency and accuracy. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide an intelligent question-answering method and system in the field of materials based on a vector database in view of the deficiencies of the prior art. In view of the special needs in the field of materials science, the question and literature processing processes are optimized, and the literature required by users can be retrieved quickly and accurately, and question responses that conform to user habits can be generated.

[0006] To solve the above technical problem, this application discloses an intelligent question-answering method in the field of materials based on a vector database, including: Identifying and extracting the question entered by the user through a large language model to obtain question keywords, determining the Chinese words in the question keywords, determining whether there are multilingual synonyms for the Chinese words, and if so, replacing the Chinese words with the multilingual synonyms to obtain a multilingual question; Retrieving a first literature list in the vector database based on the multilingual question; performing cross-matching on the multilingual question and the first literature list through a cross-lingual Reranker model to sort the first literature list to obtain a second literature list; Forming a question response based on the literature in the second literature list through the large language model.

[0007] Optionally, it further includes the step of pre - extracting data from the literature resources to obtain metadata: Perform layout prediction on the literature materials through the onnxruntime inference engine of the preset OCR model; Determine the literature type of the literature materials based on the result of the layout prediction; Identify all independent paragraphs of the literature resources, and combine the independent paragraphs with an associated relationship based on the literature type to obtain the metadata.

[0008] Optionally, the pre - extracting data from the literature resources to obtain metadata includes: Split the literature resources into multiple sub - literature sets, and allocate the sub - literatures in the sub - literature sets to multiple processes for parallel processing to obtain the metadata corresponding to the sub - literatures; Integrate the metadata corresponding to the sub - literatures to obtain the metadata of all literature resources.

[0009] Optionally, the splitting the literature resources into multiple sub - literature sets includes: Screen the literatures smaller than or equal to the preset standard literature size based on the preset standard literature size and divide them into the first sub - literature; Split the literatures larger than the preset standard literature size into multiple sub - literatures according to the preset splitting rules, and use all the split sub - literatures as the second sub - literature; Obtain multiple sub - literature sets based on the number of the first sub - literature and the second sub - literature and the number of parallel processing processes.

[0010] Optionally, it further includes the step of vectorizing and storing the metadata in a vector database, including: Pre - train the preset BCEmbedding model using multi - language and multi - domain literatures based on the XLM - RoBERTa architecture; Optimize the BCEmbedding model using a large - scale multi - language corpus in the material field, perform adaptive fine - tuning for the material field, and optimize the BCEmbedding model through parallel corpus and contrast learning techniques to obtain a cross - language BCEmbedding model at the word - level, sentence - level, and document - level; Select the cross - language BCEmbedding model at the word - level, sentence - level, and document - level to vectorize the corresponding metadata.

[0011] Optionally, it further includes the step of building an index for the vectorized metadata, including: Identify the data types of independent paragraphs in the metadata, and divide the vector data of the metadata according to the data types to obtain vector sub - data corresponding to different data types; For vector sub-data with a data volume greater than a preset value, multiple vector sub-data are obtained by dividing according to the spatial structure, logical structure or general content of the vector sub-data; Based on all vector sub-data, multiple vector sub-data sets are formed, and a separate vector index is set for each vector sub-data set to form a corresponding index file.

[0012] Optionally, the forming of the question response by the language large model based on the documents in the second document list includes: Forming a question response language based on the second document list; Based on user requirements, determine the target document in the second document list, merge partial vector sub-data of the target document to form a response document, or merge the vector sub-data of the target document after removing useless information to form a question response, or merge all vector sub-data of the target document to obtain the complete target document; Form the question response based on the question response language and the target document.

[0013] This application also discloses a material field intelligent question answering system based on a vector database, including: A question preprocessing module, configured to identify and extract the question input by the user through a language large model to obtain question keywords, determine Chinese words in the question keywords, determine whether there are multilingual synonyms for the Chinese words, and if so, replace the Chinese words with the multilingual synonyms to obtain a multilingual question; A document retrieval module, configured to retrieve a first document list in the vector database based on the multilingual question; perform cross-matching on the multilingual question and the first document list through a cross-language Reranker model to sort the first document list to obtain a second document list; A question response module, configured to form a question response by the language large model based on the documents in the second document list.

[0014] Beneficial effects: This application constructs a comprehensive and integrated RAG platform, which includes a complete process from document management, OCR processing, text chunking, vectorization to retrieval and answer generation. This application develops an efficient and stable service architecture, which can process a large number of materials science documents while ensuring the response speed and stability of the system. Particularly worth mentioning is that this application optimizes the document processing process for the special needs of the materials science field, enabling the system to accurately identify and generate text that conforms to human reading habits, laying a solid foundation for subsequent retrieval and analysis. Description of the Drawings

[0015] The following further specifically describes the present application in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present application will become clearer.

[0016] Figure 1 It is a flowchart of an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 2 It is a flowchart of extracting metadata obtained from an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 3 and Figure 4 It is a schematic diagram of a specific example of metadata extraction in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 5 It is a flowchart of multi-process processing in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 6 It is a system architecture diagram of multi-process processing in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 7 It is a flowchart of splitting literature resources into multiple sub-literature collections in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figures 8 to 15 It is a schematic diagram of the table structure in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 16 It is a flowchart of vectorizing metadata in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 17 It is an architecture diagram of the BCEmbedding model in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 18 It is a flowchart of building an index for metadata in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 19 It is a schematic diagram of building an index for metadata in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 20 It is a flowchart of S300 in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 21 It is an overall inference architecture diagram of the large model in an embodiment of the intelligent question-answering method in the material field based on a vector database of the present application; Figure 22 It is a structure diagram of an embodiment of the intelligent question-answering system in the material field based on a vector database of the present application; Figure 23A structural schematic diagram of a computer device suitable for implementing the embodiments of the present invention is shown. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0019] The technical solutions provided by the present application are mainly implemented using large model technologies. Here, a large model refers to a deep learning model with a large number of model parameters, which usually can include hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a foundation model. Through large-scale pre-training of the large model with unlabeled corpora, a pre-trained model with more than hundreds of millions of parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0020] It should be noted that in actual applications, the large model can be fine-tuned with a small number of samples for the pre-trained model, enabling the large model to be applied to different tasks. For example, the large model can be widely applied in fields such as Natural Language Processing (NLP), computer vision, and speech processing. Specifically, it can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), and image generation. It can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0021] It should be noted that without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the accompanying drawings and in combination with the embodiments.

[0022] It should be noted that in one or more embodiments of this application, the Retrieval-Augmented Generation (RAG) system refers to an artificial intelligence mechanism that combines retrieval and generation, used to improve the performance of language models in information retrieval and knowledge-intensive tasks. First, the retrieval component retrieves information or document fragments related to the input query from a large knowledge base or dataset, and then these retrieved contents are used to assist the generation component in generating a more accurate and information-rich output.

[0023] The vector database is mainly used to store and process vector data, that is, various types of data such as text, images, and audio are converted into vector form through specific algorithms for storage and management, facilitating efficient similarity search and analysis.

[0024] To solve the problems existing in the prior art, according to one aspect of this application, this embodiment discloses an intelligent question-answering method in the field of materials based on a vector database. As Figure 1 shown, in this embodiment, the method includes: S100: Identify and extract the keywords of the question input by the user through a large language model, determine the Chinese words in the question keywords, determine whether there are multilingual synonyms for the Chinese words, and if so, replace the Chinese words with the multilingual synonyms to obtain a multilingual question. S200: Retrieve a first list of documents from the vector database based on the multilingual question; rank the first list of documents by cross - matching the multilingual question and the first list of documents through a cross - language Reranker model to obtain a second list of documents. S300: Generate a question response based on the documents in the second list of documents through the language large model.

[0025] This application constructs a comprehensive and integrated RAG platform, which includes a complete process from document management, OCR processing, text chunking, vectorization to retrieval and answer generation. This application develops an efficient and stable service architecture that can process a large number of materials science documents while ensuring the response speed and stability of the system. Notably, this application optimizes the document processing flow according to the special needs of the materials science field, enabling the system to accurately identify and generate text that conforms to human reading habits, laying a solid foundation for subsequent retrieval and analysis.

[0026] In an alternative embodiment, as Figure 2 shown, the method further includes the step of pre - extracting metadata from the document resources: S010: Perform layout prediction on the document materials through the onnxruntime inference engine of a preset OCR model; S020: Determine the document type of the document materials based on the result of the layout prediction; S030: Identify all independent paragraphs of the document resources, and combine the independent paragraphs with an associated relationship based on the document type to obtain the metadata.

[0027] When storing document materials, it is necessary to support text content extraction for more than 160,000 PDF papers. To process such a large number of PDF documents, it is first necessary to develop efficient parallel processing techniques, otherwise it will not be able to meet the requirements of document processing duration well. Moreover, the diversity of the PDF format is also an important issue, and different paper layouts, fonts, etc. need to be well - processed; in addition, a robust error - handling mechanism is required to handle any possible damaged or abnormal documents.

[0028] To solve these problems, this application introduces the onnxruntime inference engine into the OCR model for text recognition of literature materials in formats such as PDF and pictures to predict the layout of the literature materials, thereby clarifying the type of the literature materials to prepare for efficient subsequent data storage. This application uses relevant interfaces such as InferenceSession provided by the python package to transmit the ONNX format text recognized by the OCR model to onnxruntime for layout prediction. Among them, onnxruntime is a powerful and easy-to-use model inference engine with characteristics such as high performance, cross-platform, and support for multiple models, which can better meet the inference needs of the OCR model and improve PDF text recognition.

[0029] In a specific example, the training process of the OCR model includes: S1: Collect a large number of PDF documents with different layouts to ensure data diversity. Annotate elements such as text, images, and tables in the PDF documents, and the annotation content includes information such as text position, font, size, and color.

[0030] S2: Before annotation, perform data cleaning to remove noise data, such as irrelevant headers, footers, watermarks, etc., and data augmentation: enhance the data through operations such as rotation, scaling, and translation to improve the generalization ability of the model.

[0031] S3: Build a neural network model based on big data technology and initialize the model parameters, and set the loss function.

[0032] S4: Divide the data into small batches for training, and use the gradient descent method to optimize the model parameters.

[0033] Use a learning rate scheduler to dynamically adjust the learning rate to improve the training effect.

[0034] Use regularization techniques (such as L2 regularization, Dropout, etc.) to prevent overfitting.

[0035] S5: Evaluate the model performance on the validation set and adjust the hyperparameters (such as learning rate, batch size, etc.), where the early stopping method is used to prevent overfitting, and stop training when the performance of the validation set no longer improves.

[0036] S6: Evaluate the model performance on the test set, and use metrics (such as accuracy, recall rate, F1 score, etc.) for evaluation to analyze the performance of the model on different layouts.

[0037] S7: Export the trained model to the deployable ONNX format, call the onnxruntime inference engine to improve the model inference speed, and predict the layout, literature type, and keywords of the literature materials through the onnxruntime inference engine.

[0038] S8: Deploy the model to the production environment, provide an API interface or integrate it into the application.

[0039] In an optional implementation, the OCR model recognizes the text of the literature materials and predicts the layout of the literature materials to obtain the literature type and keywords. After determining the literature type, the text of the literature materials is split according to the literature storage strategy corresponding to the literature type. The split literature materials and the corresponding keywords are vectorized and a retrieval index is constructed.

[0040] In the process of the OCR model of the present application recognizing the literature materials, the text content in Markdown format will ultimately be obtained, retaining the hierarchical structure and semantic integrity of the overall text of the paper. The specific process includes: Paragraph division: recognizing all independent paragraphs of the paper text; Title recognition: recognizing various levels of titles in the paper text; Hierarchical structure recognition: inferring the hierarchical structure directly between paragraphs through paragraphs and titles. In a specific example, Figure 3 shows a part of a PDF format paper. Through OCR recognition, each independent paragraph is recognized. The position and content of the independent paragraphs are used to infer the literature type, and then the relationship between paragraphs is inferred based on the literature type. The paragraphs are integrated to obtain the text content in Markdown format. Finally, the metadata of the output literature is as Figure 4 shown, all the text is intelligently recognized and the relationship between each paragraph and between paragraphs is inferred. Paragraphs with interrelated relationships are integrated into a complete paragraph, and finally a metadata literature with strong readability and logic is obtained.

[0041] In an optional implementation, as Figure 5 shown, the pre-data extraction of the literature resources to obtain metadata includes: S040: Split the literature resources into multiple sub-literature sets, and allocate the sub-literatures in the sub-literature sets to multiple processes for parallel processing to obtain the metadata corresponding to the sub-literatures; S050: Integrate the metadata corresponding to the sub-literatures to obtain the metadata of all literature resources.

[0042] Since the number of literature resources to be recognized is large, the single-process architecture cannot well meet the requirements of literature resource processing. Specifically, the disadvantages of the single-process architecture are: Low efficiency: The processing mode of the single process can only process one PDF literature at a time. When a large number of literatures need to be processed, the efficiency is low and the time cost is significantly high; Low resource utilization: Hardware resources such as the CPU cannot be fully utilized, resulting in a large degree of resource waste.

[0043] To overcome the deficiencies of a single process, this application uses a multi-process architecture to perform parallel processing on literature resources, which can significantly improve the efficiency of literature processing. Among them, parallel processing means that multiple processes can work simultaneously, and each process uses an OCR model to process the literature resources respectively. The processes do not affect each other, which can improve the overall processing efficiency.

[0044] In an alternative embodiment, the literature resources to be processed (e.g., PDF documents) are split into multiple subsets of documents through a hashing algorithm and distributed to multiple processes. Each process only needs to be responsible for processing the subset of documents assigned to it. Taking a PDF document as an example, the specific distribution method is described in detail as follows: When reading the documents, number them sequentially according to the reading order. Assuming there are 2000 documents in total, the document numbers are 0, 1, 2... 1999 in sequence. Assuming the total number of processes is 10, the process numbers are 0, 1,... 9 respectively. By taking the modulus of the document number with the total number of processes, the documents can be assigned to the corresponding processes. Specifically, for a document numbered 68, taking the modulus with the total number of processes, the result is 68 % 10 = 8, which means this document will be assigned to the process numbered 8, and so on.

[0045] Therefore, using a multi-process architecture can, on the one hand, improve efficiency, make full use of hardware resources such as the CPU, and greatly improve the PDF OCR processing speed. On the other hand, it can enhance stability. Even if a certain process fails, it does not affect the operation of other processes, ensuring the stability of the overall processing. At the same time, it is also easy to expand. As the hardware performance improves, the number of processes can be simply increased to further improve the processing capacity. Its overall architecture is as Figure 6 shown.

[0046] In an alternative embodiment, as Figure 7 shown, the splitting of the literature resources by S040 into multiple subsets of documents includes: S041: Screening the documents smaller than or equal to the preset standard document size based on the preset standard document size and dividing them into the first subset of documents.

[0047] S042: Splitting the documents larger than the preset standard document size into multiple sub-documents according to the preset splitting rules, and taking all the split sub-documents as the second subset of documents.

[0048] S043: Obtaining multiple subsets of documents based on the quantities of the first subset of documents and the second subset of documents and the number of parallel processing processes.

[0049] However, the above parallel document splitting scheme is more applicable when the size of the documents is relatively small. For cases where there is a large gap in the size of the documents included in the document resources, especially when the size of some documents is very large, splitting all the document resources into subsets for different processes only based on the number of documents in the document resources may lead to an unbalanced resource allocation problem due to the very large size of the documents processed by a certain process.

[0050] In order to further improve the processing efficiency of document resources, in another alternative implementation, a preset standard document size is set. Documents smaller than or equal to the preset standard document size are screened out and classified into the first subset of documents; documents larger than the preset standard document size are split into multiple subsets of documents according to the preset splitting rules, and all the subsets of documents obtained by splitting are used as the second subset of documents.

[0051] Based on the number of the first subset of documents and the second subset of documents and the number of parallel processing processes, multiple subsets of documents are obtained. That is, after summing the first subset of documents and the second subset of documents, they are evenly distributed to all parallel processing processes to obtain the subset of documents corresponding to each parallel processing process. This can be achieved through the following process: First, calculate the total number of the first subset of documents and the second subset of documents. Assume that the number of the first subset of documents is N1 and the number of the second subset of documents is N2, then the total number of documents is: Ntotal = N1 + N2 According to the number of parallel processing processes P, the total number of documents is evenly distributed to each process. The number of documents processed by each process is: Nper_process = P * Ntotal If Ntotal cannot be divided evenly by P, some processes may process one more document.

[0052] The first subset of documents and the second subset of documents are combined into a list, and then distributed according to the calculated number of documents processed by each process. Assume that the list of subsets of documents is files, then the subset of documents processed by each process i (from 0 to P - 1) is: filesi = files[i * Nper_process : (i + 1) * Nper_process] filesi = files[i * Nper_process : (i + 1) * Nper_process] For the last process, if Ntotal cannot be divided evenly by P, then it processes all the remaining documents.

[0053] If Ntotal < P * Ntotal < P, some processes may not be assigned any documents. In this case, the number of processes can be adjusted or the documents can be redistributed.

[0054] Optionally, after determining the subset of documents corresponding to each parallel processing process, calculate the total size of all documents in the subset of documents corresponding to each process, and adjust the documents in the subset of documents corresponding to each process according to the total size until the difference in the total size of documents between different subsets of documents is no greater than the standard document size.

[0055] In an alternative embodiment, splitting a document larger than a preset standard document size into multiple sub-documents according to a preset splitting rule can be achieved through the following steps: Determine all independent regions of the document based on the document type, and determine whether the data size of each independent region is greater than the standard document size; If so, split the content of the independent region according to the standard document size to obtain multiple second sub-documents, and set a first-level label associated with the document, a second-level label associated with the content of the independent region, and a third-level label corresponding to the divided sub-documents for each split second sub-document; If not, merge or separately form a second sub-document according to the document size of the independent region, and set a first-level label associated with the document and a second-level label associated with the content of the independent region for the second sub-document.

[0056] It should be noted that if the data size of each independent region is not greater than the standard document size, the data size of the independent region may also be very small and not suitable as an independent second sub-document.

[0057] Therefore, when the data size of an independent region is not greater than the standard document size, different independent regions can be merged so that the merged data size is not greater than the standard document size. By merging overly small independent regions, the number of second sub-documents can be reduced, simplifying the process of document allocation by parallel processing processes.

[0058] Optionally, after extracting the document type and metadata of the documents in different subsets of documents through multiple processes, the documents can be stored in a relational database to facilitate the subsequent construction of a vector database based on the stored documents. A large number of documents can be better managed, and metadata can be efficiently managed at the same time. In the relational database, the unique identifier of the document, such as the document name or document number, can be stored in association with the document.

[0059] In an alternative embodiment, for the sub-documents after process handling, multiple sub-documents belonging to the same document can be restored according to the label levels of the sub-documents. For example, for sub-documents with three-level labels, all the sub-documents corresponding to the three-level labels are merged in the order of the three-level labels. After merging the three-level labels, all the sub-documents with two-level labels are obtained. After merging the sub-documents with two-level labels in the order of the two-level labels, a complete document of the document materials corresponding to the first-level label can be obtained. After merging the keywords of all the sub-documents, they are stored in association with the complete document.

[0060] In a specific example, the system uses MySQL as the storage engine for metadata and session data. The advantages of MySQL are as follows: (1) Mature and stable: MySQL is one of the most widely used relational databases in the industry, with a mature community ecosystem and rich development tools. Its stability and reliability have been widely verified. (2) High performance: MySQL provides multiple storage engines and optimized query mechanisms, capable of efficiently processing large-scale data access and meeting the performance requirements of the system. (3) Easy to expand: MySQL supports a horizontally scalable architecture, and the processing capacity can be improved by adding server nodes to meet the growing data storage and processing needs of the system.

[0061] Specifically, this application designs 8 database tables for storing system-related data, including: knowledge base, document, document content block, session, conversation, user feedback, etc. The detailed table structures and table creation statements of the database tables will be listed in detail below.

[0062] 1) Knowledge base, used to store all the knowledge bases in the current system. The table structure is as Figure 8 shown.

[0063] The corresponding table creation statement: CREATE TABLE `knowledgebase` ( `id` bigint(20) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing id', `user_id` bigint(20) unsigned NOT NULL COMMENT 'User id, corresponding to the id in the user table', `kb_name` varchar(255) NOT NULL COMMENT 'Knowledge base name', `deleted` tinyint(1) NOT NULL DEFAULT '0' COMMENT 'Whether deleted, 0 not deleted, 1 deleted', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`) ) ENGINE=InnoDB AUTO_INCREMENT=3 DEFAULT CHARSET=utf8mb4 COMMENT='Knowledge base table' 2) Literature, used to store the literature list in the knowledge base, the table structure is as Figure 9 shown.

[0064] The corresponding table creation statement: CREATE TABLE `file` ( `id` bigint(20) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing id', `kb_id` bigint(20) unsigned NOT NULL COMMENT 'Knowledge base id, corresponding to the id in the knowledgebase table', `file_name` varchar(255) NOT NULL COMMENT 'Literature name', `file_name_new` text, `file_status` varchar(255) NOT NULL DEFAULT 'gray' COMMENT 'Literature status, red, gray, green', `deleted` tinyint(1) NOT NULL DEFAULT '0' COMMENT 'Whether deleted, 0 not deleted, 1 deleted', `file_size` bigint(20) unsigned NOT NULL DEFAULT '0' COMMENT 'Literature size', `content_length` bigint(20) unsigned NOT NULL DEFAULT '0' COMMENT 'Length of the literature content', `chunk_size` bigint(20) unsigned NOT NULL DEFAULT '0' COMMENT 'Size of the literature chunk', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`), KEY `kb_id` (`kb_id`), KEY `file_name` (`file_name`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Literature table' 3) Literature content chunk, used to store the content chunks of the literature, the table structure is as Figure 10 shown.

[0065] The corresponding table creation statement: CREATE TABLE `document` ( `id` bigint(20) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing id', `kb_id` bigint(20) unsigned NOT NULL COMMENT 'Knowledge base id, corresponding to the id in the knowledgebase table', `file_id` bigint(20) unsigned NOT NULL COMMENT 'Literature id, corresponding to the id in the file table', `file_name` varchar(255) NOT NULL COMMENT 'Literature name', `chunk_id` bigint(20) unsigned NOT NULL COMMENT 'Document block ID. For the same document, the IDs of different blocks are incremented.', `faiss_docstore_id` varchar(64) NOT NULL COMMENT 'faiss_docstore_id', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`), KEY `kb_id` (`kb_id`), KEY `file_id` (`file_id`), KEY `file_name` (`file_name`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Table for document chunks' 4) Session, used to store session history. The table structure is as follows Figure 11 shown.

[0066] Table creation statement: CREATE TABLE `session` ( `id` int(10) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing ID', `deleted` tinyint(1) NOT NULL DEFAULT '0' COMMENT 'Whether deleted', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Session table' 5) Conversation, used to store the specific conversation content in the session. The table structure is as follows Figure 12 shown.

[0067] Table creation statement: CREATE TABLE `conversation` ( `id` int(10) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing ID', `session_id` bigint(20) NOT NULL DEFAULT '0' COMMENT 'Session ID', `content` varchar(5000) NOT NULL DEFAULT '' COMMENT 'Conversation content', `role` varchar(20) NOT NULL DEFAULT '' COMMENT 'Conversation role', `extra` text, `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`), KEY `idx_session_id` (`session_id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Conversation content table' 6) User, stores user information, and the table structure is as follows Figure 13 shown.

[0068] Table creation statement: CREATE TABLE `user` ( `id` bigint(20) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing id', `user_name` varchar(255) NOT NULL COMMENT 'User name', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='User table' 7) User feedback, stores negative feedback information of users, and the table structure is as follows Figure 14 shown.

[0069] Table creation statement: CREATE TABLE `feedback` ( `id` int(10) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing ID', `conversation_id` bigint(20) NOT NULL DEFAULT '0' COMMENT 'Conversation content ID', `like` tinyint(4) NOT NULL DEFAULT '0' COMMENT 'Like', `dislike` tinyint(4) NOT NULL DEFAULT '0' COMMENT 'Dislike', `comment` varchar(200) NOT NULL DEFAULT '' COMMENT 'Remarks', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`), KEY `idx_conversation_id` (`conversation_id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Conversation content feedback form' 8) Query log. Users store query logs for easy tracing of various technical indicators. The table structure is as Figure 15 shown.

[0070] Table creation statement: CREATE TABLE `user_query_log` ( `id` int(10) unsigned NOT NULL AUTO_INCREMENT COMMENT 'Auto-incrementing ID', `session_id` bigint(20) NOT NULL DEFAULT '0' COMMENT 'Session ID', `conversation_id` bigint(20) NOT NULL DEFAULT '0' COMMENT 'Conversation ID', `query` varchar(200) NOT NULL DEFAULT '', `logid` varchar(100) NOT NULL DEFAULT '' COMMENT 'logid', `create_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP COMMENT 'Creation time', `update_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP COMMENT 'Update time', PRIMARY KEY (`id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Request log'.

[0071] In an alternative embodiment, as Figure 16 shown, the method further includes the step of vectorizing and storing the metadata in a vector database, including: S061: Pre-train a preset BCEmbedding model using multilingual and multi-domain literature based on the XLM-RoBERTa architecture; S062: Optimize the BCEmbedding model using a large-scale multilingual corpus in the materials field, perform adaptive fine-tuning for the materials field, and optimize the BCEmbedding model through parallel corpus and contrastive learning techniques to obtain a cross-lingual BCEmbedding model at the word level, sentence level, and document level; S063: Select a cross-lingual BCEmbedding model at the word level, sentence level, and document level to vectorize the corresponding metadata.

[0072] In a RAG system based on a vector library, the vectorization process of literature is particularly important because the vectorization effect will directly affect the literature retrieval effect. Generally, the evaluation criteria for a good retriever are: 1. Recall as many relevant text fragments as possible required by the user's question; 2. The more relevant and helpful fragments for answering the question should be in the front positions; 3. Filter out low-quality text fragments.

[0073] This application selects the cross - language BCEmbedding model to vectorize metadata, realizes the alignment of BCEmbedding spaces between different languages by optimizing the BCEmbedding model, and supports cross - language retrieval. At the same time, it generates BCEmbeddings at the word - level, sentence - level, and document - level, supporting semantic matching at different granularities. Specifically, this application pre - trains the BCEmbedding model using multi - language and multi - domain literature based on the XLM - RoBERTa architecture, supporting major languages such as English and Chinese. And it uses a large - scale multi - language corpus in the field of materials to optimize the BCEmbedding model, conducts adaptive fine - tuning for the field of materials, and improves the model's ability to understand professional terms. The accuracy of the BCEmbedding model's vectorization is further improved through parallel corpus and contrastive learning techniques. Figure 17 Shows the architecture of a dual - encoder of a BCEmbedding model.

[0074] In a specific example, documents, research reports, academic papers, patent texts, etc. in multiple languages (such as English, Chinese, Japanese, German, etc.) in the field of materials science are collected as a corpus. For example, multi - language literature on "nanomaterials" is downloaded from well - known academic databases, covering aspects such as the preparation methods, performance characteristics, and application fields of materials.

[0075] Taking the multi - language literature of nanomaterials as an example, use these corpora to fine - tune the BCEmbedding model: Perform pre - processing tasks such as cleaning, tokenization, and part - of - speech tagging on the collected multi - language texts. For example, for Chinese texts, use a Chinese tokenization tool (such as Jieba) to split sentences into words; for English texts, use tools such as NLTK for tokenization and part - of - speech tagging.

[0076] Input the pre - processed multi - language corpus into the BCEmbedding model, and set appropriate training parameters (such as learning rate, number of training epochs, etc.) for fine - tuning training. During the training process, the model will learn the professional vocabulary and language expressions in the field of materials, thereby making adaptive adjustments to the model. For example, during the training process, the model will gradually learn the representations of professional vocabulary such as "nanoparticle" and "graphene" in different languages.

[0077] In the same language, randomly replace some words in a sentence, and then let the model judge whether the sentences before and after the replacement are similar; in different languages, compare the embedding representations of words in parallel corpora to enable the model to learn more discriminative word embeddings. For example, in English text, replace "ceramic" with "plastic", let the model judge the change in the semantics of the sentence, and compare the embedding representations of the English "ceramic" and the Chinese "陶瓷".

[0078] Randomly select sentences in the same language or different languages ​​as positive and negative examples, and let the model learn to distinguish between positive examples (sentences with similar semantics) and negative examples (sentences with different semantics). For example, select some sentences from multilingual literature in the field of materials, use sentences describing similar material properties as positive examples, and sentences describing different material properties as negative examples, so that the model can learn the semantic relationship between sentences and optimize the sentence-level cross-language BCEmbedding model.

[0079] Different language documents on the same topic are used as positive examples, and documents on different topics are used as negative examples, so that the model can learn the semantic similarity between documents. For example, multilingual documents on "corrosion protection of metal materials" are used as positive examples, and documents on "synthesis of polymer materials" are used as negative examples, so that the model can understand the semantic relationship between documents in different languages ​​at the topic level, thereby optimizing the document-level cross-language BCEmbedding model.

[0080] In an optional embodiment, if Figure 18 As shown, the method further includes the step of building an index for the vectorized metadata, including: S071: Identify the data type of the independent paragraph in the metadata, and divide the vector data of the metadata according to the data type to obtain vector sub-data corresponding to different data types.

[0081] S072: For vector quantum data whose data volume is greater than a preset value, divide the vector quantum data into multiple vector quantum data according to the spatial structure, logical structure or content outline of the vector quantum data.

[0082] S073: forming a plurality of vector sub-data sets based on all the vector sub-data, and setting a separate vector index for each vector sub-data set to form a corresponding index file.

[0083] In an alternative embodiment, Faiss is selected as the vector database. Faiss is a library for efficient similarity search and dense vector clustering developed by Meta's fundamental artificial intelligence research group. It implements algorithms for fast search in vector sets of any size. Faiss is written in C++ and has a complete Python / numpy wrapper. It also supports running on GPUs. The vector database can use the HNSW algorithm for literature retrieval.

[0084] Since the final vector data in the vector database is close to 6 million, serial processing of retrieval will cause serious performance problems. Specifically: (1) Slow writing speed: When the number of vectors is huge, writing new data into a single index file becomes very time-consuming. This is because Faiss needs to reorganize and balance the entire index structure; (2) Slow loading speed: A large single index file takes longer to be fully loaded into memory, which will greatly extend the system startup time; (3) High memory pressure: A huge index file may consume a large amount of memory, affecting the overall performance of the system; (4) Difficult to update: Locally updating or modifying a large single index is usually difficult and time-consuming; (5) Poor scalability: As the data volume increases, the performance of a single index drops sharply.

[0085] Therefore, this application designs a format of multiple index files. By dividing a single large index file into multiple smaller index files, obvious advantages can be brought: (1) Parallel processing: Multiple smaller index files can be loaded in parallel, significantly reducing the overall loading time; (2) More flexible memory management: Only part of the index can be loaded as needed, thus making more efficient use of memory; (3) Faster writing speed: Adding new data to smaller index files is usually faster than updating a large index; (4) Easier to update and maintain: Individual indexes can be updated or rebuilt independently without affecting other indexes; (5) Better scalability: As the data volume increases, new index files can be easily added. As Figure 19 shown, the specific implementation process is roughly as follows: (1) Data segmentation: According to the division rules, the whole set of data is divided into multiple subsets; (2) Create multiple indexes: Create a separate Faiss for each data subset; (3) Parallel processing: Implement a parallel loading mechanism to load multiple index files simultaneously; (4) Data merging: Finally, merge multiple index files into one instance for retrieval.

[0086] When splitting data, the entire set of data is divided into multiple subsets, that is, all vector data is divided into subsets of vector data with different indices. In a specific example, the division rule during data splitting can be based on the data type of the vector data. After obtaining the metadata of the literature through the OCR model for literature type and data extraction from the literature resources, independent paragraphs with associations in the metadata are merged into one independent paragraph, and then the data type of the independent paragraph can be identified. For example, for journal literature, the metadata includes data types such as author information, keywords, abstracts, and the main text. The vector data of the entire journal can be divided according to the data type to obtain vector sub-data corresponding to different data types. Furthermore, for huge vector sub-data with a data volume greater than a preset value, it can also be divided according to the spatial structure, logical structure, or general content of the vector sub-data. For example, for vector sub-data of the main text data type, it can also be determined whether there are multiple independent paragraphs in the main text with the same chapter label. If so, the vector sub-data of the same chapter is further divided according to the chapter label to obtain new vector sub-data; or, it can also be determined the general content of each independent paragraph in the main text. If the general content described by multiple independent paragraphs is similar, the multiple independent paragraphs are divided into the same vector sub-data; furthermore, multiple independent paragraphs with continuous spatial structures can also be divided into one vector sub-data according to the data volume of each independent paragraph. Thus, a large literature can be divided into multiple vector sub-data through at least one of the above division rules, and multiple sets of vector sub-data are formed, and a separate vector index file is set for each set of vector sub-data.

[0087] In an alternative embodiment, as Figure 20 shown, the S300 forms a question response based on the literature in the second literature list through the language large model, including: S310: Form a question response language based on the second literature list.

[0088] S320: Based on the user's needs, determine the target literature in the second literature list, merge partial vector sub-data of the target literature to form a response literature, or merge the vector sub-data of the target literature after removing useless information to form a question response, or merge all vector sub-data of the target literature to obtain the complete target literature.

[0089] S330: Form the question response based on the question response language and the target literature.

[0090] Preferably, all vector sub-data obtained by dividing a large literature can be configured in different sets of vector sub-data. Then, when retrieving and loading the literature vector sub-data, different sets of vector sub-data can be processed by different processes, and the retrieval efficiency of vector data can be improved through parallel processing processes.

[0091] After the vector data of large-scale documents is partitioned to obtain vector sub-data in this application, an index associating the vector sub-data with the original document is set. Then, after the vector sub-data of the target document for answering the question is retrieved, all the vector sub-data can be located through the index, and the original target document can be obtained by merging some or all of the vector sub-data corresponding to the target document.

[0092] In this process, in order to improve the flexibility of users to view retrieval results, based on the target document required by the user, only some of the vector sub-data of the target document answering the user's question can be merged and then a question reply can be formed through a language large model. Or, based on the user's needs, the vector sub-data of the target document with useless information removed can be merged and then a question reply can be formed through a language large model. Or, based on the user's needs, all the vector sub-data of the target document can be merged to obtain the complete target document, and then a question reply can be formed through a language large model. This application can select different ways to view the target document based on the user's needs, improving the loading efficiency and resource utilization rate.

[0093] Optionally, when data correction or update of large-scale documents is required, the corresponding vector sub-data can be directly updated. Correcting or updating a single vector sub-data is much easier than correcting or updating a large-scale document.

[0094] In an optional implementation manner, the system for implementing the intelligent question-answering method in the material field further includes an interaction terminal. The interface for the interaction terminal to interact with the user realizes the purpose of the user retrieving relevant information through a dialogue method. At the same time, the source of the document is provided for each retrieval result, and its relevance to the query is explained. In order to collect user feedback, a negative feedback function is developed on the interface, and the information collected will be used for subsequent system optimization. In order to manage the session more conveniently, the system supports functions such as session deletion and aggregation by time.

[0095] In an optional implementation manner, for the question input by the user, the following processing is performed: The question keywords are identified and extracted from the user input through a language large model, the Chinese words in the question keywords are determined, and whether there are multilingual synonyms for the Chinese words is determined. If so, the Chinese words are replaced with the multilingual synonyms to obtain a multilingual question. Based on the multilingual question, a first document list is retrieved from the vector database; the first document list is sorted through cross-matching the multilingual question and the first document list by a cross-lingual Reranker model to obtain a second document list. A question reply is formed through a language large model based on the documents in the second document list.

[0096] Among them, it should be noted that the cross - language Rerank model is used for the second ranking in the two - stage retrieval of retrieving responses. Using the cross - attention model of the cross - language Rerank model allows for in - depth interaction between the question and the literature, capturing more fine - grained semantic matching information. Different from the first - stage language large - model's independent document scoring and ranking to obtain the first literature list, the Rerank model considers the entire candidate list and captures the relative importance among documents.

[0097] In a specific example, to fully guarantee the Q&A effect, this application selects the 72B large - model. Qwen - 72B is trained based on 3T tokens of high - quality data and can generate accurate, coherent, and insightful answers and summaries based on the retrieved information. It can not only explain complex materials science concepts but also integrate information from multiple sources to provide comprehensive knowledge support for users.

[0098] To ensure the inference effect, in terms of model deployment, the vLLM framework is adopted for the inference service of the language large - model and combined with FastChat to finally form a complete LLM interface service. vLLM is a highly optimized open - source LLM service framework that can effectively improve the service throughput on the GPU. Its features are as follows: (1) Efficient memory management: Through the PagedAttention algorithm, vLLM realizes efficient management of the KV cache, reduces memory waste, and optimizes the running efficiency of the model. (2) High throughput: vLLM supports asynchronous processing and continuous batch processing requests, significantly improving the throughput of model inference and accelerating the text generation and processing speed. (3) Ease of use: vLLM is seamlessly integrated with HuggingFace models, supports a variety of popular large - language models, simplifies the process of model deployment and inference, and is compatible with the OpenAI API server. (4) Distributed inference: The framework supports distributed inference in a multi - GPU environment. Through model parallel strategies and efficient data communication, it enhances the ability to process large models.

[0099] FastChat is an open platform for training, serving, and evaluating chatbots based on large language models. It provides state-of-the-art model training and evaluation code, such as Vicuna, MT-Bench, etc. With FastChat, it is easy to train efficient and accurate chatbots to meet the needs in various scenarios. In practical applications, the FastChat+LLM solution can be adopted to achieve the local deployment of large language models. At the same time, the training and evaluation code provided by FastChat can be used in combination with the natural language processing capabilities of LLM to build an efficient and accurate chatbot. Additionally, the robot can be customized and optimized according to actual needs to meet the requirements in different scenarios. Moreover, FastChat also provides a Web UI and an OpenAI-compatible RESTful API, facilitating the deployment and management of a distributed multi-model service system. Multiple chatbots can be integrated into one system to provide users with more comprehensive and efficient services.

[0100] The final overall inference architecture is as Figure 21 shown. Specifically, to start the large model inference service, the following 3 steps are required: (1) Start the controller: python -m fastchat.serve.controller (2) Start the vLLM inference service: python -m fastchat.serve.vllm_worker --model-path Qwen / Qwen-72B-Chat --trust-remote-code --tensor-parallel-size 2 --gpu-memory-utilization 0.98 --dtype bfloat16 (3) Start the api service: python -m fastchat.serve.openai_api_server --host localhost --port 8000 In an alternative implementation, considering the particularity of the materials field, a set of complex structured language model prompts for the materials research field is adopted in the retrieval large model. Through experimental verification, these prompts can well guide the large model to provide answers in the materials field, better stimulating its potential capabilities in the materials field. Specifically, by structuring the relevant background information to follow specific patterns and rules, it becomes more convenient and effective for the large model to understand the information.

[0101] For the rag part, the prompts are as follows: # Role: The large model IDM Alpha, especially well - versed in various fields of materials science and materials engineering, especially in the field of metallic materials. It is familiar with knowledge in various materials fields and has a lot of basic scientific research knowledge, including but not limited to materials science, metallic materials, materials engineering, etc. It is also familiar with many basic knowledge, such as density functional theory, fundamentals of materials science, fundamentals of materials engineering, mechanics of metallic materials, materials physics, solid state physics, statistical physics, etc. It can also answer the confusion of users in study, life, and emotion.

[0102] ## Profile: - author: You are the large model IDM Alpha, especially well - versed in various fields of materials science and materials engineering, especially in the field of metallic materials. It is familiar with knowledge in various materials fields and has a lot of basic scientific research knowledge, including but not limited to materials science, metallic materials, materials engineering, etc. It is also familiar with many basic knowledge, such as density functional theory, fundamentals of materials science, fundamentals of materials engineering, mechanics of metallic materials, materials physics, solid state physics, statistical physics, etc. It can also answer the confusion of users in study, life, and emotion - description: The large model IDM Alpha, especially well - versed in various fields of materials science and materials engineering, and can also answer the confusion of users in study, life, and emotion.

[0103] ## Goals: Combined with the user's question and reference information, select the most relevant ones from the provided reference information to answer the user's question and provide accurate knowledge in the field of materials.

[0104] ## Constrains: (1) The reference information may be useful or may not be useful.

[0105] (2) The answer must be faithful to the original text, the answer should be sufficient, do not fabricate randomly, do not provide false information; if the reference information is irrelevant to the question, give a useful, effective, and true answer to the question.

[0106] (3) Reply in the same language as the user's question. If the user's question is in Chinese, answer in Chinese; if the user's question is in English, answer in English.

[0107] ## Skills: (1) Well - versed in various knowledge in the field of materials scientific research, especially in the field of metallic materials.

[0108] (2) Accurately understand the user's question and answer the user's question according to the question and reference information.

[0109] ## Workflows: (1) Carefully understand the user's problem.

[0110] (2) Answer the user’s questions based on the user’s questions and reference information.

[0111] (3) Make sure your answers are accurate, concise, and authoritative. Do not make them up.

[0112] (4) If the reference information is not relevant to the question, please provide a useful, valid, and true answer to the question.

[0113] (5) Ensure that the language of the answer is consistent with the language of the question, for example, "你好" is Chinese, while "hello" is English.

[0114] ## Initialization: As a [Role], possess [Skills], comply with [Constrains] according to the relevant information provided, then start work according to [Workflows], and complete the task according to [Goals].

[0115] Reference Information: ``` {context} ``` My questions: ``` {question} ``` The models in the question-and-answer system of this application usually use self-supervision / autoregression and other methods to train models with huge parameters on massive data, but its algorithm principle still mainly uses deep learning technology, so the output content of the large model may contain sensitive information. Since the training corpus of the large model is usually composed of crawled data on the Internet, the data on the network inevitably contains some sensitive information that is not suitable for users to view, and the model will remember this sensitive information after training on this data, resulting in the output of the model during use. Therefore, it is necessary not only to improve the security of the model from the algorithm principle, but also to require a more reasonable and comprehensive large model security governance method. This application mainly solves the problem of sensitive information through the following aspects: (1) Comprehensive data cleaning to ensure the security of large pre-trained models. By filtering out harmful data and sensitive information such as privacy data in Internet data, the sensitivity of model output content can be effectively reduced. In the initial stage of the pre-trained model, the data needs to be properly cleaned and filtered.

[0116] (2) Optimize the language model in a reinforcement learning manner based on human feedback, that is, the way of fine-tuning the model after pre-training, which is also an important step in aligning the pre-trained model with human values. Specifically, in the reinforcement learning stage, a large amount of manually annotated data is used to train the AI system model. The reward model generates reward signals (Reward Signals) of different intensities according to the output results of the AI model and the annotated data to guide the AI model to converge in the expected direction, and a safer AI system model is obtained after training. The effect of the model trained under this technology depends to a large extent on the scale and quality of the manually annotated data, which helps to solve the problem of outputting sensitive information.

[0117] (3) For reinforcement learning based on AI feedback, the source of the feedback information of the reward model is partially or completely changed from human feedback to being automatically provided by the proxy model. The proxy model is a model that has been pre-trained to meet safety standards and is aligned. It serves as a "supervisor" to provide feedback signals for the reinforcement training of the new model.

[0118] This system aims to help students retrieve and understand relevant research literature more efficiently and overcome the trouble of information overload. By combining advanced retrieval strategies and large language models, the system can quickly locate the most relevant literature resources and generate concise summaries and insights. This not only greatly reduces the time investment of students in literature research, but also helps them better grasp the research frontiers and stimulate innovative thinking. In a rapidly evolving field such as materials science, the application of this system will provide strong support for the study and research of college students and lay a solid foundation for their in-depth research and innovation in the field of materials science in the future.

[0119] By integrating these advanced technologies, our RAG system can efficiently process and utilize a vast amount of materials science literature, providing comprehensive support for college students from literature management, information retrieval to knowledge generation. The system not only efficiently improves the learning and research efficiency, but also stimulates innovative thinking and helps students better grasp the frontier development in the field of materials science.

[0120] In summary, in terms of system construction, this application has successfully built a comprehensively integrated RAG platform, which includes a complete process from literature management, OCR processing, text chunking, vectorization to retrieval and answer generation. This application has developed an efficient and stable service architecture that can handle a large number of materials science literatures while ensuring the response speed and stability of the system. Notably, this application has optimized the literature processing process according to the special needs of the materials science field, enabling the system to accurately identify and generate texts that conform to human reading habits, laying a solid foundation for subsequent retrieval and analysis. The intelligent Q&A system for the materials field based on the vector database of this application can well support the rapid indexing and retrieval of high-dimensional vectors, greatly improving the retrieval efficiency and accuracy. To further improve the quality of retrieval results, a dedicated Rerank model is introduced. Based on the preliminary retrieval results, this model combines the knowledge of the materials science field and context information to refine the ranking of the retrieval results, ensuring that the most relevant information is presented first.

[0121] Based on the same principle, this embodiment also discloses an intelligent Q&A system for the materials field based on the vector database. As Figure 22 shown, the system includes a question preprocessing module 11, a literature retrieval module 12, and a question answering module 13.

[0122] The question preprocessing module 11 is used to identify and extract the question input by the user through a language large model to obtain question keywords, determine the Chinese words in the question keywords, determine whether there are multilingual synonyms for the Chinese words, and if so, replace the Chinese words with the multilingual synonyms to obtain a multilingual question. The literature retrieval module 12 is used to retrieve a first literature list from the vector database based on the multilingual question; and sort the first literature list by cross-matching the multilingual question and the first literature list through a cross-lingual Reranker model to obtain a second literature list. The question answering module 13 is used to form a question answer based on the literatures in the second literature list through the language large model.

[0123] Since the principle of this system for solving problems is similar to the above method, the implementation of this system can refer to the implementation of the method and will not be elaborated here.

[0124] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0125] In a typical example, the computer device specifically includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method executed by the client as described above, or when the processor executes the program, it implements the method executed by the server as described above.

[0126] The following refers to Figure 23 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.

[0127] As Figure 23 shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate tasks and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0128] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and speakers; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that the computer program read from it can be installed in the storage section 608 as needed.

[0129] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product that tangibly includes a computer program on a machine-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611.

[0130] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0131] For convenience of description, the above devices are described by function as various units respectively. Of course, when implementing the present application, the functions of the respective units can be implemented in one or more software and / or hardware.

[0132] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0133] This application may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. This application may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.

[0134] Each embodiment in this specification is described in a progressive manner, and for the identical or similar parts among the embodiments, reference may be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference may be made to the partial description of the method embodiments for the relevant parts.

[0135] The above description is only for the embodiments of this application and is not intended to limit this application. For those skilled in the art, various modifications and changes can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. An intelligent question-answering method in the field of materials based on a vector database, characterized in that, Including: Identifying and extracting the problem in the user input through a language model to obtain problem keywords, determining the Chinese words in the problem keywords, determining whether there are multilingual synonyms for the Chinese words, and if so, replacing the Chinese words with the multilingual synonyms to obtain a multilingual problem; Retrieving a first literature list from a vector database based on the multilingual problem; sorting the first literature list through cross-matching the multilingual problem and the first literature list by a cross-lingual Reranker model to obtain a second literature list; Forming a problem response based on the literatures in the second literature list by the language model.

2. The intelligent Q&A method in the field of materials based on a vector database according to claim 1, wherein Further including the step of pre-extracting data from the literature resources to obtain metadata: Performing layout prediction on the literature materials through the onnxruntime inference engine of a preset OCR model; Determining the literature type of the literature materials based on the result of the layout prediction; Identifying all independent paragraphs of the literature resources, and combining the independent paragraphs with an associated relationship based on the literature type to obtain the metadata.

3. The intelligent Q&A method in the field of materials based on the vector database according to claim 2, wherein The pre-extracting data from the literature resources to obtain metadata includes: Splitting the literature resources into multiple sub-literature sets, and allocating the sub-literatures in the sub-literature sets to multiple processes for parallel processing to obtain the metadata corresponding to the sub-literatures; Integrating the metadata corresponding to the sub-literatures to obtain the metadata of all literature resources.

4. The intelligent Q&A method in the field of materials based on a vector database according to claim 3, wherein The splitting the literature resources into multiple sub-literature sets includes: Screening the literatures smaller than or equal to a preset standard literature size based on a preset standard literature size and dividing them into the first sub-literatures; Splitting the literatures larger than the preset standard literature size into multiple sub-literatures according to a preset splitting rule, and taking all the split sub-literatures as the second sub-literatures; Obtaining multiple sub-literature sets based on the number of the first sub-literatures and the second sub-literatures and the number of parallel processing processes.

5. The intelligent question-answering method in the field of materials based on a vector database according to claim 1, characterized in that, Further including the step of vectorizing and storing the metadata in a vector database, including: Pre-training a preset BCEmbedding model using multilingual and multi-domain literatures based on the XLM-RoBERTa architecture; Optimizing the BCEmbedding model using a large-scale multilingual corpus in the materials field, performing adaptive fine-tuning for the materials field, and optimizing the BCEmbedding model through parallel corpus and contrast learning techniques to obtain a cross-lingual BCEmbedding model at the word level, sentence level, and document level; Selecting the cross-lingual BCEmbedding model at the word level, sentence level, and document level to vectorize the corresponding metadata.

6. The intelligent Q&A method in the field of materials based on a vector database according to claim 5, wherein, Further including the step of building an index for the vectorized metadata, including: Identifying the data types of independent paragraphs in the metadata, and dividing the vector data of the metadata according to the data types to obtain vector sub-data corresponding to different data types; For the vector sub-data with a data volume greater than a preset value, dividing it into multiple vector sub-data according to the spatial structure, logical structure, or general content of the vector sub-data; Form multiple vector sub-data sets based on all vector sub-data, set a separate vector index for each vector sub-data set, and form a corresponding index file.

7. The intelligent Q&A method in the field of materials based on the vector database according to claim 6, characterized in that, The forming of the question response by the language large model based on the documents in the second document list includes: Form a question response language based on the second document list; Based on the user's requirements, determine the target documents in the second document list, merge some of the vector sub-data of the target documents to form a response document, or merge the vector sub-data of the target documents after removing useless information to form a question response, or merge all the vector sub-data of the target documents to obtain the complete target document; Form the question response based on the question response language and the target document.

8. An intelligent Q&A system in the field of materials based on a vector database, characterized in that, Include: A question preprocessing module, configured to identify and extract the question input by the user through a language large model to obtain question keywords, determine the Chinese words in the question keywords, determine whether there are multilingual synonyms for the Chinese words, and if so, replace the Chinese words with the multilingual synonyms to obtain a multilingual question; A document retrieval module, configured to retrieve a first document list from a vector database based on the multilingual question; perform cross-matching on the multilingual question and the first document list through a cross-lingual Reranker model to sort the first document list to obtain a second document list; A question response module, configured to form a question response by the language large model based on the documents in the second document list.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the intelligent question-answering method in the field of materials based on a vector database according to any one of claims 1 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the intelligent question-answering method in the field of materials based on a vector database according to any one of claims 1 to 7.

Citation Information

Cited By

  • Planning-guided adaptive verifiable multilingual question answering method and system

    CN122222018A