Multi-vector knowledge base construction method and system for large model information retrieval enhancement, equipment and medium
By building a multi-vector knowledge base, the large model answers incomplete questions when processing scientific and technological intelligence data, and realizes efficient and accurate information retrieval, which improves the performance and scope of application of the large model.
Patent Information
- Application Number
- CN202510080645.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing large-model information retrieval methods give incomplete answers when processing scientific and technological intelligence data, resulting in inefficient retrieval and a single knowledge base construction method, which cannot fully cover multi-dimensional information in the field of scientific and technological intelligence.
By building a multi-vector knowledge base, obtaining and classifying the inventory of scientific and technological intelligence data, creating a vector database corresponding to each category, and vectorized the classified text data to form a multi-vector knowledge base to support the information retrieval enhancement of the big model.
It significantly improves the accuracy and efficiency of the scientific and technological intelligence model in the information retrieval process, enhances its ability to process specific categories of scientific and technological data, and expands its scope of application.
Smart Images

Figure CN119988610A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large model retrieval enhancement, and in particular to a method, system, device and medium for constructing a multi-vector knowledge base for large model information retrieval enhancement. Background Art
[0002] With the explosive growth of scientific and technological intelligence data, how to quickly and accurately retrieve valuable information from massive data has become an urgent problem to be solved. In related technologies, large model information retrieval methods have improved retrieval efficiency to a certain extent, but the knowledge base construction method is single and lacks consideration of the diversity of scientific and technological intelligence data. It cannot fully cover the multi-dimensional information in the field of scientific and technological intelligence, resulting in incomplete answers given by large models when processing specific categories of scientific and technological data, and at the same time resulting in low retrieval efficiency. Summary of the invention
[0003] The purpose of this application is to provide a method, system, device and medium for constructing a multi-vector knowledge base enhanced by large-model information retrieval, which can improve the accuracy and efficiency of large-model scientific and technological intelligence in the information retrieval process.
[0004] To achieve the above objectives, this application provides the following solutions:
[0005] In a first aspect, the present application provides a method for constructing a multi-vector knowledge base with enhanced large-model information retrieval, comprising:
[0006] Obtain stock data of scientific and technological intelligence;
[0007] Classifying the stock data of scientific and technological intelligence to obtain multiple categories of scientific and technological intelligence text data;
[0008] Create a vector database corresponding to each category in the knowledge base directory of the large model;
[0009] The scientific and technological intelligence text data of each category is stored in the vector database of the corresponding category to obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
[0010] Optionally, obtain scientific and technological intelligence stock data, including:
[0011] Access historical science and technology intelligence data;
[0012] The historical scientific and technological intelligence data is cleaned to obtain the scientific and technological intelligence stock data.
[0013] Optionally, the vector database includes an original text folder and a vector folder; the original text folder is used to store text data, and the vector folder is used to store vectors.
[0014] Optionally, the scientific and technological intelligence text data of each category is stored in the vector database of the corresponding category, specifically including:
[0015] Segment the scientific and technological intelligence text data of each category respectively to obtain the segmented scientific and technological intelligence text data of each category;
[0016] Store the intelligence text data after segmentation for each category in the original text folder of the vector database of the corresponding category;
[0017] The scientific and technological intelligence text data after segmentation for each category is vectorized to obtain the scientific and technological intelligence vector for each category;
[0018] The scientific and technological intelligence vector of each category is stored in the vector folder of the vector database of the corresponding category.
[0019] Optionally, a text segmenter is used to segment each category of scientific and technological intelligence text data separately.
[0020] Optionally, the large model information retrieval enhanced multi-vector knowledge base construction method further includes:
[0021] Obtaining incremental scientific and technological intelligence data at regular intervals and determining the categories of the incremental scientific and technological intelligence data;
[0022] Performing text segmentation on the incremental data of scientific and technological intelligence to obtain a plurality of incremental text data, and storing the plurality of incremental text data in an original text folder of a corresponding vector database;
[0023] Each incremental text data is vectorized to obtain an incremental vector, and the incremental vector is stored in a vector folder of a vector database of a corresponding category.
[0024] Optionally, the large model information retrieval enhanced multi-vector knowledge base construction method further includes:
[0025] A knowledge information base is created under the knowledge base directory of the large model; the knowledge information base stores the number, release time and file category of each file in the vector database.
[0026] In a second aspect, the present application provides a multi-vector knowledge base construction system for large model information retrieval enhancement, comprising:
[0027] Data acquisition module, used to obtain stock data of scientific and technological intelligence;
[0028] A classification module is used to classify the scientific and technological intelligence stock data to obtain multiple categories of scientific and technological intelligence text data;
[0029] The vector library creation module is used to create a vector database corresponding to each category under the knowledge base directory of the large model;
[0030] The knowledge base determination module is used to store the scientific and technological intelligence text data of each category into the vector database of the corresponding category respectively, and obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
[0031] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned large-model information retrieval enhanced multi-vector knowledge base construction method.
[0032] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned large-model information retrieval enhanced multi-vector knowledge base construction method.
[0033] According to the specific embodiments provided in this application, this application has the following technical effects:
[0034] The present application provides a method, system, device and medium for constructing a multi-vector knowledge base enhanced by large-model information retrieval. Aiming at the characteristics of the field of scientific and technological intelligence, vector databases of multiple categories are constructed to support the efficient training and application of large scientific and technological intelligence models, which significantly improves the performance and scope of application of large scientific and technological intelligence models, and effectively improves the accuracy and efficiency of large scientific and technological intelligence models in the information retrieval process. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 This is an application environment diagram of a multi-vector knowledge base construction method for large model information retrieval enhancement in one embodiment of the present application;
[0037] Figure 2 A schematic diagram of a process for constructing a multi-vector knowledge base with enhanced large-model information retrieval provided in one embodiment of the present application;
[0038] Figure 3 This is a schematic diagram of the processing process of stock data in one embodiment of the present application;
[0039] Figure 4This is a schematic diagram of the structure of a multi-vector knowledge base in an embodiment of the present application;
[0040] Figure 5 This is a schematic diagram of the incremental data processing process in one embodiment of the present application;
[0041] Figure 6 A schematic diagram of the functional modules of a multi-vector knowledge base construction system for large-model information retrieval enhancement provided in one embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0043] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0044] The multi-vector knowledge base construction method for large model information retrieval enhancement provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, or it can be integrated on the server 104, or it can be placed on the cloud or other servers. The terminal 102 can send the scientific and technological intelligence stock data to the server 104, and the server 104 builds a multi-vector knowledge base after receiving the scientific and technological intelligence stock data. The server 104 can feedback the multi-vector knowledge base to the terminal 102. In addition, in some embodiments, the multi-vector knowledge base construction method enhanced by large model information retrieval can also be implemented by the server 104 or the terminal 102 alone.
[0045] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0046] In an exemplary embodiment, Figure 2As shown, a method for constructing a multi-vector knowledge base with enhanced large model information retrieval is provided. The method is executed by a computer device, and specifically can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, and the steps include the following steps 201 to 204.
[0047] Step 201, obtaining the stock data of scientific and technological intelligence.
[0048] Specifically, Figure 3 As shown, firstly, historical science and technology intelligence data is obtained, and then the historical science and technology intelligence data is cleaned to obtain science and technology intelligence stock data.
[0049] The stock data of scientific and technological intelligence mainly refers to historical data. The scientific and technological information of previous years also plays an important role in the scientific and technological intelligence analysis and judgment of the scientific and technological intelligence big model. For the historical data of scientific and technological intelligence, the main operation used is the vector library initialization operation driven by the scientific and technological intelligence big model.
[0050] Among them, data cleaning includes removing duplicate data, removing zero-byte files, cleaning up garbled characters, etc. to ensure data quality.
[0051] Step 202, classify the stock data of scientific and technological intelligence to obtain scientific and technological intelligence text data of multiple categories.
[0052] Based on the characteristics of wide coverage and diverse content of scientific and technological intelligence data, this application has constructed a vector database of multiple categories. Through the integration and analysis of existing scientific and technological intelligence data, as well as repeated discussions and research with experts and researchers in the field of intelligence, 17 categories of scientific and technological intelligence have been designed. The specific categories and code numbers are as follows:
[0053] ["Science and Technology Strategy","kjzl"];
[0054] ["Technology Frontier","kjqy"];
[0055] ["Biomedicine and Big Health","swyy"];
[0056] ["new generation information technology","xxjs"];
[0057] ["Advanced Manufacturing","xjzz"];
[0058] ["new material","xcl"];
[0059] ["new energy","xmy"];
[0060] ["fintech","jrkj"];
[0061] ["intelligent society","znsh"];
[0062] ["intelligent transportation","znjt"];
[0063] ["environmental protection","hjbh"];
[0064] ["cultural tourism","whlv"];
[0065] ["beauty news","mqdt"];
[0066] ["innovation ecology","cxst"];
[0067] ["Popular Science","kxpj"];
[0068] ["transuranic nuclides","cyhs"];
[0069] ["Policies and Regulations","zcfg"].
[0070] This application is based on the above 17 categories and driven by the big model of scientific and technological intelligence to construct 17 vector knowledge bases to promote the retrieval enhancement generation of the big model of scientific and technological intelligence and improve the retrieval efficiency of the big model of scientific and technological intelligence.
[0071] Specifically, based on the characteristics, attributes and content of the scientific and technological intelligence stock data, category labels are added to the scientific and technological intelligence stock data, and the scientific and technological intelligence stock data are divided into 17 categories.
[0072] Step 203, creating a vector database corresponding to each category under the knowledge base directory of the large model.
[0073] In an exemplary embodiment, Figure 4 As shown, the vector database includes an original text folder and a vector folder. The original text folder is used to store text data, namely original text 1 to original text n, and the vector folder is used to store vectors.
[0074] Further configure the system knowledge base parameters of the scientific and technological intelligence model, and set the above 17 category codes as the system knowledge base name of the scientific and technological intelligence model: SYSTEM_KNOWLEDGE_BASE:-cxst-cyhs-hjbh-jrkj-kjqy-kjzl-kxpj-mqdt…. This setting configures the knowledge base parameters of the scientific and technological intelligence model as a multi-category vector library and binds the name of the multi-vector database.
[0075] Step 204, respectively store the scientific and technological intelligence text data of each category into the vector database of the corresponding category to obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
[0076] Specifically, 17 folders named with category code numbers are created under the knowledge base directory of the scientific and technological intelligence model to store scientific and technological intelligence data of corresponding categories, and the classified scientific and technological intelligence texts are copied to the category folders corresponding to the knowledge base directory.
[0077] In an exemplary embodiment, step 204 includes the following steps (1) to (4):
[0078] (1) Segment the scientific and technological intelligence text data of each category to obtain the segmented scientific and technological intelligence text data of each category. Specifically, a text segmenter is used to segment the scientific and technological intelligence text data of each category.
[0079] Taking into account the particularity of scientific and technological intelligence text data and the processing capabilities of the scientific and technological intelligence model itself, through comparative testing, the Chinese Recursive Text Splitter text splitter is used, and the file segmentation size is determined to be 500 characters. This can improve the efficiency of the scientific and technological intelligence model in processing scientific and technological intelligence text data while maintaining the semantic integrity of the data.
[0080] (2) The intelligence text data after segmentation for each category is stored in the original text folder of the vector database of the corresponding category.
[0081] (3) The scientific and technological intelligence text data after segmentation for each category is vectorized to obtain the scientific and technological intelligence vector for each category.
[0082] In an exemplary embodiment, a suitable vector embedding model and a similarity search clustering vector library (FacebookAI Similarity Search, FAISS) are set to build a vector database. The vector library initialization operation is called cyclically for each category of the vector database in turn. The science and technology intelligence model will automatically build a vector library based on the 17 category numbered folders under the knowledge base directory and update its corresponding information database.
[0083] (4) Store the scientific and technological intelligence vector of each category in the vector folder of the vector database of the corresponding category.
[0084] In an exemplary embodiment, the method for constructing a multi-vector knowledge base enhanced by large model information retrieval further includes step 205 of creating a knowledge information base under the knowledge base directory of the large model. The knowledge information base stores the number, release time and file category of each file in the vector database.
[0085] In order to support the construction of multi-vector knowledge base in information retrieval enhancement of the large-scale scientific and technological intelligence model, it is necessary to customize and modify the knowledge information base of the large-scale scientific and technological intelligence model. The main customization content includes adding kq_filecid, kq_pubtime and kq_type fields in the knowledge information base. The three fields mainly store the number, release time and file category of the file in the vector library, which is convenient for the large-scale scientific and technological intelligence model to use when generating retrieval enhancement.
[0086] In another exemplary embodiment, Figure 2 and Figure 5 As shown, the method for constructing a multi-vector knowledge base enhanced by large model information retrieval also includes the following steps 206 to 208.
[0087] Step 206, periodically acquiring incremental scientific and technological intelligence data, and determining the category of the incremental scientific and technological intelligence data.
[0088] Step 207, performing text segmentation on the incremental data of scientific and technological intelligence to obtain a plurality of incremental text data, and storing the plurality of incremental text data in an original text folder of a corresponding vector database.
[0089] Step 208 , respectively perform vectorization processing on each incremental text data to obtain an incremental vector, and store the incremental vector in a vector folder of a vector database of a corresponding category.
[0090] The scientific and technological intelligence data in this application is in the form of TXT text, involving existing data and dynamically updated data. Therefore, this application constructs a multi-vector knowledge base from both stock data and incremental data.
[0091] Information in the field of science and technology changes with each passing day. For the new data added every day, this application regularly updates the vector database, mainly by calling the large model interface to upload files to the vector database, to achieve incremental updates of the multi-vector knowledge base.
[0092] First, customize the update file interface and add category parameters. Since the traditional large model only has one vector database, you can directly update the default vector database when updating files. In the science and technology intelligence large model of multiple vector knowledge bases, you need to determine which vector database to add files to based on the category, so you need to customize the interface and add category parameters.
[0093] Then, the incremental files with category labels are received. When submitting a file to the science and technology intelligence model, it is necessary to first determine the category of the file and pass the category label along with the file content to the science and technology intelligence model. The file category determination method can be manually marked or the science and technology intelligence model can automatically identify the file category.
[0094] Finally, the corresponding category vector database is automatically updated according to the category label. After receiving the newly added file, the file segmentation is performed and the text vectorization model is called to perform vectorization operations on a single file, and the vectorized file is stored in the corresponding vector database.
[0095] In view of the characteristics of the scientific and technological intelligence field, this application innovatively constructs vector databases of multiple categories to support the efficient training and application of scientific and technological intelligence large models, significantly improving the performance and scope of application of scientific and technological intelligence large models, and can effectively improve the accuracy and efficiency of scientific and technological intelligence large models in the information retrieval process, providing strong support for scientific and technological innovation and industrial development.
[0096] Based on the same inventive concept, the embodiment of the present application also provides a large model information retrieval enhanced multi-vector knowledge base construction system for implementing the large model information retrieval enhanced multi-vector knowledge base construction method involved above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more large model information retrieval enhanced multi-vector knowledge base construction system embodiments provided below can refer to the limitations of the large model information retrieval enhanced multi-vector knowledge base construction method above, and will not be repeated here.
[0097] In an exemplary embodiment, Figure 6 As shown, a large model information retrieval enhanced multi-vector knowledge base construction system is provided, including: a data acquisition module 601, a classification module 602, a vector library creation module 603 and a knowledge base determination module 604.
[0098] Among them, the data acquisition module 601 is used to obtain the stock data of scientific and technological intelligence.
[0099] The classification module 602 is used to classify the scientific and technological intelligence stock data to obtain multiple categories of scientific and technological intelligence text data.
[0100] The vector library creation module 603 is used to create a vector database corresponding to each category under the knowledge base directory of the large model.
[0101] The knowledge base determination module 604 is used to store the scientific and technological intelligence text data of each category into the vector database of the corresponding category, and obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
[0102] In an exemplary embodiment, the large model information retrieval enhanced multi-vector knowledge base construction system further includes an information base construction module 605. The information base construction module 604 is used to create a knowledge information base under the knowledge base directory of the large model. The knowledge information base stores the number, release time and file category of each file in the vector database.
[0103] In another exemplary embodiment, the large model information retrieval enhanced multi-vector knowledge base construction system further includes: an incremental data acquisition module 606 , a text segmentation module 607 and a vectorized storage module 608 .
[0104] The incremental data acquisition module 606 is used to periodically acquire incremental scientific and technological intelligence data and determine the category of the incremental scientific and technological intelligence data.
[0105] The text segmentation module 607 is used to perform text segmentation on the incremental data of scientific and technological intelligence to obtain multiple incremental text data, and store the multiple incremental text data in the original text folder of the corresponding vector database.
[0106] The vectorized storage module 608 is used to perform vectorized processing on each incremental text data to obtain an incremental vector, and store the incremental vector in a vector folder of a vector database of a corresponding category.
[0107] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0108] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0109] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0110] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0111] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0112] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0113] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0114] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for constructing a multi-vector knowledge base enhanced by large model information retrieval, characterized in that: The method for constructing a multi-vector knowledge base enhanced by large model information retrieval includes: Obtain stock data of scientific and technological intelligence; Classifying the stock data of scientific and technological intelligence to obtain multiple categories of scientific and technological intelligence text data; Create a vector database corresponding to each category in the knowledge base directory of the large model; The scientific and technological intelligence text data of each category is stored in the vector database of the corresponding category to obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
2. The method for constructing a multi-vector knowledge base with enhanced large model information retrieval according to claim 1, characterized in that: Obtain the stock data of scientific and technological intelligence, including: Access historical science and technology intelligence data; The historical scientific and technological intelligence data is cleaned to obtain the scientific and technological intelligence stock data.
3. The method for constructing a multi-vector knowledge base with enhanced large-model information retrieval according to claim 1, characterized in that: The vector database includes an original text folder and a vector folder; the original text folder is used to store text data, and the vector folder is used to store vectors.
4. The method for constructing a multi-vector knowledge base with enhanced large model information retrieval according to claim 3 is characterized in that: The scientific and technological intelligence text data of each category is stored in the vector database of the corresponding category, including: Segment the scientific and technological intelligence text data of each category respectively to obtain the segmented scientific and technological intelligence text data of each category; Store the intelligence text data after segmentation for each category in the original text folder of the vector database of the corresponding category; The scientific and technological intelligence text data after segmentation for each category is vectorized to obtain the scientific and technological intelligence vector for each category; The scientific and technological intelligence vector of each category is stored in the vector folder of the vector database of the corresponding category.
5. The method for constructing a multi-vector knowledge base with enhanced large model information retrieval according to claim 4, characterized in that: A text segmenter is used to segment each category of scientific and technological intelligence text data.
6. The method for constructing a multi-vector knowledge base enhanced by large model information retrieval according to claim 3 is characterized in that: The method for constructing a multi-vector knowledge base enhanced by large model information retrieval also includes: Obtaining incremental scientific and technological intelligence data at regular intervals and determining the categories of the incremental scientific and technological intelligence data; Performing text segmentation on the incremental data of scientific and technological intelligence to obtain a plurality of incremental text data, and storing the plurality of incremental text data in an original text folder of a corresponding vector database; Each incremental text data is vectorized to obtain an incremental vector, and the incremental vector is stored in a vector folder of a vector database of a corresponding category.
7. The method for constructing a multi-vector knowledge base with enhanced large model information retrieval according to claim 1, characterized in that: The method for constructing a multi-vector knowledge base enhanced by large model information retrieval also includes: A knowledge information base is created under the knowledge base directory of the large model; the knowledge information base stores the number, release time and file category of each file in the vector database.
8. A large model information retrieval enhanced multi-vector knowledge base construction system, applied to the large model information retrieval enhanced multi-vector knowledge base construction method according to any one of claims 1 to 7, characterized in that: The large model information retrieval enhanced multi-vector knowledge base construction system includes: Data acquisition module, used to obtain stock data of scientific and technological intelligence; A classification module is used to classify the scientific and technological intelligence stock data to obtain multiple categories of scientific and technological intelligence text data; The vector library creation module is used to create a vector database corresponding to each category under the knowledge base directory of the large model; The knowledge base determination module is used to store the scientific and technological intelligence text data of each category into the vector database of the corresponding category respectively, and obtain a multi-vector knowledge base for information retrieval enhancement by the scientific and technological intelligence large model.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-vector knowledge base construction method for large model information retrieval enhancement according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a multi-vector knowledge base enhanced by large model information retrieval according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent process generation method and system based on data driving
CN118297275A