Knowledge data storage and management method and device, equipment and storage medium
By preprocessing and vectorized conversion of knowledge data, combining data popularity and storage strategies with storage performance indicators, the inefficiency of Milvus database in dynamic data updates and diversified data management is solved, and efficient data updates and retrieval is achieved.
Patent Information
- Application Number
- CN202510139836.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
When Milvus database updates dynamic data and diversified data management needs, there are problems such as low storage efficiency, slow system response speed and low retrieval efficiency.
A knowledge data storage and management method is adopted. By obtaining the knowledge data to be stored, extracting text information, preprocessing and vectorized conversion, the data is stored at the corresponding storage level of the Milvus database based on the data popularity and storage performance indicators, and an incremental update mechanism is adopted during dynamic updates to avoid retraining or batch update of the entire knowledge base.
It improves the efficiency of dynamic data updates, reduces query delays, improves retrieval efficiency, realizes intelligent storage resource allocation and multimodal data processing, and supports the scalability and real-time nature of the system.
Smart Images

Figure CN120086304A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data retrieval, and in particular, to a method, device, equipment, and storage medium for knowledge data storage and management. Background Art
[0002] With the rapid growth and increasing complexity of professional knowledge, knowledge management systems are facing unprecedented challenges. The emergence of a vast amount of knowledge has made how to efficiently store, manage, and quickly retrieve the core problem in the current knowledge management field. Although traditional relational databases perform excellently in the management of structured data, they often fall short when dealing with unstructured data. Such data includes text, images, audio, video, etc., which usually have high dimensions, non-linearity, and irregularity, and are difficult to be effectively organized and managed through traditional row-column storage structures. In addition, with the continuous expansion of the data scale and the increasing requirements of users for real-time performance, traditional databases have exposed performance bottlenecks in processing real-time queries and dynamic updates, and cannot meet the needs of efficient retrieval and flexible management.
[0003] In recent years, vector-based retrieval technologies have received extensive attention. In particular, the emergence of the Milvus database provides a new solution for processing high-dimensional vector data. Milvus is a vector database designed specifically for the storage and retrieval of unstructured data, capable of supporting fast similarity search for massive data, and showing significant advantages especially in the scenarios of text, image, and audio data. By vectorizing the data, the Milvus database realizes the effective storage and fast query of unstructured data, greatly improving the retrieval efficiency.
[0004] However, the method for knowledge data storage and management based on the Milvus database still has some problems to be solved urgently when facing dynamic data updates and diverse data management requirements. For example, the existing knowledge storage systems store knowledge data uniformly, which will result in slow system response speed and low retrieval efficiency when users actually retrieve. Summary of the Invention
[0005] In order to help solve the problem that in the scenario of dynamic data updates in the Milvus database, the existing systems often need to retrain or batch update the entire knowledge base when inserting new data, resulting in low storage efficiency, this application provides a method, device, equipment, and storage medium for knowledge data storage and management.
[0006] In a first aspect, this application provides a method for knowledge data storage and management, adopting the following technical solution: The method is applied to a knowledge data storage and management system, and the knowledge data storage and management system includes a Milvus database. The method includes:
[0007] Obtain the knowledge data to be stored, and extract the text information of the knowledge data to be stored;
[0008] Preprocess the text information and generate standardized text information;
[0009] Convert the standardized text information into data in vector format and generate knowledge vector data;
[0010] Divide the Milvus database into several storage levels according to a preset standard, and calculate the storage performance indicators of each storage level;
[0011] Calculate the data popularity of the knowledge vector data, and store the knowledge vector data in the corresponding storage level of the Milvus database according to the data popularity and the storage performance indicators.
[0012] In a specific feasible implementation, the preprocessing of the text information and generating standardized text information includes:
[0013] Use a tag cleaning function to remove redundant information in the text information and generate text redundant information removal;
[0014] Use a word segmentation tool to perform word segmentation on the text redundant information removal and generate standardized text information.
[0015] In a specific feasible implementation, the converting the standardized text information into data in vector format and generating knowledge vector data includes:
[0016] Adopt the Sentence-BERT model to convert the standardized text information into data in vector format and generate knowledge vector data;
[0017] The calculation method of the data converted into vector format includes:
[0018] Vectorize(W) = S-BERT(W);
[0019] Wherein, S-BERT() represents the Sentence-BERT model, Vectorize() represents the knowledge vector data after vector conversion, and W represents the set of words after word segmentation processing.
[0020] In a specific feasible implementation, the calculation method of the storage performance indicator of each storage level includes:
[0021] P(L i ) = αS i + βC i ;
[0022] Wherein, Li denotes the i-th storage hierarchy, P(L i ) represents the storage performance metric of the i-th storage hierarchy, S i represents the storage speed of the i-th storage hierarchy, C i represents the storage cost of the i-th storage hierarchy, and α and β represent constants.
[0023] In a specific feasible implementation, the calculation method of the data heat of the knowledge vector data includes:
[0024]
[0025] where T represents the knowledge vector data for which the data heat needs to be calculated, w k represents the weight of the k-th storage hierarchy, I k represents the storage performance metric of the k-th storage hierarchy, m represents the total number of knowledge vector data for which the data heat is calculated, and H(T) represents the data heat of the knowledge vector data.
[0026] In a specific feasible implementation, after storing the knowledge vector data into the corresponding storage hierarchy of the Milvus database, it further includes:
[0027] When new knowledge data is obtained, calculate the similarity between the new knowledge data and the knowledge data already stored in the Milvus database, and compare the calculated similarity with a preset threshold;
[0028] If the calculated similarity is greater than the preset threshold, do not store the new knowledge data, or update the knowledge data already stored in the Milvus database with the new knowledge data;
[0029] If the calculated similarity is not greater than the preset threshold, store the new knowledge data after processing into the corresponding storage hierarchy in the Milvus database;
[0030] Establish a deduplication log and record deduplication information in the deduplication log.
[0031] In a specific feasible implementation, after storing the knowledge vector data into the corresponding storage hierarchy of the Milvus database, it further includes:
[0032] When a user query request is received, convert the user request into a vector representation and generate a vector query request;
[0033] Extract keywords according to the vector query request and generate a keyword set;
[0034] Perform a query operation in the Milvus database according to the keyword set and store historical query records;
[0035] Use a feedback collection function to collect user feedback information, and dynamically adjust the similarity calculation weight according to the historical query records and the user feedback information.
[0036] In a second aspect, the present application provides a knowledge data storage and management device, which adopts the following technical solutions: The device is applied to a knowledge data storage and management system, and the knowledge data storage and management system includes a Milvus database. The device includes:
[0037] A data acquisition module, configured to acquire knowledge data to be stored and extract text information of the knowledge data to be stored;
[0038] A data processing module, configured to preprocess the text information and generate standardized text information;
[0039] A vector conversion module, configured to convert the standardized text information into data in vector format and generate knowledge vector data;
[0040] A storage division module, configured to divide the Milvus database into several storage levels according to a preset standard and calculate storage performance indicators for each storage level;
[0041] A data storage module, configured to calculate the data popularity of the knowledge vector data and store the knowledge vector data in the corresponding storage level of the Milvus database according to the data popularity and the storage performance indicators.
[0042] In a third aspect, the present application provides a computer device, which adopts the following technical solutions: It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as any one of the above knowledge data storage and management methods, is stored on the memory.
[0043] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solutions: It stores a computer program capable of being loaded and executed by the processor, such as any one of the above knowledge data storage and management methods.
[0044] In summary, the present application has the following beneficial technical effects:
[0045] 1. High dynamic data update efficiency. The setting of the incremental update mechanism can quickly process newly added knowledge data, solving the low-efficiency problem of the traditional method that requires retraining or batch updating of the entire knowledge base; the dynamic knowledge update realizes efficient knowledge update and seamless replacement, greatly improving the real-time performance and update efficiency of the system.
[0046] 2. Low query latency and high retrieval efficiency. Leveraging the high-dimensional vector retrieval ability of the Milvus database, fast similarity search for a large amount of knowledge data is achieved; by introducing metadata indexing and dynamic popularity optimization strategies, the query path is optimized and retrieval latency is reduced. Whether in a small knowledge base or an ultra-large-scale knowledge base, it can respond quickly to meet the requirements of efficient retrieval.
[0047] 3. Intelligent storage resource allocation. The adaptive storage mechanism dynamically adjusts storage resources according to the usage frequency of knowledge and data popularity, stores hot knowledge on high-performance storage media, and cold knowledge on low-cost media, realizing intelligent allocation and utilization of resources, which not only improves system performance but also reduces operation and maintenance costs.
[0048] 4. Support for multi-modal data. This application can process multiple unstructured data types such as text, images, and audio simultaneously. By using vectorization technology to uniformly represent different data types, the generality and scalability of the system are greatly improved, meeting the increasingly diverse requirements of data types in knowledge management.
[0049] 5. Strong scalability. This application fully considers the expansion ability of the system and supports horizontal expansion and cluster deployment; as the amount of knowledge data grows, the storage capacity and computing power can be flexibly expanded to ensure that system performance is not affected by the growth of data scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the method for storing and managing knowledge data in an embodiment of the present application;
[0051] Figure 2 is a schematic framework diagram of the knowledge data storage and management system based on the Milvus database in an embodiment of the present application;
[0052] Figure 3 is a schematic diagram of the knowledge data storage and management device in an embodiment of the present application;
[0053] Figure 4 is a schematic diagram for embodying a computer device in an embodiment of the present application.
[0054] Reference numerals: 301, data acquisition module; 302, data processing module; 303, vector conversion module; 304, storage division module; 305, data storage module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following further Figures 1-4 describes the present application in detail.
[0056] The embodiment of the present application discloses a knowledge data storage and management method, which can improve the retrieval efficiency of large-scale knowledge bases, the flexibility of data updates and the expansion capacity of the system. With the rapid growth and complexity of professional knowledge, knowledge management systems are facing unprecedented challenges. The emergence of massive knowledge has made how to efficiently store, manage and quickly retrieve the core problem in the current field of knowledge management. Although traditional relational databases perform well in the management of structured data, they are often unable to cope with unstructured data. Such data includes text, images, audio, video, etc., which are usually high-dimensional, non-linear and irregular, and are difficult to be effectively organized and managed through traditional row and column storage structures. In addition, with the continuous expansion of data scale and the improvement of users' real-time requirements, traditional databases have exposed performance bottlenecks in processing real-time queries and dynamic updates, and cannot meet the needs of efficient retrieval and flexible management.
[0057] In recent years, vector-based retrieval technology has received widespread attention, especially the emergence of the Milvus database, which provides a new solution for processing high-dimensional vector data. Milvus is a vector database designed for unstructured data storage and retrieval. It can support fast similarity search of massive data, especially in the scenarios of text, image and audio data. By vectorizing the data, the Milvus database realizes the effective storage and fast query of unstructured data, greatly improving the retrieval efficiency. However, the knowledge data storage and management method based on the Milvus database still has some problems to be solved when facing dynamic data updates and diversified data management needs.
[0058] For example, existing knowledge storage systems store knowledge data in a unified manner, which will result in slow system response and low retrieval efficiency when users actually search. In the scenario of dynamic data updates, existing systems often need to retrain or batch update the entire knowledge base when inserting new data, resulting in low storage efficiency. The problem of data redundancy is particularly common in traditional storage methods, especially repeated or similar knowledge fragments occupy a large amount of storage resources, further increasing the management burden of the system. In addition, as the amount of data gradually increases and the types of data become more complex, the query delay problem of traditional methods becomes more and more obvious, directly affecting the user experience. In order to help improve the efficiency of knowledge data storage and retrieval efficiency, the present application provides a knowledge data storage and management method.
[0059] Reference Figure 1 , the method comprises the following steps:
[0060] S10, acquiring knowledge data to be stored, and extracting text information of the knowledge data to be stored.
[0061] Specifically, to obtain the data that needs to be stored in the Milvus database, the data source may include two data sources. One is to obtain it from a multi-source database. For example, the data of existing multi-source databases such as Westlaw and LexisNexis databases can be continuously collected through API interfaces and web crawler technologies (such as Scrapy), and the required knowledge data and case data can be obtained from the multi-source database. The data types include text, pictures and video formats. Another data source is to continuously obtain the required data directly from websites and other places by establishing automated data crawling tasks. By setting up automatic crawling, data can be continuously obtained to ensure continuous updating of data. The data types also include text, pictures and videos. When crawling data, the system hardware resources can be used to improve the efficiency of data crawling, thereby ensuring the timeliness and completeness of document acquisition.
[0062] After obtaining the knowledge data to be stored, text information is extracted from the obtained knowledge data. It should be noted that when extracting text information, the named entity recognition model NER(T) of the natural language processing tool spaCy needs to be used for named entity recognition. The identified named entity set can be expressed as:
[0063] E={e 1 , e 2 , ..., e n};
[0064] Among them, E represents the set of named entities, e i Represents the i-th named entity.
[0065] Taking into account that there are different professional terms in different fields, when acquiring knowledge data from different data sources, the proper nouns in the acquired data are identified so that in subsequent practical applications, users can accurately retrieve the required proper nouns when searching, and avoid inaccurate search results caused by the segmentation of proper nouns as much as possible.
[0066] S20, preprocessing the text information and generating standardized text information.
[0067] Specifically, preprocess the extracted text information. The purpose of preprocessing is to remove redundant information and unify the data format for storage in the database and retrieval in practical applications. The data formats obtained from different data sources are different. For example, the data crawled through websites and the like includes data in HTML and XML formats. Use the markup cleaning function MarkupClean(T)=c(T) to remove redundant information in the text information and generate text with redundant information removed; it can also be understood that the HTML or XML tags are removed using the markup cleaning function MarkupClean(T)=c(T) and converted into pure text information. After obtaining the text with redundant information removed, that is, pure text information, use a word segmentation tool to perform word segmentation on the text with redundant information removed and generate standardized text information. Specifically, use the Jieba word segmentation tool to segment the pure text information, convert it into a set of processable words, and perform part-of-speech tagging and word frequency statistics. The text information after word segmentation is the standardized text information, which is converted into a unified text information format for subsequent processing. The set of words obtained after word segmentation can be expressed as:
[0068] W = {w 1 , w 2 , …, w n};
[0069] where W represents the set of words after word segmentation, and w i represents the i-th word after word segmentation.
[0070] S30. Convert the standardized text information into data in vector format and generate knowledge vector data.
[0071] Specifically, convert the standardized text information after word segmentation into vector representation through the Sentence-BERT model to generate knowledge vector data. Among them, after converting into vector representation, a unique ID can be generated for each piece of knowledge vector data and metadata such as data source, creation time, knowledge category, label, etc. can be attached. The type and format of the metadata can be user-defined. For example, only the data category and label can be added to the data, or the data source and creation time can be added, or nothing can be added. There is no restriction here, and users can set it according to actual needs. The core principle of the Sentence-BERT model is to encode sentences through the pre-trained BERT model to generate sentence-level representation vectors, which can capture the semantic and structural information of sentences, thereby realizing in-depth understanding and comparison of sentence semantics.
[0072] The calculation method for converting into data in vector format can be expressed as:
[0073] Vectorize(W) = S-BERT(W);
[0074] Among them, S-BERT() represents the Sentence-BERT model, Vectorize() represents the knowledge vector data after vector conversion, and W represents the set of words after word segmentation processing.
[0075] S40. Divide the Milvus database into several storage levels according to a preset standard, and calculate the storage performance indicators of each storage level.
[0076] Specifically, divide the Milvus database into several storage levels according to a preset standard. For example, it can be divided into several storage levels according to categories, time, or the correlation between knowledge data. The storage performance indicators of each storage level are different, and different types of knowledge data are stored in different storage levels. Among them, the calculation method of the storage performance indicator of each storage level can be expressed as:
[0077] P(L i ) = αS i + βC i ;
[0078] Among them, L i represents the i-th storage level, P(L i ) represents the storage performance indicator of the i-th storage level, S i represents the storage speed of the i-th storage level, C i represents the storage cost of the i-th storage level, and α and β represent constants.
[0079] S50. Calculate the data popularity of the knowledge vector data, and store the knowledge vector data in the corresponding storage level of the Milvus database according to the data popularity and storage performance indicators.
[0080] Specifically, calculate the data popularity of the knowledge vector data, and store the obtained knowledge data in the corresponding storage level according to the data popularity and the performance indicators of the storage levels divided by the storage system itself. For example, newer knowledge data can be stored in the fast retrieval layer, while older knowledge can be stored in the low-cost storage layer; when the popularity of some knowledge data decreases, the system can automatically migrate it from the high-performance storage layer to the lower-cost storage layer to reduce the storage cost; on the contrary, when the popularity of some knowledge data increases, it can be migrated from the low-cost storage layer to the high-performance storage layer to ensure the query efficiency. Through the dynamic adaptive storage method, the storage resources are dynamically adjusted according to the usage frequency and popularity of the knowledge data. The hot knowledge is stored in the high-performance storage level, and the cold knowledge is stored in the lower storage level, realizing the intelligent allocation and utilization of resources. This can not only improve the system performance but also reduce the operation and maintenance cost. Among them, the calculation method of the data popularity of the knowledge vector data can be expressed as:
[0081]
[0082] Among them, T represents the knowledge vector data for which the data heat needs to be calculated, and w k represents the weight of the k-th storage level, and I k represents the storage performance index of the k-th storage level. m represents the total number of knowledge vector data for calculating the data heat, and H(T) represents the data heat of the knowledge vector data.
[0083] In the solution of this application, by utilizing the high-dimensional vector retrieval ability of the Milvus database, fast similarity search for a large amount of knowledge data is realized. By introducing the strategies of metadata index and dynamic heat optimization of the storage method, the query path is further optimized, and the retrieval latency is reduced, enabling fast response in both small knowledge bases and ultra-large-scale knowledge bases to meet the high-efficiency retrieval requirements. In addition, in this application, the adaptive storage mechanism dynamically adjusts the resource storage according to the usage frequency and data heat of the knowledge data, stores the hot knowledge on the high-performance storage medium, and stores the cold knowledge on the low-cost medium, realizing the intelligent allocation and utilization of resources, which not only improves the system performance but also reduces the operation and maintenance costs.
[0084] In one embodiment, considering that when there is new knowledge data, the traditional method needs to retrain or batch update the entire knowledge base, resulting in low efficiency of updating the storage. Therefore, after storing the knowledge vector data in the corresponding storage level of the Milvus database, the following steps can also be executed:
[0085] First, when new knowledge data is obtained. For example, users can regularly check the updates of legal databases to obtain newly added legal provisions and cases, and obtain new knowledge data through database docking or automatic data scraping. Calculate the similarity between the new knowledge data and the knowledge data already stored in the Milvus database, and compare the calculated similarity with a preset threshold. If the calculated similarity is greater than the preset threshold, it indicates that the newly added knowledge data has a high similarity with the knowledge data already stored in the database. Then the system can automatically choose not to store the new knowledge data, or use the new knowledge data to update the knowledge data already stored in the Milvus database. If the calculated similarity is not greater than the preset threshold, it indicates that the newly added knowledge data has a low similarity with the already stored knowledge data, and there is no relevant knowledge data stored in the database. Then the new knowledge data is processed and stored in the corresponding storage level in the Milvus database. The processing method is also the processing method used in this application. First, extract the text information in the data, then perform preprocessing to unify it into a standard data format, and then perform vector conversion on the standardized data to convert it into knowledge vector data. Then determine which storage level the heat of the knowledge vector data belongs to and perform storage operations at the corresponding storage level. After the newly added knowledge data is stored in the Milvus database, a deduplication log is established, and the deduplication information is recorded in the deduplication log.
[0086] In the solution of this application, by comparing the similarity between new and old knowledge, new knowledge data can be quickly processed, avoiding the inefficiency of the traditional method that requires retraining or batch updating of the entire knowledge base. In addition, in a dynamic knowledge environment, efficient knowledge update and seamless replacement are achieved, greatly improving the real-time performance and update efficiency of the system.
[0087] In one embodiment, considering that in actual queries, due to different weight calculations in the system similarity calculation, the final retrieval results may be inaccurate. Therefore, after storing the knowledge vector data in the corresponding storage level of the Milvus database, the following steps can also be performed:
[0088] First, when a user query request is received, convert the user request into a vector representation and generate a vector query request. Extract keywords based on the vector query request and generate a keyword set. The extracted keyword set can be expressed as K = {k 1 , k 2 , …, k n}; Subsequently, query operations are performed in the Milvus database according to the keyword set. Each time a query is retrieved, the historical query records are recorded and stored. Finally, a feedback collection function is used to collect user feedback information, and the similarity calculation weight is dynamically adjusted according to the historical query records and user feedback information. Among them, the weight adjustment function can be expressed as: AdjustWeight(W,F), where W represents the initial weight set and F represents the user feedback information.
[0089] In the solution of this application, according to each historical query record and the feedback of the results obtained from the actual query retrieval by the user, the weight of the similarity calculation is dynamically adjusted to adjust the finally retrieved results, improving the accuracy of the retrieval results.
[0090] Refer to Figure 2 , which is a schematic diagram of the knowledge data storage and management system based on the Milvus database in the embodiment of this application. First, knowledge data can be obtained from the existing database through the API interface and web crawler technology, or data from places such as websites can be automatically crawled. After obtaining the data, the obtained data is cleaned, different data formats are converted into a unified data format, the data is standardized, and then converted into a vectorized data representation. Then, the vectorized data is stored according to the data popularity and different storage levels of the Milvus database. In actual applications, after receiving the user's query instruction, the retrieval is completed through the intelligent query and retrieval module.
[0091] Figure 1 It is a schematic flowchart of the knowledge data storage and management method in an embodiment. It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows; unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders; and Figure 1 at least a part of the steps in
[0092] Based on the above method, the embodiment of this application also discloses a knowledge data storage and management device.
[0093] Refer to Figure 3 , the device includes the following modules:
[0094] The data acquisition module 301 is used to acquire the knowledge data to be stored and extract the text information of the knowledge data to be stored;
[0095] The data processing module 302 is used to preprocess the text information and generate standardized text information;
[0096] The vector conversion module 303 is used to convert the standardized text information into vector-formatted data and generate knowledge vector data;
[0097] The storage division module 304 is used to divide the Milvus database into several storage levels according to a preset standard and calculate the storage performance indicators of each storage level;
[0098] The data storage module 305 is used to calculate the data popularity of the knowledge vector data and store the knowledge vector data into the corresponding storage level of the Milvus database according to the data popularity and the storage performance indicators.
[0099] In one embodiment, the data processing module 302 is specifically used to remove redundant information in the text information by using a tag cleaning function and generate text redundancy-removed information; use a word segmentation tool to perform word segmentation processing on the text redundancy-removed information and generate standardized text information.
[0100] In one embodiment, the vector conversion module 303 is specifically used to convert the standardized text information into vector-formatted data by using the Sentence-BERT model and generate knowledge vector data; the calculation method of the data converted into vector format includes:
[0101] Vectorize(W) = S-BERT(W).
[0102] Among them, S-BERT() represents the Sentence-BERT model, Vectorize() represents the knowledge vector data after vector conversion, and W represents the set of words after word segmentation processing.
[0103] In one embodiment, the calculation method of the storage performance indicator of each storage level in the storage division module 304 includes:
[0104] P(L i ) = αS i +βC i ;
[0105] Among them, L i represents the i-th storage level, P(L i ) represents the storage performance indicator of the i-th storage level, S i represents the storage speed of the i-th storage level, C iRepresents the storage cost of the i-th storage hierarchy, where α and... represent constants.
[0106] In one embodiment, the calculation method of the data popularity of the knowledge vector data in the data storage module 305 includes:
[0107]
[0108] Among them, T represents the knowledge vector data for which the data popularity needs to be calculated, w k Represents the weight of the k-th storage hierarchy, I k Represents the storage performance index of the k-th storage hierarchy, m represents the total number of knowledge vector data for which the data popularity is calculated, and H(T) represents the data popularity of the knowledge vector data.
[0109] In one embodiment, the data storage module 305 is further configured to, when obtaining new knowledge data, calculate the similarity between the new knowledge data and the knowledge data already stored in the Milvus database, and compare the calculated similarity with a preset threshold; if the calculated similarity is greater than the preset threshold, then do not store the new knowledge data, or update the knowledge data already stored in the Milvus database with the new knowledge data; if the calculated similarity is not greater than the preset threshold, then process the new knowledge data and store it in the corresponding storage hierarchy in the Milvus database; establish a deduplication log and record the deduplication information in the deduplication log.
[0110] In one embodiment, the data storage module 305 is further configured to, when receiving a user query request, convert the user request into a vector representation and generate a vector query request; extract keywords according to the vector query request and generate a keyword set; perform a query operation in the Milvus database according to the keyword set and store the historical query record; collect user feedback information using a feedback collection function and dynamically adjust the similarity calculation weight according to the historical query record and the user feedback information.
[0111] The knowledge data storage and management device provided by the embodiments of the present application can be applied to the knowledge data storage and management method provided in the above embodiments. For related details, refer to the above method embodiments. The implementation principle and technical effects are similar and will not be elaborated here.
[0112] It should be noted that when the knowledge data storage and management device provided in the embodiments of the present application performs knowledge data storage and management, only the above-mentioned division of each functional module / functional unit is used as an example. In actual applications, the above functions can be assigned to different functional modules / functional units according to needs, that is, the internal structure of the knowledge data storage and management device is divided into different functional modules / functional units to complete all or part of the functions described above. In addition, the implementation manner of the knowledge data storage and management method provided in the above method embodiments and the implementation manner of the knowledge data storage and management device provided in this embodiment belong to the same concept. For the specific implementation process of the knowledge data storage and management device provided in this embodiment, please refer to the above method embodiments, which will not be elaborated here.
[0113] The embodiments of the present application also disclose a computer device.
[0114] Specifically, as Figure 4 shown, the computer device can be a desktop computer, a laptop computer, a palm computer, a cloud server, and other computer devices. The computer device may include, but is not limited to, a processor and a memory. Among them, the processor and the memory can be connected through a bus or other means. Among them, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, graphics processing units (GPUs), embedded neural network processors (NPUs) or other dedicated deep learning coprocessors, discrete gates or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0115] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the methods in the above method embodiments. The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0116] The embodiments of the present application also disclose a computer-readable storage medium.
[0117] Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the methods in the above method embodiments are implemented. Those skilled in the art can understand that to implement all or part of the processes in the methods of the above embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0118] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications without creative contributions to this embodiment according to needs after reading this specification, but as long as it is within the scope of the claims of the present invention, it is protected by the patent law.
Claims
1. A method for storing and managing knowledge data, characterized in that: The method is applied to a knowledge data storage and management system, the knowledge data storage and management system includes a Milvus database, and the method includes: Acquire knowledge data to be stored, and extract text information of the knowledge data to be stored; Preprocessing the text information and generating standardized text information; Converting the standardized text information into data in vector format and generating knowledge vector data; Dividing the Milvus database into several storage levels according to preset standards, and calculating the storage performance index of each storage level; The data heat of the knowledge vector data is calculated, and according to the data heat and the storage performance index, the knowledge vector data is stored in a storage level corresponding to the Milvus database.
2. The method according to claim 1, characterized in that: The preprocessing of the text information and generating standardized text information includes: Using a tag cleaning function to remove redundant information in the text information, and generating text de-redundant information; A word segmentation tool is used to remove redundant information from the text and perform word segmentation processing to generate standardized text information.
3. The method according to claim 1, characterized in that: The step of converting the standardized text information into data in a vector format and generating knowledge vector data comprises: The standardized text information is converted into vector format data using the Sentence-BERT model, and knowledge vector data is generated; The calculation methods for data converted into vector format include: Vectorize(W) = S-BERT(W); Among them, S-BERT() represents the Sentence-BERT model, Vectorize() represents the knowledge vector data after vector conversion, and W represents the word set after word segmentation processing.
4. The method according to claim 1, characterized in that: The calculation method of the storage performance index of each storage level includes: P(L i )=αS i +βC i ; Among them, L i represents the i-th storage level, P(L i ) represents the storage performance index of the i-th storage level, S i represents the storage speed of the i-th storage level, C i represents the storage cost of the i-th storage level, and α and β represent constants.
5. The method according to claim 1, characterized in that: The calculation method of the data heat of the knowledge vector data includes: Among them, T represents the knowledge vector data whose data heat needs to be calculated, w k represents the weight of the kth storage level, I k represents the storage performance index of the kth storage level, m represents the total number of knowledge vector data for calculating data heat, and H(T) represents the data heat of the knowledge vector data.
6. The method according to claim 1, characterized in that: After storing the knowledge vector data in the storage level corresponding to the Milvus database, the method further includes: When new knowledge data is acquired, the similarity between the new knowledge data and the knowledge data stored in the Milvus database is calculated, and the calculated similarity is compared with a preset threshold; If the calculated similarity is greater than a preset threshold, the new knowledge data is not stored, or the new knowledge data is used to update the knowledge data stored in the Milvus database; If the calculated similarity is not greater than a preset threshold, the new knowledge data is processed and stored in a corresponding storage level in the Milvus database; A deduplication log is created, and deduplication information is recorded in the deduplication log.
7. The method according to claim 1, characterized in that: After storing the knowledge vector data in the storage level corresponding to the Milvus database, the method further includes: When receiving a user query request, convert the user request into a vector representation and generate a vector query request; Extracting keywords according to the vector query request and generating a keyword set; Perform a query operation in the Milvus database according to the keyword set, and store historical query records; A feedback collection function is used to collect user feedback information, and the similarity calculation weight is dynamically adjusted according to the historical query records and the user feedback information.
8. A knowledge data storage and management device, characterized in that: The device is applied to a knowledge data storage and management system, the knowledge data storage and management system includes a Milvus database, and the device includes: A data acquisition module (301), used to acquire knowledge data to be stored, and extract text information of the knowledge data to be stored; A data processing module (302), used for preprocessing the text information and generating standardized text information; A vector conversion module (303), used to convert the standardized text information into data in vector format and generate knowledge vector data; A storage partitioning module (304), used to partition the Milvus database into a plurality of storage levels according to a preset standard, and calculate a storage performance index of each storage level; The data storage module (305) is used to calculate the data heat of the knowledge vector data, and store the knowledge vector data in the storage level corresponding to the Milvus database according to the data heat and the storage performance index.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.