Knowledge segmentation retrieval system based on nltk

By developing a knowledge segmentation search system based on nltk, the problems of low accuracy, insufficient processing capabilities of long texts and coarse granularity in knowledge base search are solved, and the precise segmentation and efficient retrieval of long texts are achieved, which improves the efficiency and quality of knowledge management.

CN120123494AInactive Publication Date: 2025-06-10SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510268191.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems in the knowledge base search with low search accuracy, insufficient long text processing capabilities, and a coarse granularity of knowledge fragments.

Method used

Develop a knowledge segmentation search system based on nltk, including a knowledge upload segmentation module, a knowledge search module, a knowledge segmentation query module and a knowledge deletion module. The system realizes accurate segmentation and efficient retrieval of long text through intelligent segmentation, combination of multiple search methods, structured query and optimized data storage.

Benefits of technology

It realizes precise segmentation and efficient retrieval of long texts in the knowledge base, improves the efficiency and quality of knowledge management, and enhances user experience and system universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123494A_ABST
    Figure CN120123494A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge segmentation retrieval system based on nltk, and belongs to the technical field of text segmentation and retrieval, and the system comprises a knowledge uploading and segmenting module which is used for receiving and analyzing docx, pdf and txt files, segmenting the files into knowledge segments, and establishing an index; the knowledge retrieval module is used for receiving a query request of a user, quickly positioning the knowledge fragment with the highest similarity through the constructed index, and returning a retrieval result; the knowledge slice paragraph query module is used for providing a paragraph content query service after segmentation and carrying out sorting and returning according to paragraph serial numbers; and the knowledge deleting module is used for deleting a single knowledge file or the whole knowledge base and ensuring the consistency and integrity of the data. According to the method, accurate segmentation and efficient retrieval of long texts in the knowledge base can be realized, the knowledge management efficiency and quality are greatly improved, and convenient and accurate knowledge services are provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text segmentation and retrieval based on natural language processing tools, and particularly to a knowledge segmentation and retrieval system based on nltk. Background Art

[0002] With the rapid development of information technology, the construction and utilization of knowledge bases have become key requirements in multiple fields such as education, scientific research, and enterprises. Knowledge bases contain a large amount of text information. How to efficiently and accurately retrieve the knowledge fragments required by users has become an urgent problem to be solved. Traditional information retrieval systems usually rely on keyword matching, and this method has the following limitations when dealing with long texts and complex queries:

[0003] Low retrieval accuracy: Keyword-based retrieval systems often cannot accurately understand the semantics of user queries, resulting in a large amount of irrelevant information in the retrieval results.

[0004] Insufficient ability to process long texts: Long texts are quite common in knowledge bases, but traditional retrieval systems often cannot effectively segment long texts when dealing with them, resulting in low retrieval efficiency.

[0005] Coarse granularity of knowledge fragments: Traditional retrieval systems usually return the entire document or paragraph instead of knowledge fragments with independent meanings, and users need to screen through a large amount of information by themselves.

[0006] To solve the above problems, researchers have begun to explore the use of natural language processing (NLP) technology to improve the accuracy and efficiency of text retrieval. Among them, nltk (Natural Language Toolkit) is a Python toolkit widely used in the NLP field, which provides rich text processing functions such as word segmentation, part-of-speech tagging, and syntactic analysis. Although nltk has significant advantages in text processing, there is still a lack of a systematic method in the existing technologies that can effectively combine the advanced text processing capabilities of nltk with the segmentation and retrieval requirements of knowledge bases.

[0007] Therefore, it is necessary to research and develop a knowledge segmentation and retrieval system based on nltk to achieve precise segmentation and efficient retrieval of long texts in knowledge bases, and improve the efficiency and quality of knowledge management. Summary of the Invention

[0008] To solve the above technical problems, the present invention provides a knowledge segmentation and retrieval system based on nltk, which can effectively solve the problems of low accuracy, insufficient ability to process long texts, and coarse granularity of knowledge fragments existing in the prior art in knowledge base retrieval.

[0009] The technical solution of the present invention is:

[0010] A knowledge segmentation and retrieval system based on NLTK, including

[0011] a knowledge upload and segmentation module, a knowledge retrieval module, a knowledge segmented paragraph query module, and a knowledge deletion module.

[0012] Among them,

[0013] The knowledge upload and segmentation module can quickly upload, parse, and segment docx, pdf, and txt files, and perform intelligent segmentation according to the text length and paragraph characteristics to ensure that important information is properly processed during the segmentation process.

[0014] The knowledge retrieval module provides four retrieval methods, combining traditional text retrieval and vector retrieval based on deep learning, which improves the accuracy and efficiency of retrieval. This module uses an efficient similarity calculation algorithm to ensure the quick return of retrieval results, and automatically adjusts the score range according to user feedback and system performance to adapt to retrieval needs in different scenarios.

[0015] The knowledge segmented paragraph query module provides a query service for the content of the segmented paragraphs, allowing users to filter according to multiple attributes, which improves the accuracy of the query. This module returns the query paragraph content in a structured form and provides metadata information of the text, enhancing the user experience.

[0016] The knowledge deletion module supports deleting individual knowledge files and the entire knowledge base, and ensures data consistency and integrity. This module implements a batch deletion function, improves the convenience of operation, and retains the log records of deletion operations for subsequent data recovery and audit tracking.

[0017] Furthermore,

[0018] The knowledge upload and segmentation module includes:

[0019] A file storage and support function that supports file version control and data backup;

[0020] A file segmentation strategy that performs intelligent segmentation according to the text length and paragraph characteristics and retains the logical structure of the document;

[0021] A text parsing function that uses a parsing library and technology to improve the accuracy and efficiency of parsing;

[0022] A paragraph recognition function that recognizes special symbols for paragraph segmentation and retains non-text elements such as charts and formulas.

[0023] Among them

[0024] File Storage and Support: Use the Minio object storage service, which supports high-concurrency reading and writing and three mainstream file types: docx, pdf, and txt; adopt a load balancing and redundant storage mechanism;

[0025] File Splitting Strategy: Before splitting, download the target file from the Minio storage bucket and use a file reading mechanism; based on the text length and paragraph features, add a function to recognize the document structure; in addition, it also has the ability to intelligently recognize non-text elements such as charts and formulas;

[0026] Text Parsing Function: Distinguish the format according to the file name suffix and perform specialized text parsing for different formats; use parsing libraries and technologies for different formats of files; when processing PDF files, recognize and retain the original format and layout;

[0027] Paragraph Recognition: Recognize special symbols for paragraph splitting; adopt different paragraph recognition strategies for different file types;

[0028] Index Construction and Storage: The split content is associated with a unique identifier to form an index and uploaded to the vector knowledge base; at the same time, the split paragraph text, knowledge base information, and paragraph numbers are stored in the MySQL database for subsequent retrieval and query; when constructing the index, a combination of inverted index and forward index is used; at the same time, through transaction management and logging, the consistency and traceability of the data upload and storage process are ensured.

[0029] Furthermore,

[0030] The knowledge retrieval module includes:

[0031] Several retrieval methods, combining traditional text retrieval and vector retrieval based on deep learning;

[0032] Similarity matching function, using several similarity calculation methods and machine learning models for weight adjustment;

[0033] Score normalization processing, introducing a dynamic threshold adjustment mechanism to adapt to retrieval requirements in different scenarios.

[0034] Among them,

[0035] Retrieval Methods: Provide four retrieval methods, including bge, es, bge_es_reranker, and es_reranker, to meet retrieval requirements in different scenarios; on the basis of the original retrieval methods, add a semantic retrieval function based on deep learning, which can more accurately understand the user's query intention;

[0036] Similarity matching: Retrieve the top ten knowledge paragraphs with the highest similarity through the content of the uploaded file paragraphs and URL information; adopt a similarity calculation algorithm to ensure the rapid return of retrieval results; at the same time, filter out the retrieval results with the same URL as the uploaded file; adopt a similarity calculation method, combined with a machine learning model for weight adjustment to optimize the sorting of retrieval results.

[0037] Score normalization: Normalize the scores, and the calculation formula is: normalization_score = (score - min_val) / (max_val - min_val) to ensure the consistency and comparability of scores; introduce a dynamic threshold adjustment mechanism to automatically adjust the score range according to user feedback and system performance to meet the retrieval needs in different scenarios.

[0038] Furthermore,

[0039] The knowledge fragment paragraph query module includes:

[0040] Data query function, which realizes combined query of several conditions to improve the accuracy of query;

[0041] Information return function, which displays the text content and its metadata information, and shows the logical relationship between paragraphs through a visual interface.

[0042] Among them

[0043] The knowledge fragment paragraph query module includes:

[0044] Data query: Utilize the paragraph information stored in the MySQL database to query and sort according to the paragraph numbers; adopt optimized query statements and index strategies; realize combined query of multiple conditions, allowing users to filter according to several attributes;

[0045] Information return: Return the retrieved paragraph content in a structured form for easy viewing and analysis by users; when returning the query results, not only display the text content, but also provide the metadata information of the text, and show the logical relationship between paragraphs through a visual interface to enhance the user experience.

[0046] Furthermore,

[0047] The knowledge deletion module includes:

[0048] Batch deletion function, allowing multiple files or knowledge bases to be deleted at one time;

[0049] Permission control mechanism to ensure that only authorized users can perform deletion operations;

[0050] Audit and backup function, retaining the log records of deletion operations for easy data recovery and audit tracking.

[0051] Among them

[0052] Deletion operation: Support deleting a single knowledge file and the entire knowledge base; In the deletion operation, the corresponding text paragraphs of the vector knowledge base and the MySQL data will be deleted to ensure data consistency; At the same time, a permission control mechanism is introduced to ensure that only authorized users can perform the deletion operation;

[0053] Deletion judgment: When deleting a single knowledge file, first judge whether the file exists in MySQL; If it does not exist, it is considered that the file has been deleted and the deletion is returned successfully; If it exists, the deletion operation is performed, and at the same time, the corresponding data in the vector knowledge base is checked to ensure data integrity and consistency; Before deletion, detailed auditing and backup are performed to ensure the reversibility of the deletion operation; For the deleted files, the system retains log records for subsequent data recovery and audit tracking.

[0054] The beneficial effects of the present invention are

[0055] Through the collaborative work of the above four modules, the present invention can achieve precise segmentation and efficient retrieval of long texts in the knowledge base, greatly improving the efficiency and quality of knowledge management. The present invention has the following effects:

[0056] 1. Innovative text segmentation strategy: Combining text length and paragraph features to achieve intelligent segmentation and improve segmentation accuracy.

[0057] 2. Efficient retrieval algorithm: Combining multiple retrieval methods to improve the accuracy and efficiency of retrieval.

[0058] 3. Optimized data storage and query: Adopting an efficient data storage and query mechanism to ensure the stability and query speed of the knowledge base. Description of the drawings

[0059] Figure 1 is a schematic diagram of the working process of the knowledge upload and segmentation module;

[0060] Figure 2 is a schematic diagram of the working process of the knowledge retrieval module;

[0061] Figure 3 is a schematic diagram of the working process of the knowledge segment query module;

[0062] Figure 4 is a schematic diagram of the working process of the knowledge deletion module. Specific implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0064] The present invention provides a knowledge segmentation and retrieval system based on NLTK, and the specific objectives are as follows:

[0065] Improve retrieval accuracy: By leveraging the natural language processing capabilities of NLTK, deeply understand the semantics of user queries, achieve more accurate knowledge matching, and reduce interference from irrelevant information.

[0066] Optimize long text processing: Design an effective text segmentation method, use NLTK to accurately segment long texts, and improve the processing capacity and efficiency of the retrieval system.

[0067] Refine the granularity of knowledge fragments: Segment the text in the knowledge base into knowledge fragments with independent meanings, making the retrieval results more refined and facilitating users to directly obtain the required information.

[0068] Enhance the user experience: Through the above technical means, enhance the user experience in the knowledge retrieval process and reduce the time and energy investment of users in screening information.

[0069] Enhance system versatility: Ensure that the system can adapt to knowledge bases of different types and scales, and provide efficient knowledge retrieval services for fields such as education, scientific research, and enterprises.

[0070] In summary, the objective of the present invention is to develop an efficient and accurate knowledge segmentation and retrieval system by combining the natural language processing technology of NLTK, so as to promote the automation and intelligence of knowledge management and meet the high-efficiency and accuracy requirements of modern society for knowledge retrieval.

[0071] The present invention includes four core modules. The following is a detailed elaboration of the technical solutions, implementation methods, and innovation points of each module.

[0072] 1. Knowledge upload and segmentation module:

[0073] This module aims to achieve the rapid upload, parsing, and segmentation of text in the knowledge base. The technical solutions are as follows:

[0074] File Storage and Support: The Minio object storage service is adopted. This service supports high-concurrency reading and writing, ensuring the stability and scalability of the knowledge base. The system supports three mainstream file types, namely docx, pdf, and txt, ensuring the diversity and compatibility of the knowledge base and meeting the needs of different users. The integration of the Minio object storage service not only supports mainstream file types but also takes into account file version control and data backup to ensure the stability and data security of the knowledge base. The system employs a load balancing and redundant storage mechanism to cope with the risks of high-concurrency access and data loss.

[0075] File Segmentation Strategy: Intelligent segmentation is performed based on text length and paragraph features (such as line breaks, first-line indentation, title numbers, blank lines, etc.). Before segmentation, the system downloads the target file from the Minio bucket and reduces the time consumption during the upload and download processes through an efficient file reading mechanism. Based on text length and paragraph features, the system adds a function for document structure recognition, such as tables of contents and chapter titles, making the segmentation more in line with the logical structure of the document. In addition, the system also has the ability to intelligently identify non-text elements such as charts and formulas to ensure that these important pieces of information are properly handled during segmentation.

[0076] Text Parsing: The format is distinguished according to the file name suffix, and specialized text parsing is performed for different formats. For example, for docx files, the system uses a specialized parsing library to extract the text content; for PDF documents, the system considers their layout characteristics, traverses the content of each page, identifies the layout and structure of the text, and determines line breaks by parsing the text content, the size and position of the first and last characters of each line to ensure the correct segmentation of the text. For files of different formats, the system uses a variety of parsing libraries and technologies, such as Apache Tika, PDFBox, etc., to improve the accuracy and efficiency of parsing. Especially when processing PDF files, the system can identify and retain the original format and layout of the text, providing users with more accurate segmentation results.

[0077] Paragraph Recognition: The system can recognize a variety of special symbols (such as \n\n, \n, \r\n, \r, \t, 、 , \par, <w:p>Perform paragraph segmentation for different types of documents (such as etc.). The system adopts different paragraph recognition strategies to ensure the accuracy of the segmentation results. The system not only recognizes traditional paragraph delimiters but also adds the recognition of special paragraph structures in Chinese texts, such as dialogues and quotations, making the segmented text fragments more semantically complete.

[0078] Index construction and storage: The segmented content is associated with a unique identifier (UUID) to form an index and uploaded to the vector knowledge base. At the same time, the segmented paragraph text, knowledge base information, and paragraph numbers are stored in a MySQL database for subsequent retrieval and query. The system adopts efficient database operations to ensure the real-time and consistency of data storage. When constructing the index, the system combines inverted index and forward index to improve the flexibility and speed of retrieval. At the same time, the system ensures the consistency and traceability of the data upload and storage processes through transaction management and logging.

[0079] 2. Knowledge retrieval module:

[0080] This module is responsible for implementing efficient retrieval of the vector knowledge base, and its technical solutions are as follows:

[0081] Retrieval methods: Provide four retrieval methods, including bge, es, bge_es_reranker, and es_reranker, to meet the retrieval needs in different scenarios. These retrieval methods combine traditional text retrieval and deep learning-based vector retrieval, improving the accuracy and efficiency of retrieval. On the basis of the original retrieval methods, the system adds a deep learning-based semantic retrieval function, which can more accurately understand the user's query intention and improve the relevance of retrieval.

[0082] Similarity matching: Retrieve the top ten knowledge paragraphs with the highest similarity through the uploaded file paragraph content and URL information. The system adopts an efficient similarity calculation algorithm to ensure the rapid return of retrieval results. At the same time, filter out the retrieval results with the same URL as the uploaded file to ensure the diversity and accuracy of retrieval results. The system adopts multiple similarity calculation methods, such as cosine similarity and Jaccard similarity, and combines machine learning models for weight adjustment to optimize the sorting of retrieval results.

[0083] Score Normalization: Since the similarity score ranges returned by different retrieval methods are different, the system will perform normalization on the scores. The calculation formula is: normalization_score = (score - min_val) / (max_val - min_val), to ensure the consistency and comparability of the scores. This step is crucial for users to understand the retrieval results and conduct subsequent analysis. To improve the accuracy of score normalization, the system introduces a dynamic threshold adjustment mechanism, which automatically adjusts the score range according to user feedback and system performance to adapt to the retrieval needs in different scenarios.

[0084] 3. Knowledge Segment Paragraph Query Module:

[0085] This module aims to provide query services for the segmented paragraph content. The technical solution is as follows:

[0086] Data Query: Utilize the paragraph information stored in the MySQL database to query and sort according to the paragraph numbers. The system adopts optimized query statements and indexing strategies to ensure the rapid response of query operations. The system implements combined multi-condition queries, allowing users to filter according to multiple attributes such as file name, author, date, etc., improving the accuracy of the query.

[0087] Information Return: The system returns the queried paragraph content in a structured form for easy viewing and analysis by users. The returned results include paragraph text, paragraph numbers, file information to which they belong, etc., providing users with a comprehensive knowledge query service. When returning the query results, the system not only displays the text content but also provides metadata information of the text, such as source, creation time, etc., and visualizes the logical relationships between paragraphs through a visual interface to enhance the user experience.

[0088] 4. Knowledge Deletion Module:

[0089] This module is responsible for the deletion operation of files in the knowledge base. The technical solution is as follows:

[0090] Deletion Operation: Supports deleting a single knowledge file and the entire knowledge base. During the deletion operation, the system will execute the deletion of the corresponding vector knowledge base text paragraphs and MySQL data to ensure data consistency. The system implements a batch deletion function, allowing users to delete multiple files or knowledge bases at one time, improving the convenience of the operation. At the same time, the system introduces a permission control mechanism to ensure that only authorized users can perform deletion operations.

[0091] Deletion Judgment: When deleting a single knowledge file, the system first checks whether the file exists in MySQL. If it does not exist, it is considered that the file has been deleted and a successful deletion is returned. If it exists, the deletion operation is executed, and the corresponding data in the vector knowledge base is checked to ensure data integrity and consistency. Before deletion, the system conducts a detailed audit and backup to ensure the reversibility of the deletion operation. For deleted files, the system retains log records for subsequent data recovery and audit tracking.

[0092] During operation, the operation steps are as follows:

[0093] Upload the knowledge file to the system;

[0094] The system automatically splits the file and builds an index;

[0095] The user enters a query request;

[0096] The system retrieves the knowledge fragment with the highest similarity according to the query request;

[0097] Return and display the retrieval results;

[0098] The user deletes the knowledge file or knowledge base as needed.

[0099] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.< / w:p>

Claims

1. A knowledge segmentation and retrieval system based on nltk, characterized in that: include The knowledge upload and segmentation module is used to receive and parse docx, pdf and txt files, segment the files into independent knowledge fragments, and create indexes; The knowledge retrieval module is used to receive user query requests, quickly locate the knowledge fragment with the highest similarity through the constructed index, and return the retrieval results; The knowledge slicing paragraph query module is used to provide segmented paragraph content query services and sort and return them according to the paragraph sequence number; The knowledge deletion module is used to delete a single knowledge file or the entire knowledge base and ensure the consistency and integrity of the data.

2. The system according to claim 1, characterized in that The knowledge uploading and segmenting module includes: File storage and support functions, supporting file version control and data backup; Document segmentation strategy: intelligent segmentation based on text length and paragraph features, while retaining the logical structure of the document; Text parsing function uses parsing library and technology to improve the accuracy and efficiency of parsing; The paragraph recognition function recognizes special symbols for paragraph segmentation and retains non-text elements such as charts and formulas.

3. The system according to claim 2, characterized in that File storage and support: Adopts Minio object storage service, supports high concurrent reading and writing, supports three mainstream file types: docx, pdf and txt; adopts load balancing and redundant storage mechanism; File segmentation strategy: Before segmentation, the target file is downloaded from the Minio bucket and read through the file reading mechanism; based on the text length and paragraph characteristics, the document structure recognition function is added; in addition, it also has the ability to intelligently recognize non-text elements such as charts and formulas; Text parsing function: distinguish formats according to file name suffixes, and perform specialized text parsing for different formats; use parsing libraries and technologies for files of different formats; identify and retain the format and layout of the original text when processing PDF files; Paragraph recognition: recognize special symbols for paragraph segmentation; use different paragraph recognition strategies for different file types; Index construction and storage: The segmented content is associated with a unique identifier to form an index, which is then uploaded to the vector knowledge base. At the same time, the segmented paragraph text, knowledge base information, and paragraph sequence number are stored in the MySQL database for subsequent retrieval and query. A combination of inverted index and forward index is used when building the index. At the same time, transaction management and logging are used to ensure the consistency and traceability of the data upload and storage process.

4. The system according to claim 1, characterized in that The knowledge retrieval module comprises: Several search methods, combining traditional text search and deep learning-based vector search; Similarity matching function, using several similarity calculation methods and machine learning models for weight adjustment; Score normalization is performed and a dynamic threshold adjustment mechanism is introduced to adapt to retrieval requirements in different scenarios.

5. The system according to claim 4, characterized in that Retrieval method: Four retrieval methods are provided, including bge, es, bge_es_reranker, and es_reranker, to meet the retrieval needs in different scenarios; on the basis of the original retrieval method, a semantic retrieval function based on deep learning is added to more accurately understand the user's query intention; Similarity matching: retrieve the top ten knowledge paragraphs with the highest similarity through the uploaded file paragraph content and URL information; use similarity calculation algorithm to ensure the rapid return of search results; at the same time, filter out search results with the same URL as the uploaded file; The similarity calculation method is used, combined with the machine learning model for weight adjustment to optimize the ranking of search results; Score normalization: The scores will be normalized using the formula: normalization_score = (score-min_val) / (max_val-min_val) to ensure consistency and comparability of the scores. A dynamic threshold adjustment mechanism is introduced to automatically adjust the score range based on user feedback and system performance to meet retrieval needs in different scenarios.

6. The system according to claim 1, characterized in that The knowledge slice paragraph query module includes: Data query function, realizes query based on several conditions and improves the accuracy of query; The information return function displays the text content and its metadata information, and shows the logical relationship between paragraphs through a visual interface.

7. The system according to claim 6, characterized in that The knowledge slice paragraph query module includes: Data query: Use the paragraph information stored in the MySQL database to query and sort according to the paragraph sequence number; use optimized query statements and index strategies; implement multi-condition combination query, allowing users to filter according to several attributes; Information return: The queried paragraph content is returned in a structured form for easy viewing and analysis by users. When returning query results, not only the text content is displayed, but also the metadata information of the text is provided, and the logical relationship between paragraphs is displayed through a visual interface to enhance the user experience.

8. The knowledge segmentation and retrieval system according to claim 1, characterized in that: The knowledge deletion module includes: Batch deletion function allows deleting several files or knowledge bases at one time; Permission control mechanism to ensure that only authorized users can perform deletion operations; Audit and backup functions keep log records of deletion operations to facilitate data recovery and audit tracking.

9. The system according to claim 8, characterized in that Deletion operation: supports deleting a single knowledge file and the entire knowledge base. In the deletion operation, the corresponding vector knowledge base text paragraphs and MySQL data will be deleted to ensure data consistency. At the same time, the permission control mechanism is introduced to ensure that only authorized users can perform the deletion operation. Deletion judgment: When deleting a single knowledge file, first determine whether the file exists in MySQL; if not, the file is considered to have been deleted and a deletion success message is returned; if it exists, the deletion operation is performed and the corresponding data in the vector knowledge base is checked to ensure data integrity and consistency; Before deletion, detailed audit and backup are performed to ensure the reversibility of the deletion operation; for deleted files, the system retains log records to facilitate subsequent data recovery and audit tracking.

Citation Information

Cited By

  • SVN cross-modal retrieval method and system based on knowledge base

    CN120372065A

  • A knowledge base-based SVN cross-modal retrieval method and system

    CN120372065B

  • Knowledge base construction method and device, electronic equipment and storage medium

    CN120872980A