Data processing system and method based on vector database and large model
Through the vector database and the data processing system of large models, the problem of semantic information extraction of PPT content is solved, the automatic knowledge extraction and efficient retrieval of PPT files is realized, professional answers are generated, and the utilization rate of internal documents of the enterprise is improved.
Patent Information
- Application Number
- CN202510649393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional methods cannot effectively extract semantic information of PPT content, and the information is fragmented and the retrieval efficiency is low, making it difficult to quickly obtain the required knowledge.
The data processing system based on vector database and large model is adopted, and the document content is converted into content vectors through the content vectorization module. The vector database module is used to build an index. The knowledge extraction module calls the large model for matching and optimization. The relevant recommended module sets a similarity threshold to output relevant results.
It realizes automatic knowledge extraction and efficient retrieval of PPT files, generates professional answers, reduces the time cost of users to obtain information, and improves the utilization rate of internal documents of the enterprise.
Smart Images

Figure CN120492643A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data processing system and method based on a vector database and a large model. Background Art
[0002] Traditional methods cannot effectively extract the semantic information of PPT content. Information is fragmented and retrieval efficiency is low, making it difficult to quickly obtain the required knowledge.
[0003] Traditional text extraction techniques often only capture the textual content within PowerPoint presentations, while overlooking the crucial information contained in multimedia elements like images, charts, and animations. Furthermore, PowerPoint files typically contain multiple slides with complex logical relationships between them, making it difficult for traditional methods to establish effective connections, leading to gaps in information retrieval. Summary of the Invention
[0004] In view of this, the present invention proposes a data processing system and method based on a vector database and a large model, aiming to solve the problems existing in the current technology.
[0005] In one aspect, the present invention provides a data processing system based on a vector database and a large model, comprising: A content vectorization module, configured to extract document content based on a natural language processing algorithm and convert the document content into a content vector; A vector database module, configured to construct an index based on the content vector corresponding to the document content to obtain an index vector database; A knowledge extraction module is used to call the big model to vectorize the query question to obtain a question vector, match the question vector with the index vector database, obtain the document content corresponding to the most relevant content vector, and optimize it through the big model as a decision support result output; The related recommendation module is used to pre-set a minimum similarity threshold and output the document content corresponding to the content vector whose vector similarity between the question vector and the index vector database is higher than the minimum similarity threshold as a related result.
[0006] In some embodiments of the present application, the document content includes at least text, images, audio, and video of a PPT file.
[0007] In some embodiments of the present application, the natural language processing algorithm includes text preprocessing, feature extraction and vectorization.
[0008] In some embodiments of the present application, the content vectorization module converts the document content into a content vector through a word embedding model, and the content vector is a high-dimensional semantic vector.
[0009] In some embodiments of the present application, the knowledge acquisition module calculates the similarity between each content vector and the question vector through a nearest neighbor algorithm to obtain a vector similarity list, and obtains the content vector corresponding to the highest similarity according to the vector similarity list as the most relevant content vector.
[0010] In some embodiments of the present application, the relevant results do not include the decision support results.
[0011] In some embodiments of the present application, the related recommendation module is also used to obtain the document ID and document storage path of the document content corresponding to the related results, and establish a document link for the related results based on the document ID and document storage path and integrate the output.
[0012] In some embodiments of the present application, the knowledge acquisition module performs semantic analysis and summary on the most relevant content vector obtained through matching using a large model to generate the decision support result.
[0013] Compared with the existing technology, the beneficial effects of the present invention are: first, the present invention realizes the automatic knowledge extraction and efficient retrieval of PPT files, which can not only quickly locate relevant content, but also generate professional answers through semantic understanding, avoiding manual repeated searching and screening in traditional technologies; secondly, through efficient semantic retrieval and intelligent knowledge extraction, the system significantly reduces the time cost for users to obtain the required information, thereby improving the utilization rate of internal enterprise documents; in addition, the system can generate automated decision support solutions based on user questions, and attach relevant document links for further reference, providing strong support for the enterprise's knowledge management and business decision-making.
[0014] On the other hand, the present application also provides a data processing method based on a vector database and a large model, comprising the following steps: S1. Extract document content according to a natural language processing algorithm and convert the document content into a content vector; S2. Building an index based on the content vector corresponding to the document content in step S1 to obtain an index vector database; S3. Calling the big model to vectorize the query question to obtain a question vector, and matching the question vector with the index vector database in step S2 to obtain the document content corresponding to the most relevant content vector, and optimizing it through the big model as a decision support result output; S4. Preset a minimum similarity threshold, and output the document content corresponding to the content vector whose vector similarity between the question vector in step S3 and the index vector database is higher than the minimum similarity threshold as a relevant result.
[0015] It can be understood that the data processing method based on a vector database and a large model in the present application has the same beneficial effects as the above-mentioned data processing system based on a vector database and a large model, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 A structural block diagram of a data processing system based on a vector database and a large model provided by an embodiment of the present invention; Figure 2 A flowchart of a data processing method based on a vector database and a large model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the implementation regulations.
[0018] On the one hand, see Figure 1 This embodiment provides a data processing system based on a vector database and a large model, including: The content vectorization module is used to extract document content based on the natural language processing algorithm and convert the document content into content vectors; A vector database module is used to construct an index based on the content vector corresponding to the document content to obtain an index vector database; The knowledge extraction module is used to call the large model to vectorize the query question to obtain a question vector, match the question vector with the index vector database, obtain the document content corresponding to the most relevant content vector, and optimize it through the large model as a decision support result output; The related recommendation module is used to pre-set a minimum similarity threshold and output the document content corresponding to the content vector whose vector similarity between the question vector and the index vector database is higher than the minimum similarity threshold as the related result.
[0019] Specifically, in this embodiment, the vector database module preferably constructs an indexed vector database using the Chroma open-source vector database. The Chroma open-source vector database is compatible with embedding models such as text (OpenAI text-embedding-3-large), image (CLIP-ViT-L / 14), and audio (Whisper-v3), supporting cross-modal search (e.g., searching for images using text). Furthermore, the Chroma open-source vector database allows for dynamic selection of index types and similarity algorithms.
[0020] It is understood that in this embodiment, the content vectorization module extracts document content through a natural language processing algorithm and converts it into content vectors, ensuring that the vectors accurately reflect the text semantics. The vector database module uses an efficient data structure to achieve fast retrieval and storage. The knowledge extraction module uses a pre-trained language model, preferably the ChatGLM-7B model, to improve question understanding and matching accuracy. The related recommendation module dynamically adjusts the similarity threshold to adapt to the needs of different application scenarios, improving the relevance and accuracy of recommendations.
[0021] In a specific embodiment of the present application, the document content includes at least text, images, audio, and video of the PPT file.
[0022] Specifically, in this embodiment, by embedding a model in the Chroma open source vector database, text, images, and videos are mapped to the same vector space to achieve unified management of multimodal data.
[0023] In a specific embodiment of the present application, the natural language processing algorithm includes text preprocessing, feature extraction and vectorization.
[0024] It can be understood that in this embodiment, text preprocessing includes steps such as word segmentation and stop word removal, feature extraction is used to select meaningful words or phrases, and vectorization is to convert these features into numerical vectors for easy calculation.
[0025] In a specific embodiment of the present application, the content vectorization module converts the document content into a content vector through a word embedding model, and the content vector is a high-dimensional semantic vector.
[0026] Specifically, in this embodiment, when document content is converted into content vectors through a word embedding model, dynamic embedding is preferably performed through a bge-large-zh model.
[0027] It is understandable that the content vectorization module in this embodiment converts document content into high-dimensional semantic vectors through a word embedding model, allowing the document content to be better understood and processed by computers. By learning from large amounts of text data, the word embedding model can capture the complex semantic relationships between words and encode them into a vector space. The high-dimensional semantic vector not only retains the basic meaning of the words, but also reflects the subtle differences between words in different contexts, thus providing strong support for subsequent text analysis tasks (such as classification, clustering, retrieval, etc.). In addition, the high-dimensional semantic vector can also be used to calculate the similarity between sentences or documents.
[0028] In a specific embodiment of the present application, the knowledge acquisition module calculates the similarity between each content vector and the question vector using a nearest neighbor algorithm to obtain a vector similarity list, and obtains the content vector corresponding to the highest similarity according to the vector similarity list as the most relevant content vector.
[0029] Specifically, the question vector is a numerical vector representation of the user's question, allowing direct comparison between the question vector and the content vector to assess their similarity. The vector similarity list is a collection of similarities between all content vectors and the question vector, calculated using the nearest neighbor algorithm. The list is sorted by similarity to facilitate quick identification of the vector most relevant to the question. The nearest neighbor algorithm is a machine learning algorithm used to find the point (i.e., content vector) closest to the query point (i.e., question vector). Similarity is determined by comparing the distances between vectors based on Euclidean distance or other distance metrics.
[0030] It is understood that in this embodiment, the knowledge acquisition module uses the nearest neighbor algorithm to calculate the similarity between the content vector and the question vector, generating a vector similarity list. Then, by analyzing this list, the content vector with the highest similarity is selected as the most relevant answer.
[0031] Specifically, in this embodiment, the nearest neighbor algorithm is preferably FAISS, which uses the IndexFlatIP index structure to perform precise vector matching. The IndexFlatIP index structure uses the inner product (Inner Product) calculation method to normalize the content vector and convert it into a unit vector, ensuring that the inner product calculation is equivalent to cosine similarity. Dynamic embedding models (such as bge-base-zh, E5 series, MiniLM) are used to generate semantic vectors to improve retrieval relevance. After vectorization, all content vectors are directly added to the FAISS index in batches. The question vector is also normalized during query, and index.search is used for Top-K similarity retrieval. In small-scale scenarios, it is directly used to run on the CPU without the need to build complex indexes or perform quantization compression.
[0032] As you can see, FAISS provides an efficient method for storing and querying large numbers of vectors, while IndexFlatIP achieves precise matching through inner product calculations. Vector normalization ensures the accuracy of calculation results, while the dynamic embedding model enhances semantic understanding. In small-scale scenarios, FAISS can be run directly on the CPU, simplifying deployment.
[0033] In a specific embodiment of the present application, the relevant results do not include decision support results.
[0034] It is understandable that since the decision support result has the highest similarity with the question vector and has been output first, it will not be output repeatedly in the related results.
[0035] In a specific embodiment of the present application, the related recommendation module is also used to obtain the document ID and document storage path of the document content corresponding to the related results, and establish document links of the related results based on the document ID and document storage path and integrate the output.
[0036] Specifically, the document ID in this embodiment is a unique identifier for each document, used to quickly locate and retrieve specific documents, ensuring data consistency and accuracy. A document link is a hyperlink pointing to the document's storage path, allowing users to directly access the document's content by clicking on it. Document links make documents easier to access and share.
[0037] It can be understood that the relevant recommendation module in this embodiment can efficiently establish document links by obtaining document IDs and document storage paths, thereby achieving rapid access and integrated output of relevant documents, which not only improves the system's response speed, but also enhances the user experience, enabling users to find and use relevant information more easily.
[0038] See Figure 2 On the other hand, this embodiment also provides a data processing method based on a vector database and a large model, comprising the following steps: S1. Extract document content based on natural language processing algorithms and convert the document content into content vectors. S2. Build an index based on the content vector corresponding to the document content in step S1 to obtain an index vector database; S3. Call the big model to vectorize the query question to obtain a question vector, and match the question vector with the index vector database in step S2. The document content corresponding to the most relevant content vector is obtained and optimized through the big model as a decision support result output; S4. Preset a minimum similarity threshold, and output the document content corresponding to the content vector whose vector similarity between the question vector and the index vector database in step S3 is higher than the minimum similarity threshold as the relevant result.
[0039] It can be understood that the data processing method based on a vector database and a large model in the present application has the same beneficial effects as the above-mentioned data processing system based on a vector database and a large model, and will not be repeated here.
[0040] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0041] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0042] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A data processing system based on a vector database and a large model, characterized in that: include: A content vectorization module, configured to extract document content based on a natural language processing algorithm and convert the document content into a content vector; A vector database module, configured to construct an index based on the content vector corresponding to the document content to obtain an index vector database; A knowledge extraction module is used to call the big model to vectorize the query question to obtain a question vector, match the question vector with the index vector database, obtain the document content corresponding to the most relevant content vector, and optimize it through the big model as a decision support result output; The related recommendation module is used to pre-set a minimum similarity threshold and output the document content corresponding to the content vector whose vector similarity between the question vector and the index vector database is higher than the minimum similarity threshold as a related result.
2. The data processing system based on vector database and large model according to claim 1, characterized in that: The document content includes at least text, images, audio, and video of the PPT file.
3. The data processing system based on vector database and large model according to claim 1, characterized in that: The natural language processing algorithm includes text preprocessing, feature extraction and vectorization.
4. The data processing system based on vector database and large model according to claim 1, characterized in that: The content vectorization module converts the document content into a content vector through a word embedding model, and the content vector is a high-dimensional semantic vector.
5. The data processing system based on vector database and large model according to claim 1, characterized in that: The knowledge acquisition module calculates the similarity between each of the content vectors and the question vector using a nearest neighbor algorithm to obtain a vector similarity list, and obtains the content vector corresponding to the highest similarity according to the vector similarity list as the most relevant content vector.
6. The data processing system based on vector database and large model according to claim 1, characterized in that: The relevant results do not include the decision support results.
7. The data processing system based on vector database and large model according to claim 5, characterized in that: The related recommendation module is further configured to obtain the document ID and document storage path of the document content corresponding to the related results, and to establish document links of the related results based on the document ID and document storage path and integrate and output them.
8. The data processing system based on vector database and large model according to claim 1, characterized in that: The knowledge acquisition module performs semantic analysis and summary on the most relevant content vector obtained through matching through a large model to generate the decision support result.
9. A data processing method based on a vector database and a large model, characterized in that: A data processing system based on a vector database and a large model as described in any one of claims 1 to 8, comprising the following steps: S1. Extract document content according to a natural language processing algorithm and convert the document content into a content vector; S2. Building an index based on the content vector corresponding to the document content in step S1 to obtain an index vector database; S3. Calling the big model to vectorize the query question to obtain a question vector, and matching the question vector with the index vector database in step S2 to obtain the document content corresponding to the most relevant content vector, and optimizing it through the big model as a decision support result output; S4. Preset a minimum similarity threshold, and output the document content corresponding to the content vector whose vector similarity between the question vector in step S3 and the index vector database is higher than the minimum similarity threshold as a relevant result.