Intelligent question answering and text portable processing engine based on large model miniaturization technology

Through the combination of the quantitative framework and the model management scheduling engine, the miniaturization of large models is achieved, solving the limitations of large models on resource-constrained devices, and achieving efficient and intelligent Q&A and text processing on edge devices.

CN120104724APending Publication Date: 2025-06-06JIANGSU JINLING TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411874635.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Large models consume higher resources, limiting their application on resource-constrained or edge devices.

Method used

Through the quantitative framework, the floating-point number parameters in the model are converted into low-precision integer parameters, and the model loading and resource management are controlled on demand through the model management and scheduling engine to achieve miniaturization of large language models.

Benefits of technology

With limited hardware resources, efficient processing of large language models is achieved, solving the problem that the computing resources on traditional portable devices are limited and cannot fully utilize the large model, and providing intelligent question-and-answer and text processing services for accessing private domain data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104724A_ABST
    Figure CN120104724A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent question and answer and text portable processing engine based on a large model miniaturization technology, which is characterized in that a large model miniaturization processing unit converts floating-point number parameters in a model into low-precision integer parameters, and controls model loading and occupied resource management as required; the large-model miniaturization processing unit is connected with the large-model miniaturization application, and the large-model miniaturization application comprises an artificial intelligence question and answer module which enables a user to ask questions through natural languages and obtain accurate answers. And the natural language to SQL module is used for converting the natural language request of the user into a structured query language and returning a query result. And the RAG retrieval module is used for training private knowledge and providing retrieval questions for an internal knowledge base. And an auxiliary reading module based on a large model provides intelligent text understanding and abstract generation functions. According to the method, the problem that efficient question answering and text processing cannot be carried out by fully utilizing a large model due to the fact that computing resources are limited on traditional portable equipment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of intelligent intelligence and relates to an intelligent question-answering and text portable processing engine based on large model miniaturization technology. Background Art

[0002] In the field of artificial intelligence technology, large language models have been widely used in various fields, ranging from natural language processing to intelligent dialogue systems to information retrieval. However, although large models with higher parameters perform well in Internet applications, their high resource consumption limits their application on resource-constrained or edge devices. Therefore, it is particularly important to promote the development of large language model miniaturization technology. Summary of the invention

[0003] 1. Technical problems to be solved: Large models have high resource consumption, limiting their application on resource-constrained or edge devices.

[0004] 2. Technical solution: In order to solve the above problems, the present invention provides an intelligent question-answering and text portable processing engine based on large model miniaturization technology, including a large model miniaturization processing unit, the large model miniaturization processing unit converts floating-point parameters in the model into low-precision integer parameters through a quantization framework, and controls model loading and occupied resource management on demand through a model management scheduling engine. The large model miniaturization processing unit is connected to a large model miniaturization application and controls the large model miniaturization application, and the large model miniaturization application includes: Artificial Intelligence Question and Answer Module: enables users to ask questions in natural language and get accurate answers.

[0005] The natural language to SQL module converts the user's natural language request into structured query language and returns the query results.

[0006] RAG retrieval module implements the training of private knowledge and provides retrieval questions for the internal knowledge base.

[0007] The auxiliary reading module based on the large model provides intelligent text comprehension and summary generation functions.

[0008] The artificial intelligence question-and-answer module specifically provides a dialogue window for an online chat room, in which users communicate with the system or ask questions in natural language. The system generates corresponding answers based on the local large language model pre-training data to interact with the user and obtain accurate and detailed answers.

[0009] Furthermore, the artificial intelligence question-answering module does not require an Internet connection and relies on the natural language dialogue capabilities of a local large model.

[0010] Furthermore, the query results in the natural language to SQL module support returning chart modal information and support multiple display modes such as line charts, bar charts, and pie charts.

[0011] Furthermore, the training of private knowledge in the RAG retrieval module is specifically as follows: pre-processing data through text data extraction, text segmentation, and text vectorization steps to complete the pre-training of private knowledge in an internal specific domain.

[0012] Furthermore, after the user inputs a question, vector calculation, text recall, and rearrangement operations are used to retrieve information related to the question from the internally trained private knowledge base model, and the question and retrieval results are sent to the large model to generate an answer.

[0013] Furthermore, the large model-based auxiliary reading module uses text data extraction and prompt words to identify the summary and abstract generation of documents of tens of thousands of words, and uses the prompt words for generating outlines to convert the main content of the document into a mind map, deeply understand the text content, and capture the key information therein.

[0014] Furthermore, the large model-based auxiliary reading module assists users in further understanding the text, converts the entire content into multiple QA pairs using prompt words, and then provides a dialogue format to provide users with question-and-answer interaction with the document content.

[0015] 3. Beneficial effects: The present invention solves the problem that traditional portable devices cannot fully utilize large models for efficient question answering and text processing due to limited computing resources. The present invention can be widely used in government, sensitive industries and other fields to provide users with intelligent question answering and text processing services that access private domain data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a logical framework diagram of the present invention. DETAILED DESCRIPTION

[0017] The present invention is described in detail below with reference to the accompanying drawings.

[0018] An intelligent question-answering and text portable processing engine based on large model miniaturization technology, comprising a large model miniaturization processing unit, The large model miniaturization processing unit adopts the large model quantization deployment technology. On the one hand, it converts the floating-point parameters in the model into low-precision integer parameters through the quantization framework, and reduces the calculation complexity to achieve calculation acceleration, and realizes large language model parameter compression and precision reduction. On the other hand, it controls model loading and resource management on demand through the self-developed model management scheduling engine. On the whole, it reduces the storage space and computing resources required for the model while maintaining the model's performance on the task, and realizes efficient processing of language models under limited hardware resources.

[0019] like Figure 1 As shown, the large model miniaturization processing unit is connected to the large model miniaturization application and controls the large model miniaturization application, and the large model miniaturization application includes: Artificial Intelligence Question and Answer Module: enables users to ask questions in natural language and get accurate answers.

[0020] The natural language to SQL module converts the user's natural language request into structured query language and returns the query results.

[0021] RAG retrieval module implements the training of private knowledge and provides retrieval questions for the internal knowledge base.

[0022] The auxiliary reading module based on the large model provides intelligent text comprehension and summary generation functions.

[0023] In one embodiment, the artificial intelligence question-answering module provides a dialogue window similar to an online chat room, where users can talk to the system in natural language and ask questions. The system generates corresponding answers based on the pre-trained data of the local large language model to interact with the user, allowing the user to ask questions in natural language and obtain accurate and detailed answers. This process does not require networking and is completely dependent on the natural language dialogue capabilities of the quantized and accelerated local large model.

[0024] The natural language to SQL module allows users to make requests in the form of natural language, such as asking "How many players are there in the Chinese Basketball Association (CBA)?" The system can convert these requests into structured query language (SQL) so that the system can understand and process them.

[0025] In one embodiment, the natural language to SQL module uses natural language processing and database query technologies to convert the user's language request into a machine-executable query statement, thereby achieving accurate and efficient retrieval and analysis of data. At the same time, the query results support returning multi-modal information such as charts, and can support multiple display modes such as line charts, bar charts, and pie charts.

[0026] In one embodiment, the RAG retrieval module combines the retrieval enhancement generation (RAG) technology and the natural language processing (NLP) technology, pre-processes the data through steps such as text data extraction, text segmentation, and text vectorization, and completes the pre-training of internal specific domain private knowledge. After the user inputs the question, the system uses vector calculation, text recall, and rearrangement operations to quickly and accurately retrieve information related to the question from the internally trained private knowledge base model, and sends the question and the retrieval results to the large model to generate answers, providing efficient retrieval questions and answers for the internal knowledge base.

[0027] In one embodiment, the large model-based auxiliary reading module implements functions based on a large language model that supports long texts, uses text data extraction and prompt words to identify summaries and abstracts of documents with tens of thousands of words, and uses prompt words to generate outlines to convert the main content of the document into a mind map, deeply understand the text content, and capture the key information. In addition, the module can also assist users in further understanding the text, using prompt words to convert the entire content into multiple QA pairs, and then provide a dialogue format to provide users with question-and-answer interactions with the document content.

[0028] The present invention uses a large model miniaturization technology to implement intelligent question-answering and text processing methods on portable hardware devices. This solves the problem that on traditional portable hardware devices, due to limited computing resources, large models cannot be fully utilized for efficient question-answering and knowledge processing. The portable processing engine of this patent can be widely used in government, sensitive industries and other fields, providing users with intelligent question-answering and text processing services that access private domain data.

Claims

1. An intelligent question-answering and text portable processing engine based on large model miniaturization technology, characterized by: The system comprises a large model miniaturization processing unit, which converts floating point parameters in the model into low-precision integer parameters through a quantization framework, and controls model loading and resource management on demand through a model management scheduling engine. The large model miniaturization processing unit is connected to a large model miniaturization application and controls the large model miniaturization application. The large model miniaturization application comprises: AI Q&A module: enables users to ask questions in natural language and get accurate answers; The natural language to SQL module converts the user's natural language request into a structured query language and returns the query results; RAG retrieval module, which implements the training of private knowledge and provides retrieval questions and answers for the internal knowledge base; The auxiliary reading module based on the large model provides intelligent text comprehension and summary generation functions.

2. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 1, characterized in that: The artificial intelligence question-and-answer module specifically provides a dialogue window for an online chat room, in which users communicate with the system or ask questions in natural language. The system generates corresponding answers based on the local large language model pre-training data to interact with the user and obtain accurate and detailed answers.

3. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 2, characterized in that: The artificial intelligence question-answering module does not require an Internet connection and relies on the natural language conversation capabilities of a local large model.

4. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 1, characterized in that: The query results in the natural language to SQL module support returning chart modal information and support multiple display modes such as line charts, bar charts, and pie charts.

5. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 1, characterized in that: The training of private knowledge in the RAG retrieval module is specifically as follows: pre-processing data through text data extraction, text segmentation, and text vectorization steps to complete the pre-training of private knowledge in a specific internal field.

6. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 5, characterized in that: After the user enters a question, we use vector calculation, text recall, and rearrangement operations to retrieve information related to the question from the internally trained private knowledge base model, and send the question and retrieval results to the large model to generate an answer.

7. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 1, characterized in that: The large model-based auxiliary reading module utilizes text data extraction and prompt words to identify summaries and abstracts of documents of tens of thousands of words, and utilizes prompt words for generating outlines to convert the main contents of the document into mind maps, thereby deeply understanding the text content and capturing key information therein.

8. The intelligent question-answering and text portable processing engine based on large model miniaturization technology as claimed in claim 7, characterized in that: The large model-based auxiliary reading module assists users in further understanding the text, converts the entire content into multiple QA pairs using prompt words, and then provides a dialogue format to provide users with question-and-answer interaction with the document content.