Mobile terminal information extraction system and method based on OCR (Optical Character Recognition) and local large language model
By combining OCR with local large language models, using image preprocessing, multi-scale feature extraction, text optimization and quantum computing acceleration modules, the problem that OCR and large language models in the prior art cannot effectively process complex images and multimodal data, and efficient and accurate information refinement is achieved.
Patent Information
- Application Number
- CN202510325686.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
AI Technical Summary
When the prior art processes multimodal data, complex images and large-scale data, OCR and large-language models cannot be effectively combined, the computing efficiency is inefficient, and traditional computing architectures are difficult to meet the real-time response requirements.
Combining OCR and local large language model, accurate image-to-text transformation and intelligent screening of key information are achieved through image preprocessing, multi-scale feature extraction, text optimization and information screening, and quantum computing acceleration modules.
It significantly improves data processing speed and path optimization capabilities, ensures the efficiency of image data processing and text understanding accuracy, and improves the accuracy and efficiency of information refining.
Smart Images

Figure CN120279557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing and intelligent analysis, and specifically to a mobile information extraction system and method based on OCR and a local large language model. Background Art
[0002] In modern society, with the rapid development of information technology, we are faced with a large amount of text and image data every day. From emails, social media to business documents, contract agreements, almost every industry is filled with a vast amount of digital content. How to efficiently extract, process, and understand this information has become the key to improving work efficiency and decision-making quality.
[0003] In the prior art, OCR (Optical Character Recognition) technology is widely used to convert the text in images into editable text. OCR can extract the text information in the image, making it possible to digitally process the data. In the field of text analysis, large language models (such as GPT and BERT, etc.) have also been widely used, and they can effectively analyze the semantics of the text, perform semantic understanding, and generate text summaries. Based on these technologies, many enterprises and research institutions have been able to efficiently perform document analysis, information extraction, and data summarization, greatly improving the speed and accuracy of information processing.
[0004] However, the prior art still has some deficiencies when facing multi-modal data, complex images, and large-scale data; firstly, traditional OCR technology can only achieve text recognition and lacks the ability to deeply understand the content of the image, which results in a significant reduction in recognition effect in the case of blurred images and complex layouts; secondly, although large language models can perform semantic analysis and generate summaries, they cannot directly process image data, so the processing effect is often not ideal in scenarios where images and text are combined; in addition, existing text optimization and information screening technologies mostly rely on rules or keyword statistics, unable to intelligently screen out the most valuable information for the task, and the screening efficiency is low when facing a large amount of data; more importantly, when processing large-scale data, the computing power and real-time response ability of traditional computing architectures are difficult to meet the requirements, and the computing bottleneck severely limits the processing efficiency and system performance. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a mobile information extraction system and method based on OCR and a local large language model, which solves the problems that OCR and large language models in the prior art cannot be effectively combined to process complex images and multi-modal data, and the computing efficiency is low in large-scale data processing.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A mobile information extraction system based on OCR and a local large language model, comprising: An OCR recognition module for converting input image data into text data; A text optimization and information screening module for optimizing the text data output by OCR; A large language model analysis module for semantic understanding of the optimized text; A quantum computing acceleration and optimization module for optimizing the information transfer process between OCR and the large language model.
[0007] Preferably, the OCR recognition module includes: An image preprocessing unit for denoising, enhancing, and edge detecting the input image, specifically including: Performing multi-band decomposition on the input image using wavelet transform to remove noise; Extracting the edge information of the image through an edge detection algorithm to identify the text area; A multi-scale feature extraction unit for extracting image features at multiple scales through a convolutional neural network and wavelet transform; A text recognition unit for converting the extracted image features into text data.
[0008] Preferably, the text optimization and information screening module includes: A high-order entropy calculation unit for calculating the high-order entropy of the text data output by OCR and screening key information according to the entropy value, specifically including: Calculating the entropy value for each paragraph in the text and screening out the key information in the text based on the entropy value; Using information entropy to reduce redundant and irrelevant data and only retaining the parts that contribute significantly to the target task; An information gain analysis unit for performing information gain analysis on the text and selecting the parts that contribute significantly to the target task.
[0009] Preferably, the large language model analysis module includes: A semantic understanding unit for semantic understanding of the text according to the locally deployed large language model; A text summarization unit for generating structured or unstructured text summaries, specifically including: Calculating the weights of each part in the text through a self-attention mechanism; Calculating the core information of the text according to the weights and generating a structured or unstructured summary.
[0010] Preferably, the quantum computing acceleration and optimization module includes: A quantum Fourier transform unit for accelerating the frequency domain conversion of information; A quantum search algorithm unit for quickly finding the optimal data transmission path in the data stream.
[0011] Preferably, the image preprocessing unit is used to perform image denoising on the input image, and the denoising process includes: Performing multi-band decomposition on the input image using wavelet transform to remove noise; Extracting the edge information of the image through an edge detection algorithm to identify the text area.
[0012] Preferably, the multi-scale feature extraction unit includes: Performing layer-by-layer convolution operations on the input image through a convolutional neural network to extract local features at different scales; Extracting high-frequency and low-frequency information in the image through wavelet transform to facilitate the extraction of text features.
[0013] Preferably, the high-order entropy calculation unit of the text optimization and information screening module is used to calculate the high-order entropy of the text data output by OCR, and the high-order entropy calculation method includes: Calculating the entropy value for each paragraph of the text, and screening out the key information in the text based on the entropy value; Using information entropy to reduce redundant and irrelevant data, and only retaining the parts that contribute significantly to the target task.
[0014] Preferably, the semantic understanding unit of the large language model analysis module is used to perform semantic understanding on the text using a large language model based on the Transformer structure and generate a text summary, and the text summary generation method includes: Calculating the weights of each part in the text through the self-attention mechanism; Calculating the core information of the text based on the weights and generating a structured or unstructured summary.
[0015] Preferably, the mobile information extraction method based on OCR and the local large language model includes the following steps: S1. Transmitting the input image data to the OCR recognition module; S2. Performing image preprocessing on the image data, including denoising, enhancement, and edge detection processing; S3. Using the OCR recognition module to extract image features and convert them into text data; S4. Performing high-order entropy calculation and information gain analysis on the text data output by OCR, and screening out the most informative text parts; S5. Transmitting the optimized text to the large language model analysis module for semantic understanding and generating a summary; S6. During the transmission process between OCR and the large language model, using quantum computing to accelerate information flow and optimize the data transmission path; S7. Output the generated text summary or extracted key information.
[0016] The present invention provides a mobile information extraction system and method based on OCR and a local large language model. It has the following Beneficial effects: 1. Through the technical solution of combining the quantum computing acceleration and optimization module, the present invention significantly improves the data processing speed and path optimization ability. The application of the quantum Fourier transform and the quantum search algorithm enables more efficient information flow, reduces the processing delay commonly found in traditional computing methods, and is particularly prominent in large data scenarios.
[0017] 2. By deeply integrating OCR with the large language model, the present invention realizes the accurate conversion of images to text. This innovative integration overcomes the disadvantages in the prior art that OCR cannot effectively understand the content and the large language model cannot process images, ensuring the efficiency of image data processing and the accuracy of text understanding.
[0018] 3. The technical solutions of high-order entropy calculation and information gain analysis in the text optimization and information screening module effectively screen out key information and eliminate redundant content. This technological innovation makes the information extraction more accurate through an intelligent screening process, improving the processing efficiency of the subsequent analysis module. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the system architecture diagram of the present invention; Figure 2 is the module architecture diagram of the OCR recognition module of the present invention; Figure 3 is the module architecture diagram of the text optimization and information screening module of the present invention; Figure 4 is the module architecture diagram of the large language model analysis module of the present invention; Figure 5 is the module architecture diagram of the quantum computing acceleration and optimization module of the present invention; Figure 6 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the specification of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Please refer to the attached Figure 1 - attached Figure 5, the embodiments of the present invention provide a mobile information extraction system based on OCR and a local large language model, including: An OCR recognition module for converting the input image data into text data; The implementation of the OCR module is not limited to text extraction, but also involves image preprocessing, multi-scale feature extraction, and final text recognition. To ensure efficient and accurate text recognition under various complex conditions (such as low-quality images, handwritten text, complex layouts, etc.), the OCR module adopts multiple advanced technical means, including wavelet transform, convolutional neural network (CNN), recurrent neural network (RNN), etc.
[0022] Generally, the recognition accuracy of the OCR system is affected by the quality of the input image. Especially when the image contains noise or has low quality, the recognition effect will drop significantly. To solve this problem, the present invention uses an image preprocessing unit to perform multi-step processing on the input image. These steps include denoising, enhancement, and edge detection, etc., to ensure that the image can provide a clear input for subsequent OCR recognition.
[0023] In some embodiments, the wavelet transform method is used for image denoising processing. Wavelet transform can decompose the image into multiple frequency bands, removing high-frequency noise while retaining low-frequency information (such as the background). Specifically, the wavelet transform formula is: Where: W ψ (x,a,b) is the output result of the wavelet transform. It represents the convolution result of the image signal f(t) and the wavelet basis function ψ at position b and scale a; f(t) is the signal of the input image; is the wavelet basis function, a is the scale factor, and b is the position parameter; is the normalization coefficient.
[0024] In this way, wavelet transform can effectively remove noise and highlight the key information in the image.
[0025] The edge detection algorithm further strengthens the text area of the image, making the subsequent text recognition more accurate. Edge detection helps to extract the edge information of the image, making the contour of the text area clearer and enhancing the recognition degree of the text in the image.
[0026] Specifically, after the image preprocessing is completed, the image will enter the multi-scale feature extraction unit. The goal of multi-scale feature extraction is to extract the features of the image from different scales to adapt to text areas of different sizes and resolutions. In this embodiment, the combination of convolutional neural network (CNN) and wavelet transform is used for multi-scale feature extraction.
[0027] A convolutional neural network is a deep learning model that processes images through multiple convolutional layers to automatically extract features from the images. During this process, the CNN performs convolutional operations using convolutional kernels of different sizes to extract features at different scales in the image. Specifically, after the input image passes through the convolutional layers (each layer using a convolutional kernel of a different size), feature maps at different scales are obtained.
[0028] In addition, the wavelet transform technique is adopted to further extract high-frequency and low-frequency information from the image. The high-frequency information is mainly used to capture details in the image (such as text edges), while the low-frequency information can provide the overall structure of the image. This multi-scale feature extraction method can effectively improve the recognition rate of the detailed parts in the image, especially in the context of complex image backgrounds.
[0029] After the image feature extraction is completed, the text recognition unit will convert the extracted features into editable text data. The key to text recognition lies in processing the character sequence in the image, and this serialization process is best achieved through a recurrent neural network (RNN). The RNN is a neural network suitable for processing sequence data and is used in OCR to process the features extracted from the image and convert them into a sequence of characters or words.
[0030] The process of text recognition can be expressed by the following formula: where: is the text data output by OCR recognition (the predicted text content); CNN(I) represents the image features extracted through the convolutional neural network, and I is the input image; RNN is the recurrent neural network used to convert the features extracted by CNN into text output.
[0031] Specifically, the convolutional neural network (CNN) will first extract features from the input image I to obtain the feature map F(I), and then these feature maps are passed to the recurrent neural network (RNN), which generates the final text output according to the text sequence structure in the image. Through its internal memory mechanism, the RNN can effectively process the sequence information in the image and recognize the text content.
[0032] In addition, to improve the recognition accuracy, the text recognition unit also uses an end-to-end training method and trains through a large-scale dataset. In this way, the model can not only automatically learn the features in the image but also learn the sequential relationship between characters, improving the accuracy of text recognition.
[0033] In this embodiment, the OCR recognition module can effectively improve the accuracy of text recognition through multi-level processing of image preprocessing, feature extraction, and text recognition. Image preprocessing optimizes the input image through wavelet transform and edge detection, extracting key information; multi-scale feature extraction ensures that text at different resolutions and scales can be correctly recognized through CNN and wavelet transform; the text recognition part combines CNN and RNN, which can not only process local features but also smoothly process sequence information, improving the stability and accuracy of text recognition.
[0034] Through this comprehensive processing, the OCR recognition module can maintain high accuracy in complex application scenarios, such as low-quality images, handwritten texts, different languages, complex layouts, etc., and has good processing efficiency in applications with high real-time requirements.
[0035] In this embodiment, the OCR recognition module can effectively improve the accuracy of text recognition and adapt to various complex scenarios through the combination of image preprocessing, wavelet transform, convolutional neural network, and recurrent neural network. The image preprocessing part ensures the improvement of the input image quality, the feature extraction part guarantees the extraction of detailed information in the image, and the text recognition part enhances the recognition accuracy through the deep learning model. The entire OCR module can not only recognize ordinary printed text but also handwritten text, text in different formats and languages, and has broad application prospects.
[0036] This solution can be compatible with various different input image types, not limited to clearly printed text, but also able to process complex images such as low-quality, blurred, and handwritten ones, greatly broadening the application scope of OCR technology.
[0037] The text optimization and information screening module is used to optimize the text data output by OCR; The text optimization and information screening module effectively reduces redundant content through high-order entropy calculation and information gain analysis, ensuring that the most relevant information is retained, thereby improving the efficiency and quality of information extraction.
[0038] In this embodiment, the text optimization and information screening module includes two core units: the high-order entropy calculation unit and the information gain analysis unit. These two units work together to effectively screen the text after OCR recognition, enabling subsequent analysis to focus on important information. First, the system performs high-order entropy calculation on the text to evaluate the amount of information in each section and screens out the most important part according to the entropy value. Then, the system further uses information gain analysis to identify the text parts that contribute significantly to the target task.
[0039] In some embodiments, the text optimization and information screening module first processes the text data output by OCR using the high-order entropy calculation unit. High-order entropy is an information-theoretic method used to measure the complexity or uncertainty of information. In the present invention, high-order entropy is used to calculate the "information density" of each text paragraph, thereby screening out the parts with high information value.
[0040] The high-order entropy calculation formula is as follows: H(X) = -∑ i p(x i ) log p(x i ); Where: H(X) is the information entropy of text paragraph X, representing the complexity or amount of information of this paragraph of text. The higher the entropy value, the more complex the content of this paragraph of text and the greater the amount of information; a lower entropy value indicates less information in this paragraph of text; p(x i ) is the probability distribution of the i-th word or character in text X, representing the frequency of occurrence of this word or character in this paragraph of text; log p(x i ) is the logarithmic probability of the i-th word in the text, used to measure the amount of information. In information theory, the lower the probability, the larger the corresponding logarithm value, thus bringing more "information".
[0041] The purpose of high-order entropy calculation is to evaluate the relative complexity of text paragraphs. A higher entropy value usually means that this paragraph of text contains more information and can provide more valuable content for subsequent tasks (such as information extraction, text summarization, etc.).
[0042] After high-order entropy calculation, the text optimization and information screening module further uses the information gain analysis unit to screen the text. Information gain is an important measure of the importance of a certain feature for a classification task. In the present invention, the goal of information gain analysis is to optimize the text content according to the contribution degree of text paragraphs to the final task (such as information extraction or summary generation). A high value of information gain means that this paragraph of text contains information that makes a great contribution to the task.
[0043] The information gain calculation formula is: Where: IG(D, A) is the information gain, representing the contribution size of feature A to dataset D; H(D) is the entropy of dataset D, representing the overall uncertainty in the dataset; represents the weighted average entropy of dataset D under different values of feature A, v is the value of feature A, D v is the subset divided according to feature v, |D v | is the size of this subset, |D| is the size of the original dataset; H(D v ) represents that when feature A is v, dataset D vThe entropy represents the complexity or uncertainty of the subset.
[0044] Information gain helps identify the most informative paragraphs by weighting different parts of the text. Through this method, the system can remove low-information and irrelevant parts, ensuring that only the content that contributes significantly to the task is retained for subsequent analysis.
[0045] After high-order entropy calculation and information gain analysis, the text optimization and information screening module outputs optimized text data. These data are screened to retain the most important and relevant information, removing irrelevant or redundant content to ensure that the subsequent large language model can perform semantic analysis and information extraction more efficiently.
[0046] Specifically, the optimized text reduces irrelevant words, sentences, or paragraphs while strengthening the key points that contribute to the task. In this way, the subsequent large language model can perform effective semantic analysis and generate the required key information or summary when the input data is more concise and accurate.
[0047] Through high-order entropy calculation and information gain analysis, the text optimization and information screening module can significantly improve the information extraction efficiency of the system. This module can efficiently screen out the most valuable information from large-scale OCR outputs, reduce redundancy, and ensure the conciseness and high quality of the text. This optimization method is crucial for improving the efficiency and accuracy of the subsequent processing modules (such as the large language model analysis module).
[0048] For example, in some embodiments, when processing text containing a large amount of irrelevant information or noise, the system can quickly identify and remove this irrelevant content. By reducing redundancy, the information optimization module ensures that the subsequent analysis module can focus on the most informative parts of the text, thereby improving the efficiency and quality of information extraction.
[0049] In some embodiments, the text optimization and information screening module can also be used in combination with other natural language processing techniques. For example, the system can combine named entity recognition (NER) to further enhance the text screening process. Named entity recognition can help further filter out content irrelevant to the target task by identifying key information such as people's names, place names, and organization names in the text.
[0050] In addition, in complex text scenarios, information gain analysis can be used in combination with deep learning models to understand the key information of the text from a higher dimension. This deep learning method can not only rely on traditional frequency statistics during text optimization but also intelligently judge which parts are most important for the task through a trained model.
[0051] The text optimization and information screening module can not only efficiently screen out the key information in the text through high-order entropy calculation and information gain analysis, but also eliminate redundant content while retaining valid information. Through the optimization of this module, the subsequent large language model can perform accurate and fine-grained semantic analysis based on high-quality input, improving the overall performance and effect of the information extraction system.
[0052] The large language model analysis module is used to perform semantic understanding on the optimized text; The large language model analysis module is responsible for further performing semantic understanding, information extraction, and summary of key information on the optimized text data. This module not only performs intelligent analysis on the text, but also generates high-quality structured or unstructured summaries to provide accurate data support for subsequent applications.
[0053] In this embodiment, the large language model analysis module mainly analyzes the text content through the self-attention mechanism and the Transformer-based model, and combines the pre-trained large language model (such as BERT, GPT, etc.) to complete the tasks of semantic understanding and key information extraction.
[0054] Generally, the main task of the semantic understanding unit is to input the optimized text into the locally deployed large language model for in-depth semantic analysis of the text. In some embodiments, the large language model based on the Transformer structure can handle long-distance dependencies in the text and generate accurate semantic representations. The core of the Transformer architecture is the self-attention mechanism, which understands the semantic structure of the text by calculating the relevance of each part of the text.
[0055] The formula for the self-attention mechanism is: where: Attention(Q, K, V) represents the output of the self-attention mechanism, which is a weighted sum of the input vectors, and the weights are calculated from the relevance between the query vector (Q) and the key vector (K); Q is the query vector (Query), representing the features of the current text unit. It is usually obtained through the embedding (Embedding) and linear transformation of the input text; represents the inner product of the query vector (Q) and the key vector (K) divided by the square root of its dimension; K is the key vector (Key), representing the features of other text units. Like the query vector, the key vector is also obtained through the input text embedding and linear transformation; V is the value vector (Value), representing the feature information associated with the key vector. The value vector contains the specific information of the text content and is the key to understanding the text semantics; d kis the dimension of the key vector. This parameter controls the scale when calculating the attention weights, preventing the inner product from being too large or too small, which may affect the gradient calculation; the softmax function is used to transform the inner product result into a probability distribution, making the sum of all weights equal to 1.
[0056] By calculating the relative importance between text units through the self-attention mechanism, the model can intelligently identify the core part of the text and effectively extract the meaning of each text part for the entire sentence or paragraph.
[0057] After semantic understanding is completed, the large language model will then generate a structured or unstructured text summary based on the weights of each text part. Specifically, the text summary generation module relies on the self-attention mechanism to calculate the importance of each part in the text and generates a concise text summary based on these calculation results. This process can be described by the following formula: Where: is the text summary, representing the generated structured or unstructured summary. This summary contains the key information in the text and is a simplified version of the original text; w i is the weight of the i-th text unit, representing its importance when generating the summary. The weight is calculated by the self-attention mechanism and reflects the contribution of this unit to understanding the entire text; x i is the content of the i-th text unit, usually words, phrases or sentences in the text. Each unit represents a part of the information in the text.
[0058] Through the above calculations, the system can perform weighted processing on each text unit according to its weight and finally generate a concise text summary. In practical applications, the text summary can be a structured report or an unstructured description, depending on the requirements of the task.
[0059] In addition to generating summaries, the large language model analysis module is also responsible for extracting key information from the text, such as named entities, time, location, etc. In some embodiments, the model extracts these key data points from the text through named entity recognition (NER) technology.
[0060] The process of named entity recognition can be represented by the following formula: Where: is the set of extracted entities, representing the named entities recognized from the text T, such as person names, place names, organization names, etc.; NER(T) is the named entity recognition (NER) model that processes the text T to identify and extract the named entities in the text.
[0061] Through this process, the system can extract task-related key information from the optimized text, providing necessary data support for subsequent analysis or applications. For example, the system can identify dates, locations, events, etc. in the text, thus providing data for applications such as building timelines, generating reports, or sentiment analysis.
[0062] The large language model analysis module can provide efficient and accurate results for subsequent information extraction tasks by performing in-depth semantic understanding of the optimized text and generating structured or unstructured summaries. The self-attention mechanism and Transformer architecture can help the model identify key information in the text, enhancing the system's performance when dealing with long or complex texts.
[0063] Through this module, the system can not only generate concise summaries but also extract various types of key information from the text, such as named entities, sentiment colors, text classification, etc. The generation of text summaries and information extraction enable the system to provide valuable information outputs in multiple application scenarios, such as document management, market analysis, news summaries, etc.
[0064] As an extended technical solution, the large language model analysis module can also be used in combination with other deep learning technologies. For example, it can be combined with image-text multimodal learning, enabling the system to not only process text information but also perform multimodal data analysis by integrating image information. The combination of these technologies will greatly expand the application scenarios of the present invention, thus enabling the system to have stronger analysis capabilities and adaptability.
[0065] In addition, the model can be fine-tuned according to the needs of different tasks. In some specific applications, the system can be further optimized according to specific domains (such as medical, legal, financial, etc.) to provide more professional and accurate information extraction.
[0066] Through the input of optimized text, this module can provide higher-quality text summaries or key information extraction, providing users with accurate reports, analysis, or data support.
[0067] The quantum computing acceleration and optimization module is used to optimize the information transfer process between OCR and the large language model; This module utilizes the advantages of quantum computing to accelerate information flow, optimize the data transmission path, and improve the system operation efficiency. Quantum computing can effectively reduce the latency of data processing and information transmission, thus ensuring the efficient collaboration between the OCR recognition module, the large language model analysis module, and the text optimization and information screening module.
[0068] In this embodiment, the quantum computing acceleration and optimization module mainly accelerates information processing and optimizes the calculation path through the Quantum Fourier Transform (QFT) and quantum search algorithms. These quantum technologies provide an acceleration mechanism for information flow and greatly improve the computing efficiency of the system. Especially when dealing with big data, they show obvious advantages.
[0069] The Quantum Fourier Transform is an important algorithm in quantum computing. It can efficiently perform frequency domain conversion on a quantum computer and has a significant acceleration effect. The Quantum Fourier Transform accelerates the calculation of the traditional Fourier Transform through quantum parallelism and can complete the Fourier Transform with exponential time complexity.
[0070] The mathematical formula of the Quantum Fourier Transform is as follows: where: QFT|x> is the state after the Quantum Fourier Transform, representing the frequency domain representation of the input quantum bit |x> after the Quantum Fourier Transform; |x> is the state of the input quantum bit, usually the data to be processed, represented as a binary state containing the information to be processed; is the normalization coefficient to ensure that the transformed quantum state conforms to the probability constraint. N is the total number of quantum bits, representing the dimension of quantum computing; represents the summation over all possible output states |k>. k is the output index, representing the transformed frequency domain state; e 2πikx / N is the phase factor in the Fourier Transform, where k is the current index, x is the input data, and N is the number of quantum bits. This term modulates the frequency domain data and determines the frequency domain characteristics of the data.
[0071] The Quantum Fourier Transform accelerates the calculation of the traditional Fourier Transform through quantum superposition states and interference effects, significantly improving the information processing speed. Especially in frequency domain processing and signal analysis, the Quantum Fourier Transform can significantly accelerate the data conversion speed through quantum parallel processing, thereby improving the efficiency of the system in large-scale data processing.
[0072] Quantum search algorithms, such as the Grover algorithm, are mainly used to quickly search for the optimal solution in a large amount of data. In classical computing, searching requires traversing all possible paths, while quantum search algorithms utilize quantum superposition and interference effects to search for the target data in a more efficient way. The computational complexity of the Grover algorithm is significantly lower than O(N) of the classical search algorithm.
[0073] The mathematical formula of the quantum search algorithm is: where: G(f) is the output of the quantum search algorithm, representing the search result; is a normalization coefficient to ensure that the search results comply with the quantum probability constraints; denotes the summation over all states |x> in the search space, where x1 represents each state in the search and N1 is the size of the search space; (-1) f(x) is a decision function for marking the target data. f(x) is a Boolean function used to determine whether the currently searched state x meets the target condition. If f(x) is true, this term is -1; otherwise, it is 1.
[0074] Function and significance of the formula: The quantum search algorithm accelerates the search process through quantum superposition and interference effects. Classical algorithms require a time complexity of O(N) to traverse all possible paths, while the quantum search algorithm can find the target path within time. This characteristic of quantum computing enables the information flow to converge quickly when processing large-scale data, greatly improving the data processing efficiency.
[0075] Through the quantum Fourier transform and the quantum search algorithm, the quantum computing acceleration and optimization module can significantly improve the processing speed of the information flow and optimize the data transmission path. Specifically, the quantum Fourier transform performs efficient frequency-domain conversion on the data, accelerating information processing; while the quantum search algorithm optimizes the data stream transmission by parallelly searching for the optimal path, reducing the search time in traditional computing.
[0076] Improving the information flow processing speed: Quantum computing can process a large amount of data in a shorter time, significantly enhancing the speed of information transfer. Especially in large-scale data processing scenarios, quantum computing can reduce the processing time and improve the system response speed.
[0077] Optimizing the data transmission path: Through the quantum search algorithm, the system can quickly find the optimal data transmission path among a large number of data streams, thus avoiding the inefficient path search problem in traditional computing methods.
[0078] Reducing the computing latency: The parallel processing ability of quantum computing enables the data processing process to be no longer restricted by the serialized computing in traditional computing architectures, reducing the computing latency and improving the working efficiency of the entire system.
[0079] As an extended technical solution, the quantum computing acceleration and optimization module can also be used in combination with other high-performance computing architectures. For example, in combination with the acceleration computing architectures based on GPU (Graphics Processing Unit) or TPU (Tensor Processing Unit), the combination of quantum computing and traditional computing can, while ensuring the advantages of quantum computing, avoid the limitations of quantum computing in specific scenarios. In this way, the advantages of traditional computing and quantum computing can be fully utilized to further improve the computing efficiency.
[0080] In addition, this module can also be combined with classical computing methods to adopt a hybrid computing model, processing some simple tasks in the traditional computing framework and complex tasks in the quantum computing framework, so as to balance computing accuracy and efficiency.
[0081] The quantum Fourier transform improves the efficiency of the subsequent information analysis process by accelerating the frequency domain conversion; the quantum search algorithm optimizes the data transmission path and reduces the processing time in the information flow. The application of these quantum computing technologies not only improves the response speed of the system but also enables it to maintain high operating performance in big data scenarios.
[0082] The mobile information extraction method based on OCR and the local large language model described below can be correspondingly referred to the mobile information extraction system based on OCR and the local large language model described above.
[0083] Please refer to the atta Figure 6 , the present invention also provides a mobile information extraction method based on OCR and the local large language model, including the following steps: S1. Transmit the input image data to the OCR recognition module; S2. Perform image preprocessing on the image data, including denoising, enhancement, and edge detection processing; S3. Use the OCR recognition module to extract image features and convert them into text data; S4. Perform high-order entropy calculation and information gain analysis on the text data output by OCR, and screen out the most informative text part; S5. Transmit the optimized text to the large language model analysis module for semantic understanding and generate a summary; S6. During the transmission process between OCR and the large language model, use quantum computing to accelerate information flow and optimize the data transmission path; S7. Output the generated text summary or the extracted key information.
[0084] The method of this embodiment can be used to implement the above system embodiment, and its principle and technical effect are similar, so they will not be elaborated here.
[0085] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A mobile information extraction system based on OCR and a local large language model, characterized in that, It includes: An OCR recognition module for converting the input image data into text data; A text optimization and information screening module for optimizing the text data output by OCR; A large language model analysis module for semantic understanding of the optimized text; A quantum computing acceleration and optimization module for optimizing the information transmission process between OCR and the large language model.
2. The mobile information extraction system based on OCR and local large language model according to claim 1, wherein The OCR recognition module includes: An image preprocessing unit for denoising, enhancing, and edge detecting the input image, specifically including: Performing multi-band decomposition on the input image using wavelet transform to remove noise; Extracting the edge information of the image through an edge detection algorithm to identify the text region; A multi-scale feature extraction unit for extracting image features at multiple scales through a convolutional neural network and wavelet transform; A text recognition unit for converting the extracted image features into text data.
3. The mobile information extraction system based on OCR and local large language model according to claim 1, characterized in that The text optimization and information screening module includes: A high-order entropy calculation unit for calculating the high-order entropy of the text data output by OCR and screening key information according to the entropy value, specifically including: Calculating the entropy value for each paragraph in the text and screening out the key information in the text based on the entropy value; Using information entropy to reduce redundant and irrelevant data and only retaining the parts that contribute significantly to the target task; An information gain analysis unit for performing information gain analysis on the text and selecting the parts that contribute significantly to the target task.
4. The mobile information extraction system based on OCR and local large language model according to claim 1, wherein, The large language model analysis module includes: A semantic understanding unit for semantic understanding of the text according to the locally deployed large language model; A text summarization unit for generating structured or unstructured text summaries, specifically including: Calculating the weights of each part in the text through a self-attention mechanism; Calculating the core information of the text based on the weights and generating a structured or unstructured summary.
5. The mobile information extraction system based on OCR and local large language model according to claim 1, wherein The quantum computing acceleration and optimization module includes: A quantum Fourier transform unit for accelerating the frequency domain conversion of information; A quantum search algorithm unit for quickly finding the optimal data transmission path in the data stream.
6. The mobile information extraction system based on OCR and local large language model according to claim 1, characterized in that The image preprocessing unit is used for image denoising processing of the input image, and the denoising processing includes: Performing multi-band decomposition on the input image using wavelet transform to remove noise; Extracting the edge information of the image through an edge detection algorithm to identify the text region.
7. The mobile information extraction system based on OCR and local large language model according to claim 2, wherein The multi-scale feature extraction unit includes: Performing layer-by-layer convolution operations on the input image through a convolutional neural network to extract local features at different scales; Extracting high-frequency and low-frequency information in the image through wavelet transform for easy extraction of text features.
8. The mobile information extraction system based on OCR and local large language model according to claim 1, characterized in that, The high-order entropy calculation unit of the text optimization and information screening module is used for calculating the high-order entropy of the text data output by OCR, and the high-order entropy calculation method includes: Calculating the entropy value for each paragraph in the text and screening out the key information in the text based on the entropy value; Using information entropy to reduce redundant and irrelevant data and only retaining the parts that contribute significantly to the target task.
9. The mobile information extraction system based on OCR and local large language model according to claim 1, wherein The semantic understanding unit of the large language model analysis module is used for semantic understanding of the text using a large language model based on the Transformer structure and generating a text summary, and the text summary generation method includes: Calculate the weights of each part in the text through the self-attention mechanism; Calculate the core information of the text according to the weights and generate a structured or unstructured summary.
10. A mobile information extraction method based on OCR and a local large language model, characterized in that Use the mobile information extraction system based on OCR and local large language model according to any one of claims 1-9, comprising the following steps: S1. Transmit the input image data to the OCR recognition module; S2. Perform image preprocessing on the image data, including denoising, enhancement and edge detection processing; S3. Use the OCR recognition module to extract image features and convert them into text data; S4. Perform high-order entropy calculation and information gain analysis on the text data output by OCR, and screen out the most informative text parts; S5. Transmit the optimized text to the large language model analysis module for semantic understanding and generate a summary; S6. During the transmission process between OCR and the large language model, use quantum computing to accelerate information flow and optimize the data transmission path; S7. Output the generated text summary or the extracted key information.
Citation Information
Patent Citations
Big data processing method based on quantum computing
CN108320027A
Text abstract generation method and device, equipment and storage medium
CN116304005A
Text abstract generation method and device based on large language model, equipment and medium
CN118069830A
Robot path planning method and device based on quantum search and reinforcement learning, and medium
CN118999570A
Cited By
Aspect-level sentiment analysis method based on multiple grammar and multiple frequencies
CN120805937A
Geological exploration handwritten revealing board OCR (optical character recognition) method and system
CN121527784A